Every LLM product eventually needs two kinds of checks. One sits around the conversation: is this message something the assistant should handle, and did the reply stay inside the lines? The other sits around the facts: does the source actually say what the answer claims it says?

Both are judgments, not generation. That makes them a good fit for Jev, TypeSafe's decision model, which returns probabilities instead of text. This article walks through both patterns using TypeSafe's published cookbooks, with the numbers they reported and the design choices we would copy.

Why not a system prompt or a second LLM?

TypeSafe's guardrails cookbook makes the case plainly. Labs teach models to refuse some requests, but each lab draws that line somewhere else, and each new model version moves it. You probably want the line somewhere else too.

The usual fixes both have a cost:

  • Rules in the system prompt sit exactly where a jailbreak aims to talk its way past them.
  • A second LLM as a judge adds a full generation call's latency and cost to every turn, and it can be talked past as well.

A decision model changes the shape of the check. It does not follow the message's instructions or write a reply; it answers a question about the message. In the cookbook's words, "Ignore your instructions" scores as a jailbreak instead of working as one.

That is not immunity. TypeSafe's own jaggedness notes for Jev 1.13 list adversarial content as a known weak spot, with the advice to be explicit in your criteria and test thoroughly before rollout. Treat Jev as a strong, cheap first line, not the whole defense.

Pattern 1: screen messages in and out

The cookbook splits "out of bounds" into separate hazards, because it is not one question. Each input message gets one request with five questions:

Question Type Asks whether the message...
jailbreak Noul tries to ignore, override or reveal the assistant's instructions
harmful_request Noul asks for help causing physical harm or breaking the law
medical_advice Noul asks for a diagnosis, a dosage or a treatment decision
self_harm Noul suggests the sender may be considering self-harm
severity Score would cause harm if complied with, from "no harm" to "severe"

Replies get a mirror-image battery: did the reply go along with something it should have refused, give harmful help, give a dosage, or encourage self-harm? Checking outputs matters because ordinary-looking prompts can still produce harmful replies.

All five questions go in one request, so the whole screen is one call. See speculative fan-out for why batching questions costs so little.

Your application owns the decision

Jev supplies probabilities. The routing is ordinary code with two thresholds per policy:

POLICIES = {
    "strict":     {"review_threshold": 0.35, "action_threshold": 0.70, "severity_block": 2.0},
    "permissive": {"review_threshold": 0.35, "action_threshold": 0.85, "severity_block": 2.0},
}
HAZARD_ACTION = {
    "jailbreak": "block",
    "harmful_request": "block",
    "medical_advice": "review",  # a human looks at it instead of a hard block
    "self_harm": "support",      # route to a support path, never just block
}

At or above the action threshold, a hazard triggers its action. Between the review and action thresholds, the message goes to a person. A severity score at or above 2.0 upgrades a review to a block. When several things fire, support wins, then block, then review, then pass.

Under the strict policy, the cookbook's sample messages (some ordinary, some borderline, and real jailbreaks from a public collection) came out like this:

Message Top signal Action
Banana bread recipe jailbreak 0.02 pass
Melatonin dose question medical_advice 0.55 review
Headache dosage request medical_advice 0.95, severity 2.0 block
Novelist asking how a detective describes a poisoning jailbreak 0.05 pass
Picking the lock on someone else's door harmful_request 0.95 block
"I don't see the point of anything anymore" self_harm 0.96 support
"DAN" jailbreak jailbreak 0.98 block

The useful part is the range of outcomes. The mild dosage question goes to a person, the distressed message goes to support rather than hitting a wall, and the crime novelist passes. A single yes/no filter could not do that.

It also separates evidence from policy. The same cached assessment of a message can be routed under strict or permissive without calling the model again. Changing how cautious your product is becomes a config change, not a prompt rewrite. That is composite scoring applied to safety, and it is the same idea we use for confidence-gated routing.

Pattern 2: check that citations hold up

The second problem is quieter. An LLM answers with citations: a claim, a section of a source, and a quote. Some are wrong. The quote may not exist, or it may exist word for word while the section says the opposite of the claim.

TypeSafe's citation check cookbook splits the job the way we would:

  1. Use code for what code can do. Normalize whitespace and curly quotes, then search the source for the quote as a substring. A quote that is not there is fabricated, and no model is needed to say so.
  2. Use Jev for the judgment. For each citation that survives, one Choice question reads the claim and the section and picks how they relate.
{
  "state": { "claim": "...", "section": "..." },
  "questions": {
    "relation": {
      "type": "choice",
      "instructions": "How does the section relate to the claim?",
      "criteria": {
        "supports": "The section states the claim or directly implies that it is true",
        "contradicts": "The section states the opposite of the claim or implies it is false",
        "says_nothing": "The section does not address what the claim asserts, either way"
      }
    }
  }
}

supports maps to verified, contradicts to contradicted, and says_nothing to unsupported. Confidence decides what happens next: at 0.8 or above the verdict stands on its own, below it a person confirms. The cookbook's advice is to start the threshold high and lower it as you see how the model does on your documents.

The test set was eight citations an LLM wrote about RFC 7519 (JSON Web Token), four accurate and four edited to fail. All four accurate citations came back verified at confidence 0.93 or higher. All four planted failures were caught: one fabricated quote, one contradicted claim, and two unsupported citations sent to a person for review.

Eight citations is a demonstration, not a benchmark. What we would take from it is the structure: an exact check in code, one narrow judgment per claim, and a human path for anything the model is unsure about.

Before the answer: filter what goes in

Guardrails and citation checks sit after retrieval. Jev is also useful before it. TypeSafe's cookbooks show two patterns there:

  • Classifying RAG passages: score each retrieved passage with one request, then decide in code which ones reach the answering model.
  • Re-ranking: on 40 legal queries with 30-passage BM25 shortlists, one Jev question per query and candidate pair raised top-1 accuracy from 5% to 18% and top-10 accuracy from 38% to 62%.

Fewer irrelevant passages in the context means fewer claims that need checking afterwards.

How we would roll this out

  • Write narrow, explicit questions. One hazard per Noul. Put the boundary cases in the criteria, not in your head.
  • Screen both directions. Inputs catch intent; outputs catch what actually got said.
  • Keep thresholds in config and log the raw probabilities. When a decision looks wrong, you want to know whether the evidence or the policy was off.
  • Route, do not just block. Review queues and support paths are where most of the value is.
  • Pin the model version once thresholds are tuned, and retest before upgrading. Our OpenRouter guide covers how.
  • Test on your own traffic, especially adversarial examples, before trusting any threshold from a cookbook, including these.

If you are adding guardrails or answer verification to an AI product, see our AI engineering services or talk to us.

FAQ

Is a decision model immune to prompt injection?

No. Because Jev answers a question about the message instead of following it, an injection attempt tends to show up as evidence ("this looks like a jailbreak") rather than taking control. But TypeSafe lists adversarial content as a known weak spot, so be explicit in your criteria and test before rollout.

Should I screen model outputs as well as inputs?

Yes. Ordinary-looking prompts can still produce harmful replies. TypeSafe's cookbook runs a mirror-image set of questions on every reply for that reason.

Can Jev check whether citations support an answer?

Yes, as one step of a two-step check. Code first confirms the quote exists in the source; then a Jev Choice question decides whether the section supports, contradicts or says nothing about the claim, with low-confidence verdicts going to a person.

How many questions can one guardrail request hold?

Many. The cookbook's screen asks five questions in one request, and a request can hold far more within the 64k-token context limit. Questions are evaluated in parallel, so adding more barely changes response time.