An automated system needs two things from a model: an answer, and a reason to believe it. Most LLM integrations only get the first. The answer arrives in the same assured tone whether the model is right or guessing, and the code acts on it either way.
Jev gives you the second half explicitly. Every Choice and Score answer carries a probability distribution and a confidence value, and every Noul answer is itself a probability. This article explains the difference between the two numbers and how to turn them into routing rules. If you haven't met the question types yet, read Choice, Score and Noul first, or start with what Jev is.
Probability: what the answer is
For a Choice, probabilities is the distribution across your options; choice is simply the one with the highest probability. For a Score, it is the distribution across your levels. For a Noul, the single noul value is the probability that the answer is yes.
Jev is trained for calibrated probabilities. TypeSafe's AI primer explains the goal: across many predictions, outcomes assigned 0.8 should happen about 80% of the time. That is a property of groups of predictions. It is not a promise about any single answer, which is exactly why the next number matters.
Confidence: whether the answer is worth acting on
confidence collapses the shape of the distribution into one number from 0 to 1. Probability concentrated on one outcome means high confidence; probability spread across several means low confidence. TypeSafe computes it for you on every Choice and Score answer, and still returns the full probabilities so you can use a different measure if your domain needs one. The Confidence page covers the details.
A real example from TypeSafe's Choice docs makes the difference concrete. The ticket reads: "Shoes arrived two weeks late and in the wrong size. Also I see two charges of $120 on my card. What are you going to do about this?"
| Question | Top answer | Probability | Confidence |
|---|---|---|---|
| Which team should handle this? | returns |
0.61 (billing 0.35) | 0.42 |
| Why is the customer returning it? | wrong_size |
1.0 | 1.0 |
| What does the customer want to happen? | refund |
0.40 (replacement 0.34, exchange 0.24) | 0.20 |
| What is the customer's tone? | frustrated |
0.84 | 0.76 |
Look at the third row. An integration that reads only the top answer would issue a refund. But the customer never says what they want: the double charge suggests money back, the wrong size suggests a swap. A confidence of 0.20 is Jev saying "I don't know", and that is the most useful thing it could say. The first row tells a similar story: this ticket genuinely involves more than one team.
Three paths, not two
TypeSafe's recommended starting pattern splits confidence into three ranges:
- High confidence: act. The model has a clear read; proceed without a person.
- Medium confidence: proceed with caution. Ask the user to confirm, flag for review, or gather more information first.
- Low confidence: don't act. Route to a human, ask for clarification, or fall back to a different system, such as a reasoning model.
For Noul answers, the equivalent is a band around 0.5. TypeSafe's Noul page suggests 0.5 as the cut when yes and no are equally cheap to act on, a higher bar when a false yes is expensive (paging someone, issuing a refund), and a lower one when a missed yes is expensive (a safety flag). Values in the middle can go to a person instead of either code path.
Thresholds scale with risk
A threshold is not one number per system; it is one number per action. TypeSafe's confidence-gated routing pattern uses a voice banking example:
action = response.answers["intent"]
if action.confidence < 0.6:
route_to_human(message) # genuinely unsure: don't guess
elif action.choice == "check_balance":
show_balance(account_id) # low stakes: a wrong screen is recoverable
elif action.choice == "approve_transfer":
if action.confidence > 0.85:
confirm_then_execute(account_id)
else:
ask_user_to_confirm(account_id)
Reading a balance at 0.6 is fine; the worst case is an unnecessary read-out. Moving money needs more than 0.85, or the system asks first. The code, not the model, encodes your risk tolerance.
How we set thresholds at Dryhurst
Our internal Jev tools all follow one rule: Jev decides, a person or an LLM writes, and when Jev isn't sure it doesn't get the last word. The concrete cut-offs we use:
- Text triage. A Choice or Score confidence below about 0.6, or a Noul between about 0.35 and 0.65, counts as "not sure". Those items get read and decided by a person or a bigger model, and we say when that happened. See triage at a fraction of a cent.
- Model routing. Our Claude Code model router only hands a message to a smaller model when Jev is at least 60% confident about its size. Anything less stays on the main model, which is the safe default.
- Skill picking. Our skill picker loads a skill only at 60% or more, and the question always offers "none of these".
These are starting points, not truths. The right numbers depend on what a mistake costs you, and they should be tuned on your own data.
Keep thresholds honest over time
Three habits keep confidence gating trustworthy in production:
- Pin the model version you tuned on. Aliases like
jev-latestmove when a new release ships, so answers can change without a change on your side. TypeSafe's Models page recommends pinning a versioned ID when you have tuned thresholds against it, and every response reports which model answered. - Don't carry thresholds across question types. A cut tuned on a Noul does not transfer to a Choice. A Choice is relative (which option wins); a Noul is absolute (is this true at all).
- Log the full distribution, not just the answer. When an escalation rate drifts,
probabilitiestells you whether the inputs changed or your options stopped fitting them.
If you are designing escalation paths for an AI feature, or want an outside view on where your thresholds should sit, that is the kind of work our AI consulting practice does. Get in touch.
FAQ
What is the difference between confidence and probability in Jev?
Probability tells you how likely each answer is; the top one becomes the choice or drives the score. Confidence summarizes how concentrated that distribution is, which tells your code whether the answer is clear enough to act on.
What confidence threshold should I use?
It depends on the cost of a wrong action. TypeSafe's examples use a floor around 0.5 to 0.6 for anything automatic and a much higher bar, such as 0.85, for risky actions like moving money. Start conservative and tune on your own data.
Does the Noul primitive return a confidence value?
No. A Noul returns only noul, the probability of yes. Treat values near 0.5 as uncertain and send a middle band to review, with the band's edges set by what each kind of mistake costs.
Where should low-confidence answers go?
To whatever is safest for that action: a person, a clarifying question to the user, or a slower reasoning model with the full context.