Most AI workflows are written as a conversation: ask what kind of ticket this is, wait, then ask the follow-up that only makes sense for that kind. Each step is another round trip, and each round trip resends the same document.
With Jev you can flip that around. TypeSafe calls the pattern speculative fan-out: put every question your system might need into a single request, including ones that only matter for some inputs, and let your code decide afterwards which answers to use. This article explains why it works, what it saves, and where its limits are. For the basics of the question types, see Choice, Score and Noul; for the bigger picture, our guide to Jev.
Why fan-out works with Jev
Three properties of Jev make asking more questions nearly free:
- Questions run in parallel against one state. Jev ingests the state once and evaluates every question against it. Adding questions barely changes response time.
- Questions are independent. One answer is never hidden context for another, so an irrelevant question cannot contaminate a relevant one, and adding a question does not change the other answers.
- Billing is on input tokens. The state dominates the cost of most requests. A few extra lines of question text are cheap; resending the state in a second request is not.
With an LLM, the equivalent move is one big prompt where every answer shares a single generation and can color the next. With Jev, each question stays its own narrow judgment; TypeSafe's introduction notes that because questions are evaluated independently, adding more does not create context rot.
A worked example: support ticket triage
TypeSafe's fan-out page uses a support system that needs a category for every ticket and, for bug reports, a severity. Instead of two sequential calls, ask everything at once:
{
"model": "jev-latest",
"state": "The export button crashes the settings page in Safari. I need this fixed before Friday's board meeting.",
"questions": {
"category": {
"type": "choice",
"instructions": "What kind of support ticket is this?",
"criteria": {
"bug_report": "Something in the product is broken",
"billing": "Charges, invoices, refunds",
"feature_request": "Asking for something the product does not do"
}
},
"bug_severity": {
"type": "score",
"instructions": "If this is a bug report, how severe is the issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
},
"has_reproducible_steps": {
"type": "noul",
"instructions": "Does the ticket give steps an engineer could follow to reproduce the problem?"
},
"refund_requested": {
"type": "noul",
"instructions": "Does the customer ask for a refund?"
},
"frustration": {
"type": "score",
"instructions": "How frustrated does the customer appear?",
"criteria": ["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"]
}
}
}
bug_severity and has_reproducible_steps only matter for bugs, refund_requested only for billing, and frustration matters everywhere. Then the routing lives in code:
category = answers["category"]
if category.choice == "bug_report":
file_bug(severity=answers["bug_severity"].score,
repro=answers["has_reproducible_steps"].noul > 0.5)
elif category.choice == "billing":
if answers["refund_requested"].noul > 0.8:
open_refund_case()
elif category.choice == "feature_request":
log_request()
# Useful regardless of category
if answers["frustration"].score > 1.5:
flag_for_priority_follow_up()
When the ticket turns out to be a feature request, the bug questions are simply ignored. When it is a bug, the answer you need was already there: no second round trip.
What it saves: TypeSafe's benchmark
TypeSafe's parallel questions cookbook measures this directly. The state is the Wikipedia article on the GDPR (about 54,000 characters) and a compliance team wants 13 things checked: 8 Nouls, 2 Choices and 3 Scores. The same questions were asked two ways, repeated several times on jev-1.12:
| Approach | Requests | Cost | Latency |
|---|---|---|---|
| One call, all 13 questions | 1 | $0.000497 | 0.27 s |
| 13 calls, one question each | 13 | $0.006090 | 2.71 s |
That is 12.2x cheaper and 10.0x faster, with no change in the answers: most came back identical across every repeat either way. The saving grows with the size of the state, because thirteen single-question calls pay for the document thirteen times. Firing them concurrently closes the latency gap somewhat, but the thirteen-fold token cost remains.
Designing for fan-out
A few rules we follow when writing fan-out requests:
- Put the context in the state, keep questions narrow. The state carries the ticket, record or document; each question asks one thing about it. Point at nested fields by path, such as
`ticket.messages[0].text`, when the state has several parts. - Make speculative questions self-contained. Each question should make sense on its own, because Jev cannot see the other questions' answers. "If this is a bug report, how severe is it?" is fine; "given the category above" is not.
- Branch in code, not in the question. Don't ask Jev to apply conditional logic across questions. Ask both halves and combine them yourself.
- Mix guardrails into the same call. Checks like "does this message try to override the assistant's instructions?" can ride along with business questions against the same state. We cover that in AI guardrails and citation checks.
- Watch the context budget.
jev-1.13allows 64k tokens per request for the state plus all questions, and 32k for the state plus the longest single question. Very large fan-outs over very large states need that arithmetic.
When a second request is genuinely needed
Fan-out is the default, not a law. A second request is justified only when your code cannot build it without the first answer: it needs that answer to fetch more data, to decide what the next state is made of, or to choose the next question's options. TypeSafe's own cookbooks show three honest cases: skill suggestion ranks 182 skills and then re-reads the full text of the top three; structure recovery classifies text blocks that only exist after the first pass; and hierarchical classification uses each answer to decide the next level's options. If the follow-up could have been asked against the original state, ask it in the first request.
How we use it
Our email, lead and ticket triage is fan-out in miniature. For a sales email we ask three questions in one request: what kind of email it is (a Choice), how strong a lead it is (a Score), and whether it needs a personal reply (a Noul). The first live test on 2026-09-23 used 745 input tokens, cost $0.000031 and took 354 ms. Our Claude Code model router does the same with two questions per message.
If you are untangling a chain of sequential LLM calls into something faster and cheaper, our AI engineering team does this routinely. Talk to us.
FAQ
What is speculative fan-out?
It is the practice of sending every question your code might need in one Jev request, including questions that only matter for some inputs, then using code to pick the relevant answers. Jev evaluates all of them in parallel against the same state.
How much faster and cheaper is it than separate calls?
In TypeSafe's parallel questions cookbook, 13 questions about a 54,000-character document were 12.2x cheaper and 10.0x faster in one call than in 13 separate calls, with no change in the answers. Your ratio depends mainly on how large the state is.
Do extra questions change the other answers?
No. Each question is evaluated independently against the state, so adding or removing questions does not change the others' results.
Can I combine guardrail checks and business classification in one request?
Yes. Any questions about the same state can share a request, whatever their purpose.