Over the last few years, the default way to add AI to a product has been to send a prompt to a large language model and read what comes back. That default holds even when the product does not need any text at all. Classify this ticket. Is this message a refund request? Which of these twelve tools should run? We ask a text generator, then ask it to format its answer as JSON, then write code to parse and validate the JSON.

There is now a different kind of model built for exactly those moments. TypeSafe calls it a System One model, and Jev is the first one. This article compares the two kinds of model and gives a practical test for when not to generate text.

Two different contracts

The simplest way to see the difference is the contract each model offers your code.

An LLM takes text and produces text. You can constrain its format, but underneath it is generating a response for someone to read. The output is open-ended by design, which is its strength for writing, conversation and reasoning.

A decision model takes a state (your text and context) and a set of typed questions, and returns typed answers: a choice from options you define, a score on a scale you define, or the probability that a statement is true. It does not write replies, produce code or explain itself. Every answer is constrained to the options you supplied, so it cannot return a value outside your schema.

The name comes from Daniel Kahneman's Thinking, Fast and Slow: System 1 thinking is fast and intuitive, System 2 is slower and more deliberate. A decision model is built for the fast, focused judgment.

Decision model (Jev) Generative LLM
Output Typed answers with probabilities Text, optionally formatted as JSON
Answer space Only the options or levels you define Open-ended
Uncertainty Calibrated probabilities, plus confidence for Choice and Score Not a calibrated output by default
Billing Jev: $0.042 per million input tokens, output free Input and output tokens, on most APIs
Many questions about one input One request, evaluated in parallel Usually one prompt per question, or one long prompt
Good at Routing, classification, scoring, checks Writing, conversation, multi-step reasoning, code

Why training matters

The difference is not only in the interface. TypeSafe's AI primer describes three ways to post-train a pretrained model:

  • RLHF (reinforcement learning from human feedback) trains models to produce responses people prefer. It turned pretrained models into chatbots.
  • RLVR (reinforcement learning with verifiable rewards) produced reasoning models: strong at tasks like mathematics, but slower and more expensive.
  • RLCD (reinforcement learning for calibrated decisions) is TypeSafe's approach: return decisions and calibrated probabilities instead of text.

TypeSafe's argument is that preference training can reward sycophancy and confident-sounding hallucinations, because an answer can be compelling to a person without being reliable enough for unattended automation. Calibration is the opposite target: across many predictions, answers given a probability of 0.8 should be right about 80% of the time.

That last point is about groups of predictions, not a promise about any single answer. But it is what lets software act on the number. A calibrated probability can drive a threshold. A confident-sounding sentence cannot.

When not to generate text

Our rule of thumb: if the output of the AI step is going to be read by code rather than a person, question whether it should be text at all.

Reach for a decision model when your code needs to:

  1. Route a request to one of a fixed set of destinations, and know how sure the routing is. Our Claude Code model router is one example.
  2. Pick one item from a list, such as a tool, a skill or a category. See our agent skill picker.
  3. Score something on a rubric (urgency, lead quality, risk) and branch or sort on the number. See triage at a fraction of a cent.
  4. Check whether a statement is true of a document or message before acting, as in guardrails and citation checks.
  5. Replace a fragile prompt that asks an LLM to "return JSON" with a call that returns typed values by construction.

Keep an LLM, or plain code, when you need:

  • Writing. Replies, summaries, articles and code are generation. Jev will not do them.
  • Conversation. Multi-turn chat with a person.
  • Extended reasoning. If a judgment needs many steps or weighs many independent factors, either decompose it into narrow questions and combine the answers in code, or hand it to a reasoning model.
  • Open-ended values. When the answer space is not bounded, a decision model cannot pick from it. When it is bounded, TypeSafe's advice is to turn extraction into a Choice over the candidates.
  • Arithmetic, counting and dates. Those belong in code. TypeSafe's jaggedness notes are explicit about it.
  • Non-text input. Jev takes text only today. Convert images, audio and documents to text first.

One common confusion is worth heading off: Jev is not a replacement for the model behind a coding agent. There is no setting that turns Claude Code or Cursor into a Jev-powered agent. You use your coding agent as usual to write software that calls Jev where a decision is needed.

Code in control, models where they help

TypeSafe describes three ways to build software with AI:

  • Traditional code: a large decision tree built from simple, reliable primitives.
  • Agents: a model reads instructions and chooses its own next step. Powerful with a person watching, but every loop is another chance to go off the rails.
  • AI-powered software: code owns the workflow and handles the deterministic work. The model appears only where the system needs common sense over unstructured text, and each AI task is kept small and constrained.

Decision models are built for the third. It is also the architecture we prefer as a studio: the workflow is ordinary, testable code, and AI is a component inside it rather than the thing in charge.

Using both together

The systems we like best do not choose between the two. They put each where it is good.

  • Decide first, write second. Jev classifies and routes; an LLM drafts a reply only for the items that need one. This is the rule behind our internal Jev tools: Jev decides, a person or an LLM writes.
  • Escalate by confidence. Let the decision model handle the confident majority and send the uncertain slice to a person or a reasoning model. See confidence vs probability.
  • Cascade. TypeSafe's structured data extraction cascade has a small model extract, Jev verify, and a reasoning model take only what fails verification, aiming for most of the big model's quality at a fraction of the cost.
  • Batch the checks. Ask every question about one input in a single request. TypeSafe's parallel questions cookbook found batching 13 questions into one call 12.2x cheaper and 10.0x faster, with no change in answers. More in speculative fan-out.

For a sense of scale from our own tools, measured on 2026-09-23: a three-question triage request on one sales email took 354 ms and cost $0.000031; a skill pick over 16 skills took about 340 to 460 ms and cost about $0.00004.

If you are reviewing an AI feature that generates text only so that code can parse it, that is usually the first place we look. See our AI consulting and engineering services or contact us.

FAQ

What is the main difference between a decision model and an LLM?

A decision model returns typed answers (a choice, a score or a probability) constrained to options you define, with no text generation. An LLM generates text, which suits writing, conversation and reasoning.

Why are decision models cheaper for classification?

They do not generate text, so there are no output tokens to pay for. Jev is priced on input tokens only, at $0.042 per million, and many questions about the same input can share one request.

Can a decision model replace an LLM?

No. It replaces the parts of a system where an LLM was being used to make a decision. Anything that needs writing, conversation or open-ended reasoning still needs a generative model.

Is Jev more accurate than an LLM?

That depends on the task, and you should test on your own data. What Jev offers is calibrated probabilities and confidence, so your code can tell a sure answer from an unsure one and route accordingly.