Most AI features start the same way: a prompt asks a large language model to "return JSON", and a parser hopes the answer comes back in the right shape. It works until it doesn't. The model wraps the JSON in commentary, invents a category you never offered, or states a wrong answer with total confidence, and your code has no way to tell.

Jev is built for the moment your code needs a judgment, not a paragraph. It is TypeSafe's flagship model and the first of what TypeSafe calls System One models. At Dryhurst we build with it because it removes the fragile layer between "the model decided something" and "our code acted on it".

This guide covers what Jev is, what it returns, how it differs from an LLM, and where it earns its place in a product.

What a System One model is

The name comes from Daniel Kahneman's Thinking, Fast and Slow: System 1 thinking is fast and intuitive, System 2 is slow and deliberate. A System One model makes the fast, focused kind of judgment a knowledgeable person makes in a few seconds given the right context.

Like an LLM, Jev understands natural language. Unlike an LLM, it does not write replies, produce code, or explain its reasoning. You define the possible answers up front, and Jev returns a typed decision plus the probabilities behind it. The System One page puts it plainly: typed decisions and probabilities rather than generated text.

How Jev works: state plus typed questions

Every request has two parts:

  1. State: the content to judge. A string, a JSON object, or an array of text: a support ticket, a chat log, a record from your database, a policy. Jev currently accepts text only, so images, audio and video need converting to text or structured fields first.
  2. Questions: a map of typed questions you name. Each has a type, instructions (the question itself) and, for most types, criteria (the possible answers).

Here is a real request shape against the HTTP API:

POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer $TYPESAFE_API_KEY
Content-Type: application/json
{
  "model": "jev-latest",
  "state": "My running shoes arrived in the wrong size. Can I swap them for a size 10?",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "returns": "Exchanges, wrong or damaged items",
        "shipping": "Delivery status, delays, lost packages",
        "billing": "Charges, invoices, payment problems"
      }
    },
    "is_urgent": {
      "type": "noul",
      "instructions": "Does this message convey urgency or time-sensitivity?"
    }
  }
}

The response comes back under the same ids you chose. The ids themselves are never sent to the model, so the full meaning of each question has to live in instructions and criteria.

{
  "model": "jev-1.13.0",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "returns",
      "confidence": 1.0,
      "probabilities": { "shipping": 0.0, "returns": 1.0, "billing": 0.0 }
    },
    "is_urgent": { "type": "noul", "noul": 0.12 }
  },
  "usage": { "input_tokens": 328, "output_tokens": 34 }
}

The values are illustrative, modeled on the example on TypeSafe's Choice page. The point is the shape: no prose to parse, and every answer is constrained to the options you supplied.

The three primitives: Choice, Score and Noul

TypeSafe calls its question types primitives, because you compose them in code the way you compose ordinary software building blocks. We cover them in depth in Choice, Score and Noul: the three question types, but in short:

Primitive Question it answers Returns
Choice Which of these options? choice, probabilities, confidence
Score Which level on this scale? score, legend, probabilities, confidence
Noul Is this statement true? noul, the probability of yes

A Choice routes a ticket to a team. A Score rates bug severity on levels you describe. A Noul answers "does this message request a refund?" as a number between 0 and 1.

Probabilities and confidence

Jev is trained with what TypeSafe calls RLCD, reinforcement learning for calibrated decisions. The goal is that probabilities mean something: across many predictions, outcomes given 0.8 should happen about 80% of the time. The AI primer is careful to say this describes groups of predictions, not a guarantee about any single answer.

Choice and Score answers also carry confidence, a number from 0 to 1 that summarizes how concentrated the probability distribution is. That second number is what makes Jev safe to automate with: the answer tells your code what, confidence tells it whether to act. We go deeper in Confidence vs probability: when to let Jev act and when to escalate.

How Jev differs from an LLM

LLM Jev
Output Generated text Typed answers and probabilities
Answer space Open, even when you ask for JSON Only the options or levels you define
Uncertainty Implicit, often sounds confident either way Explicit probabilities, plus confidence on Choice and Score
Billing Input and output tokens Input tokens only; output is free
Best at Writing, reasoning, conversation, code Fast, narrow judgments your code consumes

Jev is also not a replacement for the model behind a coding agent. TypeSafe says so directly on its coding agents page: there is no setting that turns Claude Code or Cursor into a Jev-powered agent. You use your coding agent to write software that calls Jev where a decision is needed. For the broader argument, see Decision models vs LLMs: when not to generate text.

Four patterns that make it useful

TypeSafe documents a handful of patterns. These four cover most of what we build:

  • Speculative fan-out. Put every question your code might need in one request, including ones that only matter for some inputs, and let code ignore the irrelevant answers. Questions run in parallel against the same state, so extra ones barely move response time. See Speculative fan-out: ask every question in one call.
  • Confidence-gated routing. Act automatically above a threshold, escalate to a person or a reasoning model below it, with stricter thresholds for riskier actions.
  • Composite scoring. Split a broad judgment into atomic Scores and combine them with weights you own in code. When priorities change, you change a coefficient, not a prompt.
  • Intent routing. Classify incoming requests and send each to deterministic code, a specialist LLM, or a human.

What TypeSafe's own benchmarks show

TypeSafe publishes cookbooks with the numbers behind them. A few that stand out:

  • Batching. Asking 13 questions about the GDPR Wikipedia article in one call instead of 13 was 12.2x cheaper and 10.0x faster, with no change in the answers (measured on jev-1.12).
  • Re-ranking. On 40 legal queries, one question per query-candidate pair raised top-1 accuracy from 5% to 18% and top-10 accuracy from 38% to 62%.
  • Skill selection. Across 488 requests with Claude Haiku 4.5 and 182 skills from Nous Research's Hermes catalog, a Jev suggestion cut wrong skill loads from 16.8% to 7.3%, and loading a skill when none fit from 9.8% to 4.0%.

Treat these as TypeSafe's results on TypeSafe's test sets. The docs themselves say to validate performance in your own domain.

Limits worth knowing

TypeSafe publishes a jaggedness page for jev-1.13, which we respect. Known weak spots include counting, arithmetic and date comparison (do those in code), very literal reading of instructions, large states full of irrelevant detail, and generating values rather than choosing among them. Other practical limits from the Models page:

  • Context: 64k tokens per request, 32k for the state plus the longest question.
  • Price: $0.042 per million input tokens; output tokens are free.
  • Language: English is the primary training language; test other languages on your own content.
  • Data: TypeSafe states Jev is not trained on customer requests or responses.

How Dryhurst builds with Jev

We follow one rule: Jev decides, a person or an LLM writes. Three tools we built and use internally, measured on 2026-09-23:

  • A Claude Code model router that asks Jev how big each request is and hands small jobs to smaller Claude models. It adds roughly 350 to 530 ms per message.
  • A skill picker that names at most one agent skill for a request, in about 340 to 460 ms and around $0.00004 per pick on 16 skills.
  • Text triage for emails, leads and tickets. The first live test, three questions on one sales email, used 745 input tokens, cost $0.000031 and took 354 ms.

When Jev isn't sure, it doesn't get the last word: below our confidence cut-offs, a person or a bigger model makes the call.

If you are weighing where a decision model fits in your product, our AI engineering services cover exactly this, or start a conversation.

FAQ

What is a System One AI model?

A System One model makes fast, structured judgments that software can use directly. It understands natural language like an LLM but returns typed decisions and probabilities instead of generated text. Jev is TypeSafe's first System One model.

Does Jev generate text or explanations?

No. Jev returns a choice, a score, or a probability for each question you ask, constrained to the answers you defined. If your product needs prose, pair Jev with an LLM that writes.

How is Jev billed?

Per input token. Output tokens are free. The Models page lists jev-1.13 at $0.042 per million input tokens.

Can I use Jev as the model inside Claude Code or Cursor?

No. Jev is not a chat or code-completion model. Keep your coding agent's LLM and call Jev from the software you build wherever it needs a structured decision.