The usual way to get structured data out of an LLM is to ask for JSON and validate what comes back. The model is still generating text; the structure is a request, not a guarantee.

Jev works the other way round. Every question you send is one of three typed primitives, and every answer is constrained to the options or levels you defined. The model returns a probability distribution over your answers, never a value outside them. If you are new to Jev, start with what Jev is and how System One models work; this article goes one level deeper into the question types themselves.

The shape every question shares

Each question sits in a questions map under an id you choose, and has:

  • type: choice, score or noul.
  • instructions: the question itself, written in full. The id is never sent to the model, so refund_requested means nothing to Jev unless the instructions say it.
  • criteria: the possible answers. Required for Choice and Score, optional for Noul.

All three types can share one request and one state. Each is evaluated independently, so one answer never becomes hidden context for another.

Choice: one option from a set

Use a Choice when the answer is one of a known set of options with no order between them: which team handles a ticket, what kind of document this is, which intent a message carries. criteria is a map from option name to description, and both are sent to the model, so write descriptions that separate the options.

{
  "department": {
    "type": "choice",
    "instructions": "Which team should handle this?",
    "criteria": {
      "returns": "Exchanges, wrong or damaged items",
      "shipping": "Delivery status, delays, lost packages",
      "billing": "Charges, invoices, payment problems"
    }
  }
}

For a wrong-size shoe complaint, TypeSafe's Choice page shows this answer:

{
  "type": "choice",
  "choice": "returns",
  "confidence": 1.0,
  "probabilities": { "shipping": 0.0, "returns": 1.0, "billing": 0.0 }
}

choice is the option with the highest probability, probabilities covers every option, and confidence summarizes how peaked that distribution is. A ticket that mentions both a wrong size and a missing refund would split probability between returns and billing, and confidence would drop.

Two practical notes. Add an other or none_of_these option when your list might not cover every input, so Jev is never forced to pick. And remember a Choice is relative: it tells you which option wins, not whether any of them is a good fit.

Score: a position on levels you describe

Use a Score when the answer sits on a spectrum you can describe in steps: bug severity, customer frustration, how formal an outfit is. criteria is an ordered array of level descriptions, low to high, between two and ten levels. Levels are numbered by position starting at 0.

{
  "bug_severity": {
    "type": "score",
    "instructions": "How severe is the reported issue?",
    "criteria": [
      "Cosmetic; no impact to functionality",
      "Broken or degraded feature, but workaround exists",
      "Blocking issue; no workaround exists"
    ]
  }
}

For the report "The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari", the Score page shows:

{
  "type": "score",
  "score": 1.43,
  "confidence": 0.35,
  "legend": {
    "0": "Cosmetic; no impact to functionality",
    "1": "Broken or degraded feature, but workaround exists",
    "2": "Blocking issue; no workaround exists"
  },
  "probabilities": { "0": 0.0, "1": 0.57, "2": 0.43 }
}

The score is the probability-weighted mean of the level numbers: 1 x 0.57 + 2 x 0.43 = 1.43. It is a position, not a level, and it can fall between two levels. Here the model is split between "workaround exists" and "no workaround", which is exactly right for a bug that only blocks Safari users, and the low confidence says so.

Read probabilities alongside the score. A score of 1.0 can mean all probability on level 1, or half on level 0 and half on level 2. Those are very different situations, and confidence is what tells them apart.

Noul: the probability that something is true

Use a Noul for a clean yes or no where the probability itself is the useful signal: does this message request a refund, does it contain personal data, is the customer asking for a human. The answer is a single field, noul, between 0 and 1. There is no separate confidence: the number is the answer and the certainty in one.

{
  "is_human_escalation": {
    "type": "noul",
    "instructions": "Is the customer asking for a human agent?"
  },
  "is_repeat_contact": {
    "type": "noul",
    "instructions": "Has the customer contacted support about this before?",
    "criteria": {
      "true": "Mentions a prior attempt, ticket, or that they have asked before",
      "false": "No sign of any previous contact"
    }
  }
}

On "I have asked three times now. Can I please just talk to a real person?", TypeSafe's Noul page reports 0.99 for the first question and 0.93 for the second. Its recorded jev-1.13.0 answers for the human-agent question on other messages are just as telling: "How do I reset my password?" gets 0.07, "I need this sorted today, whatever it takes." gets 0.26 (urgent, but never asks for a person), and "Are you a bot?", which hints at wanting a human without asking, lands at 0.40.

The trap: a Noul near 0.5 means yes and no are about equally likely. It does not mean "medium". If you want to measure how strong a candidate's Python is, use a Score with defined levels, not a Noul asking whether they are "strong in Python".

Choosing between them

If the answer is... Use Your code does...
One of several unordered options Choice switch on choice
A position on a described scale Score Threshold or rank on score
Yes or no Noul if noul > threshold

When two types seem to fit, pick the one your code can act on most directly. A Choice between refund, rebook and information maps straight onto three code paths.

Composing primitives in code

The real power comes from asking several narrow questions and combining the answers yourself. TypeSafe calls this composite scoring. Ticket priority, for example, can be three Scores asked in one request: severity, frustration, and how much the report gives an engineer to work with.

# Normalize each Score to 0..1 by dividing by its top level, then weight.
severity = answers["severity"].score / 2
frustration = answers["frustration"].score / 2
actionability = answers["actionability"].score / 2

priority = 0.5 * severity + 0.3 * frustration + 0.2 * actionability

The weights are yours (these are illustrative). When your team disagrees with the ranking, you change a coefficient and rerun, without touching the questions. Because the questions share one request, adding a dimension barely changes response time; see speculative fan-out for why.

We use exactly this approach in our email, lead and ticket triage, and a Choice sits at the heart of our Claude Code model router. If you want a second pair of eyes on how to decompose a judgment in your product, that is core to our AI engineering work. Start a conversation.

FAQ

What is a Noul in Jev?

A Noul is Jev's yes/no question type. It returns noul, the probability that the answer is yes, as a number from 0 to 1. Near 1 is a strong yes, near 0 a strong no, and near 0.5 means the model is uncertain.

Why doesn't a Noul return a confidence value?

Because the probability already carries the certainty. A value close to 0 or 1 is a confident answer; a value close to 0.5 is not. Your code decides the thresholds, including a middle band that goes to a person.

Can Choice, Score and Noul questions share one request?

Yes. You can mix all three types in one call against the same state. Each question is evaluated independently and returns its answer under the id you chose.

Why is a Score fractional?

The score is the probability-weighted average of the level numbers, so it can land between levels. Round it when you need one outcome, or use it directly to rank items.