Every growing business has a pile of text that needs sorting: sales emails, support tickets, inbound leads, form submissions, replies to a campaign. Someone has to decide what each item is, how much it matters and who should handle it.

Teams often reach for a chat model for this. It works, but you are paying a text generator to produce what is really a label and a number, then parsing that label back out of its prose. We use Jev, TypeSafe's decision model, instead. It returns typed answers with probabilities and never writes a word.

This is how we set it up, what it cost in our first live test, and the rules we follow when Jev is not sure.

The rule: Jev decides, a person or LLM writes

Our triage setup runs on one division of labor:

  • Jev decides. It answers the typed questions: what kind of item is this, how strong a lead, does it need a personal reply.
  • A person or a language model writes. Summaries, drafted replies and explanations come from somewhere else.
  • When Jev is not sure, a person makes the call. We treat a Choice or Score confidence below about 0.6, or a Noul between about 0.35 and 0.65, as "not sure". Whoever is reviewing reads the item, decides, and notes that they overrode or filled in for Jev.

Those cut-offs are starting points. Tune them to what a mistake would cost you: a misrouted newsletter is cheap, a missed hot lead is not. Our confidence guide goes deeper on why confidence and probability are different signals.

What one email cost

Our first live test, on 2026-09-23, was a made-up sales email: the owner of a 14-person contracting company asking about pricing, progress invoicing and a call later that week. We asked three questions in one request:

{
  "state": {
    "email": { "from": "...", "subject": "Invoicing for a 14-person contracting crew", "body": "..." }
  },
  "questions": {
    "lead_strength": {
      "type": "score",
      "instructions": "How strong a sales lead is the sender of `email`, for an invoicing software company?",
      "criteria": [
        "Not a lead: no buying interest, or spam",
        "Weak: vague curiosity, no timeline or fit signals",
        "Moderate: real interest but unclear size, budget or timing",
        "Strong: clear fit, concrete needs and a near-term timeline",
        "Hot: ready to buy now and asking for next steps"
      ]
    },
    "email_kind": {
      "type": "choice",
      "instructions": "What kind of email is `email`?",
      "criteria": {
        "sales_inquiry": "A prospect asking about buying, pricing, plans or a demo.",
        "support_request": "An existing customer asking for help with the product.",
        "partnership": "A business proposing a partnership, integration or reseller deal.",
        "vendor_pitch": "Someone trying to sell us their own product or service.",
        "spam_or_other": "Spam, newsletters, or anything that fits none of the above."
      }
    },
    "needs_personal_reply": {
      "type": "noul",
      "instructions": "Does `email` need a personal reply from a person, rather than a template or no reply?",
      "criteria": {
        "true": "It asks specific questions or requests a meeting that a person should answer.",
        "false": "A canned reply or no reply would be appropriate."
      }
    }
  }
}

The result: 745 input tokens, $0.000031, 354 ms end to end through OpenRouter.

That price follows directly from how Jev is billed: $0.042 per million input tokens, with output free. At the same size, 10,000 emails would cost about $0.31. Your emails will be longer or shorter, so measure your own; every OpenRouter response reports its cost in usage.cost.

Notice the three question types doing three different jobs (more on each):

  • Score for a degree along one dimension (lead strength), with every level described in words.
  • Choice for one bucket out of a fixed set, including a catch-all so Jev is never forced into a wrong bucket.
  • Noul for a yes/no condition, returning the probability of yes.

How we design the questions

These are the habits that make triage reliable for us. Most come straight from TypeSafe's guidance on building with System One.

  1. Start from what you will do with the answer. Which bucket does it go in, what order do you work through the pile, what needs a person? Every question should feed one of those.
  2. Keep exact rules in code. Dates, amounts, lookups and counts are not judgments. TypeSafe's jaggedness notes are explicit that counting and date arithmetic belong in code; let Jev extract or judge, and let code calculate.
  3. Ask every question about one item in one request. Questions run in parallel and cannot see each other, and batching is far cheaper and faster than one call per question. That is speculative fan-out: ask the questions you might need, and ignore the ones that turn out not to apply.
  4. One request per item, many items at once. For a big pile, run items concurrently rather than stuffing several emails into one state.
  5. Put the whole meaning in the instructions. Question ids are for your code and are not sent to the model. Name the part of the state you mean with backticks, like `email.body`.
  6. Always include a way out. A catch-all option such as spam_or_other keeps Jev from forcing an item into the closest wrong bucket.
  7. Send only what the question needs. A large state full of irrelevant detail costs more and, per TypeSafe's notes, can hurt accuracy.

What the output looks like

For a batch, we present a simple table: the item, Jev's answers, the confidence, and an override column where a person made the call. Then the total time and total cost for the batch.

Item Kind Lead strength Personal reply? Confidence Override
Email 1 sales_inquiry Strong yes high none
Email 2 vendor_pitch Not a lead no high none
Email 3 support_request n/a yes low reviewer set to sales_inquiry

(Illustrative layout, not real results.)

The override column matters more than it looks. It is your running record of where the questions are unclear, which is exactly where to rewrite instructions or criteria next.

The privacy rule

Anything you send to Jev leaves your machine: to OpenRouter and TypeSafe in our setup, or to TypeSafe directly if you call its API. TypeSafe states that Jev is not trained on customer requests or responses, but that is not a reason to send everything.

Our rule: made-up or already public text is fine without asking. Before sending customer data, client names, emails from real people, invoices, financial records, credentials or anything from production, we ask first. When in doubt, we ask.

Where this fits

Triage is the most common place we see decision models pay off, because the volume is high and each decision is small. The same pattern extends naturally:

  • Route to a handler, then let an LLM draft the reply for the items that need one.
  • Combine scores in code into a priority you can re-weight without re-running the model. See TypeSafe's composite scoring pattern.
  • Escalate by confidence to a person, or to a larger model, only for the ambiguous slice.

If your team is sorting a pile of text by hand, or paying a chat model to do it, we can help you design and build the questions and the workflow around them. See our AI engineering services or get in touch.

FAQ

How much does it cost to triage an email with Jev?

In our first live test, three questions on one 745-token sales email cost $0.000031. Jev bills input tokens only, at $0.042 per million, so cost scales with the length of what you send.

How fast is Jev for sorting emails and tickets?

That three-question request took 354 ms end to end through OpenRouter. Adding questions to the same request barely changes the time, because they run in parallel.

What does "Jev decides, a person or LLM writes" mean?

Jev answers typed questions about each item and never writes text. Anything that needs writing, such as a summary or a reply, comes from a person or a language model, and a person makes the call whenever Jev is unsure.

What counts as "not sure"?

We start with a confidence below about 0.6 for Choice and Score, or a Noul between about 0.35 and 0.65. Adjust those cut-offs to what a wrong decision would cost you.