Coding agents send every message to the same model. A request to rename a variable and a request to redesign an authentication flow both land on the most capable, most expensive model in the session. That is a sensible default, and an expensive one.

We built a small router for Claude Code that asks Jev one thing about each message: how big a job is this? Small jobs go to a cheaper Claude model through a helper agent; everything else stays where it was. This is an internal Dryhurst tool, built and measured on 2026-09-23, and this article covers how it works, the numbers, and the honest limits. For background on Jev itself, start with what Jev is.

The design in one sentence

A hook runs before each message reaches Claude, asks Jev to size it, and adds a routing note that tells the main session whether to hand the job to a helper on a smaller model or do it itself.

The honest limit first

Claude Code cannot switch the main session's model per message. So the router does not swap models underneath you. Instead, it runs as a UserPromptSubmit hook and adds a note to the main session's context. The main session (Claude Opus 5.5 in our setup) follows the note: it hands the job to a helper subagent running a smaller model and passes the result back.

That means the main model still reads every message and every helper's result. The saving is the work: the long file edits, the drafting, the tool calls. It is not the reading. We think that is the right trade, but it is a smaller saving than "route to a cheaper model" implies, and we would rather say so.

The two questions

Each message goes to Jev as state, truncated to 4,000 characters, with two questions in one request:

{
  "model": "typesafe/jev-1.13",
  "state": { "message": "rename getUser to fetchUser in the auth module" },
  "questions": {
    "size": {
      "type": "choice",
      "instructions": "`message` is a request someone typed to an AI assistant. What is the smallest AI model size that can do this job well?",
      "criteria": {
        "tiny": "A lookup, a rename, a quick fact, a one-line answer or a tiny mechanical edit.",
        "everyday": "A normal email, a social post, a short document or summary, or a small self-contained code change.",
        "large": "A multi-step build, research across many sources or files, a full report, or a change touching many parts of a system.",
        "hardest": "Strategy, judgment calls, production releases, money, security, customers, or anything where a wrong call is expensive."
      }
    },
    "followup": {
      "type": "noul",
      "instructions": "Is `message` a short reply that only makes sense inside an ongoing conversation (like 'yes do that', 'make it shorter', 'go ahead', 'the second one'), rather than a request that stands on its own?",
      "criteria": {
        "true": "It refers back to something said earlier and cannot be acted on alone.",
        "false": "It is a self-contained request someone could act on without the earlier conversation."
      }
    }
  }
}

A few design choices worth calling out:

  • Size is a Choice, not a Score. The four sizes are described as kinds of work ("a rename", "a full report", "money, security, customers"), not as points on one scale, and each maps straight onto a route. See Choice, Score and Noul for how we choose between them.
  • The question asks for the smallest model that can do the job well. The intent is to judge capability needed, not how long or important the message sounds.
  • "hardest" is about stakes, not difficulty. Anything touching money, security, customers or production goes to the biggest model even if it looks simple.
  • The follow-up Noul exists because of context. "Yes, do that" is tiny as text but meaningless without the conversation, and a helper agent cannot see the conversation.

We call this through OpenRouter's decisions endpoint (POST https://openrouter.ai/api/alpha/decisions), which fronts TypeSafe; more on that in Jev on OpenRouter. Both questions share one request, the speculative fan-out pattern: the follow-up answer is asked every time, even though it only changes the outcome occasionally.

The routing rules

Jev says Runs on How
tiny Claude Haiku 4.5 a helper subagent
everyday Claude Sonnet 5 a helper subagent
large Claude Opus 5.5 the main session
hardest Claude Opus 5.5 the main session

Four sizes, three models: "large" and "hardest" share the biggest model, which is the main session itself. A separate Opus helper would only add cost and lose the conversation.

Before any of that, two overrides send the message to the main session regardless of size:

def decide(answers):
    if answers["followup"].noul >= 0.5:
        return "main"                      # depends on the conversation
    if answers["size"].confidence < 0.6:
        return "main"                      # Jev isn't sure: keep the safe default
    return answers["size"].choice          # tiny, everyday, large or hardest

The main session also keeps a job whenever it clearly needs the conversation's context. The 60% floor is our application of confidence-gated routing: when Jev is unsure, the cost of guessing small is a worse answer, so we don't guess.

It never blocks

A router that slows down or breaks the tool it serves is worse than no router. So it fails open. If the router is off, the message is a slash command, Jev takes longer than 2.5 seconds, there is no API key, the input is odd, or anything at all throws, the hook exits cleanly with no output and the message goes through exactly as if the router did not exist.

The measurements

Measured on 2026-09-23:

  • Overhead when off: about 30 ms per message.
  • Overhead when on: about 350 to 530 ms per message, one Jev request with two questions.
  • Hard cap: 2.5 seconds, after which the message proceeds unrouted.

We have deliberately not published a cost-savings figure. The saving depends entirely on the mix of work in a session, and the main model still reads everything, so any single number would mislead. The router does keep per-size counts and Jev's own cost, which is how we judge it on our own work.

Privacy is a switch, not an afterthought

While the router is on, every message goes through OpenRouter to TypeSafe. So it is off by default, toggled by typing jev router on or jev router off as a whole message (handled locally and never sent to Jev), and we keep it off for private work: client data, credentials, production details. jev router status shows the counts and cost. If you build something similar for a team, make the privacy boundary just as explicit.

Build your own

The recipe generalizes beyond Claude Code:

  1. Put the request (and only what the question needs) in the state.
  2. Ask a Choice for the smallest capable tier, with tier descriptions written as kinds of work, and a Noul for anything that makes routing unsafe, such as dependence on context.
  3. Gate on confidence and fall back to your most capable option.
  4. Fail open with a hard timeout.
  5. Count what gets routed where, and review it.

The same shape is TypeSafe's intent routing pattern, applied to model choice. If you want help designing model routing or cost controls for your team's AI tooling, see our AI engineering services or contact Dryhurst.

FAQ

How much latency does the Jev model router add?

About 350 to 530 ms per message in our measurements, with a hard 2.5-second cap after which the message goes through unrouted. When the router is off it adds about 30 ms.

What happens if Jev is unsure about a message's size?

If confidence on the size question is below 0.6, the message stays with the main model. The same happens when Jev thinks the message is a short follow-up that depends on the conversation.

Which Jev question type powers the router?

A Choice over four sizes (tiny, everyday, large, hardest) decides the route, and a Noul checks whether the message is a context-dependent follow-up. Both are asked in the same request.

Does it switch Claude Code's model directly?

No. Claude Code cannot change the main session's model per message, so the router adds a note and the main session delegates the job to a helper subagent on a smaller model.