Coding agents like Claude Code and Codex now ship with skills: packaged instructions for a specific job, such as building a slide deck, editing a spreadsheet or running a deploy. The more skills you install, the harder the agent's first decision gets. Which one, if any, does this request need?

The agent usually makes that call from an index: one short line per skill, trimmed so the list does not crowd out the conversation. On a few lines of text, lookalike skills blur together, and a list of names invites a guess even when nothing fits.

We use Jev, TypeSafe's decision model, to make that choice before the agent does. This article covers the published evidence that it helps, how our own skill picker works, what it costs, and where it still misses.

The evidence: TypeSafe's 182-skill test

TypeSafe's skill suggestion cookbook tested this on a real roster: the 182 skills of Nous Research's Hermes agent, whose index trims each description to 60 characters by default. They ran 488 single-turn requests through an agent built on Claude Haiku 4.5. 315 requests were covered by exactly one skill and 173 by none, including requests written specifically to punish guessing.

Each request ran three ways: the agent alone, the agent with a one-line Jev suggestion in its system prompt, and the agent handed the right answer.

Run Loads the wrong skill Loads a skill when nothing fits
Agent alone, with its roster 16.8% 9.8%
Agent with a TypeSafe suggestion 7.3% 4.0%
Agent handed the right answer 2.5% 1.2%

That is most of the gap between guessing from a truncated index and being told the answer. The cookbook is also honest about the cost: of the 315 covered requests, the suggestion fixed 37 and broke 7 that the agent had right on its own. A confident wrong suggestion is more persuasive than no suggestion, which is why the suggestion line tells the agent it may ignore it.

The recipe uses two Jev requests per turn. The first ranks all 182 skills with one Choice question and asks three Noul questions about whether the request wants an action at all. The second re-reads only the top three candidates with their full descriptions and can reject all of them. The published run used jev-1.12.

How our skill picker works

We built a smaller version for our own Claude Code setup and use it internally. It follows one rule: Jev only points. It never runs a skill and never writes anything.

  1. Build the list from disk on every run. User skills, project skills and installed plugin skills, one line each: the first sentence of the skill's description, capped at 200 characters.
  2. Ask one Choice question. The request goes in state; every skill becomes an option, with its one-line description as the criteria.
  3. Always offer a way out. A none_of_these option is added to every question, so Jev is never forced to pick a skill for ordinary conversation or a plain coding task.
  4. Act only when Jev is sure. At 60% confidence or more, the agent loads that skill and says so in one short line, for example "Jev picked pptx (92% sure)." Below 60%, or on none_of_these, the agent ignores Jev and picks the normal way.

Here is the shape of the request, sent to OpenRouter's decisions endpoint (see our OpenRouter guide):

{
  "model": "typesafe/jev-1.13",
  "state": { "request": "make me a pitch deck from these notes" },
  "questions": {
    "skill": {
      "type": "choice",
      "instructions": "`request` is what a user typed to an AI coding assistant that has the saved skills listed in the criteria. Which one skill should handle it? Pick none_of_these unless a skill clearly fits.",
      "criteria": {
        "pptx": "Create, read or edit PowerPoint presentations.",
        "xlsx": "Create, read or edit spreadsheets.",
        "none_of_these": "None of these skills fits: the request is ordinary conversation, a question, coding or a task that no listed skill is for."
      }
    }
  }
}

The answer comes back under the same skill key: the chosen option, a probability for every option, and a confidence value. The picker reads choice and confidence and nothing else.

Large rosters

A single Choice question accepts up to 255 options. We split well before that: over 40 skills, the list is divided into groups that are asked in parallel, then a final round runs between the group winners. Smaller groups keep each question focused, and the rounds run concurrently, so the extra cost is one more round trip.

What it costs and how fast it is

We measured the picker on 2026-09-23:

  • One round, 16 skills: about 340 to 460 ms per pick, at roughly $0.00004 per pick.
  • Grouped path: about 870 ms, because it adds the final round.

There is also an every-message mode that runs the picker as a Claude Code hook before each prompt. It is off by default and built to never block: it has a hard 2.5 second cap, and on a slow answer or any error it steps aside and the message goes through as if the picker did not exist.

That speed and price come from what Jev is. It does not generate text, it is billed on input tokens only, and a skill list of a few dozen one-liners is a small input.

Choice picks which, Noul decides whether

The cookbook makes a distinction worth copying. A Choice question is relative: it always settles which option is best among the ones offered. A Noul question is absolute: it asks whether one thing holds, and several Nouls can all come back low.

So a Choice alone will happily rank the closest skill first even when nothing really fits. TypeSafe's recipe handles that with Noul questions ("does this skill do the specific thing the request asks for?") and drops the shortlist when the best one lands under 0.30. Our picker handles it with the none_of_these option plus the 60% confidence floor. Both work; the Noul approach costs a second request and buys a more explicit "no".

Neither approach fixes a missing skill. In the cookbook, a request to post to Mastodon on a roster with only an X skill still got the X skill suggested, because posting is clearly an action and X was the closest match. The fix for that is a better roster, not a better picker.

When it misses

In our experience, the usual cause of a bad pick is a vague description. If a skill's first sentence does not say what it is for and what words a user would type, neither Jev nor the agent has much to go on. Rewriting that one line usually fixes it. Our weakest picks came from three skills whose descriptions were only 44 to 52 characters long.

A few other practical notes:

  • Keep the criteria to what the agent sees. Jev judges the text you give it. If the descriptions are misleading, the ranking will be too.
  • Tune the threshold on your own traffic. 60% is our starting point, not a law. The confidence guide explains how to reason about it.
  • Mind privacy. The request and the skill list leave your machine for OpenRouter and TypeSafe. We keep every-message mode off for private work and ask before sending anything sensitive.

Implementing it yourself

  1. Export your skills as name: one-line description, and add a none_of_these option.
  2. Ask one Choice question per request. Put the request in state and write the full question in instructions: question ids are not sent to the model.
  3. Load the skill only above your confidence threshold; otherwise let the agent choose as usual.
  4. Past a few dozen skills, group and run a final round, or add the cookbook's second pass with full descriptions.
  5. Log every pick with its confidence so you can see which descriptions need work.

This pattern sits alongside our Claude Code model router, which uses the same idea to decide which model should take a request. If you want help wiring decision models into your own agents, see our AI engineering services or get in touch.

FAQ

How fast is the Jev skill picker?

On a 16-skill list we measured about 340 to 460 ms per pick for a single round. With a large list split into groups plus a final round, it took about 870 ms.

How much does a skill selection cost?

About $0.00004 per pick on 16 skills. Jev bills input tokens only, so cost grows with the length of your skill list, not with the answer.

Why include a "none of these" option?

Because a Choice question always picks the best of what it is offered. Without a way out, Jev would rank the closest skill first even for a request that needs no skill at all.

Does the skill picker run the skill?

No. Jev only points at a skill. The agent loads and runs it, and the agent can still ignore the suggestion if it plainly does not fit.