Skip to content

How Jev works

Guide

Jev is the classifier that reads along while muxcode works. It is TypeSafe’s System One model. This page explains:

  • where Jev reads a run;
  • what a sure answer changes;
  • how it sets reasoning effort one model call at a time;
  • what happens when it is slow or unsure;
  • how to steer it from Settings → Classifier.

Jev writes no code and holds no seat in a workflow. At fixed moments muxcode puts a typed question to it about text muxcode already has: a goal, a contract, a report, a command. Jev answers with a probability, or with a choice and its confidence. muxcode then compares that answer with a fixed bar.

  1. Something happens.
  2. muxcode asks Jev a typed question.
  3. Jev answers with a probability.
  4. muxcode compares it with that question’s bar.
  5. muxcode acts, or leaves things as they are.

A sure answer changes one bounded thing: a stage, a worker, an effort level, a note on a permission card. An unsure, slow or missing answer changes nothing, and the run carries on as it would without Jev. Jev can never widen a permission. Scope checks, verification and the final judgement stay separate from classification.

  • Agents panel: each run’s routing line, for example “Work package 1: no classifier reading; keep the configured worker”. Each step’s receipt shows the effort it was assigned and the effort it ran at.
  • Permission cards: a risky command’s card names what Jev read in it.
  • Settings → Classifier: every question Jev is asked, grouped by where it is asked, with calibration charts for each.

Every call has a short time limit and one retry. Past that, the question changes nothing and the run carries on as if Jev were off. The run never waits on Jev. Where a skipped answer leaves a mark, you see the default instead: a routing line that reads “no classifier reading”, or a step that kept its effort level.

A writing run meets Jev at each hand-off:

  1. Intake reads your goal and your own words.
  2. The Architect plans.
  3. Screen reads the contract.
  4. Route picks a worker.
  5. Effort is read before every model call.
  6. Verify reads the worker’s report.
  7. The judge decides.

Each answer can move the run one bounded step. None of them skips the Architect’s plan, the scope gate or the judge.

Intake. Before any model is paid, Jev reads your goal and your own messages:

  • the kind of change (logic, UI or terminal) and its risk;
  • whether you asked for a test;
  • whether the change can only be checked in the built app;
  • which paths you protected;
  • the constraints you stated.

The constraints become a checklist the judge is handed. The Architect still plans every change.

Contract screen. Jev reads every contract before a worker runs, for four defects:

  • it cannot prove its goal;
  • it cannot reach a file the goal needs;
  • it leaves out the verification you asked for;
  • it contradicts a constraint.

A sure defect sends the contract back to the Architect once. The roster muxcode ships has no separate critic: this screen, plus the Architect re-reading its own plan, does that job.

Which worker. The roster muxcode ships has one worker, so there is nothing to route. With more than one, your workers are listed from cheapest to strongest, and for each work package Jev makes one choice: the cheapest worker whose criteria the package clearly meets, or the configured worker. Only a choice made with at least 0.85 confidence moves the package. The configured worker, last in the list, keeps the package when:

  • the Architect set a reasoning minimum;
  • evidence is missing;
  • the risk is high;
  • Jev is unsure.

The Architect and the judge are never routed.

Verify and caveats. After the checks run, Jev reads the worker’s summary against the verification log. Does it claim a check the log contradicts or never ran? Does it admit to skipping work? An answer over its bar reaches the judge as a line to read first, such as “read the log before the summary”. An accept that comes with caveats is read for how much they weigh. A worker that reports it is blocked is read for why: a missing tool or runtime pauses the run for you instead of spending the retry.

A seat’s reasoning effort is not fixed for a whole step. Before each model call, once the last tool results are in, Jev reads what the seat is doing. It answers two questions:

  • the effort for the next call, from the levels the model offers;
  • a lease: how many calls that level should hold for, 1, 2, 5 or 10.

While a lease holds, Jev is not asked again. So the loop runs like this:

  1. Tool results come in.
  2. Jev reads the seat.
  3. It answers with a level and a lease.
  4. Calls run under the lease.
  5. The lease ends, and the loop starts again.

What ends a lease early:

  • you send new input;
  • a tool call fails;
  • the level is changed elsewhere;
  • a fresh context window opens, which also clears what Jev reads, just as it clears the model’s view.

Two missing answers in a row stop that seat asking for the rest of its step.

Bands: a starting effort and an “Up to”. Each seat has an effort it starts at and, optionally, an “Up to”: the furthest Jev may raise it. Jev is offered only the levels in that band. A seat whose “Up to” is its own level is fixed, and Jev is not asked about it at all.

Seat, as shipped Starts at Up to
Architect: plan, review, replan, judge (Claude Opus 5.5) high high, so fixed
Worker (Claude Opus 5.5) medium xhigh
Retry (Claude Fable 5.1) high xhigh

A seat is never taken below where it starts. The worker starts at medium because the Architect has already done the hard thinking in the contract, and the gate, the screen and the judge catch a thin attempt. When the judge rejects the work or a check fails, the retry takes it on a second model, Fable, starting at high. The worker shares the Architect’s model, so the judge that accepts the work is the model that wrote it, at a higher and fixed effort; the retry on Fable is the second opinion. Every seat, its starting effort and its “Up to” can be changed on the Agents page.

Every question has a fixed bar. A bar is high where acting wrongly would be costly: moving a worker or an effort level takes 0.85. Dropping a tool result from the context works the other way round: its probability of mattering must be under 0.1.

Bar Questions
0.90 route, bypass guard
0.85 worker route, effort
0.80 block kind, guard
0.70 defect, risk
0.60 retry, turn outcome
under 0.10 elide or prune a result

The interesting answers sit just under a bar. The Classifier page counts how many answers fall within 0.1 of each one: those are the answers that would flip if the bar moved.

Chat tasks meet Jev too, in the same read-then-compare way:

  • how a turn ended, and whether more is needed;
  • whether a plain question can go to a cheaper model for one turn;
  • whether a tool result carries an instruction, which is flagged before the model reads it;
  • in long tasks, when to remind the model to checkpoint, what the handoff to a fresh window keeps, and which tool results can be left out;
  • the order of history search results;
  • the risk named on a permission card.

The page opens with what Jev is: something happens, Jev reads it, only a sure answer acts. It then lists every question Jev is asked. The switches are grouped by what they change.

Card Switch What it changes
Jev Read along turns Jev on or off
Jev Reuse recent answers reuses Jev’s answer when the same question is asked again
Whose key pays TypeSafe key uses your own key, for any project
What Jev may change Send plain questions to a cheaper model one cheap turn for a plain question
What Jev may change Guard bypassed sessions a check on risky commands when you bypass permissions
When Jev can’t answer Stand-in model a chat model for low-stakes questions

muxcode’s own classifier key is used only for projects on a paid mux plan. A TypeSafe key of your own works for any project.

Every question has stakes:

Stakes Jev on and allowed Jev off, not allowed, or failing
High the classifier nothing changes
Low the classifier the stand-in model, if you named one

The stand-in answers low-stakes questions, such as how a turn ended, the order of search results, or a reader’s wrap-up. Its answers are uncalibrated, logged separately, and never drawn on the charts.

For each question the page draws:

  • a histogram of Jev’s answers, with the bar drawn over it;
  • how many answers acted;
  • how many sit near the bar;
  • where the outcome is known, a reliability curve and a Brier score.

An outcome is only what actually happened, for example whether the chosen worker’s first attempt was accepted. It is never a second model’s opinion. Use the charts to ask whether a bar sits in the right place. The bars themselves change only in new versions of muxcode.