LLM

Jev Measured Independently: 92–214 ms per Decision, Under 1 Cent per 8 Requests — and the 25x-vs-200x Catch

Jev Measured Independently: 92–214 ms per Decision, Under 1 Cent per 8 Requests — and the 25x-vs-200x Catch

TL;DR: Three independent measurement waves put countable numbers on the System One class for the first time: 92–214 ms latency per decision, under 1 cent for eight requests, and 75.3 points for the original on the JevBench leaderboard. On re-measurement, the marketing factor of 200x shrinks to roughly 25x. This article shows the data points, the catch, and a three-step quick test you can use to measure any router yourself.

The System One class in one sentence

Jev does not write — it decides. Instead of generating token by token, the class of System One models outputs a classification or routing decision in a single, extremely fast pass. That is exactly what it is built for: as a pre-stage in front of the expensive autoregressive model, which only kicks in when the decision calls for it.

The three measurement waves at a glance

Within two weeks, three independent sources instrumented the same class. The numbers land close together — which is rare and makes them reliable.

Measurement sourceLatency per decisionCostNote
First measurement wave (primary blog)92–214 ms< 1 cent / 8 requestscounted directly at the gateway
Isenberg re-measurement~200 ms / query18 cents for dozens of emailseveryday workflow, repeatable
Kai (public measurement)70–500 ms$42/M input, output freewidest latency band, same order of magnitude

The third column is the point: each source calculates differently, but all of them land at milliseconds and fractions of a cent per task. That turns the class from a vague promise into a measurement protocol.

The 200x catch: recalculated, roughly 25x remains

The advertised 200x leap over a classic autoregressive model does not fully survive cross-checking. The reason: the 200x usually compares the complete generation pass with the pure decision pass — including warm-up and startup effects on both sides. The recalculated, more apples-to-apples figure is about 25x, measured on open re-implementations of the same architecture. That is still a leap of an order of magnitude, but one with a footnote.

The mechanics behind it

Distilled, the mechanics consist of three building blocks:

  • Decision-centred architecture: one compact pass instead of iterative decoder loops.
  • Parallel sampling: several candidate decisions run simultaneously rather than one after another.
  • RLCD calibration: decision quality is tuned directly to the TASK distribution, not to generic next-token log-likelihood.

Caveat from review 7: parts of the setup are reminiscent of classic classifier architectures plus marketing amplification. The class is real, and so is the pillar of numbers — but how to put it all in perspective remains an open question.

JevBench and the clone wave

With JevBench, the class has its own leaderboard for the first time. The top three overall:

ModelOverall
Jev75.3
SemIF75.1
DJev74.3

On top of that comes a remarkable sign of adoption: 20–30 open re-implementations appeared within eight days — a clone ecosystem growing at that pace is the strongest signal that an architecture fills a gap in the routing stack. Laya (non-autoregressive) shows up as a second independent mention of the class, supporting the two-class taxonomy from the first runs.

Adoption as the fourth data point

The Vercel AI Gateway reports the fastest model adoption in its gateway history for Jev: around 13 percent of teams on day one, measured by the number of active teams. For comparison: the GPT-5.6 family took twice as long and Fable 5.1 six times as long — so Jev caught on considerably faster.

Cost-per-task table (copyable)

MetricValueSource
Latency per decision92–214 msfirst measurement wave
Cost per 8 requests< 1 centfirst measurement wave
Latency, re-measured~200 ms / queryIsenberg
Adoption day 1~13 % of teamsVercel AI Gateway
JevBench overall75.3 (Jev), 75.1 (SemIF), 74.3 (DJev)JevBench
Clones in 8 days20–30public count

A minimal calculation for your own stack:

python
# Cost per task for a decision pass
delay_ms = (92 + 214) / 2          # ~153 ms mean latency
cost_per_request = 0.01 / 8         # 8 requests per cent -> ~0.00125 EUR
print(f"1500 decisions: {1500 * cost_per_request:.3f} EUR, "
      f"{1500 * delay_ms / 1000:.0f} s cumulative")

The 3-point quick test for your own router

  1. Latency point: Measure 50 requests at your gateway and take the median per decision. If you land below 300 ms, the class is relevant for you.
  2. Cost point: Divide your monthly request count by 8 and multiply by the per-cent price for your region. Below five times your current router cost per call? Move on to point 3.
  3. Quality point: Compare at JevBench level: how often does the router decide differently from what the large model concludes in hindsight? Deviations below 20 percent are considered viable.

Anyone who gets through all three points within an hour has verified the pattern themselves — without the hype as a crutch.

What this means for builders

  • Decider coupling: router tasks (classification, intent, sorting) move to the fast class, while generation stays with the large model.
  • Numbers before names: every newly mentioned model only gets a slot once latency and cost per task are documented.
  • Use the clone ecosystem: with 20–30 open re-implementations, SemIF and DJev are worth a look as local alternatives with a similar profile.

FAQ

What is the difference between a System One decision and a regular model call? A classic model call generates token after token, delivering text as a by-product of the probability chain. A System One pass, by contrast, outputs a finished decision in a single, compact pass — such as a class or a routing path. That saves exactly those iteration loops that cause most of the latency in classification tasks.

How should I read the 92–214 ms per decision range? The range marks the lower and upper end of the measured medians across several independent waves, each per individual decision. The mean is around 150 ms, which puts a sequential chain of ten decisions at roughly 1.5 seconds. For interactive pipelines, that is enough headroom to run before or alongside the first full-model layer.

Why is 25x rather than 200x the more honest figure? The 200x figure compares different warm-up stages and pass depths, so startup costs on both sides inflate the difference. In a direct comparison of a complete generation pass with a decision pass on the same hardware, about 25x remains. This figure has been recalculated reproducibly and is therefore the conservative working value.

What does a day-one adoption of 13 percent of teams tell me? The metric comes from the Vercel AI Gateway and counts active teams that called the model at least once within 24 hours of release. In the context of the gateway's history, this is the fastest uptake ever measured, ahead of the well-known model families. The figure measures curiosity, not necessarily retention — for a cost-benefit estimate, the median of the following weeks is more meaningful.

How do I build the 3-point quick test into my pipeline? Create a request counter per route, log latency and cost in cents for each call, and take the median after 50 lines. Then compare the decision against the result of the large model and flag deviations. With these three values you can place any new router candidate within an hour, without needing a second reference source.

Is a clone such as SemIF or DJev worth it compared to the original? The clone wave produced 20–30 open re-implementations in eight days, the two best of which score 75.1 and 74.3 JevBench points respectively — only 0.2 to 1.0 points behind the original. That makes them fully viable for local setups with fixed hardware limits. The decision usually comes down to the licence and the available parameter count, not the score itself.

Sources

Relevant for your team? Let's talk for 30 minutes.

We sort out what of this actually works in your company — concrete, no slide marathon.

No commitment · 30 minutes · Proposal within 48 h