Provider Comparison

Jev vs. SemIF vs. DJev: The JevBench Duel in 8 Days

Jev vs. SemIF vs. DJev on JevBench: 75.3, 75.1, 74.3 — under one point apart. The switch is decided by role: calibrated original or open weights.

Reviewed by Michael Kerkhoff, as of

Definition
On 15 September 2026 TypeSafe AI released Jev — the first System One model, returning typed decisions with calibrated probabilities instead of text. Within eight days, 20 to 30 clones appeared. The vendor's own JevBench currently reads: Jev 75.3, SemIF 75.1, DJev 74.3. All three sit within one point — which is exactly why the leaderboard is not the decision. The decision axis is role: calibrated hosted original versus open weights you run yourself.
Category
Provider Comparison
Options
Jev (TypeSafe AI)SemIF & DJev (open-source clones)

Detailed Comparison

A side-by-side analysis of key factors to help you make the right choice.

Jev (TypeSafe AI) vs SemIF & DJev (open-source clones)
FactorJev (TypeSafe AI)SemIF & DJev (open-source clones)
JevBench score75.3 — top of the first benchmark built specifically for decision models, authored by the same hand as the model WinnerSemIF 75.1, DJev 74.3 — every gap under one point, below the resolution of most in-house evaluation sets
Latency and cost70–500 ms per decision; 42 $ per 1M input tokens, output free; no second infrastructure layer required Winnerruns on your own inference (4B/35B backbone); cost depends on your hardware, no hosted SLA
CalibrationObjective is calibrated probability: a 95% answer should be right about 95% of the time — the basis for automatic threshold decisions WinnerSemIF uses a plain three-class NLI classifier on the last token; confidence gradation is coarser
Openness and operationHosted API, launched as early access; no open weightsOpen checkpoints on Hugging Face — SemIF as 4B and 35B variants on a Qwen3.5 backbone, DJev in the same class; reproducible on your own hardware Winner
Ecosystem and adoptionFastest adoption in Vercel AI Gateway history: roughly 13% of teams on day one; Playground and JevBench ship from the same vendor WinnerYounger lineage but fast: the Kev line grew from 0.5B to 8B within days; clones typically land right after the original
Total Score · 0 ties4 / 51 / 5

Key Statistics

Real data from verified industry sources to support your decision.

All statistics come from verified third-party sources. Source, year, and direct link are shown on each metric.

When to Choose Each Option

Clear guidance based on your specific situation and needs.

Our Recommendation

The leaderboard barely decides this matchup — the gaps sit under one point, and that is less than most in-house evaluation sets can resolve. The axis is role: Jev is the calibrated original with a hosted API, SemIF and DJev are the open replicas for your own inference. Pick Jev when the probabilities themselves carry the logic. Calibration means a 95% answer should be right about 95% of the time — that is what makes threshold decisions automatable: code acts above the value, everything below goes to a person. At 70–500 ms response time, 42 $ per 1M input tokens and free output, no second infrastructure layer is needed. A clone is enough when the task is narrow and the hardware already exists. SemIF ships 4B and 35B as a Qwen3.5 backbone with a plain three-class NLI classifier, DJev is in the same class — if you already run Qwen3.5, you get the decision model as a by-product. The trade-offs are coarser confidence gradation and self-operation. Two hygiene notes: JevBench comes from the same hand as the model, and the price claim moved from 400x to 440x without an explained method change. Both numbers are usable but not finished — re-measure on your own task set, then switch. The field widened again in early October: Cloudflare's Clef and Clef-flash (Apache 2.0, Jev-API compatible, hosted on Workers AI) entered the Jev Decision Index above Jev with reported latencies of ~39–209 ms against Jev's ~524 ms — the clone wave reached the infrastructure layer. The role logic still decides, but treat every cross-vendor number as vendor-reported and re-measure on your own task set.

Choose Jev (TypeSafe AI) when...
  • The calibrated probabilities are themselves part of your logic: code acts above a threshold, everything below goes to a human — with no in-house inference operation.
Choose SemIF & DJev (open-source clones) when...
  • You need open weights on your own hardware: narrow routing or classification tasks on a Qwen3.5 backbone you already run, with a reproducible checkpoint instead of a hosted API.

Common questions about this comparison answered.

Frequently Asked Questions

(01)What does JevBench measure?
The first benchmark built specifically for System One decision models. It is authored by TypeSafe itself, so Jev topping it is expected rather than surprising — the scale is still useful, because the category had no dedicated evaluation before. Task composition and scoring rubric were initially available as headline numbers only.
(02)Why don't the points decide?
Because the gaps are under one point — less than many in-house evaluation sets can resolve. For the switching decision what counts is the role: calibrated hosted original versus open weights with self-operation.
(03)When is a clone enough?
When the task is narrow (routing, classification, booleans), the Qwen3.5 backbone already runs on your hardware, and a coarser confidence gradation fits your logic.
(04)When the original?
When calibration itself is the criterion, when you run no in-house inference, or when 70–500 ms response time matters in the agent loop. Hygiene detail: TypeSafe said "up to 400x cheaper" at launch and "440x" four days later, with no disclosed methodology change — re-measure both on your own task set.
(05)Where does Cloudflare Clef fit into this field?
As the infrastructure-layer answer to the clone wave: Jev-API compatible, open-weighted (Apache 2.0), hosted on Workers AI and sold with its own RL fine-tuning platform. Its self-published table puts it above Jev on eight of ten rows — but those are cross-vendor vendor-reported numbers, so re-measure on your own task set, exactly like JevBench itself.

Need help deciding?

Book a free 30-minute consultation and we'll help you determine the best approach for your specific project.

Free consultation · No obligation · Personal reply