---
type: "Comparison"
title: "Jev vs. SemIF vs. DJev: The JevBench Duel in 8 Days"
description: "Jev vs. SemIF vs. DJev on JevBench: 75.3, 75.1, 74.3 — under one point apart. The switch is decided by role: calibrated original or open weights."
resource: "https://www.contextstudios.ai/comparisons/jev-vs-semif-vs-djev"
language: "en"
generated:
  by: "process:contextstudios-md/1"
  at: "2026-10-08T20:46:20.946Z"
status: "stable"
---

# Jev vs. SemIF vs. DJev: The JevBench Duel in 8 Days

On 15 September 2026 TypeSafe AI released Jev — the first System One model, returning typed decisions with calibrated probabilities instead of text. Within eight days, 20 to 30 clones appeared. The vendor's own JevBench currently reads: Jev 75.3, SemIF 75.1, DJev 74.3. All three sit within one point — which is exactly why the leaderboard is not the decision. The decision axis is role: calibrated hosted original versus open weights you run yourself.

## Detailed Comparison

| Factor | Jev (TypeSafe AI) | SemIF & DJev (open-source clones) | Winner |
|--------|------|------|--------|
| JevBench score | 75.3 — top of the first benchmark built specifically for decision models, authored by the same hand as the model | SemIF 75.1, DJev 74.3 — every gap under one point, below the resolution of most in-house evaluation sets | Jev (TypeSafe AI) |
| Latency and cost | 70–500 ms per decision; 42 $ per 1M input tokens, output free; no second infrastructure layer required | runs on your own inference (4B/35B backbone); cost depends on your hardware, no hosted SLA | Jev (TypeSafe AI) |
| Calibration | Objective is calibrated probability: a 95% answer should be right about 95% of the time — the basis for automatic threshold decisions | SemIF uses a plain three-class NLI classifier on the last token; confidence gradation is coarser | Jev (TypeSafe AI) |
| Openness and operation | Hosted API, launched as early access; no open weights | Open checkpoints on Hugging Face — SemIF as 4B and 35B variants on a Qwen3.5 backbone, DJev in the same class; reproducible on your own hardware | SemIF & DJev (open-source clones) |
| Ecosystem and adoption | Fastest adoption in Vercel AI Gateway history: roughly 13% of teams on day one; Playground and JevBench ship from the same vendor | Younger lineage but fast: the Kev line grew from 0.5B to 8B within days; clones typically land right after the original | Jev (TypeSafe AI) |

## Key Statistics

- **75.3 · 75.1 · 74.3 — Jev, SemIF, DJev on JevBench — a one-point span across the whole class** — [JevBench reports, AI digest 19–20 September 2026](https://explainx.ai/blog/jev-playground-jevbench-launch-2026) (2026)
- **20–30 — clones of the Jev lineage within eight days — six in the first two alone** — [Latent.Space AINews 17–18 September 2026](https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in) (2026)
- **≈13% — of teams used Jev on day one via the Vercel AI Gateway — 2x GPT-5.6, 6x Fable 5.1** — [@vercel on X, 18 September 2026](https://www.latent.space/p/ainews-here-are-6-clones-of-jev-in) (2026)
- **Cloudflare released Clef and Clef-flash (27B/9B, Qwen backbone, Apache 2.0) on Workers AI in early October 2026 — Jev-API compatible and open-weighted; on Cloudflare's self-published table Clef beats Jev on eight of ten benchmark rows (e.g. API-Bank 93.11 vs 88.19, BANKING77 macro-F1 94.2 vs 79.74), with Jev ahead on When2Call and BRIGHT — all vendor-reported numbers.** — [Cloudflare blog — Introducing Clef (Oct 2026)](https://blog.cloudflare.com/clef-decision-models/) (2026)
- **Cloudflare's reported median latencies: Clef-flash 38.8 ms and Clef 209.3 ms against Jev's 524.1 ms (p95: 122.4 ms Clef-flash vs 536 ms Jev, across 43 evals); Clef adds a vision encoder and a 64k context window where Jev is text-only at 32k — treat as vendor-reported and re-measure.** — [Cloudflare blog — Introducing Clef (Oct 2026)](https://blog.cloudflare.com/clef-decision-models/) (2026)

## Choose Jev (TypeSafe AI) when...

- The calibrated probabilities are themselves part of your logic: code acts above a threshold, everything below goes to a human — with no in-house inference operation.

## Choose SemIF & DJev (open-source clones) when...

- You need open weights on your own hardware: narrow routing or classification tasks on a Qwen3.5 backbone you already run, with a reproducible checkpoint instead of a hosted API.

## Our Recommendation

The leaderboard barely decides this matchup — the gaps sit under one point, and that is less than most in-house evaluation sets can resolve. The axis is role: Jev is the calibrated original with a hosted API, SemIF and DJev are the open replicas for your own inference.

Pick Jev when the probabilities themselves carry the logic. Calibration means a 95% answer should be right about 95% of the time — that is what makes threshold decisions automatable: code acts above the value, everything below goes to a person. At 70–500 ms response time, 42 $ per 1M input tokens and free output, no second infrastructure layer is needed.

A clone is enough when the task is narrow and the hardware already exists. SemIF ships 4B and 35B as a Qwen3.5 backbone with a plain three-class NLI classifier, DJev is in the same class — if you already run Qwen3.5, you get the decision model as a by-product. The trade-offs are coarser confidence gradation and self-operation.

Two hygiene notes: JevBench comes from the same hand as the model, and the price claim moved from 400x to 440x without an explained method change. Both numbers are usable but not finished — re-measure on your own task set, then switch.

The field widened again in early October: Cloudflare's Clef and Clef-flash (Apache 2.0, Jev-API compatible, hosted on Workers AI) entered the Jev Decision Index above Jev with reported latencies of ~39–209 ms against Jev's ~524 ms — the clone wave reached the infrastructure layer. The role logic still decides, but treat every cross-vendor number as vendor-reported and re-measure on your own task set.

## Frequently Asked Questions

**Q: What does JevBench measure?**
A: The first benchmark built specifically for System One decision models. It is authored by TypeSafe itself, so Jev topping it is expected rather than surprising — the scale is still useful, because the category had no dedicated evaluation before. Task composition and scoring rubric were initially available as headline numbers only.

**Q: Why don't the points decide?**
A: Because the gaps are under one point — less than many in-house evaluation sets can resolve. For the switching decision what counts is the role: calibrated hosted original versus open weights with self-operation.

**Q: When is a clone enough?**
A: When the task is narrow (routing, classification, booleans), the Qwen3.5 backbone already runs on your hardware, and a coarser confidence gradation fits your logic.

**Q: When the original?**
A: When calibration itself is the criterion, when you run no in-house inference, or when 70–500 ms response time matters in the agent loop. Hygiene detail: TypeSafe said "up to 400x cheaper" at launch and "440x" four days later, with no disclosed methodology change — re-measure both on your own task set.

**Q: Where does Cloudflare Clef fit into this field?**
A: As the infrastructure-layer answer to the clone wave: Jev-API compatible, open-weighted (Apache 2.0), hosted on Workers AI and sold with its own RL fine-tuning platform. Its self-published table puts it above Jev on eight of ten rows — but those are cross-vendor vendor-reported numbers, so re-measure on your own task set, exactly like JevBench itself.

