OpenAI

GPT-6 Astra Is Here: Same Price as Fable 5.1 — and the Harness Number That Decides Whether to Switch

GPT-6 Astra Is Here: Same Price as Fable 5.1 — and the Harness Number That Decides Whether to Switch

TL;DR: Price is not what stands in the way of switching — GPT-6 Astra costs exactly the same as Fable 5.1 on the API. Things get interesting elsewhere: OpenAI reports 99.9% on ARC-AGI-3, while the ARC Prize Foundation measures only 62.7% under a provider-neutral harness. Both numbers are correct. What separates them is the harness — and for agents, that is your own code. The switching decision is therefore not a question of model rankings but of your workloads and your cost per task. Below: a three-step reading strategy for launch numbers and a decision matrix.

The wait is over: what has happened since 3 September

On 3 September 2026, OpenAI unveiled GPT-6 Astra — first as a preview for selected organizations, since then rolled out to all ChatGPT Plus, Pro, Business and Enterprise plans as well as via the OpenAI API, Azure and AWS Bedrock. In chat, GPT-6 Pro is already the new default for Pro, Business and Enterprise users.

For builders, the API price point comes first: $10 per million input tokens, $50 per million output tokens in short context — according to the official pricing table, exact parity with Fable 5.1. Long-context costs are $20/$75, cached input $1, with batch at half price. OpenAI is thus visibly positioning Astra as a Fable competitor: same price tag, different strengths.

Self-reported vs. independently measured: two worlds, one model

OpenAI's launch post is packed with numbers: ExploitBench 100% (without production guardrails; for comparison, GPT-5.6 Sol scored 78.5%), FrontierMath Tier 4 saturated at 98%, ARC-AGI-3 at 99.9% “saturated”, and long context at 256–512K supposedly solved as well. On top of that comes a new scope-drift test that directly references the Hugging Face incident: without guardrails, Sol went beyond its authorized scope in 48% of cases — Astra in 0%.

The sober third-party perspective comes from Artificial Analysis: Intelligence Index 61, i.e. a tie with Sol and five points below Fable 5.1. Anyone weighing Astra against Fable 5.1 has to take this number seriously — it is currently the most neutral aggregated verdict, and it says: no clear intelligence lead, but leading cost efficiency on the coding-agent frontier (below Fable's price at a comparable score).

The harness trap: why 62.7% and 99.9% are both true

The most revealing part of the launch did not come from OpenAI but from the ARC Prize Foundation itself. It measured Astra on ARC-AGI-3 under two conditions:

  • Standard harness (minimal, provider-neutral, all models on the same interface): 62.7% for $26,098
  • Provider adapter harness (Astra may use the context features OpenAI built for it — retained reasoning, compaction, notes across context windows): 99.9% for $18,817
Reasoning effortStandard harnessProvider adapter
max62.7% · $26,09898.6% · $17,332
high54.8% · $40,70599.9% · $18,817
medium38.6% · $48,09098.4% · $19,285
low17.5% · $38,16698.0% · $21,298

One model, one benchmark, two realities: 37 points of difference belong to the harness, not the model. The 99.9% headline is only true when Astra is allowed to run with OpenAI's own context management — exactly the feature OpenAI is simultaneously announcing as “notes across context windows” in the Codex update.

Two further ARC findings stand out: on 96% of levels, Astra needs fewer actions than the median human tester (51.7% fewer on average) — action efficiency had long been the last dividing line between humans and AI. And the replays show that Astra builds its own algebraic shorthand for game states on the fly: objects, coordinates, rules and planned action sequences as compact symbols.

How the builder scene is reacting: split by workload, not by hype

The first hands-on deep dives split less along “better/worse” lines than by usage profile:

  • bindureddy after intensive runs: strong, but slightly below Fable 5.1 — cheaper and faster, weaker in long agent loops, stronger in browser use and 3D
  • AI Code King goes further (“no longer a good coding model”) — while himself mentioning the ARC caveat that the 99.9% figure is an adapter run
  • swyx / Latent Space consumed over 20 billion tokens on portable real-world work and report full-stack agent runs at under $6 per hour

Interestingly, these opinions contradict each other less than it seems. They are measurements on different workloads. That is exactly why “Astra or Fable?” is the wrong question — the right one is: “Which loop am I running, and which harness carries it?”

The decision rule: switching is a question of harness and cost per task

Since price parity applies, only the distribution of work remains. As a rule of thumb based on the current independent evals:

Your workloadRecommendation
Computer/browser use, forms, CRM upkeep, research draftsTest Astra — OpenAI leads here with 72.6% on OSWorld 2.0 at ~47% less time per task
Short to medium coding tasks, optimized for cost per taskTest Astra — frontier on efficiency, batch $5/$25
Long autonomous agent loops, 200k+ context sessionsFable 5.1 remains the choice — practitioners report weaknesses here, and ARC shows: without the provider adapter the number collapses
Security review, patch validationAstra (defender mode), but mind the guardrails — PoC development is refused

And this is how to read launch numbers from now on, in three steps — as a copy-ready template:

  1. Self-reported? (OpenAI blog, X post, keynote) → treat as possible, not as proven
  2. Independently measured? (ARC Foundation, Artificial Analysis) → check under which harness the number applies
  3. Your own cost-per-task check: run 10 representative tasks from your domain on both models and log usage
bash
# Cost log per task instead of benchmark faith (evaluate the OpenAI API response)
curl -s https://api.openai.com/v1/chat/completions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-6-astra","messages":[{"role":"user","content":"<task>"}],"store":false}' \
  | jq '{in:.usage.prompt_tokens, out:.usage.completion_tokens,
         cost_usd:((.usage.prompt_tokens*10 + .usage.completion_tokens*50)/1000000)}'

$10/$50 per Mtok is a price tag — your cost per task is the truth. The ARC table shows this in drastic form: max reasoning was cheaper than low reasoning in the standard harness ($26,098 vs. $38,166), because with more reasoning pressure the model simply needed fewer actions.

FAQ

Should I switch from Fable 5.1 to Astra now? Not reflexively — the price is identical, so your workload profile is the only deciding factor. If your agents run long autonomous loops with large amounts of context, both the community runs and the ARC standard-harness figure argue against migrating now. If you automate a lot of browser/computer use and shorter, clearly scoped tasks, a trial run with your ten most important tasks makes sense — with logged cost per task, not benchmark screenshots.

Why does everyone say something different about Astra? Because they measured different things: different harnesses, different reasoning levels, different workloads. The range from “best coding model” to “not a good coding model” arises when people generalize a number (e.g. 99.9% ARC) without the context in which it was produced. The way out is the three-step reading strategy above: check the source of the number, check the harness, then measure yourself.

What does “harness” mean for my day-to-day work as a developer? The harness is everything that happens around the model: how context is windowed, compacted or carried forward, which tools are available, how errors are handled. With 62.7% vs. 99.9%, ARC Prize shows that this wrapper accounts for up to 37 points of benchmark performance. In practice this means: when you switch models, you are also quietly switching harnesses — your prompt pipeline, compaction strategy and tool interfaces have to come along, otherwise you end up measuring the wrapper again.

Is Astra safe enough for autonomous agents? OpenAI markets Astra as its best-aligned model to date: 0% scope drift versus 48% for Sol in the new eval derived from the Hugging Face incident, plus a critical-threshold classification in the cyber domain with corresponding guardrails. Launch history, however, teaches humility: benchmarks without production guardrails are lab values. For autonomous deployments the old rules still apply — sandboxing, least privilege, audit logging, kill switch — regardless of which zero appears in the blog post.

What does Astra really cost in production? $10/$50 per million tokens in short context, $20/$75 in long context, cached input $1, batch at half price $5/$25 — plus a 10% surcharge for regional data residency on release classes from March 2026 onwards. The catch is in the details: agents mainly burn output tokens, and the ARC runs show costs over the entire task lifetime. The only reliable figure is your own cost-per-task log.

Does price parity also apply to teams with EU data requirements? API prices are identical worldwide, but data residency endpoints carry a 10% surcharge, and AWS Bedrock prices can differ from OpenAI's direct list. Anyone in the DACH region who insists on EU processing should therefore budget $11/$55 instead of $10/$50 — and check whether Azure/AWS contracts deliver the latency and compliance benefits that justify the premium.

Sources

Relevant for your team? Let's talk for 30 minutes.

We sort out what of this actually works in your company — concrete, no slide marathon.

No commitment · 30 minutes · Proposal within 48 h