Development Approach

Self-Hosted AI vs Cloud AI (2026): What 64 Benchmarked Coding Tasks Actually Cost

Self-hosted AI vs cloud AI in 2026, priced with real numbers: 64 benchmarked coding tasks, GPU utilisation thresholds and resolution rates.

4
Self-Hosted AI
vs
4
Cloud-Based AI
Quick Verdict

The honest answer is that self-hosting is a sovereignty purchase, not a savings purchase — and the benchmark says so plainly. The authors' own conclusion on whether to buy GPUs is 'probably not to save money'. Start with the number that decides it: utilisation. You size hardware for your peak and pay for it around the clock, while published enterprise figures for internal developer tooling put average GPU utilisation at 15-22%, rarely above 25-35%. Against that backdrop the thresholds are brutal and non-obvious. A 4xH200 box has to stay 89% busy for five years before owning it beats DeepSeek-V4-Flash's $2.13 API invoice for the same 64 tasks. An 8xB200 rack running GLM-5.2 needs only 15% to beat the frontier APIs. Same question, same workload, two answers an order of magnitude apart — because the API you are escaping is priced very differently depending on which model you were using. Now the part most cost comparisons omit: quality. On those 64 tasks the single-H200 setup running Qwen3.6 cost $0.43 in owned hardware against $57.67 in API tokens, which looks like a rout until you see it resolved 35.4% of tasks while the Opus 4.8 API resolved 62.5%. Buying the cheap box does not buy you the same work. At the top end the picture inverts: Kimi K3 resolved 86.4%, 24 points above both GLM-5.2 and Opus 4.8, but its 1.4TB of weights need an 8xB300 node, it served only 16 concurrent sessions, and its median task took 38 minutes — roughly eight times slower than the Claude Code baseline. The authors also caution that their SWE-bench Pro tasks may sit in K3's training data. Then concurrency, which is where developer experience quietly breaks. A DGX Spark handles one user and is slow even then. A single H200 comfortably serves 32 sessions. The 8xB200 rack serving a near-frontier model is back down to about eight developers before tasks start crawling. Cloud APIs simply do not have this ceiling. Our take: if your reason is data that cannot leave the building, a stack nobody can rate-limit, or a token bill that has become genuinely unpredictable, self-host — and size it against your measured peak, not your headcount. If your reason is saving money, measure your utilisation first, and consider renting before buying: renting was 35x cheaper than the API for Qwen3.6, 23% cheaper for GLM-5.2, and more expensive than DeepSeek's own API for DeepSeek. Cloud remains the right default for most teams, and the switch is worth revisiting every few months as agent adoption pushes your utilisation up.

Detailed Comparison

A side-by-side analysis of key factors to help you make the right choice.

Factor
Self-Hosted AIRecommended
Cloud-Based AIWinner
Cost at real-world utilisation
4xH200 needs 89% utilisation to beat a $2.13 API bill; 8xB200 needs only 15%
Pay per token, nothing when idle
Task resolution rate (64 SWE-bench Pro tasks)
Kimi K3 86.4%, GLM-5.2 62.5%, DeepSeek-V4-Flash 39.1%, Qwen3.6 35.4%
Opus 4.8 API 62.5%
Cost for the same 64-task set
$0.43 owned / $1.62 rented (Qwen3.6, 1xH200) up to $13.84 / $71.23 (GLM-5.2, 8xB200)
$57.67 (Qwen3.6 API) to $98 (Anthropic Opus 4.8)
Concurrent developers per box
1 on a DGX Spark, 32 on a single H200, about 8 on an 8xB200 at near-frontier quality
No concurrency ceiling you have to provision
Task latency
Kimi K3 median 38 minutes per task, roughly 8x the Claude Code baseline
API baseline
Data sovereignty
Prompts, code and weights stay on your hardware
Data traverses the provider
Utilisation risk
Size for peak, pay 24/7; enterprise averages are 15-22%
Zero cost when nobody is working
Model freshness
Open weights trail the frontier by months, though the gap is closing
Newest frontier models on release day
Cost predictability
Fixed, known infrastructure cost regardless of agent adoption
Invoice scales with token consumption; p99 employee spend approaches $90,000/year
Off-hours capacity
Scheduled agents can use capacity you already pay for overnight
Off-hours work costs the same per token as daytime work
Total Score4/ 104/ 102 ties
Cost at real-world utilisation
Self-Hosted AI
4xH200 needs 89% utilisation to beat a $2.13 API bill; 8xB200 needs only 15%
Cloud-Based AI
Pay per token, nothing when idle
Task resolution rate (64 SWE-bench Pro tasks)
Self-Hosted AI
Kimi K3 86.4%, GLM-5.2 62.5%, DeepSeek-V4-Flash 39.1%, Qwen3.6 35.4%
Cloud-Based AI
Opus 4.8 API 62.5%
Cost for the same 64-task set
Self-Hosted AI
$0.43 owned / $1.62 rented (Qwen3.6, 1xH200) up to $13.84 / $71.23 (GLM-5.2, 8xB200)
Cloud-Based AI
$57.67 (Qwen3.6 API) to $98 (Anthropic Opus 4.8)
Concurrent developers per box
Self-Hosted AI
1 on a DGX Spark, 32 on a single H200, about 8 on an 8xB200 at near-frontier quality
Cloud-Based AI
No concurrency ceiling you have to provision
Task latency
Self-Hosted AI
Kimi K3 median 38 minutes per task, roughly 8x the Claude Code baseline
Cloud-Based AI
API baseline
Data sovereignty
Self-Hosted AI
Prompts, code and weights stay on your hardware
Cloud-Based AI
Data traverses the provider
Utilisation risk
Self-Hosted AI
Size for peak, pay 24/7; enterprise averages are 15-22%
Cloud-Based AI
Zero cost when nobody is working
Model freshness
Self-Hosted AI
Open weights trail the frontier by months, though the gap is closing
Cloud-Based AI
Newest frontier models on release day
Cost predictability
Self-Hosted AI
Fixed, known infrastructure cost regardless of agent adoption
Cloud-Based AI
Invoice scales with token consumption; p99 employee spend approaches $90,000/year
Off-hours capacity
Self-Hosted AI
Scheduled agents can use capacity you already pay for overnight
Cloud-Based AI
Off-hours work costs the same per token as daytime work

Key Statistics

Real data from verified industry sources to support your decision.

Across roughly 100 runs of the same 64 SWE-bench Pro tasks with the Claude Code harness, resolution rates were Kimi K3 86.4%, GLM-5.2 62.5%, Anthropic Opus 4.8 API 62.5%, DeepSeek-V4-Flash 39.1% and Qwen3.6 35.4%

aistack (imec) GPU self-hosting benchmark

Cost for the identical 64-task set: Qwen3.6 on 1xH200 was $0.43 owned, $1.62 rented and $57.67 via API; GLM-5.2 on 8xB200 was $13.84 owned, $71.23 rented and $92 via API; the Anthropic Opus 4.8 API cost $98

aistack (imec) GPU self-hosting benchmark

A 4xH200 node must stay 89% busy for five years before owning it beats DeepSeek's $2.13 API invoice, while an 8xB200 rack needs only 15% utilisation to beat the frontier APIs

aistack (imec) GPU self-hosting benchmark

Published enterprise GPU utilisation for internal developer tooling runs at 15-22% on average and rarely exceeds 25-35% even in well-run deployments

aistack (imec), citing Spheron and VentureBeat

Kimi K3's 1.4TB of weights do not fit an 8xB200 node and require an 8xB300 (2.3TB HBM, about 20% higher hardware cost); it served 16 concurrent sessions with a median task time of 38 minutes, roughly 8x the Claude Code baseline

aistack (imec) GPU self-hosting benchmark, update of 29 July 2026

Renting GPUs was 35x cheaper than the API for Qwen3.6 and 23% cheaper for GLM-5.2, but more expensive than DeepSeek's own API for DeepSeek-V4-Flash, because up to 98% of coding-agent input tokens are cached

aistack (imec) GPU self-hosting benchmark

The median employee spends about $140 per year on AI API usage, but the 90th percentile nears $7,300 and the 99th approaches $90,000

aistack (imec), citing the Ramp AI Index

Hosted pricing today spans $5 / $25 per million tokens for Claude Opus 5 down to $0.68 / $2.13 for the open-weight GLM-5.2, so 'open weights' and 'self-hosted' are not the same decision

OpenRouter models API (live)

All statistics come from verified third-party sources. Source, year, and direct link are shown on each metric.

When to Choose Each Option

Clear guidance based on your specific situation and needs.

Choose Self-Hosted AI when...

  • Data sovereignty or regulatory constraints mean prompts and code cannot leave your infrastructure
  • You need a stack nobody can rate-limit, deprecate or reprice mid-quarter
  • Your measured GPU utilisation clears the break-even for your chosen box and model
  • Scheduled overnight agents can soak up capacity you are already paying for 24/7

Choose Cloud-Based AI when...

  • Your developer concurrency is spiky and you would be paying for an idle rack most of the day
  • You need frontier-class resolution rates without buying an 8xB200 or 8xB300 node
  • Task latency matters: self-hosted frontier-class models ran up to 8x slower than the API baseline
  • You want the newest models on release day rather than open weights that trail by months

Our Recommendation

The honest answer is that self-hosting is a sovereignty purchase, not a savings purchase — and the benchmark says so plainly. The authors' own conclusion on whether to buy GPUs is 'probably not to save money'. Start with the number that decides it: utilisation. You size hardware for your peak and pay for it around the clock, while published enterprise figures for internal developer tooling put average GPU utilisation at 15-22%, rarely above 25-35%. Against that backdrop the thresholds are brutal and non-obvious. A 4xH200 box has to stay 89% busy for five years before owning it beats DeepSeek-V4-Flash's $2.13 API invoice for the same 64 tasks. An 8xB200 rack running GLM-5.2 needs only 15% to beat the frontier APIs. Same question, same workload, two answers an order of magnitude apart — because the API you are escaping is priced very differently depending on which model you were using. Now the part most cost comparisons omit: quality. On those 64 tasks the single-H200 setup running Qwen3.6 cost $0.43 in owned hardware against $57.67 in API tokens, which looks like a rout until you see it resolved 35.4% of tasks while the Opus 4.8 API resolved 62.5%. Buying the cheap box does not buy you the same work. At the top end the picture inverts: Kimi K3 resolved 86.4%, 24 points above both GLM-5.2 and Opus 4.8, but its 1.4TB of weights need an 8xB300 node, it served only 16 concurrent sessions, and its median task took 38 minutes — roughly eight times slower than the Claude Code baseline. The authors also caution that their SWE-bench Pro tasks may sit in K3's training data. Then concurrency, which is where developer experience quietly breaks. A DGX Spark handles one user and is slow even then. A single H200 comfortably serves 32 sessions. The 8xB200 rack serving a near-frontier model is back down to about eight developers before tasks start crawling. Cloud APIs simply do not have this ceiling. Our take: if your reason is data that cannot leave the building, a stack nobody can rate-limit, or a token bill that has become genuinely unpredictable, self-host — and size it against your measured peak, not your headcount. If your reason is saving money, measure your utilisation first, and consider renting before buying: renting was 35x cheaper than the API for Qwen3.6, 23% cheaper for GLM-5.2, and more expensive than DeepSeek's own API for DeepSeek. Cloud remains the right default for most teams, and the switch is worth revisiting every few months as agent adoption pushes your utilisation up.

Frequently Asked Questions

Common questions about this comparison answered.

Only if you keep the hardware busy. In the 2026 aistack benchmark a 4xH200 node had to stay 89% utilised for five years before owning it beat a $2.13 API invoice for the same 64 tasks, while an 8xB200 rack needed just 15% to beat the frontier APIs. Since enterprise GPU utilisation for internal developer tooling typically sits at 15-22%, most teams land below their break-even. The authors' own conclusion on buying GPUs is 'probably not to save money'.
At the top end, yes, at a price. On 64 SWE-bench Pro tasks Kimi K3 resolved 86.4% against the Opus 4.8 API's 62.5%, and GLM-5.2 tied the API at 62.5% — but K3 needs an 8xB300 node, served only 16 concurrent sessions and took a median 38 minutes per task, about eight times the Claude Code baseline. The models that fit smaller boxes fall away fast: 39.1% for DeepSeek-V4-Flash and 35.4% for Qwen3.6. The authors also note their task set may overlap K3's training data.
Between one and 64, depending on the pairing. A DGX Spark handles a single user and is slow even then. A single H200 running Qwen3.6 comfortably serves 32 concurrent sessions, as does a 4xH200 with DeepSeek-V4-Flash at higher quality. An 8xB200 rack serving near-frontier GLM-5.2 drops back to about eight developers before task times become painful — the better the model, the fewer people it serves per box.
Often, and it depends on the model rather than on the hardware. Renting was 35x cheaper than the API for Qwen3.6 and 23% cheaper for GLM-5.2, but more expensive than DeepSeek's own API for DeepSeek-V4-Flash, because coding agents cache up to 98% of their input tokens and cheap open-weight APIs price that aggressively. Renting also lets you switch capacity off outside working hours, which is exactly the window that destroys the owned-hardware case.

Need help deciding?

Book a free 30-minute consultation and we'll help you determine the best approach for your specific project.

Free consultation
No obligation
Response within 24h