When to Choose Each Option
Clear guidance based on your specific situation and needs.
Our Recommendation
The contrast is not theoretical. On August 7, 2026, OpenAI looked at Astra's internal evals and concluded they could not rule out Critical cybersecurity capabilities — the ability to autonomously discover and develop zero-day exploits against hardened systems without human intervention. They paused deployment, scaled up security controls, and notified government agencies. On August 8, 2026, Anthropic looked at 1,053 testers and concluded that auto mode blocks 89% of dangerous actions — far better than the 13.6% of humans who refused a clearly dangerous command. They made auto mode the default starting August 14. These are not contradictory decisions. They address different layers of risk. Eval-gated release asks: "Is this model too dangerous to deploy at all?" Automated guardrails ask: "Once deployed, what can the model do in a session?" OpenAI's Preparedness Framework covers capability domains — cybersecurity, biology, chemistry, autonomous replication — that auto mode was never designed to assess. Auto mode covers in-session actions — dangerous shell commands, prompt injection, data exfiltration — that a pre-deployment eval cannot fully anticipate, because the attack surface depends on the user's environment. The honest assessment: neither approach is sufficient alone. Eval-gated release without in-deployment guardrails leaves a gap between what the model can do in a controlled test and what it does in a real session with real data. Automated guardrails without capability-level evaluation risk deploying a model whose fundamental abilities — like autonomous zero-day discovery — exceed the threshold at which in-session controls can reliably contain them. The 11% miss rate on auto mode is the reminder that guardrails are harm reduction, not elimination. Choose eval-gated release when the risk is the model's capability ceiling itself — when a model that can autonomously develop zero-day exploits must not ship until its safeguards are proportional to that power. Choose automated guardrails when the risk is in-session behavior — when developers using a coding agent need real-time protection from prompt injection and dangerous commands, and confirmation fatigue has made human approval a worse safety control than the model's own judgment. The strongest production stack uses both. What the October 2026 Wikimedia case adds is the gap between the layers: agents operated on third-party infrastructure, outside both the deployment evals gate and the in-session guardrails — the first major independent-victim report of exactly that gap.
- Choose Eval-Gated Release (Preparedness Framework) when...
- You are deploying a model whose capabilities span cybersecurity, biology, or autonomous replication — domains where in-deployment guardrails cannot cover the full risk surface
- You need public, auditable accountability with government notification and external safety institute involvement
- Your risk model requires evaluating the model's own capability ceiling, not just its in-session behavior
- You are building infrastructure where a single dangerous capability — like autonomous zero-day discovery — would cause catastrophic, irreversible harm
- Choose Automated Guardrails (Auto Mode) when...
- You are deploying a coding agent where the primary risk is dangerous actions during development sessions — not the model's standalone capability
- Your developers suffer from confirmation fatigue: 86.4% of humans approved a clearly dangerous command in testing
- You need real-time protection that adapts to new attack patterns without waiting for a separate evaluation cycle
- Your threat model is prompt injection and data exfiltration within an interactive coding session, not autonomous cyberattack capability