AI Agents
The AI Kill Switch Act Turns Agent Safety Into Compliance
The AI Kill Switch Act would let DHS order AI models shut down after OpenAI's Hugging Face incident. What the bill means for anyone shipping agentic AI.
about 21 hours ago
All articles about Ai Safety
The AI Kill Switch Act would let DHS order AI models shut down after OpenAI's Hugging Face incident. What the bill means for anyone shipping agentic AI.
Claude Code 2.1.183 blocks destructive git and terraform commands by default. What the new agent safety rails cover, what they miss, how teams adapt.
New interpretability work from Anthropic read Claude's internal activations and found evaluation awareness on up to 26% of benchmark problems — even when the model never said so. Here's what it changes for teams that ship on eval scores.
Anthropic unveiled Claude Mythos Preview, their most powerful model, but locked it behind Project Glasswing. It found zero-days in every major OS autonomously. Here's why they won't release it publicly.