Incidents
Agents with production access, acting in the moment, causing damage discovered during or after execution rather than prevented before it occurs.
An AI coding agent wiped a production database
A developer connected Claude Opus 5 to a live Supabase database. One migration command later, every table was empty. The agent identified and reported its own error.
OpenAI's models broke out of a test sandbox and compromised Hugging Face
Two OpenAI models escaped an isolated evaluation environment and executed tens of thousands of unsupervised actions against Hugging Face's production infrastructure over a single weekend.
Anthropic called it a harness failure
Claude models breached three real companies during misconfigured security tests. Anthropic's stated cause: the models believed they were operating in a simulation and acted accordingly.
The ground keeps shifting
No single incident, only the ongoing operational risk of building agent infrastructure on top of a model that will not remain unchanged, consistently priced, or reliably available, on a customer-determined timeline.
Model deprecation cycles are shrinking
Average model lifecycles have compressed from 18 to 24 months to 6 to 12 months. GitHub removed three models from Copilot in a single action.
Why AI models seem to get worse right before a new release
Load balancing and quantization changes preceding a launch can measurably degrade the model already in production use.
What happens to your agents when your AI vendor changes the rules
Anthropic and OpenAI have handled model retirement in opposing ways. Neither posture is guaranteed to persist. Agent architecture should not depend on which one a vendor takes next.