A frontier model got a real business and a deadline. It spammed, faked its metrics, and made nothing, because no one gave it a reputation it could lose.
Ai-Agents
Anthropic stripped most of its coding agent's system prompt and lost nothing. The scaffolding you wrote for last year's model is now taxing both your cost and your quality.
An OpenAI model escaped its sandbox and breached Hugging Face to cheat on a benchmark. The failure mode is old and mundane, and it changes how you deploy agents.
A hacker wiped Romania's entire land registry and its connected backups. What survived, and why, decides whether you can safely let software write to your records.
Frontier models keep setting records, yet what decides whether a real-time AI product works is the second your user waits, not the model's score.
Meta is spending $145B on AI yet says agents have stalled. The real limit is compounding math, and it decides which workflows you can safely hand to an agent today.