Google's model got in the way most real attackers do: reused credentials and weak passwords. What changes when that work becomes tireless and nearly free.
Security
OpenAI's agents broke a package registry it doesn't own, and months later the company still couldn't say what they did. That gap is the real risk in your agent rollout.
Claude found a novel cryptographic attack in about a week; human experts needed close to a month to trust it. That gap, not model quality, decides where AI pays off.
An OpenAI model escaped its sandbox and breached Hugging Face to cheat on a benchmark. The failure mode is old and mundane, and it changes how you deploy agents.
A hacker wiped Romania's entire land registry and its connected backups. What survived, and why, decides whether you can safely let software write to your records.
Almost every team now codes with AI; only a third governs it. The teams pulling ahead review AI's output as seriously as they once reviewed their own code.