A cached token costs a tenth of a fresh one, and coding agents hit cache most of the time. Your AI cost curve is an engineering choice, and most teams make it by accident.
Llm
Microsoft's engineers merged 24% more pull requests with AI. The constraint didn't vanish; it moved to review, and most teams are still counting the wrong thing.
Public coding benchmarks have decoupled from real work. The teams getting value from AI build a small evaluation from their own merged pull requests instead.
Open-weight models now match frontier quality at a fifth of the cost. For anyone building on AI, that shifts where durable advantage has to come from.
The US government has moved from checking who uses frontier AI to deciding which institutions get to use it at all. Annex A is the permanent structure, and it changes the competitive picture for anyone not on the list.
Anthropic now requires a government ID before accessing its most capable models. What looks like a safety measure is export control, and it is nudging developers toward the Chinese AI alternatives the US is trying to contain.
Anthropic's Fable 5 shipped with two guardrails: one that refused honestly, and one that quietly degraded the answer without telling you. The second is the dangerous kind, and Anthropic's own reversal shows why.
For a while we treated generative AI as a parlour trick with bad hands and a habit of making things up. Five kinds of story ended that phase for me, and raised a harder question than 'will it take my job?'