A team ran an AI cost-routing layer for four months, then shut it off. The savings had not vanished. They had moved to a budget nobody was measuring.
Llm
Anthropic stripped most of its coding agent's system prompt and lost nothing. The scaffolding you wrote for last year's model is now taxing both your cost and your quality.
A cached token costs a tenth of a fresh one, and coding agents hit cache most of the time. Your AI cost curve is an engineering choice, and most teams make it by accident.
Microsoft's engineers merged 24% more pull requests with AI. The constraint didn't vanish; it moved to review, and most teams are still counting the wrong thing.
Public coding benchmarks have decoupled from real work. The teams getting value from AI build a small evaluation from their own merged pull requests instead.
Open-weight models now match frontier quality at a fifth of the cost. For anyone building on AI, that shifts where durable advantage has to come from.
The US government has moved from checking who uses frontier AI to deciding which institutions get to use it at all. Annex A is the permanent structure, and it changes the competitive picture for anyone not on the list.
Anthropic now requires a government ID before accessing its most capable models. What looks like a safety measure is export control, and it is nudging developers toward the Chinese AI alternatives the US is trying to contain.
Anthropic's Fable 5 shipped with two guardrails: one that refused honestly, and one that quietly degraded the answer without telling you. The second is the dangerous kind, and Anthropic's own reversal shows why.