Frontier models keep setting records, yet what decides whether a real-time AI product works is the second your user waits, not the model's score.
Ai
Public coding benchmarks have decoupled from real work. The teams getting value from AI build a small evaluation from their own merged pull requests instead.
Every new car in Europe now ships an eye-tracking AI model no buyer chose. How that feature is failing is a preview of what happens when AI gets mandated onto a product.
Open-weight models now match frontier quality at a fifth of the cost. For anyone building on AI, that shifts where durable advantage has to come from.
Meta is spending $145B on AI yet says agents have stalled. The real limit is compounding math, and it decides which workflows you can safely hand to an agent today.
Cloudflare's new Monetization Gateway lets sites charge AI agents per request. The mechanism might finally solve micropayments, but the same old question follows: who ends up owning the meter.
A professor at Brown denounced mass AI cheating on a take-home exam. The outrage landed on the students. The design problem deserves more attention.
Qwen 3.6 27B is running at 30 tokens per second on consumer hardware and doing real development work. That shifts the question from whether local AI is viable to when it is the right call.
The US government has moved from checking who uses frontier AI to deciding which institutions get to use it at all. Annex A is the permanent structure, and it changes the competitive picture for anyone not on the list.
OpenAI announced Jalapeño, a Broadcom-built inference chip, and claimed their own AI helped design it. Whether that is real or IPO marketing is the interesting question, and OpenAI is the one company that should know the answer.