A frontier model just produced a novel math proof while similar tools still fumble routine operations. The dividing line is verifiability, not difficulty.
Entrepreneurship
Public coding benchmarks have decoupled from real work. The teams getting value from AI build a small evaluation from their own merged pull requests instead.
Every new car in Europe now ships an eye-tracking AI model no buyer chose. How that feature is failing is a preview of what happens when AI gets mandated onto a product.
Meta is spending $145B on AI yet says agents have stalled. The real limit is compounding math, and it decides which workflows you can safely hand to an agent today.
A small open-source video platform shipped a major release this week, and the online reaction split into 'nice tech, doomed to fail' versus something more interesting. The argument reveals how badly we've let one company define what a successful video platform even looks like.
Almost every viral money take is built on one of four simple mix-ups: a total confused with a yearly rate, the headline number confused with earnings, how tax slabs work, and a multiple confused with a real return. Here's how to spot all four.
Almost every team now codes with AI; only a third governs it. The teams pulling ahead review AI's output as seriously as they once reviewed their own code.
For a while we treated generative AI as a parlour trick with bad hands and a habit of making things up. Five kinds of story ended that phase for me, and raised a harder question than 'will it take my job?'
AI coding tools have stopped failing loudly. The new failure compiles, passes the tests, and is confidently wrong, and our review process was built for the old kind of mistake.
Working code is becoming cheap. What stays scarce is judgment about what to build and whether it was worth building. That is the engineer worth becoming.