Inside the exponential

If it feels like AI is speeding up, that's because it is. The big American labs are releasing new models faster than they ever have. But release dates are the boring part. What's really accelerating is what the models can do.

The clearest way to see this is to measure how much real work they can do. Several groups now try to. METR and the UK's AI Security Institute both estimate how many hours of a human programmer's work a model can do from a single prompt. GDPval takes a different approach: professional judges compare AI output against human experts across many fields. All three curves are going up, and not merely exponentially. They're rising faster than that.

Curves are abstract, so here's a concrete case. Epoch recently let Opus 4.7 work on its own for 14 hours. It built a software package they estimate would have taken human engineers somewhere between 2 and 17 weeks. The tokens cost $251. Ethan Mollick, in his own experiments, had Fable run unattended for 9 hours on software projects that would have taken a team more than a week. None of this means the models pass every test, or that they're always cheap to run. It means they're getting better, fast.

So far I've been talking about the frontier models, the smartest ones. Those come from three American companies: Anthropic, OpenAI, and Google. But there's a second tier that matters nearly as much, and it's entirely Chinese. These are open-weights models, meaning anyone can download, run, and modify them once they're released. That makes them cheap. They lag the frontier by 6 to 12 months, but they're climbing their own exponential. You can see it in AA-Briefcase, a test that simulates a multi-week consulting engagement: the closed American models on one curve, the open Chinese ones on a parallel curve a step behind.

Graphs only get you so far, though. They hide how jagged the frontier is. A model that can do two weeks of engineering can still stumble on things that look easy, and the open-weights models in particular often do worse in practice than on benchmarks. The only way to know what a model is good at is to try it on the things you care about and judge the results carefully. Mollick's version of this is a test where models have to build an interactive simulation of a harbor town evolving over time. You can play with the results here. What's striking is how much the models differ in the things benchmarks don't capture: taste, design, judgment. And as tasks get longer, those become exactly the things that matter most.

Which brings me to the real point. As models can do longer tasks, the way people use them changes. Until recently the normal way to work with AI was as a collaborator. You asked for something, checked it, then asked for the next step. With enough attention and careful prompting you could steer a model through a long, complicated job. Call this the chatbot way of working.

It still works, and lots of people still do it. But increasingly it's not where the valuable work happens. A system that can run for hours, notice its own mistakes, and fix them doesn't need you sitting beside it. And unlike a chatbot, an agent comes with machinery: a harness that gives the model tools and a place to act, and apps built for that purpose, like Claude Code or OpenAI's Codex. The harness matters. A good one makes an already good model noticeably better.

So work is shifting from collaborating with a chatbot to handing things off to agents. OpenAI and a group of economists published a study of how this happened inside OpenAI itself, and the speed is remarkable. The interesting part is that it wasn't just engineers. Legal, HR, and the other non-technical functions adopted agents at nearly the same rate. OpenAI is an unusual company, but that's what makes it a useful canary. It probably shows what the rest of work will look like a little later.

Work at OpenAI increasingly looks like managing AI. A quarter of the staff have at least four agents running at once in any given week. And once the coding is done by a model inside a harness, everyone becomes a coder of sorts. It turns out they're not bad at it. A separate study of Claude Code users found that software engineers succeeded at coding tasks at about the same rate as people from other professions.

What predicted success wasn't the user's profession. It was their expertise. The more someone knew about a domain, the more they got done in it with Claude Code, and, more interestingly, the more useful output they got from each prompt. An expert's prompt is simply worth more.

That's the shift. The chatbot era was non-experts using AI to fill gaps in what they knew. The agent era is experts using AI to get work done. And the skill that matters most isn't prompting. It's management. The best way to work with agents is to think of yourself as their manager.

One more thing, about what it feels like to be on an exponential. The defining property of an exponential is that each step is bigger than the last. If your organization wrote its AI strategy any time before the winter of 2025, it described a system that could do a couple of hours of work with a high error rate. A few months later a single prompt gets you sixteen hours or more. On paper this is one smooth curve. From the inside it feels like a series of jolts. We're bad at feeling exponentials, and right now we're living inside one.

I think this explains the turbulence around AI better than the usual story about hype. AI isn't a real cybersecurity threat until suddenly it is, and then governments improvise policy at the highest level. Markets don't believe AI can undermine a business model until suddenly it can, and then stocks swing wildly. People read these lurches as signs of an immature field that will settle down eventually. I don't think it will, not soon. The instability isn't noise. It's what happens when institutions that move at the speed of people, or worse, of committees, try to track a curve that doesn't move at human speed at all. And for as long as the exponential lasts, the gap keeps widening.