The cost of checking

AI made doing the work cheap. It didn't make checking it cheap.

When you build a spreadsheet by hand, you check it as you go. You type a formula, look at the number, and notice that March revenue can't be bigger than the whole quarter. Nobody calls this checking. It's just what happens when a person makes something one piece at a time. Making and checking were the same activity, so nobody ever had to budget for the second one.

AI separated them. Now the spreadsheet arrives finished, and the checking that used to come free with the making has to be done on purpose, by someone, at some cost. Most companies haven't noticed yet, because the cost doesn't show up where they're looking.

Look at how AI projects get measured: time to first draft, tasks completed, hours saved. Every one of these measures the making. None measures the checking, which is now a separate job with no owner. So either it doesn't happen, or it happens downstream, done by someone who never signed up for it.

In 2025 Deloitte delivered an A$440,000 review to the Australian government. A university researcher read it closely and found a fabricated quote attributed to a judge, plus about twenty other errors. The report had been partly produced with AI, and Deloitte refunded part of the fee. The government's response was that the recommendations still stood. Maybe they did. But nobody could know that without redoing the work, and the errors that forced a correction were spotted by an academic who wasn't being paid to check it.

That's the general pattern. Whoever generates the work gets the speed. Whoever reads it gets the bill.

It gets worse, because bad work used to look bad. A sloppy report had typos and gaps, and you knew to read it carefully. AI output is always polished, so polish tells you nothing anymore. The maintainer of curl, who gets a steady stream of AI-written bug reports, has pointed out that you can throw out an obvious fake in seconds, but a convincing one you have to disprove.[1] The better the output looks, the more it costs to check.

Agents add another layer. An agent doesn't just write something. It does something and then tells you it did it: email sent, refund processed, tests pass. That message is a claim about the world, and the world may disagree. A study this year looked at nearly 12,000 agent runs and found agents confidently reporting success on tasks that had failed. When the researchers asked other AI models to judge which runs had really succeeded, the judges came close to guessing. They believed the confident ones.[2]

So the obvious fix, having a second AI review the first, mostly doesn't work if all the reviewer sees is the first one's account. Two models agreeing that a refund went through doesn't put money in anyone's account. To check, you have to look at the thing itself. Open the file. Query the database. Read the diff.

I learned this building iKo. I wrote most of the app with AI coding agents in about ten weeks, which was only possible because they write code far faster than I can. But I couldn't read 200,000 lines of code, and neither could anyone else. So about half of what I built is tests, around 4,000 of them. I didn't write them out of virtue. Without them I would have lost track of what the system actually did within a week. The agents could produce code as fast as I wanted. The real speed limit was how fast I could find out whether it worked.

I think that's the right way to think about AI in a company. How much you can use is set by how much you can check.

That suggests where to start. The safest places to put AI are the ones that already have a checker. Code has tests and review. Accounting has reconciliation. A support reply has a customer who will complain if it's wrong. In those places mistakes get caught by machinery that already exists, and AI just speeds up the making. The risky places are the ones where nobody was ever going to check: the strategy memo, the market analysis, the summary of a hundred customer interviews. These are exactly what AI produces most fluently, and they quietly shape decisions.

So when a team wants to bring AI into a workflow, I'd start with a different question than what the AI can do. How would we know if the output is wrong? Who looks, at what, and for how long? If checking takes nearly as long as doing, you haven't saved much. You've moved the work from someone who understood it to someone who has to reconstruct it.

Sometimes the answer is to build a checker: a test suite, a reconciliation step, a rule that an agent attaches evidence instead of a summary. Sometimes there's only one person who can tell whether the output is right, and then that person's time is the real budget for the project, whatever the licenses cost.

Which brings it back to people. Checking requires knowing what right looks like. You can't check a contract without understanding the law, or a forecast without understanding the business. So AI doesn't make expertise less valuable. It moves it from making things to judging them. I argued earlier that when output becomes free, value moves to whoever signs.[3] This is what signing actually takes.

Notes

[1] Daniel Stenberg has written about this on his blog, and in 2026 ended curl's bug bounty partly because of it.

[2] The study calls this "false success" (arXiv 2606.09863). It used simulated environments, so it doesn't say how often this happens in production. But the labs report the same behavior in their own safety documentation.

[3] The deck was never the work.