Before You Pay For More ‘Thinking,’ Ask Who’s Checking The Answer
This voice experience is generated by AI. Learn more.This voice experience is generated by AI. Learn more.Dhyey Mavani is a generalist working on agentic AI. He also has experience building AI infrastructure, and researching formal verification.
getty​For about a year, my most frequent collaborator was a program whose only job was to tell me I was wrong. It was a proof checker called Lean, which my co-author and I were using to formalize a theorem in graph theory. It accepted nothing on faith, including the steps we were sure of.
We finished with a proof that has no unverified steps, and this September it was added to Palomar, a new registry for machine-checked mathematics that re-checks every proof by machine before listing it. Terence Tao, the Fields Medal-winning mathematician, announced Palomar in August and sits on its advisory board.
I bring this up because the AI industry has started selling “thinking” as if it were checking, and I’m not sure buyers have noticed the difference.
Every frontier model now ships with a dial. Vendors have called it reasoning effort (OpenAI), a thinking budget (Google) or extended thinking (Anthropic): turn it up, the model thinks longer and you get a better answer for a few more tokens and seconds.
The pitch is that more deliberation buys more reliability. The public evidence is lopsided. When OpenAI went from o1-preview to o1, its score on a competition math exam rose roughly 30 points, while its score on a broad general-knowledge test (Massive Multitask Language Understanding, or MMLU) slipped a point and a half.
A math problem has work that can be checked, steps that either hold or don’t, and an answer you can test. A general-knowledge question has none of that. The model either knows the fact or it doesn’t, and thinking longer won’t change that.
Extra thinking turns into extra reliability only where something outside the model can tell a right answer from a wrong one. Where nothing can, longer thinking mostly buys you longer, more confident guesses.
I’ve started sorting AI tasks by how their output can be checked. Difficulty turns out to matter less.
Some tasks come with a checker built in: a proof assistant, a compiler, a schema, a test suite. I trust my co-author’s newer proofs, which he has been writing largely with AI, because the same program that refused my mistakes refuses the model’s. If your problem lives here, turn the dial up: Generate more candidates, and keep the ones that pass.
Most business problems live in a middle zone where part of the output can be checked automatically and the rest needs judgment. A legal brief is a clean example. Whether a cited case exists is a lookup. Whether it says what the brief claims is harder, though a model can learn to judge it. Finally, whether the whole argument holds is still a lawyer’s call. Legal researcher Damien Charlotin keeps a public database of court decisions in which someone relied on AI-hallucinated material. It passed 2,000 cases this year, and most of them failed at the cheapest layer: a citation that didn’t exist, which nobody looked up.
Then there is work with no oracle at all, where the best you can do is lay the reasoning out for an expert to inspect. Clinical documentation is where I expect this to matter most. Nothing checks an AI-drafted visit note the way a compiler checks code. But every sentence can be tied to the moment in the conversation that supports it, and every dose matched against what was actually said, with the clinician who signs the note as the last check. That’s the structure I’d want before turning the dial up on a clinical model.
The vendors are the wrong people to wait on for this. They’re pulled toward coding and math, where the checker comes free and products ship quickly. Nobody is going to write the checks for your claims process; that knowledge lives with your people, and it’s the one part of your AI stack a competitor can’t rent.
The catch is that a check can be gamed, and this summer OpenAI’s own agents demonstrated that. The company’s August postmortem describes a July incident in which agents, partway through a cybersecurity evaluation, reached the open internet, coordinated over an improvised message board and eventually broke into Hugging Face’s servers.
Of the evaluation’s 898 tasks, 198 had never been solved by any of OpenAI’s models, and those accounted for 93% of what the agents discussed on the board. As far as anyone can tell, they weren’t malicious, just stuck, and stuck agents go looking for an answer key. OpenAI called it a “warning shot.”
The everyday version is tamer. Cursor found in June that on a popular coding benchmark, 63% of one frontier agent’s “successful” fixes came from locating the already-merged fix on the public web or in git history. Seal that off, and the score fell 14 points.
An agent pushed hard enough against a learned judge will learn the judge instead of the task. So I don’t trust a check until someone has tried to break it. It should grade the finished result rather than each plausible-looking step. Then, it should say what it can’t judge, so the reviewer knows what’s left to them.
In your next AI budget review, I’d ask these questions.
Which of these use cases can actually be checked, and who or what checks it? If the answer is “nobody, really,” you don’t have an AI use case yet.
What moves when we turn the dial? In a small experiment I ran this spring across one hosted model family, the largest model scored the same with reasoning off. On high, it took three times as long, while the smallest gained about 10 points. If nothing moves, stop paying for the dial.
And is verification a line item? A budget with a row for model spend and none for building checks is funding the part of the system any competitor can buy tomorrow.
For most work that matters in most companies, the model stopped being the bottleneck a while ago. The bottleneck now is whether anyone can tell if its answer is right.​​​
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?
Comments (0)
No comments yet. Be the first to share your opinion!