How I stopped trusting “great results.”
Yet again, the AI told me it had solved my problem. It had graded my Wheat Cents perfectly. To prove it, it produced the charts. As usual, the claim quickly crumbled under pressure.
The Convincing Failure
This is the part of building with an AI that nobody warns you about. Not that it fails, but that it “succeeds” so relentlessly, so cheerfully, with such a well-organized summary of everything that improved that it SEEMS correct. Reams of charts and graphs and numbers to “prove” it’s right. An assistant that has never once entertained the possibility that it might be wrong.
The evidence was rarely fabricated. That’s one reason why it’s so convincing. Like P-hacking, it was correct and valid data. It’s just that it had a goal and ran hundreds of experiments until one of the results met that goal. The chance of that goal appearing by accident is less than 1 in 100. Guess what? If you run 200+ experiments, the chances of at least one false positive are very high.
The AI measured on the Wheat Cents it had to work with, scored on the slice that worked and framed against whichever comparison flattered it most. Every number in the summary was defensible on its own. The whole picture was wrong.
Lies can be more easily caught than a true number measured on the wrong coins looking exactly like good news. News you desperately want to hear.
So I started scrutinizing results more and more carefully. For months I read through results asking where a number came from, what it was measured against, what quietly got left out. This method works. It is also the most tiring thing I have ever done, and it does not scale. When I got tired, and really wanted a result to be true, that’s exactly when things slipped through and later became problems.
You cannot audit your way out of a problem that produces new evidence faster than you can read it.
Building the Gate
So I stopped arguing about evidence afterward and moved the argument to before. I built myself a “bureaucrat.” A system that forced me to go through an onerous series of tests before the result could be accepted.
The clear goal and measurement now gets set first. Before any work starts on an improvement to CoinAILyzer, we write down what would count as success: which coins, measured how, against what. Then the work happens. When we think we have an answer, we record the methodology, the results, and the images used. Then it runs against a set of carefully vetted coins the model has never seen.
And here is the piece I did not expect to build, and the reason I am writing this at all. Those coins are walled off so that I cannot see them either.
The Person I Didn’t Trust
This is because I found the AI wasn’t the only one so desperate to get passing results that it would tweak the data and measurements to pass. I was guilty of it as well. If I can see the test set I will drift toward it. Not deliberately. I just found I’d make a hundred small choices over a hundred days with those particular coins somewhere in the back of my mind, and one morning you have a model tuned to an exam instead of a model that grades coins.
A test you can see is a test you eventually teach to. But you cannot move a bar that was written down before the attempt.
Bureaucracy as a Feature
Of the improvements that reach that gate, roughly four in five fail. Usually abysmal failure, not a near miss. Things where the numbers ahead of time look really solid. But then they get tested against coins I’ve never seen and they don’t work at all.
The bigger change is that because we know the gate is coming, we now test far harder before we get there. A huge percentage, 98%+, of experiments get killed by our own preliminary checks first. Our own screening is nowhere near as good as the real gate. But it means the largest thing the gate produces is not the errors it catches. It is the work that never gets proposed, because we already went and checked.
That process has radically improved how I build CoinAILyzer.
Anyone in a corporation has seen it. The manager takes the numbers and spins them for themselves in the best possible light. Their project is months behind. That gets only a minor mention while they spend an hour highlighting the one portion just completed (and fail to mention it hasn’t been tested yet).
The machine only does this faster, and with better charts.
I call the whole apparatus the ML bureaucrat, because that is exactly how it feels. I built a system whose entire job is to make me fill out form after form before it will let me do the thing I had already decided to do. I have spent a career fighting bureaucracy, most of which exists so that no one has to be accountable for a judgment. This is my one exception, and the reason is narrow. When a broken AI and a working AI produce identical outward signals, process is the only instrument left. Twenty years of experience cannot see it. And you certainly cannot trust the confident summary somebody handed you, least of all when that somebody is you.
Most of what dies at that gate stays dead. A few things get fixed and pass later. A surprising number get picked up later for another purpose. Many do not, which means a great deal of confident, plausible, well evidenced improvement was never real, and I would have shipped nearly all of it.
I built the gate to keep the machine honest.
It has spent most of its time keeping me honest instead.

Leave a Reply