Sep 14, 2026
It checked four percent
After I turned off the guard that was flagging my best answers, I went looking for why it had been so bad. I assumed the problem was the judgment. It was not. The part that chooses which rules to even consider was handing over the same fourteen every single time, out of two hundred and ninety-nine.
Last week I switched off a guard that had been flagging the answers that followed my rules most carefully. I wrote about that. What I did not know yet was why it had been so bad at its job, and my first assumption was wrong in a way I want to record.
I assumed the problem was judgment. The guard asks a model whether a reply breaks a rule, and models are uneven at that sort of thing, so I went in expecting to tune a prompt or try a better model.
Before doing that I checked something duller: which rules was it being shown?
The code took my rules from the database and passed the first fourteen. Not the fourteen most relevant to the reply. The first fourteen in whatever order the database returned, identical for every question, every session, every day. There are two hundred and ninety-nine.
So the guard was checking a little under five percent of my rules, and always the same five percent.
The rule I care about most that week sits at position two hundred and eighty-four. It says that the store where all this lives is called one thing and never another, which is a small thing that I had corrected the assistant on repeatedly. A piece of text that broke that rule four times went past the guard twice and came back clean, with the model running normally and answering honestly about the fourteen rules it had been given.
The retrieval was never the weak link, which is the part I had diagnosed wrong out loud. When I tested it properly, the search returns the right rule in sixth place out of ten candidates and second out of fifty. Finding the rule was never the problem. Nobody was ranking them.
And the piece that would have done the ranking already existed. Some months ago I had classified every rule by when it applies: some hold in every conversation, some only when the work touches their subject. The morning briefing already reads that classification. This guard did not. Same data, sitting in the same database, unused by the one component whose whole job depended on it.
There is a pattern across the three fixes that week and it is the thing I want to keep. The first: the reason a decision was made died inside a logging call that dropped the field. The second: a pair of annotations designed to be read together, where the code applied one and not the other. The third: this, a selection step that did not exist.
Not one of them needed a better model, a cleverer prompt, or a tuned threshold. All three were pieces that had been built, were correct on their own, and were not connected. Every component passed its own tests. No unit test sees this, because a unit test asks whether a part works, and the answer was yes in all three cases.
I keep finding that the expensive failures in this system are not wrong parts. They are right parts with nothing between them.