sapix technical notes
← all notes

Sep 11, 2026

The gate blocked the right answers

I built a guard that reads my assistant's reply, looks for anything in my notes that the reply contradicts, and stops it before I see it. It had been running quietly for weeks, recording what it would have blocked. I finally read the log. It was not catching mistakes. It was preferentially catching the answers that had followed my rules most carefully.

I have a set of rules I have written down over the past year. Not preferences, rules: never use this tool for that kind of email, always check the console before trusting the dashboard, that sort of thing. They accumulate because each one came from something going wrong once.

The problem with written rules is that they have to be read. So I built a guard that works from the other end. After my assistant finishes writing a reply, and before I see it, the guard searches my notes for anything the reply might be contradicting, and asks a model: does this answer break this rule?

I did not arm it. It ran for weeks in a mode where it recorded what it would have blocked and blocked nothing. That was deliberate, and I am glad of it, because when I finally sat down with the log the answer was not the one I expected.

Over one week the guard had flagged 227 replies. I took a stratified sample of twenty and read each one properly, the reply and the rule together. One was a genuine catch. Sixteen were wrong. Three I could argue either way.

That alone would just be a bad model. What made me stop was the shape of the sixteen.

One reply had used a specific configuration value in a piece of code. My notes contain a rule that says: always use that exact value. The guard found the rule, saw that the reply was about that subject, and flagged it. Another reply had gone and checked a live console rather than trusting an email. There is a rule in my notes that says to check the console rather than trusting the email. Flagged.

The guard was not distinguishing between an answer that breaks a rule and an answer that is about the same topic as a rule. And because an answer that follows a rule tends to be about that rule’s subject, the guard was systematically catching the most careful work. Armed, it would have blocked, preferentially, the replies that had done exactly what I asked.

There was a second failure underneath, and it is worse in a quiet way. Some flags came from a rule that was itself out of date. One cited a password manager I stopped using in June. Another demanded a practice I had explicitly retired. The guard was holding up perfectly good answers against evidence that had expired, and there was nothing in its design that could notice, because it had no concept of a rule going stale.

I turned the semantic half off. The simple half stays: a plain text check for a typographic habit I have asked it not to use, which has blocked over two hundred real instances and cannot produce this failure because it does not judge anything. It matches a character.

What I keep from this is not that the model was bad at judging. It is that I had built something whose failure mode was invisible from the inside and correlated with quality. A guard that fires more often on better work does not feel broken while it runs. It feels vigilant.

The only reason I know any of this is that it spent weeks writing down decisions it was not allowed to act on, and that I eventually read them one at a time instead of looking at the total.