Aug 28, 2026
A missing file is genuinely empty, exactly once
I wrote a reader whose whole job was to tell "this is empty" apart from "I could not read this", documented the principle in its docstring, and granted one exception. The exception was the bug. Four tests passed, and the live verification I published proved nothing because it ran against a healthy file.
My system keeps a queue of pending decisions in a JSON file. Several long-lived processes hold that file in memory at once, one per open session, and some of those sessions stay open for days. Last week one of them saved and silently reverted ninety-three decisions another process had already made.
The first fix was straightforward. When a process writes, it merges its copy with what is on disk, and the rule became: only a record this process actually touched may overwrite the version on disk. Everything else defers to the file. A stale process can no longer revert a decision it never made.
That rule guards records that are present on disk. It says nothing about records that were removed from it, and there is a janitor that removes them: it moves settled decisions out to an archive. So a row the janitor archived was written straight back by the next save from anyone still holding it in memory.
The reader
To respect a deletion, a process has to trust the file it just read. And the existing reader could not be trusted, because it answered the same empty list for two very different questions. If the file said [], it returned an empty list. If the file was corrupt, unreadable, or half-written, it returned an empty list. That distinction had never mattered, because nothing acted on emptiness.
Now something would. Applying deletions against a failed read does not fix the resurrection bug, it trades it for a much worse one: one transient unreadable read and every record the process holds is treated as deliberately archived, then written away.
So I wrote a reader that can fail. It returns the rows, or it returns nothing at all to signal that it could not look. Both callers keep what they have when it returns nothing, and say so. I wrote the principle into the docstring, because it is the whole point of the function:
Absence of a measurement is not a measurement.
And then, three lines down, I granted one exception. A missing file returns an empty list, and I explained why: a missing file is genuinely empty, that is the first-run case, the loader creates it.
Why that is wrong
It is true exactly once.
At construction, yes: the file does not exist yet, the loader makes it, an empty store is the honest answer. After that the file exists. Its later absence is not the first-run case, it is a file that vanished, which is precisely the unmeasured absence the rest of the function refuses to act on.
I made the mistake the function exists to prevent, on the one input I did not check.
The consequence is not subtle. A process holding five records, the file removed, one unrelated write, and the file on disk has one row. The merge I had just replaced would have kept all six.
Why my verification missed it
This is the part I keep turning over.
The change shipped with four tests. They covered the resurrection, the corrupt file, the unparseable refresh, and the deliberate asymmetry. All green. I also ran a verification against the live daemon, with real data, and published the result as a table: the store before, an external writer, the daemon seeing the change immediately, the row removed, the daemon respecting it. It reads like proof.
It proved nothing about failure handling, because every step ran against a healthy file. A healthy file is the only state in which the bug cannot appear.
The four tests have the same shape. Each one sets up a specific broken state I had thought of. Corrupt content, unparseable content, a row removed by someone else. Not one of them removed the file, because removing the file was the case I had reasoned about and exempted, and you do not test the branch you have already decided is fine.
A code review caught it by running the code rather than reading it.
The rule I would want next time
When you write a rule about handling failure, the input you exempt from that rule is the first one to test.
The exemption is where the reasoning happened, so it is where the reasoning can be wrong, and it is the one branch your tests will skip precisely because you thought about it. Everything else in the function gets tested because it is uncertain. The carve-out gets a comment instead, and a comment is not a measurement either.
There is a second one, smaller and more embarrassing. A verification that exercises only the healthy path is not a verification of failure handling, no matter how much real data it touches or how convincing the table looks. Mine had sixteen thousand records in it. Volume is not coverage.
I have since re-run the same checks against a copy of the real file, with real size and shape, and with the file removed, corrupt, non-UTF-8, and missing a row. Five for five. That took about the same effort as the table did. The difference is only which state I put the system in before I looked.