Sep 24, 2026
The backup waited for a click
One morning the nightly backup had not happened, and two other jobs had sat idle for an hour and a half. The external drive they read was mounted and healthy. What held them was a permission dialog on my screen, raised overnight for a different program, and each of them moved within a second of my click.
When a nightly job fails, I expect to find a failure: a crash, a full disk, an error somewhere in the log. This morning had none. The backup that runs at two had used its whole two-hour allowance and been killed with nothing to show for it. Two jobs that start at a quarter to eight were still running at nine, at zero CPU, and between them they held two of the three slots the scheduler allows at once, so everything behind them started late. The external drive they all read was mounted, and a listing from my own terminal came back instantly.
The assistant sampled one of the stuck processes. It sat inside the call that opens a directory, waiting on the kernel, and had been for over an hour. A throwaway job, started by the system the way the background services are, stopped on the same line. So it was neither the jobs nor the drive. Something in the operating system was holding the directory reads those jobs made on that drive, and letting mine through.
Its first explanation was a missing permission: the database where the system keeps these grants seemed to say the program the jobs run under had never been allowed onto removable drives. It had, since May. The query had been cut at thirty rows, and the row that said so sat further down. I could have spent the morning granting a permission that was already there.
What the screen held was a dialog. The assistant’s command-line client updates itself almost every day, and the operating system treats each version as a new program, so it asks again whether that program may read files on a removable drive. I have answered that question sixty-one times since June. A scheduled job starts the client four times a day with a one-word prompt, and at ten the night before it was the first thing to run the new version. At startup the client reads the skills it can load, and five of them were links onto the external drive. The dialog went up with nobody in front of it, and the one-word prompt waited eleven hours.
I clicked Allow at 9:12:44. One of the stuck jobs finished at 9:12:45, the other a minute later, once it had done its work. The night before, the same dialog had appeared for the previous version; I answered it within half an hour, and nothing else got caught behind it.
There were two fixes, and neither was clicking faster. The links are gone: those skills reach the assistant through Sapix when a task calls for them, so the scheduled run has no reason to touch the drive. And the jobs that got stuck, along with the others that read the drive the same way, now give it thirty seconds to answer. On silence the backup fails with a line saying the drive did not answer and that a permission dialog is the likely reason; the other jobs skip the drive, say so, and carry on with everything else. The assistant tested it with a read built never to return: the backup failed after thirty seconds instead of two hours, and wrote nothing.
A wait with no deadline hands your schedule to whatever you are waiting on, and last night that was a dialog nobody could see. Waiting looked exactly like working: no error, no log line, just a process asleep. The deadline matters less than what happens when it runs out. Thirty seconds of silence now ends in a sentence that tells me where to look, which is more than two hours told me.
The part I have not finished is why a question about one program held reads made by another. Not every read that night was held, either: a job at a quarter to six listed the top of the same drive without trouble. In July a hang with the same look cleared only after a reboot, and I now suspect it was a dialog too. I cannot prove either yet.