Chapter 7 · The judgment gap and the paradox
What Tom read on Friday morning
Tom Brindle read one letter at Priya's demonstration, for an account he has covered since 2019, and said he would have been happy to receive it.
He meant it, and he was the right person to ask, and the letter was good. Correctly formatted, right invoice number, firm without being unpleasant, better than the template Ostermill had been sending since 2011. Tom knows more about how his customers receive things than anyone in the building, and his reaction was the most informed reaction in the room.
It told everybody one thing. The system can write a collections letter.
That is the entire content of what Tom learned, and it is the smallest of the things anyone needed to know that morning.
The letter is the output. The job is the decision to send it, and those two are produced by different capabilities that arrive inside the same artifact wearing the same clothes. Fluency comes out of the box. Judgment does not, and the box is very much better at the first than the second.
Ruth's forty-one holds a week are the whole job, and not one of them shows up in the quality of a letter. If the agent had drafted to the family shop five days after the owner's death, the letter would have been just as well written, Tom would have read it just as approvingly, and the only person in Ohio who would have known something was wrong was not in the room.
Why the good letter is worse than a bad one
The two are inversely reassuring, which is the part that makes this dangerous rather than merely difficult.
A badly written letter makes a room suspicious. Somebody says the tone is off, somebody else says it does not sound like us, and the project slows down while people look more carefully. That suspicion is useful. It is a free error-detection system that costs nothing and works.
A beautifully written letter to an account that should never have been contacted produces no suspicion at all, because every visible signal is excellent. The only thing wrong with it is invisible from inside the artifact.
In ordinary software this problem is not the tester's problem, and it is worth being precise about why, because the whole testing discipline rests on it. If an invoice total renders as four hundred and ten dollars, the arithmetic ran. You can inspect the output and learn something real about the system.
Here a correct-looking output is evidence about the language and nearly nothing about the world. The letter cites invoice 4471, four hundred and ten dollars, forty-one days past terms, and every one of those facts is right. Whether the letter should exist is a different question, and reading it more carefully does not extract the answer, because the answer is not in it.
Only somebody who knows that the owner died in July looks at that letter and sees the problem.
Which is the reason a person stays in this story after the agent arrives, and also the reason the person has to be a specific person rather than any competent reviewer. What is needed here is not general diligence. Part of it is twenty-two years of one company's accounts, and that part can be written down by anybody willing to sit with Ruth long enough to ask. The rest is that she knew the owner had died, which was never anybody's job to record, and which nobody would have thought to ask her for.
The paradox
Which brings the argument to the thing that makes this hard, in a way nobody can be blamed for.
The standard answer to everything above is to put a person in the loop. The agent proposes, Ruth approves, and Ruth catches what it gets wrong. Ostermill launches exactly that way, and this book recommends it.
It contains a mechanism that quietly undoes it, and the mechanism is not carelessness.
What Ruth does when she reads a draft is not one act but two. She checks whether the letter is any good, which she can do from the letter. And she asks whether it should have been written at all, which she cannot, because that answer is not in the letter. It is in what she knows about the account.
The second act is the expensive one, and a low error rate is what retires it. Give her two hundred drafts a week of which two are wrong, and she learns, correctly, that this queue is almost always fine. So she keeps doing the cheap act and quietly stops doing the expensive one. She is not reading carelessly. She is reading the letter instead of the account, and the letter is the one place these errors never appear.
The better the agent gets, the less she finds. The less she finds, the less often she looks past the letter. The less often she looks past it, the more the two get through.
Supervision degrades in proportion to how well the thing being supervised performs.
There is no version of that which resolves by asking people to concentrate. The automation literature has been describing this shape since the early nineteen eighties, and it is not confined to one field.
Priya's version of this is one line. A check that almost never fires is not a check. It is a habit.
She had a specific thing in mind when she wrote it, and it had nothing to do with agents.
Ostermill's ERP has a duplicate-invoice flag. It has had one since the 2014 project, it works correctly, and it fires on about forty invoices a month. Thirty-eight of those forty are legitimate re-bills, partial shipments invoiced twice on purpose, or the same part number on two lines of one order. The flag is right to fire. It is doing exactly what it was built to do.
The receivables team clears it in a batch on Thursday afternoons, four or five seconds an invoice, because after twelve years everyone knows what that flag is worth.
In 2023 a supplier's system hiccuped and issued the same invoice twice, six weeks apart, for fourteen thousand dollars. The flag fired. It fired on the Thursday, in the batch, between a partial shipment and a two-line order, and it was cleared in about four seconds by somebody competent who had cleared close to five hundred of them that year and had never once found a real one.
Ostermill paid it twice. The duplicate was caught the following March by the supplier, who mentioned it apologetically, and the money came back.
Nobody was disciplined and nobody should have been. The flag worked. The person worked. The system produced a fourteen-thousand-dollar error out of two components that were each performing exactly as designed, and the mechanism that produced it was the base rate.
This is why the paradox cannot be solved by choosing better people or asking them to care more. The two fields that have studied it longest, aviation and clinical monitoring, spent decades learning that lesson expensively. Both arrived at the same conclusion from opposite directions: you cannot train vigilance into somebody watching a system that is almost always right, because what they are learning from the system is true. The response in both fields was to change the work rather than the worker. Fewer alarms with more meaning behind each one. Alarms that state what they think is wrong rather than that something is. And, in the places that got furthest, measuring how fast people clear them and treating a falling number as the leading indicator it is.
Ostermill has a duplicate-invoice flag and no idea how fast anyone clears it.
What follows from it
Two things follow, and they are why Part Two gives the supervisor's half of Design as much space as the agent's.
The first is that a human in the loop is not a safety measure. It is a design surface, and the safety comes from how it is built rather than from the human being present. A confirmation box asking whether to proceed transfers liability without transferring understanding. A queue showing only the agent's conclusion trains approval, and faster than anybody expects.
The second is that Ruth's new job has to be designed against the paradox rather than in ignorance of it. That means showing her the strongest reason the agent might be wrong rather than only what it decided. It means making a reject as cheap as an approve, because a system where rejecting takes four clicks and approving takes one is a system with an opinion. It means measuring how long she actually spends and treating a falling number as a warning rather than as efficiency.
And it means accepting a cost nobody puts in a business case. Supervision that works is slower than supervision that does not, and the version that catches things will always look worse on every throughput metric than the version that rubber-stamps. Part Two makes that trade explicit, in the phase where the review time is budgeted, and it is the moment where the number a supervisor is given stops being a preference and becomes a design constraint.