Back matter

Sources and status notes, by chapter

### Chapter 7 · The judgment gap and the paradox · sources

The vigilance argument. The shape described here, that detection degrades on a task where almost nothing is ever wrong, is the classic finding of the automation literature rather than a claim of ours. Lisanne Bainbridge, Ironies of Automation, Automatica 19(6), 1983, pages 775 to 779, DOI 10.1016/0005-1098(83)90046-8, is the canonical statement of it and of the related point that automating the easy parts of a task leaves a person with the hard remainder and less practice at it. The chapter states the shape and deliberately quotes no rate, because published vigilance decrements vary widely with the task and a borrowed number would be precision this argument has not earned.

Aviation and clinical monitoring. Both are named as domains that responded by redesigning the work rather than by exhorting the worker. That is a characterization of how those fields moved, not a citation of a specific intervention, and it is stated at that level on purpose.

Evidence tier. The duplicate-invoice flag, the fourteen-thousand-dollar double payment and the supplier who returned the money are Ostermill. They are illustrative composite, not a documented incident, and they carry the shape of a failure class rather than evidence for it. Nothing in the chapter's argument rests on them; they are there because the argument is easier to see with a concrete case attached.

### Phase 1 · Decide · sources

This phase cites nothing, and that is the note. Decide is the one chapter in Part Two built entirely out of the case. No study, benchmark or published finding is invoked in it, because the question it answers is not what is known about agents in general; it is what one company's own week of work turns out to contain when somebody counts it. The reader who wants the research is pointed to Part One, which carries it, and to Design's note, which carries the automation-complacency literature this phase's kill number anticipates.

No number in this phase is external. The $8,400, the $310 and the $1,150 are all constructed, on three different bases, which Exhibit B's own note says. The $1,150 is an all-in figure: a few cents of model compute, ninety seconds of loaded review on every draft, and a rework tail. The only market-priced input is the compute, it is the smallest term in it, and it is the only one that falls, which is exactly what the chapter's argument turns on. The figure is the arithmetic of a decision taken at one moment, and the point of the passage survives the number changing, because the argument is about which side of the ledger is measured and which is judged.

Evidence tier. Ostermill does not exist. Chapter 2 declares the composite and states the trade both ways. Nothing in this phase is offered as a documented incident, and the work reconstruction in Exhibit A is a constructed artifact in the shape of a real one, not a transcribed record.

Numbers that must reconcile with the case file. The week of 3 March is 312 accounts: 220 written to, 41 held, 38 escalated, 13 written off. The 41 break down 17, 9, 6, 5, 4, and the five with disputes are the five whose dispute field is empty, which is the whole of the argument against a deterministic filter. The alternatives are $8,400 for clerk time, $310 for the module and $1,150 for the agent, and the three are not on one basis, which Exhibit B's own note says. The prototype run is twenty cases, twelve drawn from the holds: the judge writes to 7 of the 12 and is confident on 5, the sorter writes to all 12 and catches nothing, the triager writes to 3 and names a signal on the other 9. The stopping rule was more than 3 of the 12, written before the run, and the triager met it at exactly 3.

Deliberately absent. No payback period, no year-one total, no cumulative fixed cost. The chapter argues that a ratio built from a measured numerator and a judged denominator hides which side each number came from, so it does not then compute one.

### Phase 2 · Design · sources

Automation complacency and the irony of automation. Lisanne Bainbridge, Ironies of Automation, Automatica 19(6), 1983, pages 775 to 779, DOI 10.1016/0005-1098(83)90046-8. The chapter names her and the year in the text. Her argument is that automating the tractable parts of a task leaves the operator the intractable remainder, less practice at it, and a monitoring job people are poorly suited to, which is the mechanism under this phase's approval moment and under the one-day-a-week clause that arrives three phases later.

Warning fatigue. The chapter states that people dismiss interrupting warnings at high rates even when the warning blocks their work. This is the well-replicated alert-fatigue finding from clinical decision support, and the chapter deliberately quotes no override rate, because published rates range across a wide band and depend heavily on the alert class. The argument needs only the direction, and the direction is not in dispute.

Evidence tiers. Every artifact in this phase is Ostermill and therefore designed teaching material: both briefs, the wall sort, the approval screen, the measured revocation time. They are internally consistent and were checked to reconcile. They are not findings, and the eleven-minute revocation in particular is a number chosen to make a point about rehearsal, not a number anybody measured.

Numbers that must reconcile with the case file. 220 sends a week at a 90-second budget is 5.5 hours, which is the arithmetic the interruption budget exists to answer. Four walls and one fence, with the fence in its own table. R. Vaughn receives the ordinary queue and the wall stops; M. Ellery receives at most 25 items a week.

### Phase 3 · Prove · sources

Evidence tiers are load-bearing. Everything Ostermill does is a designed composite and is declared as such in the front matter. Everything below is external and documented.

A note on how numbers are used in this chapter. The prose teaches with arithmetic the reader can do in their head and check without trusting anyone: nine in ten, eight times over, is worse than half. Published figures appear once, unnamed and unquantified, only to establish that the effect is measured rather than merely argued. Precise numbers live in the exhibit, where they belong, because an exhibit is a document and documents carry figures. No argument in this chapter depends on a percentage the reader has to accept on authority.

pass@k. Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. de O., et al., "Evaluating Large Language Models Trained on Code," arXiv:2107.03374, 7 July 2021. Defines pass@k as the probability that at least one of k sampled solutions is correct.

pass^k. Yao, S., Shinn, N., Razavi, P., and Narasimhan, K., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains," arXiv:2406.12045, 17 June 2024. Introduces pass^k as the probability that all k trials of the same task succeed, as a reliability measure rather than a capability measure.

"Fall by more than half on the same tasks." Supported by τ-bench above and, independently, by Mehta, S., "Beyond Accuracy," arXiv:2511.14136, 18 November 2025, which reports agent performance dropping from 60 percent on a single run to 25 percent on 8-run consistency. The chapter deliberately does not quote either figure, because the model tested in 2024 is long superseded and the benchmark now has a successor. The gap is the durable finding; the scores are not.

Judge bias. Position, verbosity and self-preference effects are documented, with Zheng et al., MT-Bench, NeurIPS 2023, arXiv:2306.05685, as the standard reference. Calibration by agreement against a multi-annotator gold set is current practice, not a proposal of ours.

The ambient scribe trial. Lukac et al., Ambient AI Scribes in Clinical Practice: A Randomized Trial, NEJM AI, November 2025, DOI 10.1056/AIoa2501000. Three-arm pragmatic randomized trial, 238 outpatient physicians, 14 specialties, run November 2024 to January 2025. Primary outcome was the change in logged time-in-note. One scribe arm fell about 9.5 percent against control and the difference was significant; the other showed no significant change. The chapter uses the contrast between the two arms rather than either arm alone, because that contrast is what the trial actually reports and it is the stronger illustration: a team holding only adoption and sentiment data could not have told the two products apart.

Composites, declared. Everything Ostermill does is designed teaching material, including the eval figures in Exhibit F and the judge's kappa. They are internally consistent and were checked to add up, and they are not findings.

### Phase 4 · Observe · sources

The METR study. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, arXiv:2507.09089. Sixteen experienced developers, 246 real issues, randomized by issue. Beforehand they forecast a 24 percent speedup; measured, they were 19 percent slower; asked afterward, they still believed they had been about 20 percent faster. The chapter rests on the persistence of that belief rather than on the size of the slowdown, and states the study's limits in the text rather than hiding them here.

Not used, deliberately. The 1.5-second per-claim figure from the medical-necessity litigation. It is derived from alleged volumes rather than from findings, and this chapter already has a measured distribution of its own that makes the same argument without borrowing an allegation.

### Phase 5 · Operate · sources

Numbers that must reconcile with the case file. Eval set: twelve at launch, thirteen after the August incident, nineteen at twelve months. Two model swaps. Rung 4 for one case class, and the rung 5 write-off authority restored for another. The forty-one from Ruth's week of March third moves to about thirty-six as records improve, and never to zero. The unset dispute flag appears three times: the near-miss in March, the revoked write-off authority in week eleven, and the reconstruction request in May.