References and Sources
Front matter and Part I: Decide
Field snapshot. Time-reallocation evidence (a few hours per week for a typical heavy AI user; closer to a day for the top tier; routine falling from roughly half the week to a third, with part of the freed time absorbed by checking the agent): synthesized across multiple 2026 industry productivity studies and presented in the text as an informed estimate, not a single measured finding.
AI literacy (Chapter 1). The four-layer stack: author framework. BIDMC / o1-preview diagnostic comparison (model 67.1% exact-or-close at triage vs 55.3% and 50.0% for two internal-medicine attending physicians, 76 ED cases): Brodeur et al., “Performance of a large language model on the reasoning tasks of a physician,” Science, April 2026 (the same study developed in full in Chapter 15). Supervision paradox preview: Bainbridge (1983) and the modern deskilling literature, fully cited in Chapters 17-18. Constitution as the agent’s behavioral contract: the term originates in Anthropic’s Constitutional AI (Bai et al., “Constitutional AI: Harmlessness from AI Feedback,” arXiv:2212.08073, 2022); “Claude’s Constitution,” Anthropic, 2023, updated with a new public-domain constitution January 2026 (anthropic.com/constitution). The operational, agent-governance reading (purpose, prohibitions, escalation, value-priority, runtime enforcement) reflects 2026 enterprise usage, e.g. “agentic constitution” as policy-as-code with ethical boundaries an agent must never cross (CIO, January 2026). The prompt/policy/guardrail/constitution distinction is the author’s framing. The note that GitHub’s Spec Kit overloads “constitution” for engineering build-discipline is drawn from the Spec Kit constitution command/template (github.com/github/spec-kit) and is flagged as a homonym, not the accepted sense.
Not a bridge (Chapter 2). Bridge operator vs architect, the Autonomy Ladder, two-channel design: author frameworks.
Not every problem (Chapter 3). Token-price decline (~10x/yr; a frontier chat model ~$36/M tokens in early 2023 to ~$1-2/M) and the architecture multiplier (the real bill is architecture, not the model’s sticker price): industry pricing data, figures stated as a range and stamped to mid-2026. Four suitability tests and earned-vs-scheduled autonomy: author frameworks. Utah Doctronic (a 2026 prescription-renewal pilot under a regulatory sandbox; NEJM Perspective NEJMp2601148): the “autonomy on a schedule” reading is the author’s interpretation, stated as such where the pilot documents do not explicitly frame it that way. Klarna (assistant reported as doing the work of ~700 agents, then rebalanced toward humans): company-reported figure (OpenAI/Klarna), labeled as a company claim, not an independent measurement. MVP house of cards: author article, “The House of Cards: Why MVP Thinking Is Breaking Healthtech” (2026).
Part II: Prototype & Collaborate
How the Work Splits (Chapter 4). The two-channel operating model, the dual brief (Human / Executable), and the work-unit shift from user story to outcome-centric spec graded by evals: author frameworks, synthesized from “The Supervisory Layer: A System-2 PM” and the spec-driven-development pattern, the clearest 2026 instance of which is GitHub’s open-source Spec Kit (github.com/github/spec-kit; “Spec-driven development with AI,” The GitHub Blog, 2026). The Executable Brief and its Appendix A template draw on that pattern’s structure, numbered testable requirements, explicit clarification markers, given/when/then acceptance, and a pre-build gate, while adding the supervisory (Channel 2) requirements the general tooling does not supply. The refund agent runs as the worked example from here through the design, operate, and human-system chapters and is worked in full in Appendix A.
Vibe coding (Chapter 5). Vibe-coding tooling and the AI-generated-code vulnerability figures: the elevated-vulnerability finding draws on Perry et al. (Stanford, 2023) and subsequent 2024-2025 code-security audits (Veracode 2025 GenAI code-security report; Cloud Security Alliance vibe-coding research note); figures vary by study and method, and the text states the range as such rather than as a single precise rate. Package-hallucination / slopsquatting (~1 in 5 recommended packages not existing): Lanyado / slopsquatting research, 2024-2025. Lovable audit (a featured app with inverted authentication exposing ~18,697 records): security-researcher disclosure and trade coverage (Vibe Graveyard, 2025-2026); carried as a category problem, not a single-vendor indictment, and as a researcher’s report rather than an adjudicated finding. Prototype-as-learning-instrument and the opportunity-solution-tree framing: author framework and the product-discovery literature.
Collaborator (Chapter 6). The six runtime ownership domains and the kill-switch-cannot-be-owned-by-the-agent’s-own-team point: author framework, drawing on 2026 agentic-risk-management standards work. ChatGPT-Health under-triage of gold-standard emergencies (~52% routed to 24-48h care rather than the emergency department): Mount Sinai (Icahn School of Medicine), “ChatGPT Health performance in a structured test of triage recommendations,” Nature Medicine 2026 (s41591-026-04297-7); a structured test of 60 clinician-authored vignettes across 16 conditions (960 interactions), not live deployment. Cross-referenced to Chapter 9. Four co-owners of the launch decision: author framework.
Part III: Design
Design the behavior (Chapter 7). System-type declaration, four runtime artifacts, consequence classification, approval-as-authorization: author chapter (v3). The placement rule for the autonomy boundary (gate only where the expected value of human judgment outweighs its cost; hand control back under low confidence) draws on Eric Horvitz, “Principles of Mixed-Initiative User Interfaces” (CHI 1999) and the decision-theoretic framework behind Microsoft’s Guidelines for Human-AI Interaction. The enforcement principle (a boundary must be enforced in the execution path, not stated in the prompt): author framework, illustrated by the PocketOS incident. Current security frameworks, OWASP Top 10 for Agentic Applications (late 2025, ASI prefix), MAESTRO (Cloud Security Alliance, 2025, seven-layer), and the CISA / Five Eyes “Careful Adoption of Agentic AI Services” (April 2026), plus the five PM security decisions: synthesized from those primary framework documents. PocketOS nine-second deletion (a flagship-model coding agent deleting a production storage volume and its same-volume backups via an over-scoped infrastructure token, after which the agent produced a written confession enumerating the safety rules it had violated): primary first-hand account by the company’s founder, Jer Crane (@lifeof_jer), X, 25 April 2026 (https://x.com/lifeof_jer/status/2048103471019434248); corroborating trade coverage (Tom’s Hardware and others). Vendors described by role rather than named in the text, deliberately. This is a separate incident from the 2025 Replit production-database deletion (Chapter 10); the two are kept distinct and not conflated.
Two kinds of HITL (Chapter 8). Program-level vs transaction-level oversight: new framework. Regulatory tiers: EU AI Act Article 14 (human oversight; Article 14(5) two-person verification for certain remote biometric ID); California SB 1120 (signed 2024, effective Jan 2025; physician determination of medical necessity, not AI alone); the CMS prior-authorization pilot; credit adverse-action reason requirements; financial-services sign-off expectations; and the Colorado AI Act (SB 24-205) right to human review. Primary law where available. Cigna PXDX claim-review allegations (physicians signing >300,000 denials over ~2 months, ~1.2s each; of appealed denials, ~90% alleged overturned): ProPublica (2023) and Kisting-Leung v. Cigna class action; litigation-stage, presented as alleged. The text separates the rapid-denial figures from the appeal-reversal rate and does not conflate the latter with an error rate.
Evals (Chapter 9). The three eval breaks, pass@k, judge-bias, state-vs-semantic validation, coverage statement, model-update-as-deployment: author chapter (v3). The compounding arithmetic (0.95 to the tenth ≈ 0.5987): exact; presented as the limiting case assuming step independence, with the text noting real systems drift from it. DAX Copilot RCT (high adoption and improved physician task-load ratings, but no statistically significant change in documentation time vs. control: DAX time-in-note -1.7%, p=0.66): “Ambient AI Scribes in Clinical Practice: A Randomized Trial,” NEJM AI 2025, a three-group pragmatic trial of 238 outpatient physicians across 14 specialties. The text’s point is that the metric the business case rested on (time saved) did not move, even as adoption and comfort did.
Part IV: Operate
Operational guardrails (Chapter 10). Platform cost-control reality (no major provider ships a hard cost kill-switch; the platform-vs-team line; pre-call vs post-call enforcement; circuit breakers; backoff; burn-rate tiers): synthesized from platform documentation and 2026 agent-ops practice. The $47k multi-agent loop (eleven days, escalating weekly spend, every span green, monthly alert as “receipt not brake”): illustrative composite assembled from documented runaway-cost patterns. The Replit production-database deletion under an unenforced code freeze (Jul 2025; Replit CEO public apology; AI Incident Database entry; documented by the user, Jason Lemkin): public record. Distinct from the PocketOS incident (Chapter 7). The widely circulated “$340k weekend” was found to have no traceable origin and is deliberately not used.
Observation (Chapter 11). The observation phase, the six instruments, platform-emits-PM-composes, the metric-is-the-design reversal, data observability: author chapter (v3). The procurement agent (340 transactions, fully confirmed, none delivered, six months): illustrative composite, assembled from documented background-failure patterns.
Silent degradation (Chapter 13). Silent degradation, the six drift vectors, the currency question, instrument half-life, the three disciplines: author chapter (v3). Epic Sepsis Model: the v1 external-validation critique (vendor-claimed discrimination in the low 0.8s; externally validated AUROC ~0.63 with low sensitivity) is Wong et al., JAMA Internal Medicine 2021. A later validation of v2 (JAMA Network Open, 2026) reported higher AUROC; the text distinguishes the v1 critique from v2 and does not claim v2 mostly alerts after clinicians act unless that specific finding is sourced at typeset. GPT-4 behavioral drift (prime-number / benchmark task, ~84% to ~51% over three months, a sibling model improving on the same task): Chen, Zaharia & Zou (arXiv 2023; Harvard Data Science Review 2024); labeled as a benchmark task in a controlled comparison, not a product metric, and noting the paper’s methodology was contested by some critics. AI model aging across many real-world datasets: peer-reviewed, 2022.
Audit trails (Chapter 14). Decision provenance, the sealed decision artifact, the five audit-trail requirements, retention-floor-vs-litigation-timeline: synthesized from EU AI Act Articles 12/19/26, GDPR Article 22, sector retention rules, and the 2026 LLM audit-trail academic framework. Estate of Lokken et al. v. UnitedHealth Group (D. Minn., filed Nov 2023; nH Predict algorithm in Medicare Advantage post-acute denials): the figures (of appealed denials, ~90% alleged overturned; ~0.2% of patients appealed) derive from the complaint and STAT News reporting; plaintiff allegations, not adjudicated facts. The text presents the ~90% as an appeal-reversal rate, alleged, and does not equate it with a system error rate.
Part V: The Human System
Change management (Chapter 16). Actor-to-supervisor transition, two-channel design, the supervisory system as second product: author frameworks. Environmental-reversion finding (high stated intent to adopt a reflective AI practice in a workshop, near-total reversion to extraction-style prompting back in normal work): 2026 field study of experienced professionals; reported qualitatively, the original small-sample percentage is deliberately not stated because the primary source could not be independently confirmed. United 173 (1978) as a catalyst for Crew Resource Management (developed by United/NASA, 1979-1981): aviation human-factors record.
Why HITL fails (Chapter 17). The Loop Test (time, skill, attention): new framework. Supervision paradox, two-stage supervision collapse, automation-expectation-vs-complacency: Bainbridge (1983) and the modern literature. ACCEPT / Budzyń colonoscopy result (adenoma detection falling after AI exposure): peer-reviewed. Cigna and UnitedHealth nH Predict figures: litigation-stage, alleged.
Skill erosion (Chapter 18). The three erosions (deskilling, never-skilling, cognitive surrender), validator-of-validator regress, the four apprenticeship conditions, the AI literacy gradient: author framework plus literature. Deskilling: Budzyń et al. / ACCEPT trial, Lancet Gastroenterology & Hepatology 2025 (unassisted adenoma detection 28.4% → 22.4% after AI exposure; observational). Never-skilling: Bastani et al., PNAS 2025 (high-school math; GPT-4 access improved practice but degraded subsequent unaided performance, guardrailed “tutor” version did not). Cognitive surrender (acceptance of incorrect AI answers reaching ~80%, 79.8% in the relevant condition, across 1,372 participants): Shaw & Nave, “Thinking, Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender,” Wharton, 2026.
The agent as team member (Chapter 19). The agent HR stack and its constituent constructs (persistent contractor with API access, the missing category, non-local cultural defaults, costume vs actor, configuration as onboarding, the missing supervisor): author frameworks. Agent-management and personnel-practice material: synthesized from 2026 agent-operations and personnel-practice writing. “Be Brief, Be Bright, Be Gone” and the polite-disagreement-in-the-wrong-country case: author article material.
Part VI: Carry the Weight
The people the agent never sees (Chapter 20). User-vs-affected-person, disparate performance, the constitutional runtime layer, moral architecture vs committee: author chapter (v3). Maria (a clinical summary compressing away risk factors): illustrative composite clinical case. The supervision-paradox-meets-regulation argument and the “AI supervisory fallacy”: cross-referenced to Chapters 17-18; BMJ Digital Health 2026 and the deskilling literature.
A day in the life (Chapter 22). Composite week, no external sources; built from the artifacts of Chapters 3-20 and the spine.
Back matter
Field Manual. Assembled from the closing exercises of each chapter. Toolkit. PM tooling landscape: surveyed mid-2026; tool names are dated illustrations. Glossary. Assembled from the book’s named constructs.