Guide
Human-in-the-Loop AI: When Human Review Actually Helps
Your team is about to let an AI system draft decisions that people used to make from scratch. Someone proposes the safeguard everyone expects: “A human stays in the loop.” The room relaxes, and the meeting moves on.
That sentence ends the discussion without answering the question behind it. Does the reviewer make the decisions better than the AI would on its own? Can the reviewer spot the errors that matter, or does a person mostly add time and a signature? And when review does help, is the gain worth the hours it consumes?
This guide shows you how to answer them before you write the policy. You compare three conditions (human alone, AI alone, and human plus AI), record six measures, and decide with a worksheet.
What human-in-the-loop AI means in this guide
Human-in-the-loop AI here means a workflow in which a person reviews, edits, approves or overrides an AI system’s output before it takes effect. The output might be a draft decision, a classification, a summary or a recommended next step. The person might check every case or only the cases routed to them.
The phrase has a second meaning in machine learning: people labeling data or rating outputs to train and improve a model. That is a different job with different measures, and this guide does not cover it.
Why “better than before” is the wrong test
Most teams judge an AI-assisted process against the process it replaces. That comparison skips a third option: letting the AI’s output stand without review. A 2024 meta-analysis in Nature Human Behaviour by Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone names the difference. Human augmentation means the combination beats the human working alone. Human-AI synergy means it beats both the human alone and the AI alone.
The researchers pooled 370 effect sizes from 106 experiments published between January 2020 and June 2023, each measuring all three conditions. Using a standardized difference (Hedges’ g), they found that combinations beat humans alone by a medium-to-large margin (g = 0.64). But combinations did worse than the better of the human or the AI alone (g = −0.23, a small difference).
Two moderators matter. Combinations lost ground in decision tasks, where people chose among fixed options, and fared better in creation tasks such as writing, though that estimate was less certain. When people alone outperformed the AI, combinations tended to beat both (g = 0.46); when the AI outperformed people, combinations tended to fall short of the AI (g = −0.54). The authors suggest that people who are less accurate than the AI are also poor at judging which of its answers are wrong.
An experiment by Ángel Alexander Cabrera and colleagues shows the pattern and its limits. Participants used a simulated AI fixed at about 73% accuracy, described to them as 90% accurate. In detecting fake hotel reviews, people alone scored about 55% and people shown the AI’s predictions about 62%, still below the AI. In classifying bird images, people alone scored about 81% and people shown the predictions about 83%; people who also saw where the AI tends to fail reached about 90%, beating both.
On satellite images, people alone were already about 90% accurate, and no AI condition reliably changed that. (These values are read from the paper’s chart.) Who is better alone predicts the direction on average, not in every case.
What older studies can and cannot tell you
The meta-analysis covers experiments published by mid-2023, many with crowdworkers and some with simulated AI. It supplies a method, not a verdict on your tool.
In a 2024 trial by Ethan Goh and colleagues, physicians given GPT-4 scored a median of 76% on diagnostic reasoning vignettes, against 74% for physicians without it. In an exploratory analysis, three runs of GPT-4 answering alone scored a median of 92%. Only measuring the AI alone exposed that gap. The trial’s formal test compared GPT-4 alone with physicians working without it, and GPT-4 scored 16 points higher.
A worked example: refund decisions
The rest of this guide follows one hypothetical case. The organization, numbers and results are invented to show the method.
An online retailer handles about 10,000 refund requests a month. Agents check the order history and shipping records, then approve, deny or escalate each request to a senior agent. The team wants an AI system to draft an approve-or-deny decision for each request, and proposes that agents review every draft before the customer sees it.
The operations lead wants three answers. Would reviewed drafts beat the current process? Would they beat unreviewed drafts? And what would review cost in agent time?
Step 1: Decide what counts as a correct decision
You cannot count errors without a standard. The team forms an audit panel of two senior agents, who decide each sampled case with the full record and no time pressure, without seeing the AI draft or knowing the condition. A third senior agent settles disagreements.
Score every condition on its final decision. For an escalated case, that is the senior agent’s decision, and the senior’s time counts toward that condition’s cost. Use later facts, such as a confirmed fraud, where they exist. The panel is still fallible, so the comparison is only as good as this standard.
Define a serious mistake before you see results: here, denying a valid refund worth more than $200, or approving a refund on an order with confirmed fraud. Track these separately, because a review that reduces minor slips while missing costly errors is not doing its main job.
Step 2: Set up three conditions
The team draws 600 requests from one month and randomly assigns each to one of two arms of 300. Random assignment makes the arms comparable, though chance still plays a part, which Step 4 deals with. The randomized controlled trial entry and the research methods guide explain the design.
- Human alone. An agent handles the request with the current process. The AI still drafts a decision, but the draft is hidden and only recorded.
- Human plus AI. An agent sees the AI draft and reviews it, exactly as the proposed policy would require.
- AI alone. No third set of cases and no live AI decisions are needed. Score the drafts in the human-plus-AI arm against the panel.
The hidden drafts in the human-alone arm check the randomization: the AI should score about the same in both arms. Test the workflow you intend to run, including whether agents see the draft before forming their own view.
Step 3: Record six measures
For each condition, record:
- Quality: the share of decisions that match the audit panel.
- Serious mistakes: decisions that meet your definition.
- Caught errors and missed correct advice: in the human-plus-AI arm, wrong drafts corrected, correct drafts changed to a wrong decision, and drafts accepted unchanged.
- Time: agent minutes per case, from the case log.
- Review burden: total agent hours at your expected volume, including senior time on escalations.
- Escalation: the share of cases sent to a senior agent and the senior minutes they used.
Read measure 3 two ways. The counts give the net effect: if overrides of correct drafts outnumber corrections, review is making decisions worse. The rates show whether reviewers can tell good drafts from bad: corrections divided by wrong drafts (the catch rate), and overrides divided by correct drafts (the false-alarm rate). These are the hit and false-alarm rates of signal detection theory; if they are close, reviewers are changing drafts almost at random.
The two readings can disagree. When the AI is usually right, correct drafts vastly outnumber wrong ones, so even a low false-alarm rate can break more answers than a good catch rate fixes.
Step 4: Read the results
Here are the invented results for the retailer.
| Measure (hypothetical) | Human alone | AI alone | Human plus AI |
|---|---|---|---|
| Cases | 300 | 300 (drafts in the human-plus-AI arm) | 300 |
| Correct decisions | 267 (89.0%) | 276 (92.0%) | 288 (96.0%) |
| Serious mistakes | 9 (3.0%) | 9 (3.0%) | 3 (1.0%) |
| Wrong drafts corrected (catch rate) | n/a | n/a | 15 of 24 (62.5%) |
| Correct drafts overridden (false-alarm rate) | n/a | n/a | 3 of 276 (1.1%) |
| Final decision same as the draft | n/a | n/a | 282 (94.0%) |
| Agent minutes per case | 6.0 | 0 | 2.5 |
| Escalated to a senior agent | 30 (10.0%) | 0 (drafts only approve or deny) | 18 (6.0%) |
Read the table in three passes.
First, compare with the best single condition. The AI alone (92.0%) beat agents working alone (89.0%). Reviewed drafts (96.0%) beat both. In Vaccaro’s terms, this is synergy, not just augmentation. My test for any intervention is the same one I use for nudges: show what changes, compared with what, and at what cost.
Second, find where the gain came from. Reviewers corrected 15 of 24 wrong drafts and broke 3 of 276 correct ones: 276 − 3 + 15 = 288. A 62.5% catch rate against a 1.1% false-alarm rate means they could tell bad drafts from good. Serious mistakes fell from 9 to 3: reviewers caught 7 of the 9 serious draft errors, and one override created a new one. Most wrong drafts involved parcels with an open carrier investigation; agents check those records routinely, and the AI’s inputs did not include them.
Third, ask whether the gain could be chance. The human-plus-AI arm gives a paired comparison: the same 300 cases before and after review, where only the changed cases carry information. For overall accuracy, 15 fixed against 3 broken makes chance an unlikely explanation by an exact sign test (McNemar’s exact test). For serious mistakes, 7 fixed against 1 broken points the same way, but the same test does not rule out chance at the 5% level. Confirm it on a larger sample of high-value requests.
Step 5: Price the review
At 10,000 requests a month, with 15 senior minutes per escalated case (also invented):
- Human plus AI: 2.5 minutes × 10,000 ≈ 417 hours, plus 600 escalations × 15 minutes = 150 hours. Total: about 567 hours.
- Human alone, the current process: 6.0 minutes × 10,000 = 1,000 hours, plus 1,000 escalations × 15 minutes = 250 hours. Total: 1,250 hours.
- AI alone: no agent time, because the drafts approve or deny and nothing is escalated.
Against the current process, reviewed drafts are more accurate and take less than half the agent time, which supports adopting AI drafts once you add the tool’s running cost. The review decision rests on the comparison with AI alone. If the sample rates hold, the extra 567 hours of review buy roughly 400 more correct decisions a month, about 200 of them serious mistakes avoided. That is about 2.8 extra hours per serious mistake avoided and 1.4 per extra correct decision.
Is that gain large enough to matter? Suppose, with invented figures, that a serious mistake costs the retailer about $250 in chargeback fees and lost goodwill, and an agent hour costs $35. A serious mistake is then worth about 7 agent hours, well above the 2.8 hours it takes to prevent one. The worksheet also counts ordinary corrections at a smaller value, but serious mistakes carry most of the total. Review earns its place here, provided the larger sample confirms the serious-mistake result.
When review adds little: low-value refunds
Now change one condition. The team runs the same evaluation on a different queue: requests under $25 for items that are cheap to replace. These are again invented results.
| Measure (hypothetical) | Human alone | AI alone | Human plus AI |
|---|---|---|---|
| Cases | 300 | 300 (drafts in the human-plus-AI arm) | 300 |
| Correct decisions | 285 (95.0%) | 294 (98.0%) | 295 (98.3%) |
| Serious mistakes | 0 | 0 | 0 |
| Wrong drafts corrected (catch rate) | n/a | n/a | 2 of 6 (33%) |
| Correct drafts overridden (false-alarm rate) | n/a | n/a | 1 of 294 (0.3%) |
| Final decision same as the draft | n/a | n/a | 297 (99.0%) |
| Agent minutes per case | 3.0 | 0 | 1.5 |
| Escalated to a senior agent | 0 | 0 | 0 |
Review moved accuracy from 98.0% to 98.3%, one net decision in 300 resting on three changed cases. That is indistinguishable from no effect. Even at face value, reviewing every draft would cost 250 agent hours per 10,000 requests for about 33 extra correct decisions. That is roughly 7.5 hours per extra correct decision, on refunds worth under $25.
Here review adds little. Let the drafts stand, audit a weekly random sample against the panel standard to catch drift, and send agents only cases that meet a defined exception rule. This arrangement is often called human on the loop: people monitor the system and handle exceptions instead of approving every output.
What changed between the queues? Not the agents or the tool. In both, the AI beat agents working alone, and both are decision tasks: the two conditions under which Vaccaro’s meta-analysis found that combinations, on average, fall short of the AI. The first queue goes against those averages, and it has a specific cause. On the cases where the AI erred, agents held evidence it never saw, the carrier records, so they were more accurate than the AI on its mistakes while less accurate overall.
In the low-value queue, agents saw the same order data as the AI, errors were rare, and each was cheap. Do not assume your queue looks like the first one; the averages point the other way.
Can your reviewers detect the errors that matter?
The carrier records point to the first condition for useful review: reviewers need evidence the AI did not use. A reviewer who sees only the draft and its explanation is re-reading the AI’s own case. Asking the model whether its draft is right is not an independent check either, and the article on AI sycophancy explains why a model’s answer can bend toward what the person asking appears to want.
Telling reviewers to be careful is no substitute; the automation bias entry explains why vigilance instructions, explanation boxes and confidence displays are not dependable fixes. Fluent output adds a newer problem. In a 2026 study of Boston Consulting Group consultants, GPT-4 access lowered the share reaching the right answer on a case built so that GPT-4 would struggle. Graders also rated the AI-assisted answers as more persuasive, even the wrong ones; the AI adoption guide covers the study in detail.
Some designs reduce overreliance on wrong advice, at a cost. In a 2021 experiment, Zana Buçinca, Maja Barbara Malaya and Krzysztof Gajos gave 199 online participants a simulated AI that was right 75% of the time. The task was choosing which ingredient to replace in a meal to cut carbohydrates. Their cognitive forcing functions made people decide before seeing the AI’s suggestion, request it, or wait 30 seconds for it.
Across all questions, people picked the right ingredient about 57% of the time with these designs and 56% with simple explanation designs (with or without a confidence display), against 42% with no AI. So the forcing designs gave no overall gain over simple explanations. Their benefit came where the AI was wrong: about 27% correct, against 8% with simple explanations and 49% with no AI. Participants also rated the most effective designs least favorably, so test a decide-first step as a variant of your human-plus-AI condition rather than assuming it works.
The UK Cabinet Office’s Mitigating ‘Hidden’ AI Risks Toolkit, published in June 2025, warns that human oversight without supporting measures “won’t be sufficient.” Reviewers may lack the expertise, time or authority to challenge AI outputs. Its three conditions make a practical checklist:
- Expertise: reviewers know the subject well enough to judge whether an output is accurate.
- Time: the workflow allows a real check, not a glance.
- Authority: reviewers can challenge an output, or the use of the tool, and escalate when needed.
The toolkit is practice guidance; your comparison supplies the test. If the errors you fear most are rare, a 300-case sample may contain few of them. In a test environment, add cases with known errors of that type and count how many reviewers catch.
Accountability is a separate requirement
Sometimes a person must make or sign a decision whatever the accuracy results show. A statute, regulation, contract or internal policy may require a named official to decide a benefit claim, or a manager to approve every denial above a threshold. That requirement tells you a person must be in the loop. It does not tell you the person improves the decision.
Keep the two questions apart. Suppose a hypothetical benefits program requires a caseworker to make each eligibility decision, and the agency introduces AI drafts. If caseworkers reviewing drafts do worse than the drafts alone, the agency cannot remove them. It can change the conditions: source documents rather than only the draft’s summary, a decide-first step, or more time on complex cases. Then it runs the comparison again, without reporting the signature as evidence of accuracy.
Try it on a new case
A city housing office tests an AI tool that flags rental-assistance applications with missing documents. Caseworkers do this check by hand, using the same uploaded files the tool reads. The office runs the three-condition comparison with an audit panel and 300 applications per arm (hypothetical numbers):
- Caseworkers alone: 255 of 300 applications classified correctly (85%).
- AI alone, scored on the drafts in the review arm: 279 of 300 (93%).
- Caseworkers reviewing the AI’s flags: 270 of 300 (90%). In that arm, caseworkers corrected 6 of the 21 wrong flags and overrode 15 of the 279 correct ones.
What does this tell the office, and what should it do next?
Answer. Reviewed results beat caseworkers alone but fall below the AI alone: 279 − 15 + 6 = 270. The rates show why. Caseworkers changed 29% of wrong flags and 5.4% of correct ones, so they can partly tell good flags from bad. But correct flags outnumber wrong ones 279 to 21, so even a 5.4% false-alarm rate breaks more flags than the catch rate fixes. Unlike the refund agents, caseworkers had no evidence the tool lacked.
Next, the office should examine the 15 overridden flags, then redesign the review. It might limit review to the kinds of flags where the audit shows the tool errs most often, which raises the share of reviewed flags that are wrong. Or it might give caseworkers a source the tool cannot read, such as the agency’s record of documents already on file. If a caseworker must legally decide each application, the review stays either way, and the office runs the comparison again.
Decide where review belongs
Go back to that meeting. “Keep a human in the loop” is a hypothesis about a particular task, tool and team. In some queues, review is the most valuable part of the system. In others, it is a cost with a signature attached.
Pick one queue where the review policy is still undecided, draw a sample, and fill in the review-evaluation worksheet before the policy is written. If you have not yet settled which tasks should get AI drafts at all, start one step earlier with AI adoption at work: choose the task before the tool.
Review-evaluation worksheet: does human review of AI output earn its cost?
Use this worksheet with the guide Human-in-the-Loop AI: When Human Review Actually Helps. It follows the guide’s five steps, then adds the detection checks, the decision rules and the decision record. Fill in one worksheet for each queue or task you are deciding about.
The web version appears inside the guide. The printable version at the end has the same fields in a compact layout.
Part A. The decision
| Field | Your answer |
|---|---|
| Task or queue being evaluated | |
| Expected volume per month | |
| What the AI produces (draft decision, classification, summary, other) | |
| Proposed review policy (every case, some cases, none) | |
| Workflow to test (reviewer sees the draft first, or decides first and then sees it) | |
| Is a human decision or signature required by law, regulation, contract or policy? Name the source. |
Part B. Step 1: what counts as a correct decision
| Field | Your answer |
|---|---|
| Reference standard: who decides each sampled case | |
| What information the standard sees (full record, time allowed) | |
| Blinded to the AI draft and to the condition? (yes/no) | |
| Who settles disagreements | |
| How escalated cases are scored (normally the senior decision-maker’s final decision, with senior time counted as cost) | |
| Later facts you will use where available (for example, confirmed fraud, upheld appeals) | |
| Serious mistake definition, written before any results |
Part C. Step 2: three conditions
| Field | Your answer |
|---|---|
| Sample size and how cases were drawn | |
| How cases are randomly assigned to the two arms | |
| Human alone arm: AI drafts generated, hidden and recorded? (yes/no) | |
| Human plus AI arm: reviewers see the draft as the real workflow would show it? (yes/no) | |
| AI alone: drafts in the human-plus-AI arm scored against the standard? (yes/no) | |
| Randomization check: AI accuracy in the human-alone arm (hidden drafts) | |
| Randomization check: AI accuracy in the human-plus-AI arm |
Part D. Steps 3 and 4: the six measures
| Measure | Human alone | AI alone | Human plus AI |
|---|---|---|---|
| Cases | |||
| 1. Correct decisions (number and %) | |||
| 2. Serious mistakes (number and %) | |||
| 3a. Wrong drafts corrected (of all wrong drafts) | n/a | n/a | |
| 3b. Correct drafts overridden to wrong (of all correct drafts) | n/a | n/a | |
| 3c. Serious draft errors caught (of all serious draft errors) | n/a | n/a | |
| 3d. New serious errors created by overrides | n/a | n/a | |
| 3e. Drafts accepted unchanged (number and %) | n/a | n/a | |
| 4. Agent minutes per case | |||
| 5. Agent hours at expected volume, including senior time (from E9 to E11) | |||
| 6. Escalated cases (number and %) and senior minutes per escalated case |
Part E. Steps 4 and 5: calculations
| Calculation | Result |
|---|---|
| E1. Best single condition: the higher of human alone and AI alone (% correct) | |
| E2. Human plus AI minus the best single condition (percentage points) | |
| E3. Check: AI correct drafts − overridden + corrected = human-plus-AI correct | |
| E4. Catch rate: 3a ÷ number of wrong drafts | |
| E5. False-alarm rate: 3b ÷ number of correct drafts | |
| E6. Changed cases, all decisions: corrected vs overridden | |
| E7. Changed cases, serious mistakes: serious errors caught vs new serious errors created | |
| E8. Optional: exact sign test (McNemar’s exact test) on E6 and on E7 | |
| E9. Agent hours per month, human alone (minutes × volume ÷ 60, plus escalation hours) | |
| E10. Agent hours per month, human plus AI (same method) | |
| E11. Agent hours per month, AI alone (any escalation or exception handling the AI-alone process still needs) | |
| E12. Extra hours of review over AI alone (E10 − E11) | |
| E13. Extra correct decisions per month, human plus AI vs AI alone | |
| E14. Serious mistakes avoided per month, human plus AI vs AI alone | |
| E15. Extra hours per extra correct decision (E12 ÷ E13) | |
| E16. Extra hours per serious mistake avoided (E12 ÷ E14; not applicable if E14 is zero or negative) | |
| E17. Your estimated cost of one serious mistake, in agent-hour equivalents | |
| E18. Your estimated cost of one ordinary wrong decision, in agent-hour equivalents | |
| E19. Value of errors avoided per month: E14 × E17 + (E13 − E14) × E18 |
Part F. Can reviewers detect the errors that matter?
| Check | Yes / no, and notes |
|---|---|
| Reviewers have evidence the AI did not use (records, documents, a second source) | |
| Reviewers are not checking the draft by asking the same model whether it is right | |
| Expertise: reviewers know the subject well enough to judge accuracy | |
| Time: the workflow allows a real check, not a glance | |
| Authority: reviewers can challenge an output or the use of the tool, and escalate | |
| Rare serious errors: known-error test cases added in a test environment and scored separately? Catch rate: | |
| Optional variant tested: reviewer decides before seeing the draft |
Part G. Decision rules
Apply these in order.
- Compare with the best single condition, not only with the current process. If human plus AI neither beats the best single condition on correct decisions (E2 is zero or negative) nor reduces serious mistakes (E14 is zero or negative), review is not earning its cost on accuracy. Go to rule 6 or rule 7.
- Check the net balance. If correct drafts overridden (3b) equal or exceed wrong drafts corrected (3a), review is making decisions worse on balance. Do not adopt it as designed. Use rule 3 to find out why.
- Check whether reviewers can tell good drafts from bad. Compare the catch rate (E4) with the false-alarm rate (E5).
- If the two rates are close, reviewers are changing drafts almost at random. Give them evidence the AI did not use, and fix expertise, time and authority (Part F). Then run the comparison again.
- If the catch rate is well above the false-alarm rate but rule 2 still fails, the problem is base rates: correct drafts are so common that a low false-alarm rate still breaks more than review fixes. Route review to the kinds of cases where the AI is wrong more often, then run the comparison again.
- Check whether the gain could be chance. If the gain rests on a handful of changed cases (E6 or E7), or the optional test in E8 does not rule out chance, treat that result as unsettled. Confirm it with a larger sample, or with known-error test cases for rare serious errors, before relying on it.
- Check whether the gain is large enough to matter. Compare the value of errors avoided (E19) with the extra hours of review (E12). If E19 is not larger than E12, the gain is too small to pay for the review; go to rule 6. Use E15 and E16 to see which errors carry the value. E16 does not apply when E14 is zero.
- If review adds little, replace full review. Let the drafts stand and audit a random sample against the reference standard on a fixed schedule to detect drift. Route to a person only the cases that meet a defined exception rule (human on the loop).
- If a person is required (Part A), keep the reviewer and record why. State that the reason is accountability. If rules 1 to 5 show the review is not improving decisions enough, fix the conditions in Part F and run the comparison again. Do not report the signature as evidence of accuracy.
Part H. Decision record
| Field | Your answer |
|---|---|
| Decision for this queue | |
| Rule or rules that decided it | |
| What would change the decision | |
| Date of the next audit or re-test |
Filled example (hypothetical)
This example uses the invented refund-queue numbers from the guide. The retailer, results, costs and times are illustrations of the method, not data from a real company or study.
Part A.
- Queue: all refund requests.
- Volume: 10,000 a month.
- AI output: a draft decision (approve or deny).
- Proposed policy: an agent reviews every draft.
- Workflow tested: the agent sees the draft first, as planned for rollout.
- Required human decision: no legal or policy requirement for this queue.
Part B. Standard: an audit panel of two senior agents, with the full record and no time pressure, blinded to the draft and the arm; a third senior agent settles disagreements. Escalated cases: scored on the senior agent’s final decision; senior time counted as cost. Later facts: confirmed fraud and chargeback outcomes where available. Serious mistake: denying a valid refund worth more than $200, or approving a refund on an order with confirmed fraud.
Part C. 600 requests drawn at random from one month; each randomly assigned to one of two arms of 300. AI drafts generated in both arms, hidden in the human-alone arm. Randomization check: the AI scored about the same in both arms.
Part D.
| Measure | Human alone | AI alone | Human plus AI |
|---|---|---|---|
| Cases | 300 | 300 (drafts in the human-plus-AI arm) | 300 |
| 1. Correct decisions | 267 (89.0%) | 276 (92.0%) | 288 (96.0%) |
| 2. Serious mistakes | 9 (3.0%) | 9 (3.0%) | 3 (1.0%) |
| 3a. Wrong drafts corrected | n/a | n/a | 15 of 24 |
| 3b. Correct drafts overridden to wrong | n/a | n/a | 3 of 276 |
| 3c. Serious draft errors caught | n/a | n/a | 7 of 9 |
| 3d. New serious errors created by overrides | n/a | n/a | 1 |
| 3e. Drafts accepted unchanged | n/a | n/a | 282 (94.0%) |
| 4. Agent minutes per case | 6.0 | 0 | 2.5 |
| 5. Agent hours at 10,000 a month | 1,250 | 0 | ≈ 567 |
| 6. Escalated cases; senior minutes each | 30 (10.0%); 15 | 0 (drafts only approve or deny) | 18 (6.0%); 15 |
Part E.
| Calculation | Result |
|---|---|
| E1. Best single condition | AI alone, 92.0% |
| E2. Human plus AI minus best single condition | 96.0 − 92.0 = +4.0 points |
| E3. Check | 276 − 3 + 15 = 288 |
| E4. Catch rate | 15 ÷ 24 = 62.5% |
| E5. False-alarm rate | 3 ÷ 276 = 1.1% |
| E6. Changed cases, all decisions | 15 corrected vs 3 overridden (18 changed cases) |
| E7. Changed cases, serious mistakes | 7 caught vs 1 created (8 changed cases); serious mistakes 9 − 7 + 1 = 3 |
| E8. Exact sign test | E6: p ≈ 0.008. E7: p ≈ 0.07, so the serious-mistake gain is not settled at the 5% level |
| E9. Agent hours, human alone | 6.0 × 10,000 ÷ 60 = 1,000 hours, plus 1,000 escalations × 15 ÷ 60 = 250 hours. Total 1,250 hours |
| E10. Agent hours, human plus AI | 2.5 × 10,000 ÷ 60 ≈ 417 hours, plus 600 escalations × 15 ÷ 60 = 150 hours. Total ≈ 567 hours |
| E11. Agent hours, AI alone | 0 (drafts approve or deny; nothing is escalated) |
| E12. Extra hours of review over AI alone | 567 − 0 ≈ 567 hours |
| E13. Extra correct decisions per month vs AI alone | 4.0% of 10,000 = 400 |
| E14. Serious mistakes avoided per month vs AI alone | (3.0% − 1.0%) of 10,000 = 200 |
| E15. Extra hours per extra correct decision | 567 ÷ 400 ≈ 1.4 hours |
| E16. Extra hours per serious mistake avoided | 567 ÷ 200 ≈ 2.8 hours |
| E17. Cost of one serious mistake (invented) | $250 ÷ $35 per agent hour ≈ 7.1 agent-hour equivalents |
| E18. Cost of one ordinary wrong decision (invented) | $30 ÷ $35 ≈ 0.9 agent-hour equivalents |
| E19. Value of errors avoided per month | 200 × 7.14 + (400 − 200) × 0.86 ≈ 1,429 + 171 = 1,600 agent-hour equivalents |
Part F.
- Independent evidence: yes, agents check carrier investigation records that the AI’s inputs do not include; most wrong drafts involved these cases.
- Not asking the model to check itself: yes.
- Expertise: yes.
- Time: 2.5 minutes per case on average, enough to open the carrier record.
- Authority: yes, agents can override and escalate.
- Rare-error test cases: planned for the confirmation run.
- Decide-first variant: not tested in this run.
Part G.
- Human plus AI beats the best single condition by 4.0 points and avoids serious mistakes. Continue.
- Overrides (3) are well below corrections (15). Review improves decisions on balance. Continue.
- A catch rate of 62.5% against a false-alarm rate of 1.1%: reviewers can tell good drafts from bad. Continue.
- The overall gain rests on 18 changed cases and is unlikely to be chance. The serious-mistake gain rests on 8 changed cases and is not settled. Confirm it on a larger sample of high-value requests, with known-error test cases.
- E19 (about 1,600 agent-hour equivalents) is well above E12 (about 567 hours). The value comes mainly from serious mistakes: ordinary corrections alone are worth about 200 × 0.86 ≈ 171, far below 567. So the decision depends on rule 4’s confirmation of the serious-mistake result.
- Not applicable to this queue.
- Not applicable: no required human decision.
Part H. Decision: require agent review of every draft in this queue during a three-month confirmation period. Rules: 1 to 5. What would change it: the confirmation sample showing no reduction in serious mistakes, or the catch rate falling toward the false-alarm rate. Next re-test: after the confirmation sample is scored.
Second queue (refunds under $25, also hypothetical).
- E2 = +0.3 points.
- E4 = 2 ÷ 6 = 33%; E5 = 1 ÷ 294 = 0.3%.
- E6 = 2 corrected vs 1 overridden (p = 1.0).
- E10 = 1.5 × 10,000 ÷ 60 = 250 hours; E11 = 0; E12 = 250 hours.
- E13 ≈ 33; E14 = 0, so E16 does not apply.
- E15 ≈ 7.5 hours.
- Each wrong decision costs under $25, so E18 is at most about 0.71, and E19 is at most about 33.3 × 0.71 ≈ 24 agent-hour equivalents, against 250 extra hours.
- Rule 4 (the gain is unsettled) and rule 5 (the gain is too small to pay for the review) point to rule 6: let the drafts stand, audit a weekly random sample and route exceptions.
Printable version
Same fields as the web version. Print on two pages; write in the blanks.
A. Decision. Queue: ________ Volume per month: ________ AI output: ________ Proposed policy: ________ Workflow tested (draft first / decide first): ________ Required human decision and its source: ________
B. Standard (Step 1). Who decides: ________ Information and time: ________ Blinded to draft and arm (Y/N): ___ Disagreements settled by: ________ Escalated cases scored by: ________ Later facts used: ________ Serious mistake definition (written before results): ________
C. Conditions (Step 2). Sample and how drawn: ________ Random assignment method: ________ Hidden drafts recorded in human-alone arm (Y/N): ___ Reviewers see drafts as in the real workflow (Y/N): ___ AI-alone drafts scored (Y/N): ___ AI accuracy, human-alone arm: ___ AI accuracy, human-plus-AI arm: ___
D. Measures (Steps 3 and 4).
| Measure | Human alone | AI alone | Human plus AI |
|---|---|---|---|
| Cases | |||
| 1. Correct decisions | |||
| 2. Serious mistakes | |||
| 3a. Wrong drafts corrected | n/a | n/a | |
| 3b. Correct drafts overridden to wrong | n/a | n/a | |
| 3c. Serious draft errors caught | n/a | n/a | |
| 3d. New serious errors created | n/a | n/a | |
| 3e. Drafts accepted unchanged | n/a | n/a | |
| 4. Agent minutes per case | |||
| 5. Agent hours at expected volume | |||
| 6. Escalated cases / senior minutes each |
E. Calculations (Steps 4 and 5). E1 best single condition: ___ E2 difference: ___ E3 check: ___ E4 catch rate: ___ E5 false-alarm rate: ___ E6 changed cases, all: ___ vs ___ E7 changed cases, serious: ___ vs ___ E8 optional test: ___ E9 hours, human alone: ___ E10 hours, human plus AI: ___ E11 hours, AI alone: ___ E12 extra hours of review: ___ E13 extra correct per month: ___ E14 serious avoided per month: ___ E15 extra hours per extra correct: ___ E16 extra hours per serious avoided: ___ E17 cost of a serious mistake (agent hours): ___ E18 cost of an ordinary wrong decision (agent hours): ___ E19 value of errors avoided: ___
F. Detection. Independent evidence: ___ Not asking the model to check itself: ___ Expertise: ___ Time: ___ Authority: ___ Known-error test cases and catch rate: ___ Decide-first variant: ___
G. Rules. (1) Beat the best single condition? (2) Overrides below corrections? (3) Catch rate well above false-alarm rate? If not, fix evidence and conditions; if yes but rule 2 fails, route review to cases where the AI errs more. (4) Gain more than chance? (5) Value of errors avoided above extra review hours? (6) If review adds little: drafts stand, scheduled random audit, exception routing. (7) Required human: keep, record the reason, fix conditions, re-test.
H. Decision. Decision: ________ Rules used: ________ What would change it: ________ Next audit or re-test: ________
References
- Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188.
- Cabinet Office, Government Communications (2025). The Mitigating ‘Hidden’ AI Risks Toolkit. GOV.UK, published June 4, 2025.
- Cabrera, Á. A., Perer, A., & Hong, J. I. (2023). Improving human-AI collaboration with descriptions of AI behavior. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW1), Article 136.
- Dell’Acqua, F., McFowland, E., III, Mollick, E., Lifshitz, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2026). Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organization Science, 37(2), 403–423.
- Goh, E., Gallo, R., Hom, J., et al. (2024). Large language model influence on diagnostic reasoning: A randomized clinical trial. JAMA Network Open, 7(10), e2440969.
- Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8(12), 2293–2303.