Glossary

Automation Bias

Updated Published 4 min read

An automated invoice checker marks a payment as ready. You approve it without comparing the bank details with the supplier’s record. The checker missed a changed account number, and so did you. The important failure was not simply that software made a mistake: its reassurance displaced a check you could have made.

What is automation bias?

Automation bias is inappropriate reliance on an automated aid in making a judgment or decision. Researchers commonly distinguish two kinds of error:

  • Commission: acting on incorrect advice. Imagine a scheduling system tells you to cancel an appointment that is still required, and you follow the recommendation despite contradictory information.
  • Omission: failing to act because the system did not alert you. In the invoice example, the absence of a warning is treated as evidence that the account details need no attention.

These are hypothetical illustrations of the distinction described in Goddard, Roudsari, and Wyatt’s systematic review. An omission is not automatically automation bias: someone could miss the same problem without any automated aid. The question is how the aid changed information seeking and the decision.

A useful system can introduce new mistakes

Suppose, in a hypothetical test, a person gets 80 of 100 decisions right without assistance. A tool helps correct 15 mistakes but also persuades them to change five correct answers to wrong ones. Their final score is 90. The tool improved overall performance and introduced five errors.

Both statements matter. Looking only for harmful switches would miss the benefit; looking only at the higher final score would hide a preventable failure. The numbers here are illustrative, not a published effect size.

The systematic review found evidence of this broader tradeoff across decision-support settings: improved overall performance can coexist with errors caused by incorrect advice. The studies used different tasks and systems, so the review does not supply a universal error rate for today’s AI products.

Why “stay alert” is incomplete advice

Checking consumes attention. If a workflow makes independent evidence hard to find, or leaves no time to inspect it, a final human approval can become little more than another click. In their review of complacency and automation bias, Parasuraman and Manzey connect reliance and monitoring problems with how attention is allocated. Both experienced and inexperienced users can be affected; simple instructions to remain vigilant are not a dependable guarantee.

Automation bias can overlap with confirmation bias when a recommendation supports what someone already believes. But accepting a recommendation that contradicts one’s initial correct answer is also possible. Agreement with the user’s prior belief is not required.

Automation bias in generative AI

The studies in Goddard and colleagues’ and Parasuraman and Manzey’s reviews mostly examined alerts, recommendations and classifications: a system offered an answer, and a person accepted or rejected it. Generative AI tools draft whole analyses, summaries and recommendations in fluent prose. The same failure applies. A person can adopt a wrong output in place of a check, and a well-written draft offers few surface cues that something is wrong.

A preregistered study of 758 Boston Consulting Group consultants, published by Fabrizio Dell’Acqua and colleagues in Organization Science, shows the tradeoff described above with a generative tool. GPT-4 access raised productivity and rated quality on tasks chosen to fall within the model’s capabilities. On a business case built so that GPT-4 would struggle, consultants without AI were right about 84.5% of the time, against about 60% and 70.6% in the two groups with access. Graders also rated the AI-assisted recommendations as more coherent and persuasive, including the incorrect ones.

The study compared access to AI with no access, and its reported outcomes are correctness, time spent and rated quality. The authors attribute the drop to workers becoming overly reliant on AI, but the comparison does not isolate that mechanism, so it is not a direct measurement of automation bias. It does show why independent checks matter for generative tools: an output that needs correcting can read as convincingly as one that does not. Asking the same model to confirm its own answer does not supply that independence; the article on AI sycophancy describes how a model’s responses can shift toward what the user appears to believe.

Make the check concrete

For the invoice workflow, “review carefully” is vague. “Compare the destination account with the independently maintained supplier record before approving a changed account” identifies an action, an evidence source, and a trigger. Re-reading the checker’s own explanation would not provide the same independence.

When evaluating an aid, examine cases where it is wrong, cases it fails to flag, and cases where it improves the decision. Record what the user decided before seeing its advice where that comparison is feasible. The review discusses design and implementation approaches that may improve reliance, but an explanation box or confidence display should not be assumed to solve the problem merely because it is present. Whether a human review step is worth requiring at all is a separate question. Answering it means comparing people alone, the system alone and people reviewing the system’s output; the guide to human-in-the-loop AI works through that comparison with an example.

The practical test is whether people can use the system’s strengths while still detecting consequential failures. Choose the checks around the particular decision, available evidence, and cost of an error.

References