Article

Do Nudges Work? What the Evidence Can Actually Tell Us

Published 5 min read

Some nudges change measured behavior. That is a much narrower claim than “nudges work” as a general approach to solving behavioral problems. A reminder before an appointment, a default contribution rate and a message about what other people do are different interventions. Their effects cannot be inferred from sharing a label.

My objection is to the jump from a striking result to a dependable design recipe. A useful answer needs to identify the intervention, the comparison, the people affected and the outcome that changed. It also needs to survive scrutiny of the studies that did not produce a striking result.

What this assessment covers

This is a selected evidence assessment, reviewed on September 9, 2026. It examines a large vaccination experiment, a comparison of government and academic nudge trials, the disputed broad choice-architecture meta-analysis, and a later review of health-related social norms messages. It is not an exhaustive systematic review of every intervention called a nudge.

The sources were selected to answer different questions: whether a specific intervention can work, what happens across routinely conducted trials, and how publication bias changes an average assembled from published research. Those questions are related, but their results are not interchangeable. The earlier essays The Death of Behavioral Economics and Bad News for Nudges present the broader criticism. This page keeps the evidence and the practical decision in view.

First define the result

A nudge changes the arrangement or presentation of choices while leaving options available and without substantially changing economic incentives. In practice, the boundary is disputed: a simpler form, a reminder and an automatic enrollment policy can all appear in nudge discussions despite changing quite different parts of a service.

“Worked” might mean that someone clicked a link, registered, completed a task, continued a behavior or experienced a benefit. A study of registration does not establish an effect on the final service delivered. A study of next week's action does not establish a new habit. Choose the outcome before interpreting the result.

A positive example: vaccination texts before an appointment

Milkman and colleagues randomly assigned 47,306 patients with upcoming primary-care appointments to usual care or one of 19 text-message interventions. Across the messages, influenza vaccination increased by an average of 2.1 percentage points over a control rate of roughly 42%. That is about a 5% relative increase, not a 5-percentage-point increase. The original report describes the interventions and analyses.

The outcome was vaccination on the appointment day or during the preceding three days, at two US health systems in 2020. Usual care included ordinary appointment reminders, so the contrast concerns additional vaccination messaging. It does not show lasting behavior change or a reduction in illness. The 2025 correction repairs image links in the supporting information; it does not revise these results.

This is evidence for an intervention delivered when an opportunity to act already existed. It gives a team a reason to investigate similar messaging in a comparable service. It does not give every message, every vaccine campaign or every nudge the same expected effect.

Routine government trials tell a less dramatic story

DellaVigna and Linos compared a comprehensive set of 126 randomized trials from two US nudge units with a separate sample of academic trials. Average effects were 1.4 percentage points in the nudge-unit sample and 8.7 points in the academic sample. The reported standard errors were 0.3 and 2.5 points, respectively.

These are different collections of interventions and populations, not the same interventions tested small and then scaled up. Defaults were excluded. Publication selection helps explain the difference, alongside differences in the trials and institutional constraints. The authors regard the government effects as meaningful given low marginal costs. My conclusion is that the larger published average is a poor default planning assumption.

Why the broad meta-analysis became a dispute

Mertens and colleagues pooled choice-architecture interventions across domains. The meta-analysis reported a positive average and variation across interventions. A formal correction removed observations from a retracted paper and repaired errors. The meta-analysis itself was corrected, not retracted.

Maier and colleagues' reanalysis questioned whether the pooled positive result survived publication-bias adjustment. Their analyses did not provide convincing support for a positive overall effect. They also reported heterogeneity: this does not imply that every individual intervention has zero effect. Their available methods handled dependent estimates imperfectly, including an analysis that selected the most precise effect from each study.

In their reply, Mertens and colleagues acknowledged publication bias and emphasized variation between interventions and settings. Their position was to identify when effects occur through better research, not to deny the bias problem.

The disagreement is a reason to distrust a single confident number for “the effect of nudging.” Bias corrections model a selection process that is not fully observed. An unadjusted average is not automatically trustworthy; an adjusted estimate is not a direct measurement of every missing study. The shared data also mean that the meta-analysis, reanalysis and reply are one evidence dispute, not three independent replications.

Social norms messages need their own assessment

A 2025 review of social norms messaging examined randomized studies of health-related behavior in people aged 16 and older in developed countries. The authors reported a small positive pooled effect before publication-bias adjustment and no clear positive average after it. Effects varied across studies.

The scope matters. The review excluded, among other topics, alcohol and cannabis interventions in school and college students, and populations in developing countries. It considered several outcomes and forms of normative messaging; it is not a test of every influence one person can have on another. Its findings challenge the routine recommendation to add “social proof” to a health message. They do not establish that social relationships or group expectations never matter.

What the selected evidence can support

EvidenceComparison and outcomeWhat to carry into a decision
Vaccination text megastudyAdditional SMS messages versus usual care; recorded vaccination around an existing appointmentA specific positive result with an immediate opportunity to act; no tested long-term benefit
Two government nudge unitsEach trial's intervention versus its control; separate comparison with an academic sampleRoutine trial effects were smaller than the academic average; costs and trial differences matter
Choice-architecture meta-analysis and reanalysisHeterogeneous studies pooled under different bias assumptionsThe broad average is disputed and cannot be treated as a dependable forecast
Health social norms reviewNormative messages versus eligible comparison conditions; varied health-related endpointsPublication-bias adjustment removes clear evidence of a positive average within this review's scope

How I would decide whether to test a nudge

Start with the behavior and the obstacle. Suppose a service needs people to attend a booked appointment. A reminder is a plausible response to forgetting. It is a poor response to appointments that conflict with work, inaccessible transport or a service people no longer want. Calling all of these “noncompliance” hides the information needed to choose a response.

Compare feasible alternatives before refining the message. Could the appointment move? Could the person complete the task remotely? Could the service remove a needless visit? This is the behavior-selection question in Behavior Market Fit. A small communication change may still be worth testing, but it should earn its place among those alternatives.

Specify a comparison and a useful threshold. If the outcome is attendance, count attendance among eligible people assigned to each version. Also examine cancellations, rebooking and staff work: an apparent improvement that creates unusable bookings may not help. Estimate costs using the actual delivery process. “Cheap” should include setup and staff time, not just the price of an extra text.

Distinguish exploration from confirmation. If a team tests many messages and selects the largest apparent winner, its first estimate may be optimistic. Confirm the promising version in another suitable comparison before treating that estimate as a forecast. Decide in advance which outcome and follow-up period will justify wider use.

The research-methods guide explains study selection. The practical-significance reference explains why a detectable difference can still be too small to matter—and why a small difference can sometimes be worth the cost.

What would change this assessment?

A new positive study would strengthen the case for its intervention to the extent that its design and setting support the conclusion. Stronger evidence for a family of nudges would require more: accessible results from credible comparisons, outcomes that matter, adequate follow-up, and successful tests beyond a single setting. Null findings and failed replications belong in that assessment too.

The practical standard is demanding but straightforward. Show what changes, compared with what, and at what cost. A clever explanation for why a message should work is not the same thing as evidence that it did.