Glossary

Goodhart’s Law

Updated Published 6 min read

A hypothetical city wants safer roads. Its road crews have long reported the share of potholes repaired within 48 hours of a complaint, and the figure has been a fair guide to how well the crews work. Then the city starts paying crew bonuses on that figure. Within a year, almost every pothole is “repaired” within 48 hours. Many of the patches crumble after the next rainstorm, and drivers complain about the same holes again.

What is Goodhart’s law?

Goodhart’s law is the observation that a measure that tracks something useful under ordinary conditions can stop tracking it once people are pressured or rewarded to move the measure. Sales figures, waiting times, test scores, and benchmark results all stand in for something harder to observe. The most quoted version is a single sentence: “When a measure becomes a target, it ceases to be a good measure.”

That sentence is not Charles Goodhart’s. It comes from the anthropologist Marilyn Strathern, writing in 1997 about audits of British universities. In “‘Improving ratings’: audit in the British University system”, she used it to describe how a good degree classification (a 2.1) loses its power to distinguish students once it becomes an expectation. She noted that the educationalist Keith Hoskin called this pattern “Goodhart’s law”, after Goodhart’s observations on monetary control.

Where the idea came from

Charles Goodhart is a British economist. He first stated the idea in a 1975 paper on problems of monetary management in the United Kingdom, prepared for a Reserve Bank of Australia conference in Sydney. He later called it a humorous throwaway line. In his own later account, it reads:

Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.

In other words, a relationship that authorities rely on to steer the economy can break down once they start using it to steer.

The social scientist Donald Campbell reached a similar conclusion independently, in work on evaluating public programs. He first published it in 1975, and restated it in a 1976 paper on assessing planned social change. The statement is now called Campbell’s law:

The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.

One of Campbell’s examples was educational testing. Achievement tests can indicate general achievement under normal teaching, he argued, but once scores become the goal of teaching they lose that value and distort teaching itself.

Campbell’s version adds that the measured activity itself can be corrupted, not only the statistic. In a 2018 note in Significance, Jeffery Rodamar traced Campbell’s version to talks in 1974 and earlier writing, and argued that Campbell has priority. The two names are now often used interchangeably.

Why measures stop working

The road crews show two ways a target damages a measure.

First, people find cheaper ways to move the number than doing the underlying work. A quick patch that fails in a month counts the same as a lasting repair. Gaming can also mean choosing easy jobs or altering records.

Second, a measure is usually only correlated with the goal. Suppose the city promotes the crews with the best 48-hour figures. Part of what made those figures high was skill, but part was luck: newer roads, lighter traffic, fewer complaints. Selecting hard on the figure also selects for that luck, so the promoted crews will tend to disappoint. David Manheim and Scott Garrabrant, in a 2018 paper categorizing variants of Goodhart’s law, call this “regressional” Goodhart.

Examples of Goodhart’s law

Production and sales targets

Campbell, summarizing studies of Soviet planning, gave the example of a nail factory. When factories were judged by the total weight of their products, a nail factory could meet its target by making only the largest nails. When judged by the number of items, it could make only the smallest. Either way, he wrote, factories overproduced unneeded items and underproduced much-needed ones.

Modern sales targets show the same pattern. Wells Fargo admitted in a 2020 settlement with the US Department of Justice that onerous sales goals and management pressure had driven misconduct from 2002 to 2016. According to the department, thousands of employees provided millions of accounts or products to customers under false pretenses or without consent. Inside the bank, many of these practices were called “gaming.”

Government targets

UK ambulance services shared a target: reach 75% of calls about possibly life-threatening emergencies within eight minutes. Gwyn Bevan and Richard Hamblin compared what happened across the UK in a 2009 study. Only in England was the target part of a public star-rating system that could damage a service’s reputation, and only in England was the target met. Services in the other UK countries missed it by large margins.

England also shows the cost of that pressure. A reanalysis of all English ambulance data by the Commission for Health Improvement, which Bevan and Hamblin report, found manual “corrections” of response times in around a third of trusts. Their distributions of recorded times showed sharp discontinuities around eight minutes, implying that calls just over the target had been recorded as eight minutes or less.

The Commission estimated that the most dramatic corrections had improved reported performance by at most 6%. Bevan and Hamblin concluded that the target regime still produced substantial improvements, of up to 20% since 1999. The same pressure produced both genuine change and gaming.

AI benchmarks and reward hacking

In a 2016 post, OpenAI researchers described an agent trained on the boat-racing game CoastRunners, whose score came from hitting targets along the route. The agent found a lagoon where it could circle and hit the same targets as they reappeared. Without finishing the course, it scored on average about 20% higher than human players. Researchers call this reward hacking or specification gaming.

The same problem appears in training language models on human preferences. A reward model learns to predict which answers people prefer, and the language model is then optimized against that reward model. Leo Gao, John Schulman, and Jacob Hilton (2023) studied this in a synthetic setup in which a fixed “gold-standard” reward model stood in for human judgment. As optimization against the proxy continued, the gold-standard score first rose and then fell, a pattern they describe as Goodhart’s law. Training on human approval may also reward agreeable answers over accurate ones; the article on AI sycophancy examines that evidence.

What Goodhart’s law does not say

Goodhart’s law does not say that every target fails or that measurement is futile. A measure under pressure can lose some information without losing all of it, as the ambulance data show. In the road example, the 48-hour figure may still have risen partly because crews really did reorganize and respond faster. The task is to find out how much of the rise is real.

It also differs from neighboring ideas. Surrogation describes decision makers coming to treat a measure as if it were the goal itself, a change in how they think. In the road example, crews patching to hit the 48-hour figure is Goodhart’s law: the measure has been corrupted. Council members who then cite the figure as proof that the roads are safer would be showing surrogation. A perverse incentive is a reward that pays people to make the underlying problem worse, which is one way, but not the only way, a measure can break down.

Using measures without being ruled by them

Several practices follow from these cases, though none removes the problem:

  • Ask how the number could rise without the goal improving. Write down the cheapest routes before the target is set.
  • Check measures against independent evidence. The city could send an inspector with no stake in the bonus to check a random sample of repairs three months later.
  • Look at distributions, not only averages. A pile-up of cases just inside a threshold, as in the ambulance times, is a warning sign of manipulation.
  • Weigh the pressure to game against the pressure to improve. Keeping a measure away from pay or rankings reduces the incentive to game it. The ambulance comparison suggests it can also reduce the pressure that produces real improvement. Decide which risk is larger for the task at hand.

Back on the city’s roads, those inspections would have shown within one season that the 48-hour figure no longer meant safer roads, while it was still worth knowing how quickly crews responded. For a worked application to bonus schemes, see the guide to employee incentives.

References