Article
How Long Does It Take to Form a Habit? What Lally et al. (2010) Actually Found
A Close Reading of the Most-Cited Habit Formation Study
If you've ever heard that “it takes 66 days to form a habit,” the number comes from Lally and colleagues’ study, published in the European Journal of Social Psychology in 2010. The answer is more specific than the slogan: 66 days was the median estimated time to reach 95% of a modeled automaticity plateau for 39 participants whose data met the researchers’ criteria. It was not a deadline for every person or every behavior.
The popular shorthand drops the very details that determine what the number means. Read the study carefully and the picture is far less encouraging for promises that repetition will make any demanding activity effortless.
Here's what the study found, what it didn't find, and what it means for anyone trying to change their behavior. (This article is a companion to our complete guide to behavior change, which covers the full evidence base.)
The Setup
Ninety-six university volunteers (mostly postgraduate students, mean age 27) were asked to choose a new healthy behavior, eating, drinking, or exercise, and perform it daily in the same context for 84 days. Examples: "eating a piece of fruit with lunch," "drinking a bottle of water with lunch," "running for 15 minutes before dinner."
Each day, participants logged onto a website and reported whether they'd performed the behavior, then completed a modified version of the Self-Report Habit Index (SRHI). The researchers used 7 of the SRHI's 12 items, the ones measuring automaticity specifically (e.g., "I do automatically," "I do without thinking," "I would find hard not to do"), excluding items about repetition history and identity. This gave an automaticity score ranging from 0 to 42.
The researchers then fit an asymptotic curve to each individual's automaticity scores over time, using Mitscherlich's law of diminishing returns. The idea: automaticity should increase rapidly at first, then slow down, eventually plateauing at a maximum. The curve parameters tell you the plateau height (how automatic it gets), the rate of change (how fast it gets there), and the time to reach 95% of the asymptote.
Behavior type breakdown: 27 chose eating, 31 chose drinking, 34 chose exercise, 4 chose other (e.g., meditation).
Finding #1: Fewer Than Half Met the Full Modeling Criteria
This is the finding that never makes it into the popular summaries.
Of 96 participants, 82 provided enough data for analysis. Fourteen did not. Among the 82:
- For 12 participants, the statistical software couldn't even fit the asymptotic curve after 100 iterations
- For 8 participants, the model produced a flat line, meaning the fitted model estimated no increase in automaticity
- For the remaining 62, the curve could be fit, but:
- 16 had poor fits (R² < 0.7)
- 4 had asymptote values below 21 (out of 42), which the researchers treated as too low to indicate habit; a score at or above 21 was not itself proof that a habit had formed
- 3 had unrealistically high modeled asymptotes (above 49 on a 42-point scale, the model was extrapolating beyond the data)
That left 39 participants, 48% of those with adequate data, for whom the asymptotic curve was a "good fit."
Let that sink in. The familiar timeline describes 39 people, not the entire sample. The researchers themselves used success/failure language in their discussion and connected inconsistency with difficulty forming habits. But they also suggested that some people whose curves could not be fitted, or whose estimated plateaus were very high, may simply have needed longer. The selection criteria are therefore not a validated test that divides everyone into permanent habit successes and failures. Participants were paid for completing the study, not for successfully changing behavior.
The researchers noted that poorer fits were associated with fewer repetitions. That supports taking consistency seriously. It does not tell us whether a person who failed the modeling criteria would never develop automaticity, or why they found the action difficult to repeat.
Finding #2: The "66 Days" Describes a Selected Group
The famous 66-day figure comes from the 39 participants with good curve fits. Among those 39:
| Parameter | Median | Range |
|---|---|---|
| R² (model fit) | 0.88 | 0.72 – 0.98 |
| Asymptote (max automaticity, out of 42) | 35 | 21 – 48 |
| Time to 95% of asymptote | 66 days | 18 – 254 days |
The estimates ranged from 18 to 254 days—roughly a fourteenfold difference. But these were modeled times, not all observed completion dates. The study lasted 84 days. Any estimate beyond that point projected the fitted curve beyond the period when researchers were collecting reports.
When someone tells you “it takes 66 days to form a habit,” they have turned a conditional estimate into a general promise. The median applies to the 39 selected cases, about 48% of the 82 people with sufficient data. It says nothing about a universal date when a behavior becomes effortless.
Finding #3: The Exercise Estimate Was Longer Than the Study
Among the 39 participants with good fits:
| Behavior Type | N | Median Time to 95% Asymptote | Quartiles (Q1:Q3) |
|---|---|---|---|
| Eating | 10 | 65 days | 35–106 |
| Drinking | 15 | 59 days | 39–75 |
| Exercise | 13 | 91 days | 44–118 |
The exercise median was 91 days, compared with 65 for eating and 59 for drinking. The authors described this as about one and a half times as long. The overall subgroup comparison was not statistically significant (p = 0.328), and they explicitly noted limited power for subgroup analysis. That makes the difference an uncertain descriptive finding, not an established rule that exercise always takes longer. Low power is a reason to avoid a firm conclusion in either direction.
Here is the critical distinction: the observation period was 84 days, while the exercise median estimate was 91 days. More than half of those 13 selected exercise participants therefore had model-estimated times beyond the observation window. That does not tell us exactly when they actually plateaued. The upper quartile of 118 days likewise describes the distribution of estimates, not a directly observed four-month follow-up. One other-behavior case brings the subgroup counts in the table to the full selected sample of 39.
Among the selected cases, median reported compliance was 86% for exercise and 93% for drinking. Compliance was the percentage of reported days on which the action was reported performed; missing reports leave uncertainty about total performance. Across the 39 cases, compliance correlated with model fit (r = 0.34, p = 0.035). These observations support the practical concern that consistency and modeling success are related. They do not isolate task complexity as the cause.
Finding #4: They Measured "Automaticity," Not "Habit"
This is the most important and most overlooked finding in the paper.
The researchers used the Self-Report Habit Index to measure automaticity, a participant's subjective sense that a behavior feels automatic, effortless, and unintentional. But automaticity and habit are not the same thing.
In the action–habit research tradition associated with Dickinson, a key test asks whether responding persists after its outcome is devalued. Other habit research examines learned cue-response associations and automaticity, including self-report measures. These approaches do not establish exactly the same thing. A questionnaire score alone cannot show how someone would respond in an outcome-devaluation experiment.
The SRHI cannot distinguish these. When a participant reports that doing 50 sit-ups "feels automatic," they could mean:
- The decision to start feels automatic: "When I finish my coffee, I automatically think 'time for sit-ups' without deliberation." This is initiatory automaticity: the cue-response link for getting started has been learned.
- The performance itself feels automatic: "I do 50 sit-ups without conscious effort or attention." This would be true motor automaticity.
- The routine feels established: "This is just what I do now. It's part of my day." This is goal-directed automaticity: a subjective sense of routine that remains tied to the person's goals and would stop if the goals changed.
The authors were fully aware of this ambiguity. In their discussion, they cited Wood and Neal (2007), who proposed that repeating complex behaviors in a consistent setting could develop goal-directed automaticity rather than habit, meaning the behavior becomes routinized and feels effortless, but remains flexible and tied to conscious goals.
The authors wrote explicitly:
"Using the SRHI in this study means that we have assessed the development of automaticity for performing these behaviours rather than specifically habit. New measures will be needed to disentangle these two forms of automaticity."
That measurement qualification matters. The authors were studying how automatic these behaviors felt; their measure could not establish whether a high score reflected habit or a flexible, goal-directed routine. The important practical question is what became easier, not which reassuring label we attach to it.
For a simple action, the cue and response may be fairly easy to identify. Exercise is more complicated: beginning a session, selecting movements, sustaining effort and deciding when to stop are different parts of an activity. Later research by Phillips and Gardner explicitly distinguishes exercise instigation from execution. Starting can become cue-driven without making the whole session effortless. Lally's study did not separately measure those parts.
Finding #5: The "50 Sit-Ups" Example Is More Ambiguous Than It Appears
The study's Figure 3 shows three example curves from participants with good model fits. One is "Doing 50 sit-ups after my morning coffee," which shows a clean asymptotic curve reaching about 35 on the 42-point automaticity scale.
This is one of the most reproduced figures in the habit formation literature, and it creates the impression that even exercise behaviors reliably become automatic. But several things are worth noting:
It is an illustration, not a success rate. Figure 3 presents three examples of good model fits. The sit-up curve demonstrates what one such trajectory looked like; it does not show that all exercise participants followed it. The full sample and selection results are needed to judge how representative it is.
The score does not identify what became automatic. A modeled score of about 35 on the 42-point scale indicates strong reported automaticity. It does not reveal whether the participant meant getting started, moving through a familiar sequence, or no longer reconsidering the plan each morning. Assigning one of those explanations to this participant would go beyond the data.
A different measure could answer a different question. A controlled test of response after outcome devaluation would investigate a property the questionnaire did not measure. Simply asking whether someone would stop after advice from a doctor would not reproduce that experiment: the advice could also change beliefs about safety, the task and its consequences.
50 sit-ups is more complex than "drinking water" but less complex than "exercising regularly." It's a bounded, specific, single-movement behavior performed in one context. It's not the same as "going to the gym" or "maintaining a varied exercise routine", the kinds of exercise goals that people actually care about. The gap between "50 sit-ups after coffee" and "a sustainable fitness practice" is enormous.
What the Study Actually Supports
To be fair to Lally et al., the study made genuine contributions:
- It killed the 21-day myth. The selected group had a median estimated time of 66 days, with substantial variation.
- A single missed opportunity did not materially disrupt reported automaticity in this study. The measured change was less than half a point, and scores recovered quickly. This is reassuring evidence about a single omission, not a test of repeated missed days or a long break.
- It showed that an asymptotic curve can fit individual-level data. Prior work (Hull, 1943, 1951) had only shown this for group averages. Lally et al. demonstrated it for individuals.
- It showed that simple behaviors in stable contexts can start to feel automatic. Drinking water after breakfast, eating fruit with lunch, these are plausible habit candidates, and the data supports that.
What the Study Does NOT Support
Here's what the popular interpretation gets wrong:
"It takes 66 days to form a habit." The median estimate was 66 days in the 39 selected cases, about 48% of those with adequate data, with estimates from 18 to 254 days. For exercise, the median was 91 days. For many participants, the "asymptote" was a model projection beyond the observation window.
"Any behavior can become a habit with enough repetition." More than half of the participants with adequate data did not meet the full modeling criteria. The study's own authors acknowledged that complex behaviors may develop goal-directed automaticity rather than true habit. Verplanken (2006), cited in the paper, had already shown that "for the same number of repetitions a simple behaviour has a higher habit score than a complex behaviour."
“A whole exercise routine will become effortless.” This study does not establish that claim. It examined self-reported automaticity for chosen actions; the exercise subgroup was small and its median modeled time exceeded the observation period. Automatic initiation remains a different question from effortless execution.
“If you just stick with it long enough, it becomes effortless.” The authors concluded that developing automaticity may require sustained self-control during the learning period. They did not establish that every complex activity eventually stops requiring attention or effort. The 84-day study cannot settle what happens to every activity over an indefinite period.
What later research adds
A 2024 systematic review of health habit formation also found variation in the time taken and limitations in the available studies. A later review can broaden the evidence; it cannot make the original 66-day estimate apply to every person or activity. For the distinction between habits, routines and effortful practices, see the habits reference.
The Implication Nobody Wants to Hear
If the most-cited habit formation study shows that:
- More than half of participants with adequate data did not meet the full modeling criteria in 84 days
- The exercise subgroup had a longer, uncertain median estimate extending beyond the study
- The measurement tool can't distinguish true habits from established routines
- And the authors themselves say they measured automaticity, not habit
...then the popular advice to "just build the habit" is far less grounded than it appears.
For simple, bounded behaviors in stable contexts, the Lally data show that actions such as drinking water and eating fruit can start to feel automatic. That is a useful finding, even though it does not establish when a fully formed habit has arrived.
But the activities many people want to sustain—exercising, writing, managing their finances—contain decisions and effortful work. They may become more familiar, and starting may require less deliberation. That is useful. It is different from making the whole activity automatic. Lally's study did not directly test writing, financial management or whether any of these activities could become effortless.
The practical implication is significant. If your target behavior is genuinely goal-directed, then the most important question isn't "how do I make this automatic?" It's "have I chosen a version of this behavior that I can sustain with ongoing conscious effort?" That shifts the focus from habit engineering to what behavioral scientists call person-behavior fit, finding the specific implementation of a behavior that works with your personality, preferences, constraints, and life circumstances.
The habit formation industry has built an empire on the assumption that any behavior can become effortless with enough repetition. The Lally study, the industry's own flagship citation, suggests otherwise.
For an exercise in choosing alternatives, use the behavior menu. For my longer argument about person-behavior fit, see the behavior-change guide. I also develop that argument in my book, Real Change: Achieve Lasting Transformation. The practical question remains the same: which version of the activity can you sustain with the effort it actually requires?
References
Daw, N. D., & O'Doherty, J. P. (2014). Multiple systems for value learning. In P. W. Glimcher & E. Fehr (Eds.), Neuroeconomics: Decision making and the brain (2nd ed., pp. 393–410). Academic Press.
Dickinson, A. (1985). Actions and habits: The development of behavioural autonomy. Philosophical Transactions of the Royal Society of London B, 308(1135), 67–78.
Hull, C. L. (1943). Principles of behavior: An introduction to behavior theory. Appleton-Century-Crofts.
Lally, P., van Jaarsveld, C. H. M., Potts, H. W. W., & Wardle, J. (2010). How are habits formed: Modelling habit formation in the real world. European Journal of Social Psychology, 40(6), 998–1009.
Verplanken, B. (2006). Beyond frequency: Habit as a mental construct. British Journal of Social Psychology, 45(3), 639–656.
Wood, W., & Neal, D. T. (2007). A new look at habits and the habit-goal interface. Psychological Review, 114(4), 843–863.