Glossary

Test Norms and Percentiles

Updated Published 4 min read

Test norms describe how a group of people scored on a test. That reference group supplies a comparison for interpreting someone else’s score: is the score common in this group, or unusually high or low? A percentile rank expresses relative standing: where a score falls among the comparison scores. It is not the percentage of questions answered correctly or the percentage of a personality trait that someone possesses.

This entry explains psychological test scores. For expectations about behavior in a group, see social norms.

If a personality report gives you “70” for Conscientiousness, a trait concerning organization and follow-through, the number alone tells you little. It could be a total from your answers, a percentile comparing you with other people, or a score converted to a 0–100 display. To see the difference, start with the answers and follow the calculation.

Start with the raw score

Imagine a personality scale: a set of statements whose answers are combined to measure one tendency. Each statement is called an item. This invented four-item example concerns organization, with responses running from 1 for a very inaccurate description of you to 5 for a very accurate one. It is not a real questionnaire or published set of norms.

Suppose your four responses are 4, 2, 3, and 2. The scoring instructions say that the second and fourth items describe less of the tendency being measured. Those two responses must be reverse keyed before they are added to the others. Otherwise, agreeing that you are organized and agreeing that you are disorganized could both raise the same score.

Reverse keying makes responses point in the same direction before they are combined. On this 1–5 scale, subtract a reverse-keyed response from 6: 1 becomes 5, 2 becomes 4, and 3 stays 3. This is the convention described in the IPIP scoring instructions; another instrument’s instructions must be checked separately.

  • First item: 4 stays 4.
  • Second item: 6 − 2 = 4.
  • Third item: 3 stays 3.
  • Fourth item: 6 − 2 = 4.

The raw total is 4 + 4 + 3 + 4 = 15, on a possible range of 4–20. That tells you how the keyed answers add up. It does not yet tell you how common the score is. Dividing 15 by 20 produces 75%, but that arithmetic does not turn the score into a percentile or “75% conscientious.”

The same score, two comparison groups

Now imagine that two different groups took the same four-item scale and that their answers were scored by the same rules. Each group contains 20 other people, each contributing one total. They are invented reference groups, chosen to make the counting easy; 20 is not a recommended sample size for developing real norms.

For each group, count how many totals fall below your 15, divide by the 20 people in that group, and multiply by 100. This expresses the count as a percentage. In this example, percentile rank means the percentage of reference scores strictly below your score. Neither group contains a total of 15, so there are no ties to handle.

  • Group A: 14 people scored lower and 6 scored higher. The percentile rank is 14 ÷ 20 × 100 = 70.
  • Group B: 5 people scored lower and 15 scored higher. The percentile rank is 5 ÷ 20 × 100 = 25.

The 70th percentile in Group A means your total exceeds 14 of its 20 scores. The 25th percentile in Group B means it exceeds only 5 of that group’s 20 scores. Your answers and total stayed the same; Group B simply contains more scores above yours. A percentile describes your position within a specified comparison, so these two results do not contradict each other.

Real tests may calculate ranks differently. Some count everyone at or below the score; others give partial weight to people with exactly the same score. A test may also estimate ranks between observed scores, called interpolation, or smooth irregularities in the observed distribution. These choices can change the reported number. Read the scoring documentation before recreating a report’s exact number. The 2014 testing Standards glossary defines percentile rank in terms of scores below the score being ranked; the example above states that convention explicitly.

Which comparison answers your question?

A test report needs to identify its reference group. Students in one class, volunteers on a website, and a sample designed to represent a country answer different comparison questions. A percentile among website volunteers should not silently become a claim about everyone your age.

The testing Standards, standards 5.8–5.11, call for clear population and sampling information, testing dates, and evidence that the norms remain appropriate. Age-specific comparisons may be useful for some questions, but the subgroup must still be suitable for the intended interpretation.

The IPIP norms guidance illustrates a local comparison: students can locate their scores within their own class’s distribution. That can answer a classroom question without claiming to describe a national population. The useful comparison depends on the question, not simply on choosing the largest available dataset.

What a percentile leaves out

A percentile tells you relative position, not the distance between scores. Imagine many people have totals close to yours. Increasing your raw score by one point could move you past several people. In a sparse part of the distribution, a larger increase might move you past very few. That is why moving from the 40th to the 50th percentile need not represent the same raw-score difference as moving from the 80th to the 90th: ranks count people, while raw-score differences compare totals.

It also does not establish a threshold for success. The median is the middle of the comparison scores. Being above it does not, by itself, show that someone is ready for a job, has mastered a skill, or needs treatment. Those conclusions require evidence about the particular decision. A rank within a group is not that evidence.

Finally, an exact-looking rank does not remove measurement uncertainty. A person might answer some items differently on another occasion, and a different sample of people could produce different norms. These are two separate sources of uncertainty: one concerns the person’s score, the other the comparison used to interpret it. To judge a small percentile difference, you need evidence about both. A universal “plus or minus five percentiles” rule would invent precision rather than explain it.

Read the number with its description

Before drawing a conclusion, find the score’s name, the exact test version, the reference group, and the interpretation the report actually supports. Keep these together. “70th percentile among this test’s reference group” says more than “a personality score of 70,” even though both contain the same number.

For the wider check, use How to Evaluate Personality Tests. Psychometrics explains how measurement, reliability, and validity fit together.