Glossary
Construct validity
What is Construct Validity?
Construct validity concerns the evidence for interpreting a score as a measure of the concept it is meant to represent. That concept is the construct: reasoning ability, for example, or a tendency to plan ahead. Researchers cannot establish what a score means simply by naming it. They examine whether the measure behaves as the theory predicts and whether another explanation could account for the results. One successful prediction does not settle what it measures.
How it works
Researchers examine several kinds of evidence when assessing construct validity:
- Convergent validity: Does the measure relate to measures of similar concepts as expected? For example, do people who score higher on one planning measure also tend to score higher on another?
- Discriminant validity: Does it distinguish the intended concept from other concepts? A proposed planning score should not simply reproduce a reading-skill score if the two are meant to measure different things.
- Nomological validity: Does the score fit the wider pattern of relationships predicted by the theory? Researchers specify what else should—and should not—be associated with the concept, then examine those predictions together.
For a depression scale, relevant questions might concern its relationship with clinical assessments, everyday functioning, or change during treatment. Which relationships are expected depends on the intended interpretation and use.
When a reliable test measures the wrong thing
Suppose a test intended to measure reasoning presents every problem in a long, difficult paragraph. People receive similar scores when they retake it. That is useful reliability evidence, but the test could consistently favor strong readers, even when their reasoning is no better.
To investigate this hypothetical case, researchers could compare performance on carefully matched reasoning problems with different reading demands, examine participants’ explanations of how they solved the items, and compare scores with independent measures of reading and reasoning. The question is whether reading skill explains results that the reasoning interpretation cannot.
Reading demand is not always unwanted. If the intended construct is reasoning from complex written material, reading belongs in the task. The problem arises when the score is interpreted more broadly than the evidence allows.
Two ways a measure can miss its target
Construct underrepresentation means leaving out important parts of the concept. A planning measure that asks only about keeping lists may miss prioritizing and adapting plans. Construct-irrelevant variance means scores also reflect something outside the intended construct, such as unnecessary reading difficulty in a nonverbal-reasoning assessment. These distinctions are developed in the testing standards.
Reliability concerns consistency under specified conditions. Construct validity asks whether the interpretation of those consistent scores is justified. Predicting a later outcome supplies another kind of evidence, but a useful prediction does not by itself identify what produced it.
Why it matters
The name of a test is a claim about its meaning. Construct validation asks whether the evidence supports that claim and whether a competing interpretation fits better.
Sources and further reading
Cronbach & Meehl (1955): Construct Validity treats construct validation as evaluating competing interpretations using a body of evidence. No single correlation or successful prediction settles the question.