Glossary

Predictive validity

Published 2 min read

What is predictive validity?

Predictive validity concerns evidence that a measurement score supports a prediction about an outcome assessed later. The later outcome is often called the criterion. A score’s relationship with that criterion contributes to the evidence for its intended use.

“This test predicts success” leaves almost everything important unspecified. Success at what, for whom, over what period, and compared with what simpler prediction? Those questions determine whether the claim is useful.

A prediction needs an outcome

Imagine a training provider considering a reasoning assessment to help identify learners who may need extra support. In this hypothetical study, everyone completes the assessment before the course, and researchers record performance on a defined practical exercise at the end.

Higher initial scores might be associated with higher later exercise scores. That supports prediction of this criterion in this sample. It does not establish a prediction of career success, explain why particular learners struggled, or show that using the test to assign support improves learning.

The last claim requires an additional evaluation of the support decision itself. The testing standards distinguish proposed interpretations and uses rather than granting a test unrestricted validity.

Predictive versus concurrent validity

Predictive evidence uses a criterion measured later. Concurrent evidence relates scores to a criterion measured at roughly the same time. Both can inform validity, but they answer different timing questions. Showing that a measure distinguishes current experts from novices does not automatically establish how well it predicts beginners’ future progress.

What to inspect in a predictive-validity study

  • The criterion: A course grade, attendance record and independently scored work sample describe different outcomes. The criterion itself needs defensible measurement.
  • The population: Evidence among people already selected into a program may not generalize to all applicants. Selection can restrict the range of scores and alter the observed relationship.
  • The comparison: Does the measure add information beyond readily available evidence, such as prior performance? This is incremental predictive value.
  • New cases: A prediction developed and evaluated on the same data can look better than it performs with new people. Separate evaluation data help test that optimism.
  • Errors and consequences: For a decision threshold, inspect missed cases and false alarms, not just an overall association.

Prediction is not causation

A score can predict an outcome without identifying a cause that can be changed. In the training example, previous education might contribute both to the assessment score and to later performance. Raising the reported score would not necessarily improve learning.

Use predictive evidence to make a bounded forecast. Use causal inference to investigate what would change the outcome. And check construct validity before assuming the score measures the characteristic in its label.