Guide
How This Site Evaluates Tests
Suppose a personality test gives similar scores when people take it twice. That is useful evidence about consistency. It does not yet tell you whether the test can help someone choose a career: that is a different claim, with a different outcome to examine. This site’s test explanations keep those questions separate, so you can see what the research supports and what remains uncertain. A universal accuracy percentage or single “best test” ranking would hide those differences.
This page describes the evidence standard used in this site’s explanations of personality tests and measurement. For a practical sequence you can apply to a test you are considering, use How to Evaluate Personality Tests.
Identify the instrument, not just the label
The Big Five is a way of organizing personality traits, not a single test. IPIP is a collection of questionnaire statements, called items, and sets of items scored together, called scales. IPIP-NEO-120 is one particular questionnaire built from that collection. A company can administer it online, calculate scores, compare them with a reference group, and write a report. Research on the questionnaire does not automatically check every one of those later steps.
That is why a profile names the exact instrument—the test being discussed—and its version. If questions are translated or changed, people may understand them differently. If scoring rules change, the same answers may produce different results. Evidence for the original version therefore needs examination before it can support the changed one.
Match the source to the claim
Original studies support descriptions of what researchers tested and found. Official instrument documentation establishes such details as version names and scoring keys: the instructions for turning answers into scores. Provider pages establish what a service says it offers. A provider’s description is not treated as an independent test of its effectiveness.
A test publisher may report its own research statistics. Those findings can be informative, but they are identified as publisher evidence. Independent replication means researchers independent of the original study investigate whether a finding holds again; repeating the publisher’s number on another website is not replication. Who collected and analyzed the data matters, and independent studies still need to be judged by what they actually did.
In the IPIP-NEO-120 profile, the original development paper supports the account of the studied instrument; the official IPIP pages support its scoring and interpretation guidance. The worked example in Test Norms and Percentiles uses invented numbers, clearly labeled as such. It is not presented as a study result.
Sources are linked near the claims they support. A reference is useful because it supports the explanation, not because a long reference list makes a page look scientific. An explanation also needs to distinguish the finding from the editorial judgment about what a reader might do next.
Ask what was actually tested
Reliability concerns consistency: for example, whether related items give a coherent score or whether scores are stable when measured again under similar conditions. Validity concerns the evidence for what a score means. A questionnaire could consistently measure something other than its intended trait, so consistency alone cannot settle the meaning.
The next question is what the proposed use requires. If two questionnaires tend to give high scores to the same people, that supports a relationship between their measurements. It does not show that taking either test improves relationships or work. That claim would need evidence about what happens when people use the results. The people studied matter too: a very large group of volunteers can still differ from the population a service wants to describe.
This approach follows the distinction between supported score interpretations and intended uses in the Standards for Educational and Psychological Testing. The standard is a basis for asking better questions, not a certificate that every test discussed here meets it.
Make the score understandable
A number is useful only if its meaning is clear. A raw score combines answers according to scoring rules. A percentile places that score relative to other people. A category, such as “high,” groups scores under a label. These are different steps, and a reader should be able to follow how one became the next. If a small score difference could reflect measurement error, the explanation should say so; a numerical margin of error requires evidence, not guesswork.
The aim is useful understanding: what a result adds, what remains unknown, and which next question would clarify it. A report can be interesting without being a prescription for a career, relationship, or routine.
Disclose commercial interests
Jason Hreha owns Twofold, a personality-assessment product. That interest is relevant when this site discusses assessments. Research on a personality model or a published questionnaire is not automatically a validation of Twofold’s implementation, reference groups, reports, or practical benefits. The same distinction applies to other providers.
Reading these explanations does not require taking a test or buying a report. Commercial interests do not supply an exception to the evidence standard.
State the review scope honestly
These explanations discuss selected studies and official materials. They are not systematic reviews—a search and assessment of the wider research using an explicit method—or checks of every service using an instrument’s name. They also do not claim review by a psychometric specialist. Stating that scope lets you distinguish an explanation of selected evidence from a more comprehensive investigation.
New evidence should change the relevant explanation when it changes the answer. A different result deserves examination of its version, sample, method, and proposed use; it should not be dismissed because an older page reached a convenient conclusion.