What reliability asks
Reliability evidence examines whether scores are sufficiently consistent across relevant items, forms, raters, or occasions. The appropriate evidence depends on how the test is designed and used.
High reliability reduces random error but does not prove that the test measures the intended construct or supports a particular decision.
Sources of inconsistency
Fatigue, distraction, unclear instructions, device problems, item sampling, scoring, and normal day-to-day variation can affect results. Practice can also change retest performance.
Standardized administration and quality control reduce avoidable variation, while confidence intervals communicate remaining uncertainty.
What to look for
Responsible test documentation identifies the population, sample, reliability method, coefficients, standard errors, and conditions. A single unexplained “accuracy” percentage is not equivalent.
Evidence should match the score being interpreted. Reliability for a total score does not automatically establish equal precision for every subscore.
Common questions
Frequently asked questions
Does a reliable IQ test always give the same score?
No. Reliability allows normal variation and is summarized statistically rather than requiring identical results.
Is reliability the same as accuracy?
No. Reliability concerns consistency; validity concerns whether evidence supports the intended interpretation and use.
References