What reliability asks

Reliability evidence examines whether scores are sufficiently consistent across relevant items, forms, raters, or occasions. The appropriate evidence depends on how the test is designed and used.

High reliability reduces random error but does not prove that the test measures the intended construct or supports a particular decision.

Sources of inconsistency

Fatigue, distraction, unclear instructions, device problems, item sampling, scoring, and normal day-to-day variation can affect results. Practice can also change retest performance.

Standardized administration and quality control reduce avoidable variation, while confidence intervals communicate remaining uncertainty.

What to look for

Responsible test documentation identifies the population, sample, reliability method, coefficients, standard errors, and conditions. A single unexplained “accuracy” percentage is not equivalent.

Evidence should match the score being interpreted. Reliability for a total score does not automatically establish equal precision for every subscore.

Common questions

Frequently asked questions

Does a reliable IQ test always give the same score?

No. Reliability allows normal variation and is summarized statistically rather than requiring identical results.

Is reliability the same as accuracy?

No. Reliability concerns consistency; validity concerns whether evidence supports the intended interpretation and use.

References

Sources and further reading