How Psychometricians Measure the Quality of Standardized Tests
Before standardized tests are deployed for high-stakes decisions, psychometricians must prove they meet strict scientific standards. Reliability, validity, and fairness serve as the three non-negotiable properties that determine whether an assessment accurately measures student ability without systemic bias.
By Nabil Faris
- Psychometricians
- Focus on statistical rigor, minimizing measurement error, and ensuring score comparability across different years and test forms.
- Equity Advocates
- Focus on eliminating cultural bias, ensuring universal design, and preventing high-stakes decisions from disproportionately harming marginalized subgroups.
- State Regulators
- Focus on federal compliance, practical deployment at scale, and defending the legal validity of graduation and licensure requirements.
Perspectives this story doesn't cover
- Students taking the assessments
- Classroom teachers administering the tests
The short answer
- Reliability measures the consistency of test scores across different administrations and forms.
- Validity ensures that an assessment actually measures the specific construct it claims to measure.
- Fairness is now considered a foundational prerequisite of validity, not just a separate statistical check.
- High-stakes decisions should never rely on a single standardized test score.
- Psychometricians use Differential Item Functioning to identify and remove biased questions before tests are finalized.
State education boards, university admissions committees, and professional licensing bodies dictate which standardized tests determine graduation, college entry, and career certification. Before these organizations deploy a new assessment to millions of students, psychometricians must prove the instrument actually measures what it claims to measure without systemic bias. They do this by evaluating three non-negotiable properties: reliability, validity, and fairness. These metrics form the scientific foundation of educational measurement, ensuring that a single exam does not unjustly derail a student's academic trajectory.[1][7]
The stakes attached to these assessments are immense, influencing school funding, teacher evaluations, and individual student futures. The American Educational Research Association explicitly warns against over-reliance on isolated metrics. In its official position statement, the organization declares that "decisions that affect individual students' life chances or educational opportunities should not be made on the basis of test scores alone." This principle requires test developers to build instruments that can withstand intense statistical and legal scrutiny before they ever reach a classroom.[3]
To standardize this scrutiny, the American Psychological Association, the American Educational Research Association, and the National Council on Measurement in Education jointly publish the Standards for Educational and Psychological Testing. Last comprehensively revised in 2014, this document serves as the definitive rulebook for test developers worldwide. It outlines the specific empirical evidence required to defend a test's use in high-stakes environments, establishing a baseline of scientific rigor that commercial publishers must meet.[1]
The first of these foundational properties is reliability, which measures the consistency and stability of test scores. If a student takes the same reading comprehension test on a Tuesday and then an equivalent version on a Thursday, assuming no new learning occurs in the interim, their scores should be nearly identical. A reliable test minimizes random measurement error, ensuring that a score reflects actual ability rather than a lucky guess, a poorly worded question, or environmental distractions.[6]
Psychometricians quantify this consistency using reliability coefficients, which are expressed on a statistical scale from 0 to 1.0. For low-stakes classroom quizzes, a coefficient of 0.70 might be acceptable to a teacher. However, for high-stakes decisions like high school graduation or medical board certification, state agencies and licensing boards typically demand a reliability coefficient of 0.85 to 0.90 or higher, leaving very little room for statistical noise.[6]
Even with perfect reliability, a test is useless if it lacks validity. Validity is universally considered the most critical property in psychometrics. It asks a fundamental question: does the test actually measure the specific construct it claims to measure? For example, a mathematics assessment laden with highly complex, convoluted vocabulary might inadvertently measure reading comprehension rather than mathematical reasoning, rendering the math scores invalid for students who are still learning the language.[1]
Even with perfect reliability, a test is useless if it lacks validity.
Modern psychometric theory emphasizes that validity is not an inherent property of the test itself, but rather of the interpretation of the scores for a specific purpose. An exam might be highly valid for placing college freshmen into introductory calculus, but entirely invalid for determining whether a high school senior should receive a diploma. Test developers must provide distinct empirical evidence for every proposed use of the scores, proving that the inferences drawn from them are scientifically sound.[1]
Historically, fairness was treated as a secondary consideration or a separate statistical check performed after a test was already built. However, the Center for Assessment notes that expectations for fairness have evolved significantly over the past two decades. Fairness is no longer viewed merely as the absence of overt bias, but as a proactive requirement for equal opportunity to demonstrate knowledge, demanding that test designers consider diverse populations from the very beginning of the drafting process.[4]
By 2017, the Buros Center for Testing highlighted a major epistemological shift in the field: fairness is now inextricably linked to validity. If a test systematically disadvantages a specific demographic subgroup—whether due to cultural references, linguistic barriers, or inaccessible formats—it is not measuring the target construct accurately for that group. Therefore, a biased test is fundamentally an invalid test, as the scores do not reflect the true ability of the marginalized test-takers.[5]
The National Council on Measurement in Education's Code of Fair Testing Practices in Education, published in 2018, operationalizes this integration. The code requires developers to design assessments that are accessible to all test-takers, including those with disabilities or diverse linguistic backgrounds, right from the initial blueprinting stage. This universal design approach minimizes the need for retroactive accommodations and ensures that the test measures the intended skill rather than the student's ability to navigate the test format.[2]
To enforce these fairness standards, psychometricians employ advanced statistical techniques like Differential Item Functioning analysis. This technique identifies specific test questions where test-takers of equal overall ability perform differently based on demographic factors such as race, gender, or socioeconomic status. If a question exhibits severe Differential Item Functioning, it is flagged for review by content experts and typically removed from the scoring pool before the test is finalized.[6]
State agencies apply these principles rigorously to comply with federal law and defend their graduation requirements in court. The South Carolina Department of Education's psychometrics division, for instance, continuously monitors state assessments for reliability and validity. They conduct extensive field testing and statistical equating to ensure that a scaled score of 400 means the exact same thing in 2026 as it did in 2023, regardless of which specific test form a student receives on exam day.[6]
Despite these rigorous controls, psychometricians acknowledge the inherent limits of educational measurement. Every test score contains some degree of measurement error. To communicate this uncertainty transparently, score reports often include a Standard Error of Measurement, which provides a confidence interval around a student's reported score, reminding stakeholders that a score is an estimate rather than an absolute truth.[1]
As state education boards and licensing agencies transition toward digital, adaptive testing formats and explore algorithmic scoring, the application of these three properties must evolve. The organizations that govern psychometric standards will next need to determine how artificial intelligence impacts the fundamental guarantees of reliability, validity, and fairness, ensuring that new technologies do not compromise the scientific integrity of high-stakes decisions.[4][7]
Jargon, explained
- Construct
- The specific skill, trait, or concept that a test is designed to measure, such as mathematical reasoning or reading comprehension.
- Reliability Coefficient
- A statistical measure from 0 to 1.0 indicating how consistently a test measures performance across different administrations.
- Differential Item Functioning
- A statistical analysis used to identify test questions that unfairly disadvantage specific demographic groups despite equal overall ability.
- Standard Error of Measurement
- An estimate of the amount of error inherent in a student's test score, providing a confidence interval around the reported result.
Sources
[1]American Psychological AssociationPsychometriciansThe Standards for Educational and Psychological Testing
Read on American Psychological Association →
[2]National Council on Measurement in EducationPsychometriciansCode of Fair Testing Practices in Education
Read on National Council on Measurement in Education →
[3]American Educational Research AssociationEquity AdvocatesAERA Position Statement on High-Stakes Testing in Pre-K–12 Education
Read on American Educational Research Association →
[4]Center for AssessmentEquity AdvocatesEvolving Expectations for Fairness in Testing
Read on Center for Assessment →
[5]Buros Center for TestingPsychometriciansFairness in Educational and Psychological Tests: Critical Issues and Solutions (2017)
Read on Buros Center for Testing →
[6]South Carolina Department of EducationState RegulatorsPsychometrics
Read on South Carolina Department of Education →
[7]Factlen Editorial TeamSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Education
See all →Assessment Reform
How Standards-Based Grading Separates Academic Mastery From Classroom Behavior
2 sources
Title IX Compliance
How the 1972 Title IX Three-Part Test Defines Compliance for Gender Equity in School Sports
6 sources
Employer Benefits
How Section 127 and SECURE 2.0 Employer Student Loan Benefits Work
6 sources
PISA Framework
The 500-Point Mean: How the OECD's PISA Score Compares Student Performance Across Nations
6 sources
Every angle. Every day.
Get Education stories with full source coverage and perspective breakdowns delivered to your inbox.




