Skip to main content
ExplainerAssessment ScienceExplainer· 6 min read· in Education

Content, Criterion, and Construct: How the Three Types of Validity Define the Quality of an Educational Assessment

An educational test is only as useful as the evidence proving it measures what it claims to measure. The classical tripartite model—content, criterion, and construct validity—provides the foundational framework for evaluating whether an assessment yields accurate and meaningful results.

By Nabil Faris

Classical Psychometricians 40%Modern Unified Theorists 40%Practical Test Developers 20%
Classical Psychometricians
Emphasize the distinct, measurable nature of content, criterion, and construct validity as separate checklists required for test development.
Modern Unified Theorists
Argue that all validity is ultimately construct validity, requiring a cohesive, argument-based approach rather than isolated statistical tests.
Practical Test Developers
Focus on the legal and practical necessity of documenting content validity to defend high-stakes exams against bias or fairness challenges.

Perspectives this story doesn't cover

  • Student Test-Takers
  • Classroom Teachers

Common questions

What is the difference between validity and reliability?

Reliability refers to the consistency of a test's results, while validity refers to whether the test actually measures what it claims to measure. A test can be highly reliable but completely invalid, but it cannot be valid without first being reliable.

Why is construct validity considered the most important?

Construct validity is the overarching concept that confirms a test is measuring the intended invisible trait, such as intelligence or reading comprehension. Modern psychometricians argue that all other forms of validity are simply evidence supporting construct validity.

How do test makers prove content validity?

Test developers use subject matter experts to map individual test questions directly to established curriculum standards, ensuring the exam covers the appropriate breadth and depth of the topic.

What is Kane's argument-based framework?

Introduced in 1992, Kane's framework requires test developers to build a logical chain of evidence supporting the specific real-world decisions made based on test scores, rather than just checking statistical boxes.

The short answer

  • Validity is the most critical metric of an educational assessment, proving the test measures what it claims to measure.
  • The classical tripartite model divides validity into three categories: content, criterion, and construct.
  • Content validity ensures the test covers the right material, while criterion validity proves the scores predict real-world outcomes.
  • Modern psychometrics, led by Kane's framework, treats all validity as construct validity, requiring a unified, argument-based approach.
  • A test can be highly reliable without being valid, but it cannot be valid if it is not first reliable.

An educational assessment’s quality is defined by its validity—the extent to which a test accurately measures the specific trait, skill, or knowledge it claims to evaluate. This quality is established through three classical pillars: content validity ensures the test covers the right material, criterion validity proves the scores predict real-world outcomes, and construct validity confirms the exam actually measures the underlying psychological or educational concept. Without these three foundational elements, a test score is merely a number disconnected from reality.[5]

The stakes for getting this right are massive. From third-grade reading benchmarks to the medical board exams that license physicians, standardized assessments dictate the allocation of billions of dollars in public funding and determine individual career trajectories. Test developers cannot simply assert that their exams work; they must provide empirical evidence. The framework for organizing that evidence relies on the tripartite model, which was first formalized in 1954 by the American Psychological Association (APA).[1]

Content validity is the most intuitive of the three types. It asks whether the items on a test adequately represent the entire domain of knowledge being assessed. If a high school geometry final consists entirely of questions about triangles, it lacks content validity because it ignores circles, polygons, and three-dimensional shapes. To establish this, test publishers rely heavily on subject matter experts who map individual test questions directly to established curriculum standards.[5]

The Massachusetts Tests for Educator Licensure (MTEL) provides a clear practical example of this mechanism in action. According to the official MTEL documentation, "Validity is the most important consideration in test evaluation." For state licensure exams, content validity is legally required to prove that the test measures the specific knowledge a teacher needs in the classroom, protecting the state against claims that the exam is arbitrary or discriminatory.[6]

The classical tripartite model divides validity evidence into three distinct categories.

Criterion validity shifts the focus from the content of the test to the outcomes it predicts. It measures how well one assessment predicts an outcome for another established measure. This is typically divided into two subcategories: predictive validity and concurrent validity. Predictive validity evaluates whether a score can forecast future performance, such as how well a college entrance exam predicts a student's first-year grade point average.[5]

Concurrent validity, by contrast, measures how well a new test correlates with an already established assessment administered at the same time. If a school district wants to replace a lengthy, expensive three-hour reading diagnostic with a cheaper 20-minute version, they must demonstrate concurrent validity. If students score similarly on both the short and long tests, the new assessment is deemed valid for that specific criterion.

Construct validity is the most complex and, according to modern psychometricians, the most critical of the three pillars. A "construct" is an abstract, theoretical concept that cannot be directly observed—such as intelligence, reading comprehension, or mathematical reasoning. Construct validity asks whether the test is actually measuring that specific invisible trait, rather than being skewed by unrelated factors.[1]

Construct validity is the most complex and, according to modern psychometricians, the most critical of the three pillars.

For example, a math test composed of highly complex word problems might inadvertently measure a student's reading comprehension rather than their mathematical ability. If a student fails the test because they could not parse the vocabulary, the test lacks construct validity as a math assessment. The APA PsycTests methodology database explicitly tracks how researchers validate these constructs using statistical tools like factor analysis and item-response theory.[1]

While the classical tripartite model remains the standard for introductory measurement courses and legal defense, the scientific consensus has evolved. In 1989, psychometrician Samuel Messick introduced a unified theory of validity, arguing that content and criterion evidence are simply subsets of construct validity. Under this unified view, every piece of evidence serves to prove that the test measures the intended construct.[3]

Validity theory has evolved from isolated statistical checks to unified, argument-based frameworks.

This unified approach was further refined in 1992 by Michael Kane, who introduced the argument-based approach to validation. As detailed in contemporary PubMed literature, Kane’s framework requires test developers to explicitly state the proposed interpretations and uses of test scores, and then build a logical, evidence-based argument to support those specific claims.[2]

Kane's framework treats test validation like a courtroom trial. The test developer is the prosecutor, presenting a chain of inferences—from the scoring of a single item to the final decision made about a student. If any link in that chain is weak, the entire validity argument collapses. This approach forces developers to consider not just the statistical properties of the test, but the real-world consequences of its use.[2][4]

The health professions have aggressively adopted this modern framework. Research published in the National Center for Biotechnology Information outlines how medical education assessment now relies on four tenets of modern validity theory, moving beyond traditional checklists. When licensing a surgeon, proving that a multiple-choice test covers the right content is insufficient; the validity argument must prove that the score translates to safe clinical practice.[3][4]

Despite these theoretical advancements, the classical tripartite model remains deeply embedded in educational policy. The Association for Education Finance and Policy notes that policymakers and school administrators still rely heavily on the distinct concepts of content and criterion validity when evaluating commercial assessments. The tripartite language is simpler to communicate to school boards, parents, and legislators than complex argument-based frameworks.

Subject matter experts play a critical role in establishing the content validity of an assessment.

It is also impossible to discuss validity without addressing its prerequisite: reliability. Reliability refers to the consistency of a test. If a student takes the same exam three times in one week and receives wildly different scores, the test is unreliable. A typical reliability coefficient must exceed 0.70 before validity can even be considered. A test can be highly reliable but completely invalid, but a test cannot be valid if it is not first reliable.[5]

The tension between classical and modern frameworks is fundamentally a debate about scope. The classical model asks whether the test has the right parts, while the modern unified model asks whether the test score justifies the decision being made. Both frameworks agree that a test is never universally valid; it is only valid for a specific purpose, with a specific population, at a specific time.[2]

As educational assessments increasingly incorporate artificial intelligence, adaptive algorithms, and gamified interfaces, the burden of proof on test developers is growing. Whether organized into three classical pillars or a single unified argument, the fundamental requirement remains unchanged: an assessment must earn the trust placed in its results through rigorous, continuous empirical validation.[4]

Why it matters

Billions of dollars in educational funding, college admissions, and professional licensure decisions hinge on standardized test scores. If an assessment lacks validity, the resulting data can misguide policy, misallocate resources, and unfairly penalize students.

Jargon, explained

Validity
The extent to which an assessment accurately measures the specific trait, skill, or knowledge it claims to evaluate.
Reliability
The consistency of a test's results over time and across different populations.
Construct
An abstract psychological or educational concept, such as mathematical reasoning or reading comprehension, that a test attempts to measure.
Criterion
An external measure or real-world outcome, such as college GPA or job performance, that a test score is expected to predict.

Sources

Source coverage

7 outlets

3 viewpoints surfaced

Classical Psychometricians 40%Modern Unified Theorists 40%Practical Test Developers 20%
  1. [1]American Psychological AssociationPractical Test Developers

    APA PsycTests Methodology Field Values

    Read on American Psychological Association
  2. [2]PubMedModern Unified Theorists

    A contemporary approach to validity arguments: a practical guide to Kane's framework

    Read on PubMed
  3. [3]PMCModern Unified Theorists

    Four tenets of modern validity theory for medical education assessment and evaluation

    Read on PMC
  4. [4]Taylor & Francis Online

    Bridging Validity Frameworks in Assessment: Beyond Traditional Approaches in Health Professions Education

    Read on Taylor & Francis Online
  5. [5]PressbooksPractical Test Developers

    Chapter 9 Validity

    Read on Pressbooks
  6. [6]MTELClassical Psychometricians

    TEST VALIDITY

    Read on MTEL
  7. [7]Factlen Editorial Team

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Education stories with full source coverage and perspective breakdowns delivered to your inbox.