Skip to main content
ExplainerHiring ComplianceExplainer· 6 min read· in Careers & Work

Content, Criterion, and Construct: The Three Types of Validity Evidence Required to Legally Defend a Hiring Test

Under federal guidelines, employers using hiring assessments that cause adverse impact must prove the tests are job-related using one of three validation methods. The choice between content, criterion, and construct validity dictates the statistical burden required to defend the selection process in court.

By Amira Darwish

Industrial-Organizational Psychologists 35%Employment Law Counsel 35%Human Resources Practitioners 30%
Industrial-Organizational Psychologists
Focuses on the psychometric rigor and statistical reliability required to prove a test predicts performance.
Employment Law Counsel
Focuses on risk mitigation and surviving a disparate impact claim under Title VII guidelines.
Human Resources Practitioners
Focuses on the practical application, cost, and scalability of testing tools across different roles.

Perspectives this story doesn't cover

  • Candidates who are screened out by algorithmic assessments
  • Test vendors who design and market off-the-shelf construct assessments

In August 1978, the Equal Employment Opportunity Commission (EEOC), the Department of Justice, and the Department of Labor jointly published a set of rules in the Federal Register that fundamentally altered how American companies hire. The Uniform Guidelines on Employee Selection Procedures (UGESP) established a strict legal standard: any hiring test that results in an adverse impact against a protected group is legally discriminatory unless the employer can prove it is strictly job-related. "The guidelines are designed to aid in the achievement of our nation's goal of equal employment opportunity without discrimination on the grounds of race, color, sex, religion or national origin," states the EEOC's official interpretation. That proof is not a vendor's promise or a marketing claim; it is a mathematical and procedural defense known as test validation.[4][5]

The requirement to validate an assessment is triggered by a specific mathematical threshold known as the "four-fifths rule." If a selection procedure screens out a protected group at a rate less than 80% of the highest-selected group, the employer faces a costly legal choice: eliminate the adverse impact, adopt an alternative procedure, or validate the test. Under 29 CFR § 1607.5, the federal government recognizes three distinct frameworks for this defense: content validity, criterion validity, and construct validity. The choice between them dictates the statistical burden the employer must carry into court, directly impacting an organization's legal exposure and hiring budget.[1][4]

Content validity is the most direct and legally straightforward defense available to an employer. It requires demonstrating that the assessment visibly samples the actual work the candidate will perform on the job. A typing test for a data entry clerk, a welding demonstration for a fabricator, or a coding challenge for a software engineer are classic examples of content-valid assessments. Because the assessment is a direct replica of the daily tasks, the legal defense is highly intuitive: the employer simply asked the candidate to perform the job they are applying for. This proximity to the actual work makes it the easiest framework to explain to a judge or jury during a disparate impact claim.[2][6]

While content validity is the easiest to defend conceptually, it still requires strict procedural rigor to survive legal scrutiny. Employers cannot simply guess what the job entails or rely on an outdated job description; content validation must be preceded by a formal, documented job analysis. Subject-matter experts must observe and record the specific duties and skills required, and the test items must link directly to those observed tasks. It is highly defensible, but its primary limitation is that it can only be used to create job-specific procedures. This makes it difficult to scale across entirely different roles within a large organization, forcing human resources teams to build bespoke assessments for every distinct position.[1][6]

The legal and statistical burden of defending a hiring test scales inversely with its proximity to the actual work.

When a test cannot directly sample the work—such as when hiring for entry-level roles that require training—employers often turn to criterion validity. This framework requires concrete statistical proof that scores on the assessment correlate with later job performance. It is generally viewed by the courts as the strongest validity evidence, but it is also the most expensive and mathematically demanding to produce. Establishing criterion validity requires a large sample size—often at least 200 candidates to achieve statistical significance—and a reliable, objective measure of job performance to correlate against the test scores. Without clean performance data, the correlation falls apart, leaving the employer exposed.[3]

When a test cannot directly sample the work—such as when hiring for entry-level roles that require training—employers often turn to criterion validity.

Criterion validity is established through two primary methodologies: predictive validation and concurrent validation. Predictive validation involves testing applicants during the hiring process, hiring them without using the scores to make the decision, and later comparing those initial scores to their six-month or one-year performance reviews. Concurrent validation, by contrast, tests current employees and correlates their scores with their existing performance metrics. Both approaches demand rigorous data collection and a performance metric that is entirely free from managerial bias or contamination. If the performance reviews used as the criterion are themselves biased, the entire validation study becomes legally indefensible.[3][6]

Construct validity is the furthest removed from the actual daily tasks of the job. It involves demonstrating that a test accurately measures an underlying psychological trait or characteristic—such as cognitive ability, leadership potential, spatial reasoning, or attention to detail—that has been judged critical to acceptable job performance. This is the framework typically used to defend off-the-shelf personality inventories and general cognitive ability tests. Because these tests measure abstract human qualities rather than concrete skills, they require a deep foundation of psychometric research to prove that the test actually captures the trait it claims to measure.[2][6]

Because construct validity measures abstract traits rather than direct job tasks, it carries the highest legal exposure for an employer. The organization must justify a complex two-step leap: first, that the test accurately measures the abstract trait, and second, that the specific trait actually predicts success in this specific role. If a vendor cannot produce the underlying validation studies, the employer holds the entirety of the legal risk. Reaching for a construct-valid instrument buys an organization statistical sophistication and scalability, but it significantly raises the burden of proof if the test disproportionately screens out protected candidates during a hiring surge.[4][6]

A selection rate for any group that falls below 80% of the highest group's rate generally triggers the requirement for formal test validation.

A common and dangerous misconception among human resources professionals is that a test can be inherently "valid" straight out of the box. Under federal guidelines, validity is a property of the test's use, not a permanent label attached to the instrument itself. A cognitive ability test might be perfectly valid for selecting financial analysts but entirely indefensible for hiring warehouse staff. The EEOC requires local validation, meaning the employer must document that the test is valid for their specific use and their specific applicant population. Relying solely on a vendor's generalized marketing claims without local data is a fast track to a Title VII violation.[2][4]

As algorithmic assessments and artificial intelligence tools flood the hiring market, the foundational rules of test validation remain unchanged. Even when the federal government updated its guidance in 2004 to address Internet-based recruiting, the core mandate held firm. The EEOC has signaled increased scrutiny on automated selection procedures, reminding employers that they cannot outsource their Title VII liability to a software vendor's proprietary algorithm. The legal defense of a hiring test still rests on the exact same three pillars established nearly fifty years ago: the employer must prove the test measures the work, predicts the outcome, or isolates the trait that matters.[5][6]

Key points

  1. The EEOC requires employers to validate any hiring test that causes an adverse impact against a protected group.
  2. Content validity defends a test by proving it directly samples the actual tasks required on the job.
  3. Criterion validity uses statistical data to prove that test scores accurately predict future job performance.
  4. Construct validity measures abstract traits like cognitive ability, carrying the highest burden of proof to link the trait to the role.
  5. Validity is a property of how a test is used for a specific job, meaning employers cannot rely solely on off-the-shelf vendor claims.

Key terms

Adverse Impact
A substantially different rate of selection in hiring or promotion that works to the disadvantage of a protected race, sex, or ethnic group.
Four-Fifths Rule
A rule of thumb used by federal agencies stating that a selection rate for any group which is less than 80% of the rate for the highest group is evidence of adverse impact.
Job Analysis
A systematic process of gathering, documenting, and analyzing information about the duties, responsibilities, and necessary skills of a specific job.
Disparate Impact
A legal doctrine under Title VII where a facially neutral employment practice disproportionately excludes a protected group and is not justified by business necessity.
Local Validation
The process of gathering evidence to prove that a selection procedure is valid for a specific job and applicant population within a particular organization.

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Industrial-Organizational Psychologists 35%Employment Law Counsel 35%Human Resources Practitioners 30%
  1. [1]Law.Cornell.EduEmployment Law Counsel

    29 CFR § 1607.5 - General standards for validity studies.

    Read on Law.Cornell.Edu
  2. [2]Recruiters LineUpHuman Resources Practitioners

    What Makes a Hiring Test Legally Compliant?

    Read on Recruiters LineUp
  3. [3]TestGeniusIndustrial-Organizational Psychologists

    Criterion Validity and Pre-Employment Testing

    Read on TestGenius
  4. [4]EEOCEmployment Law Counsel

    Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures

    Read on EEOC
  5. [5]Federal RegisterEmployment Law Counsel

    Adoption of Additional Questions and Answers to clarify and provide a common interpretation of the Uniform Guidelines on Employee Selection Procedures

    Read on Federal Register
  6. [6]Factlen Editorial TeamIndustrial-Organizational Psychologists

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Careers & Work stories with full source coverage and perspective breakdowns delivered to your inbox.