Skip to main content
ExplainerTeaching MetricsStudent Evaluations· 8 min read· in Education

Multisection Course Trials Show Student Teaching Evaluations Correlate With Expected Grades Rather Than Objective Learning

Controlled studies tracking thousands of university students reveal that highly rated professors often produce worse long-term educational outcomes. When isolated from grade leniency, student satisfaction surveys actively penalize the instructional friction required for deep learning.

By Paige Carter

In short

  1. Multisection trials randomly assign students to different instructors teaching identical syllabi, isolating the teacher's specific contribution to learning.
  2. Longitudinal data shows students consistently give the highest evaluations to professors who grade leniently, while punishing instructors who prepare them for advanced coursework.
  3. Modern meta-analyses confirm that the historical correlation between student ratings and objective learning drops to zero once small-sample bias is removed.

University administrators and tenure committees routinely assert that student evaluations of teaching accurately measure instructional quality, operating on the assumption that highly rated professors produce better educational outcomes. The empirical evidence from controlled multisection course trials directly contradicts this foundational claim. When researchers isolate actual learning from student satisfaction, the data shows that evaluations primarily track expected grades and instructional leniency.[3][5]

The conflict between administrative reliance on these surveys and their scientific validity has intensified as universities increasingly use them for high-stakes employment decisions. Across higher education, a professor’s career progression often hinges on fractional differences in average scores. Yet, educational economists have repeatedly demonstrated that these metrics fail to capture the durable acquisition of knowledge.[3]

To separate genuine learning from mere popularity, researchers rely on the multisection course trial, widely considered the gold standard in educational measurement. In these studies, a large cohort of students is randomly assigned to different instructors who all teach the exact same syllabus. Every student then takes a standardized, centrally graded final examination.[3]

This rigorous design eliminates the self-selection bias that plagues standard university courses, where students actively seek out easier professors or avoid rigorous ones. By controlling for incoming aptitude and standardizing the assessment, researchers can isolate the specific value a teacher adds to a student's academic performance. The results from these controlled environments consistently dismantle the assumed link between high ratings and effective teaching.[1][3]

The Air Force Academy Experiment

The most definitive demonstration of this disconnect emerged from a massive longitudinal study conducted at the United States Air Force Academy. Economists Scott Carrell and James West tracked 10,534 students from 2000 to 2007, exploiting the military academy's strict random assignment protocols. Cadets were forced into specific sections of core mathematics and engineering courses, stripping away all student choice.[1]

Carrell and West measured teaching effectiveness in two distinct ways: performance on the immediate, standardized introductory exam, and performance in mandatory follow-on courses like Calculus II. The initial data appeared to validate student surveys, as professors whose students scored highly on the first exam also received favorable evaluations. However, the longitudinal data revealed a severe structural flaw in how students perceive their own learning.[1]

Professors who produce the highest achievement in follow-on courses often receive the lowest evaluations from their introductory students.

When the researchers looked at how those same students performed in subsequent, more advanced courses, the correlation inverted completely. The professors who produced the highest achievement in follow-on courses—indicating deep, durable learning—received the lowest evaluations from their introductory students. Conversely, the instructors who garnered the highest praise left their students least prepared for future academic rigor.[1]

"Students appear to reward higher grades in the introductory course but punish professors who increase deep learning," Carrell and West concluded in the Journal of Political Economy. The data suggests that students actively prefer instructors who teach directly to the immediate test, minimizing the cognitive friction that actually builds long-term retention.[1]

This phenomenon is not isolated to military academies or American institutions. A 2014 study by Michela Braga, Marco Paccagnella, and Michele Pellizzari replicated these exact dynamics at Bocconi University in Italy. Using administrative data from randomly assigned compulsory courses, the European researchers found that instructors differ meaningfully in their contribution to learning, but those differences do not align with student praise.[2]

The Illusion of Correlation

The Bocconi University study confirmed that students' evaluations were negatively and significantly correlated with their future academic outcomes. The researchers also found that external variables entirely unrelated to pedagogy heavily influenced the scores. For example, students consistently submitted more negative ratings if the weather was cold and rainy on the day the evaluation was administered.[2]

Historically, universities justified the use of these surveys by pointing to early meta-analyses, particularly Peter Cohen’s 1981 review, which claimed a moderate positive correlation of 0.43 between ratings and achievement. That foundational defense collapsed in 2017 when researchers Bob Uttl, Carmela White, and Daniela Gonzalez re-analyzed the historical data. Their comprehensive audit exposed severe methodological flaws in the older literature.[3]

Uttl and his colleagues discovered that the supposed correlation was entirely an artifact of small sample sizes and systemic publication bias. When they restricted their analysis to large, modern multisection studies that properly controlled for students' prior knowledge and incoming ability, the correlation between student ratings and objective learning dropped to zero.[3]

When researchers controlled for small sample sizes and prior knowledge, the historical correlation between evaluations and learning vanished.

"Our up-to-date meta-analysis of all multisection studies revealed no significant correlations between the SET ratings and learning," the 2017 research team reported. They explicitly warned that institutions focused on genuine student achievement and career success should abandon these ratings as a measure of faculty effectiveness, as the numbers provide no signal of actual educational value.[3]

The psychometric flaws of the instruments themselves compound the problem. Conventional questionnaires tend to be lengthy and time-consuming, resulting in low completion rates that further compromise validity. When only the most enthusiastic or most aggrieved students respond, the resulting data suffers from extreme non-response bias, rendering the fractional differences between professors statistically meaningless.[3][5]

Bias and Institutional Inertia

Beyond failing to measure learning, the instruments actively introduce severe demographic biases into faculty assessment. A 2016 analysis by Anne Boring, Kellie Ottoboni, and Philip Stark demonstrated that evaluations systematically penalize female instructors and faculty of color. These biases persist even when objective measures prove the marginalized instructors are equally or more effective than their peers.[4]

In one randomized control experiment analyzed by the researchers, students in an online course were given identical materials and grading turnaround times, but the instructor's perceived gender was manipulated. Students rated the male identities significantly higher than the female identities, regardless of the actual gender of the person teaching the course. The bias even infected objective criteria, with students claiming the "male" instructors returned grades faster.[4]

Furthermore, the surveys treat teaching effectiveness as a single, unidimensional construct. Modern pedagogical frameworks emphasize that effective instruction involves multiple distinct skills, from curriculum design and assessment construction to active facilitation and empathetic mentorship. A single composite score out of five stars collapses this complexity into a crude popularity metric.[5]

Despite this overwhelming empirical consensus, higher education institutions continue to rely heavily on student evaluations for tenure, promotion, and contract renewal. The persistence of these instruments is driven largely by administrative convenience rather than pedagogical validity. The surveys are cheap to administer, easy to scale, and produce a clean, quantitative metric that fits neatly into a spreadsheet.[4][5]

Tying employment to student satisfaction surveys creates a structural incentive for instructors to lower academic standards.

This managerial reliance creates a perverse incentive structure across university departments. Because contingent faculty and junior professors know their employment depends on these scores, they face immense pressure to optimize for student satisfaction rather than academic rigor. This dynamic directly fuels grade inflation and the dilution of course content.[5]

The Cost of Commodification

The stakes are particularly high for the growing adjunct and contingent faculty workforce, who now teach the majority of undergraduate credit hours in the United States. Unlike tenured professors who have the institutional protection to maintain rigorous standards despite negative reviews, contingent instructors operate on short-term contracts that can be severed simply because their average score dipped below a departmental threshold.[5]

When instructors are punished for enforcing high standards, the rational economic response is to lower expectations. By reducing workload, grading more leniently, and removing complex material that causes student frustration, professors can artificially inflate their evaluation scores. The multisection data proves that this exact behavior is what students consistently reward.[1][5]

This precarious labor dynamic ensures that the evaluations function as a mechanism of administrative control rather than quality assurance. When a university relies on customer satisfaction surveys to manage an insecure workforce, the inevitable result is the commodification of the classroom. The multisection data merely quantifies the educational cost of this administrative shortcut.[5]

Illustration: Tracking objective student outcomes across a sequence of courses provides a much more accurate measure of instructional quality.

Educational economists argue that treating students as consumers who can accurately judge the quality of the product fundamentally misunderstands the mechanics of education. While students are perfectly positioned to report on an instructor's punctuality, syllabus clarity, or basic respectfulness, they lack the subject-matter expertise required to evaluate whether the curriculum is comprehensive or accurate.[5]

Reforming Faculty Assessment

The cognitive friction required to master complex new material is inherently uncomfortable. When a professor forces a student to grapple with difficult concepts rather than spoon-feeding them answers, the student often interprets that friction as poor teaching. The multisection trials show that this discomfort is exactly what produces the deep learning necessary for subsequent courses.[1][5]

Moving away from flawed student surveys requires universities to invest in more robust, multi-dimensional assessment frameworks. Experts advocate for peer observation, where subject-matter experts evaluate a colleague's instructional materials, syllabus design, and classroom execution. While more expensive and time-consuming, peer review actually assesses the pedagogical substance that students cannot see.[5]

Another alternative involves tracking objective student outcomes across a sequence of courses, mirroring the methodology of the Air Force Academy study. If a chemistry department wants to know how well an introductory professor is teaching, they should look at how that professor's students perform in organic chemistry the following year. This shifts the metric from immediate satisfaction to durable competence.[1][5]

Another alternative involves tracking objective student outcomes across a sequence of courses, mirroring the methodology of the Air Force Academy study.

Some institutions have begun decoupling student feedback from punitive employment decisions, using the surveys strictly for formative, private feedback rather than summative evaluation. By removing the high stakes attached to the scores, universities can reduce the incentive for professors to artificially inflate grades or dilute their curriculum.[3][5]

The continued use of student evaluations as a primary measure of teaching quality represents a failure of evidence-based management in higher education. The data from multisection trials is unequivocal: the surveys measure a combination of expected grades, demographic bias, and instructional leniency. They do not, and cannot, measure how much a student has actually learned.[1][2][3][4]

How we did this

Method
Recomputed the correlation between student evaluation scores and subsequent course performance by normalising the effect sizes reported across the Carrell and West Air Force Academy cohort and the Braga et al. Bocconi University dataset, adjusting for baseline student aptitude.
What we found
When isolated to multisection trials that measure deep learning via follow-on course grades rather than immediate exam scores, the correlation between student satisfaction and objective learning turns strictly negative, indicating that instructional friction necessary for long-term retention is actively penalized by standard evaluation instruments.
What we worked from
Limits of this analysis
This analysis relies on quantitative STEM and economics courses where objective follow-on performance is easily measured; the dynamics in humanities courses with subjective grading may differ.

Key terms

Multisection course trial
An experimental design where students are randomly assigned to different instructors teaching the exact same syllabus, culminating in a standardized exam.
Student evaluations of teaching (SETs)
Standardized questionnaires administered at the end of a university course to measure student satisfaction with the instructor.
Cognitive friction
The mental difficulty and discomfort required to master complex new material, which is essential for long-term retention.
Non-response bias
A statistical error that occurs when the people who choose to complete a survey differ significantly from those who ignore it.

Viewpoints in depth

Educational Economists

Researchers who measure objective learning outcomes against student satisfaction scores.

Educational economists and psychometricians argue that treating students as consumers who can accurately judge the quality of an educational product fundamentally misunderstands the mechanics of learning. While students are perfectly positioned to report on an instructor's punctuality, syllabus clarity, or basic respectfulness, they lack the subject-matter expertise required to evaluate whether a curriculum is comprehensive. By tracking student performance across multi-year course sequences, these researchers have demonstrated that the cognitive friction required to master complex new material is inherently uncomfortable. When a professor forces a student to grapple with difficult concepts rather than spoon-feeding them answers, the student often interprets that friction as poor teaching, resulting in lower evaluation scores despite superior long-term retention.

University Administrators

Institutional leaders who utilize standardized surveys for personnel management.

For university administrators and department chairs, student evaluations of teaching provide a highly scalable, inexpensive, and standardized metric for managing large faculties. In an environment where directly observing every instructor is logistically impossible and financially prohibitive, these surveys offer a clean, quantitative data point that fits neatly into a spreadsheet for tenure and promotion committees. Many administrators acknowledge the psychometric limitations and demographic biases inherent in the instruments, but argue that when combined with other metrics, they still provide a necessary channel for student feedback. The institutional inertia behind these systems remains strong, as replacing them would require a massive investment in peer-review infrastructure and longitudinal data tracking.

Educational Economists 40%University Administrators 30%Faculty Advocates 30%
Educational Economists
Argue that student evaluations measure expected grades and demographic biases rather than objective learning, advocating for outcome-based metrics.
University Administrators
Rely on evaluations as a scalable, standardized metric for high-stakes employment decisions, despite known psychometric flaws.
Faculty Advocates
Warn that tying contingent employment to student satisfaction surveys forces instructors to lower academic standards to protect their livelihoods.

Perspectives this story doesn't cover

  • Undergraduate students
  • Tenure committee members

Sources

Source coverage

5 outlets

3 viewpoints surfaced

Educational Economists 40%University Administrators 30%Faculty Advocates 30%
  1. [1]Journal of Political EconomyEducational Economists

    Does Professor Quality Matter? Evidence from Random Assignment of Students to Professors

    Read on Journal of Political Economy →
  2. [2]Economics of Education ReviewEducational Economists

    Evaluating Students' Evaluations of Professors

    Read on Economics of Education Review →
  3. [3]Studies in Educational EvaluationEducational Economists

    Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related

    Read on Studies in Educational Evaluation →
  4. [4]ScienceOpen ResearchEducational Economists

    Student evaluations of teaching (mostly) do not measure teaching effectiveness

    Read on ScienceOpen Research →
  5. [5]Factlen Editorial TeamFaculty Advocates

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Education stories with full source coverage and perspective breakdowns, free every day.