top of page

Be Notified of New Research Summaries -

It's Free!

Which Students Are Most Likely to Exhibit Gender Bias in Professor Evaluations?

  • Writer: Greg Thorson
    Greg Thorson
  • 2 hours ago
  • 7 min read

El Khoury (2026) asks which students are most responsible for gender bias in online professor evaluations. She combines RateMyProfessor data from business and economics instructors at Arizona State University and several other large public universities with a randomized experiment involving 288 students. Students evaluated identical teaching content delivered with either a male or female voice. Female professors received online quality ratings 0.13 to 0.17 points lower on a 1–5 scale. In the experiment, the female instructor was rated 0.11 points lower overall. Among students who had previously posted professor reviews, the penalty increased to about 0.22 points, or 5.1 percent.


Why This Article Was Selected for The Policy Scientist

This article addresses a broad issue in higher education: whether widely used information systems accurately represent teaching quality or instead transmit systematic distortions that can influence student choices. That question is increasingly important as students rely heavily on online ratings when selecting courses and instructors. El Khoury’s study is particularly timely because digital review platforms now shape decisions well beyond higher education, making selective participation and biased evaluations important concerns across many markets. The article also extends a substantial literature showing gender differences in teaching evaluations, including influential experimental work demonstrating bias even when instructional quality is held constant. The data are reasonably strong because the study combines observations from several universities with a randomized experiment involving identical instructional content. The experimental design is especially persuasive because random assignment supports a causal interpretation and provides stronger evidence than multivariate regression alone. Generalizability beyond large public universities remains uncertain, but the consistency of the observational patterns across institutions strengthens the broader relevance of the findings. The article was published in AEA Papers and Proceedings, a well-regarded outlet of the American Economic Association that disseminates current empirical research across economics.


Full Citation and Link to Article

El Khoury, S. (2026). Who is biased? Heterogeneity in gender bias in evaluations. AEA Papers and Proceedings, 116, 594–598. https://www.aeaweb.org/articles?id=10.1257/pandp.20261120


Central Research Question

The study asks who generates gender bias in online professor evaluations. A substantial literature has established that female instructors often receive lower teaching evaluations than male instructors even when measures of teaching effectiveness are similar. The more specific question addressed here is whether this bias is broadly distributed across students or concentrated among particular types of evaluators. The analysis focuses especially on students who actively participate in professor-review websites such as RateMyProfessor.com. The central hypothesis is that the ratings observed on these platforms may not represent the views of the broader student population because the students who choose to post reviews may differ systematically from those who do not.


The study therefore separates two related questions. First, do female instructors receive lower ratings than comparable male instructors on professor-review websites? Second, when instructional content is experimentally held constant, which students are most likely to penalize an instructor who is perceived to be female? The analysis considers whether bias differs according to student gender, minority status, or prior participation in professor-review websites. This distinction matters because online ratings can influence subsequent students’ beliefs about professors and affect course-selection decisions.


Previous Literature

The study builds on a well-developed literature examining gender bias in student evaluations of teaching. MacNell, Driscoll, and Hunt (2015) provide particularly relevant experimental evidence. Their research demonstrated that students could evaluate instructors differently depending on the gender they believed an instructor to have, even when the underlying teaching behavior was held relatively constant. This work strengthened the argument that observed gender gaps in evaluations cannot necessarily be interpreted as differences in teaching performance.


Boring, Ottoboni, and Stark (2016) similarly questioned whether student teaching evaluations accurately measure instructional effectiveness. Their findings contributed to a broader methodological critique of using student ratings as direct indicators of teaching quality. Boring (2017) provided additional evidence of gender bias in student evaluations, showing that female instructors could receive systematically different ratings even after accounting for factors intended to capture teaching performance.


Boring and Philippe (2021) extended this literature by examining whether an intervention designed to increase awareness of gender bias could reduce discrimination in teaching evaluations. Their work is especially relevant because it moves beyond documenting evaluation differences to considering whether information can alter biased judgments. Binderkrantz, Skorkjær, and Bisgaard (2024) also examined the role of gender in university teaching evaluations, providing evidence that evaluator and instructor gender may shape ratings. Saygin and Zhang (2025) further connected gender differences in teaching evaluations to student course enrollment, demonstrating that evaluation patterns may have consequences beyond the ratings themselves.


The present study advances this literature by shifting attention from whether gender bias exists to identifying who produces it. Rather than assuming that an average gender difference reflects a uniform tendency across students, it examines heterogeneity among evaluators and focuses particularly on the behavior of students who actually contribute ratings to online platforms.


Data

The empirical analysis combines observational data from RateMyProfessor with original experimental data. The observational component begins with professor-level RateMyProfessor data for economics and business instructors at Arizona State University. The dataset includes instructor names, inferred binary gender, department, average quality rating, average difficulty rating, and the number of reviews received. The Arizona State University sample contains 1,202 instructor profiles.


The author then extends the observational analysis to business schools at several additional large public universities, including Florida International University, Michigan State University, Georgia State University, the University of Alabama, California State University Fullerton, and Washington State University. This broader dataset contains 3,781 instructor profiles. Because RateMyProfessor profiles accumulate over many years, these observations include individuals who have taught at the institutions in different instructional capacities, including faculty, lecturers, adjunct instructors, graduate teaching assistants, visiting faculty, and others.


The experimental component consists of 288 students enrolled in introductory economics and business statistics courses. Data were collected in two waves during summer and fall 2024. Approximately 73 percent of participants were economics or business students, while 18 percent were STEM majors. Students completed an online survey outside class after being invited through their courses.


Methods

The observational analysis estimates whether female professors receive different average quality ratings from male professors. The regressions control for factors including department, institution, perceived course difficulty, and the number of ratings received. These models identify systematic associations between instructor gender and online ratings, but the study explicitly recognizes that observational differences cannot by themselves establish gender bias. Female and male instructors could differ in unobserved characteristics related to teaching, course assignments, or student populations.


The randomized experiment addresses this identification problem. Students watched a two-minute instructional video covering a basic probability concept. The video displayed only a virtual whiteboard, eliminating visual information about the instructor. Professional actors recorded male and female voice-overs using the same instructional script. Students were randomly assigned within their courses to hear either the male or female version.

Immediately afterward, students evaluated the instructor’s quality and difficulty using scales similar to those employed by RateMyProfessor. Because the instructional content was identical and perceived instructor gender was randomly assigned, differences in quality ratings can be interpreted as causal effects of perceived gender.


The author then estimates heterogeneous treatment effects by interacting the female-instructor assignment with student characteristics. These include student gender, minority status, and whether the student had ever submitted a professor review online. This design permits the analysis to identify not only the average gender penalty but also whether it is concentrated among particular groups.


Findings/Size Effects

The observational evidence shows a consistent gender gap in online professor ratings. At Arizona State University, female instructors receive quality ratings approximately 0.128 points lower than male instructors on a 1-to-5 scale. This difference corresponds to roughly 3.5 percent of the average rating. Across the additional public universities, the estimated gap is approximately 0.169 points. The consistency across institutions suggests that the pattern is not unique to Arizona State University.


The randomized experiment provides direct evidence that at least part of this difference reflects gender bias rather than differences in instructional content. On average, students assigned to the female voice-over rate the instructor approximately 0.109 points lower than students assigned to the male voice-over. Because the instructional script and presentation are otherwise identical, this difference isolates the effect of perceived instructor gender.

The most important result concerns heterogeneity in bias. The study finds little evidence that the gender penalty varies systematically by student gender. Female and male students exhibit similar responses to the instructor’s perceived gender. Similarly, minority and nonminority students do not differ significantly in the magnitude of their evaluations of the female instructor.


A much larger difference emerges according to whether students participate in professor-review websites. Among students who report never posting an online professor review, the estimated female-instructor penalty is only about 0.028 points and is statistically insignificant. Among students who have previously posted at least one review, however, the additional penalty is approximately 0.193 points. Combining these coefficients produces a total estimated penalty of approximately 0.22 points for active reviewers.


This 0.22-point reduction represents roughly 5.1 percent of the male instructor’s average rating and is statistically significant at the 1 percent level. The magnitude is therefore substantially larger among the students who actually generate the reviews that appear on online platforms. The findings imply that the gender gap observed on professor-review websites can be disproportionately generated by a relatively active subset of students rather than reflecting an equally strong bias throughout the student population.


Conclusion

The study concludes that professor-review websites may transmit gender-biased information because participation in these platforms is selective. Female instructors receive lower average ratings in observational data, and the randomized experiment demonstrates that students assign lower quality ratings to an otherwise identical instructor when the instructor is perceived to be female.


The principal contribution is the identification of who appears to generate this bias. The differences are not meaningfully concentrated among male students or among nonminority students. Instead, the strongest bias appears among students who have previously submitted online professor reviews. Students who do not participate in these platforms exhibit little measurable gender penalty, while active reviewers impose a substantially larger one.


These results distinguish the views of the broader student population from the ratings produced by the subset of students who contribute online evaluations. Because professor-review websites are widely consulted when students select courses, selective participation can give a relatively small group of reviewers disproportionate influence over the information available to other students. The study therefore demonstrates that understanding online evaluation systems requires examining not only average ratings, but also the composition and behavior of the individuals who produce them.

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
Screenshot of Greg Thorson
  • Facebook
  • Twitter
  • LinkedIn


The Policy Scientist

Offering Concise Summaries*
of the
Most Recent, Impactful 
Public Policy Research

*Summaries Powered by ChatGPT

bottom of page