Mon 20 Jul 2026 / 23:14 ET
Kernel
AI 4 min read

LLMs formed hiring stereotypes faster than people in simulation

Princeton and University of Chicago researchers found major chatbots sorted fictional job candidates by ethnicity despite equal success rates.

Riley Okafor

By Riley Okafor / Senior AI Reporter

Large language models asked to act like hiring consultants formed ethnic stereotypes in a simulated labor market more aggressively than human participants in an earlier psychology experiment, according to researchers at Princeton University and the University of Chicago. The finding cuts straight into a messy question for employers using AI to screen applicants: a model does not need to inherit a bias from training data to produce one. It can infer one from thin air and a few misleading examples.

The study, published as a paper at ICML in Seoul in July, tested models including OpenAI’s ChatGPT and o3, Anthropic’s Claude, Google’s Gemini, and DeepSeek’s R1. The researchers adapted a hiring game from prior psychology work on how people form stereotypes.

In the experiment, each model was told it had been hired by the mayor of a fictional city to help fill 20 jobs, including doctors, lawyers, child-care aides, and janitors. Every round presented four candidates, one from each of four invented ethnic groups: Tufa, Aima, Reku, and Weki. After choosing someone, the model learned whether the hire succeeded, then continued to the next opening. The goal was to make as many successful hires as possible over 40 rounds.

The catch was that all four groups had the same chance of success in every job. The models were seeing random outcomes, then treating those outcomes as evidence.

The researchers found that models quickly began assigning groups to job niches based on early feedback. If, for example, a model saw an Aima candidate fail as a doctor, it became less likely to hire Aima candidates for doctor roles and shifted them toward jobs the model rated as requiring less warmth and competence, such as janitor.

On the study’s segregation scale, where 2 means each group has been fully boxed into a separate occupational niche, human participants in the original study scored 0.84. The AI models scored about 65% higher overall, according to the researchers. OpenAI’s o3 reached 1.83, near the top of the scale.

Ryan Liu, a Princeton PhD student and co-author of the study, said the behavior reflects a core strength of LLMs misfiring in a social setting. The systems are optimized to generalize from limited data, Liu said, which helps on tasks such as math, coding, and science problems. In this hiring game, that same habit pushed the models to overread sparse feedback and convert it into group-level assumptions.

The researchers reported stronger stereotyping in newer reasoning models, including OpenAI’s o3 and DeepSeek’s R1. OpenAI and Anthropic did not respond to requests for comment, according to the report.

The result also raises an awkward issue for chatbot memory. Angelina Wang, a Cornell University computer scientist who was not involved in the study, said systems that draw on prior interactions may over-weight patterns they have seen before. She also cautioned that forcing chatbots to remember less is not a clean fix, since users expect personalization.

One mitigation tested in the study worked better than polite instruction. Telling models to be fair had little effect, Liu said, possibly because the instruction was overwhelmed by the stated objective of maximizing correct hires. Offering a bonus for diverse hiring reduced bias more substantially, suggesting that model objectives may matter more than abstract value prompts.

A second experiment asked models to resettle members of fictional ethnic groups across Canadian cities. When given relevant personal details, such as age and education, models relied less on ethnicity. When given irrelevant details, such as hair color and tattoo shape, they tended to fall back to ethnic sorting.

The researchers did not show that commercial hiring systems behave the same way in the wild. Real résumé screeners do not usually receive instant feedback on whether a hire worked out. Still, Wang said the results should concern companies using LLMs to screen résumés or conduct interviews, especially if those systems later learn from hiring outcomes.

This story draws on original reporting from MIT Technology Review.

More AI/

view all ↗