AI can become more biased than humans while choosing who gets hired; study finds |
Artificial intelligence is increasingly being trusted to help employers screen CVs, rank candidates and even conduct early-stage interviews. The promise is simple: AI can process thousands of applications quickly, reduce administrative workloads and, in theory, make hiring decisions more objective than humans. According to a recent review by MIT Technology, based on the 2024 study published in the Journal of Experimental Psychology titled “Costly Exploration Produces Stereotypes With Dimensions of Warmth and Competence”, large language models (LLMs) may not only inherit human biases from the data they are trained on but also develop entirely new stereotypes based on their own experiences. Rather than simply reflecting existing prejudice, these systems can create fresh patterns of discrimination as they attempt to optimise decision-making over time.The findings are particularly significant because businesses are rapidly integrating AI into recruitment. If these systems begin making assumptions about groups of applicants after only a handful of hiring decisions, they could gradually reinforce unfair employment practices without any explicit human instruction.The research, conducted by scientists at Princeton University and the University of Chicago, builds on earlier psychological work showing how stereotypes can emerge from repeated decision-making. It suggests that the same learning strategies that make AI effective at solving complex problems can also make it unusually prone to stereotyping job candidates.
How AI hiring models developed stereotypes during the recruitment experiment
The researchers designed a hiring simulation based on an established psychology experiment investigating how stereotypes form. Instead of human volunteers alone, they tested several leading language models, including ChatGPT, Claude and Gemini, to see how they behaved when repeatedly making hiring decisions. Each model was placed in the role of a consultant hired by the mayor of a fictional city. Its task was to recruit people for 20 different occupations, including doctors, lawyers, child-care aides and janitors.As per the MIT review, applicants belonged to four fictional ethnic groups named Tufa, Aima, Reku and Weki. Importantly, every candidate had the same underlying chance of succeeding in every job. The groups differed only in name, ensuring that any stereotypes developed during the experiment could not be explained by genuine differences in ability. During each hiring round, the model selected one applicant and was then told whether that person succeeded or failed. It repeated this process across 40 rounds while trying to maximise successful hires.
Why AI hiring systems formed stereotypes from random hiring outcomes
The crucial point was that early successes and failures occurred purely by chance. Nevertheless, the language models rapidly began assigning particular groups to particular occupations.For example, if an applicant from the fictional Aima group happened to fail as a doctor early in the experiment, the models became less willing to hire other Aima candidates for medical positions. Instead, they increasingly recommended them for lower-status occupations such as janitors, despite there being no evidence that Aima applicants were less capable.This pattern reflected a tendency to generalise from extremely limited information. Rather than continuing to explore whether other members of the same group might perform differently, the models quickly settled on assumptions about where each group “belonged.”The behaviour mirrors a concept in psychology known as the exploration-exploitation dilemma. Decision-makers constantly face a choice between trying unfamiliar options to gather more information or relying on previous experiences that appear successful. In hiring, exploring means giving different candidates opportunities, while exploiting means repeatedly choosing applicants who resemble those who succeeded before.The underlying psychological theory has been explored in previous human research. As per the study, in general, stereotypes can emerge even when decision-makers have no prejudice and when all groups possess identical abilities, simply because repeatedly relying on experience reduces the perceived cost of exploring unfamiliar options. The latest AI study suggests language models may amplify this tendency even further.
Why AI hiring models developed stronger stereotypes than humans
Perhaps the most striking finding was not simply that AI formed stereotypes, but that it did so more readily than human participants. The original psychology experiment measured the extent to which people segregated different groups into different occupations. Human participants recorded an average segregation score of 0.84.The language models performed substantially worse. According to the findings highlighted by MIT Technology Review, the models produced segregation scores roughly 65% higher than those observed in humans. OpenAI’s reasoning model o3 recorded a score of 1.83, approaching the maximum possible value on the researchers’ scale.Ryan Liu, a PhD student at Princeton University and one of the study’s co-authors, explained that the behaviour reflects what language models are fundamentally designed to do. “They really are eager to create generalisations from limited data. That’s literally a lot of what they’re optimised for.”That tendency is highly valuable when solving mathematical problems, writing computer code or identifying patterns across large datasets. AI systems are rewarded during training for recognising useful relationships from relatively few examples. However, social decisions are very different from mathematical reasoning.When evaluating people, concluding too quickly can lead to unfair assumptions that become increasingly difficult to reverse. A single negative outcome can disproportionately shape future recommendations, causing entire groups to be associated with particular occupations without any objective basis.The researchers found that newer reasoning models showed even stronger stereotyping than earlier systems. Greater reasoning ability therefore did not automatically translate into fairer decision-making.According to the study, simply instructing AI to “be fair” had little effect. Although the models understood fairness in principle, they continued prioritising strategies that appeared to maximise successful hiring. This finding echoes broader concerns in AI safety: language models may understand ethical principles yet still pursue optimisation strategies that unintentionally conflict with those values when carrying out complex tasks.The earlier psychology research helps explain why this occurs. It argues that stereotypes need not arise because decision-makers are irrational or intentionally biased. Instead, they can emerge naturally whenever individuals repeatedly rely on previous experiences while trying to maximise rewards under uncertainty.In the context of AI hiring systems, this means bias may not simply be inherited from historical datasets; it can emerge through the model’s own decision-making process as it accumulates experience.
Why AI hiring systems can create new biases instead of just inheriting them
One of the study’s most important conclusions was that the bias did not originate from the language models’ training data alone. Instead, the models developed new stereotypes through their own experiences during the hiring task. According to the researchers, this happens because large language models are designed to identify patterns quickly and often generalise from limited information, even when early results are based purely on chance.The study describes this behaviour through the exploration-exploitation dilemma, in which decision-makers favour options that appeared successful before rather than exploring unfamiliar ones. While this approach is valuable for tasks such as mathematics, coding and logical reasoning, it can produce unfair assumptions when applied to people. As per MIT, including OpenAI’s o3 and DeepSeek’s R1, showed stronger stereotyping than earlier systems. Simply instructing the models to “be fair” had little measurable effect. However, according to the study, rewarding more diverse hiring decisions significantly reduced bias, suggesting that fairness depends more on how AI systems are designed and optimised than on ethical instructions alone.
What the study means for the future of AI hiring and recruitment
Many organisations already use AI to filter CVs, rank applicants and support recruitment. Although real-world hiring provides much slower feedback than the experiment, researchers say the risk of AI developing new biases remains, particularly as models gain memory and personalisation features. Cornell computer scientist Angelina Wang warned that AI systems may “over-index” on previous experiences, making them more likely to develop stereotypes over time.According to the study, that providing relevant information about candidates, such as education or age, reduced stereotyping, whereas irrelevant details had little effect. Overall, the findings suggest AI can generate new biases through its own learning process, meaning fairness depends not only on training data but also on how AI systems are designed and deployed in hiring and other high-stakes decisions.