AI is more likely than humans to form biases when hiring

· Source: Artificial intelligence – MIT Technology Review · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Human Resources & Workforce Development, Research Methodology & Innovation · Depth: Intermediate, short

Summary

New research from Princeton University and the University of Chicago demonstrates that large language models (LLMs) can develop their own biases from experience, stereotyping job applicants more significantly than humans. In a simulated hiring game, LLMs like ChatGPT, Claude, and Gemini, hired for 20 jobs from four fictional ethnic groups. Despite all candidates being equally qualified, models quickly segregated groups into specific job niches based on early, limited feedback. For example, if an Aima candidate failed as a doctor, the model would then disproportionately hire Aimas as janitors. On a segregation scale, LLMs scored 1.83 (OpenAI's o3 model), approximately 65% higher than human participants' 0.84. This tendency arises from LLMs' optimization for rapid generalization. While instructing models to be "fair" was ineffective, offering a bonus for diverse hiring or providing relevant personal information about candidates substantially reduced bias. This has serious implications as companies deploy LLMs for resume screening.

Key takeaway

For Directors of AI/ML or HR professionals deploying LLMs for hiring, be aware that these models can develop novel biases from experience, stereotyping applicants more than humans. Your systems, optimized for generalization, may quickly segregate candidates based on limited feedback. To mitigate this, you must design explicit diversity bonuses into your LLM objectives and ensure the models receive relevant, individualized candidate information, not just group affiliations. This proactive approach is crucial to prevent unintended discriminatory outcomes.

Key insights

LLMs develop novel biases from experience, stereotyping more than humans due to their generalization optimization.

Principles

Method

Researchers simulated a hiring game where LLMs (ChatGPT, Claude, Gemini) hired for 20 jobs over 40 rounds, learning success outcomes. Candidates from four fictional ethnic groups were equally qualified.

In practice

Topics

Best for: CTO, VP of Engineering/Data, Executive, AI Scientist, AI Ethicist, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial intelligence – MIT Technology Review.