AI is more likely than humans to form biases when hiring
Summary
New research from Princeton University and the University of Chicago demonstrates that large language models (LLMs) can develop their own biases from experience, stereotyping job applicants more significantly than humans. In a simulated hiring game, LLMs like ChatGPT, Claude, and Gemini, hired for 20 jobs from four fictional ethnic groups. Despite all candidates being equally qualified, models quickly segregated groups into specific job niches based on early, limited feedback. For example, if an Aima candidate failed as a doctor, the model would then disproportionately hire Aimas as janitors. On a segregation scale, LLMs scored 1.83 (OpenAI's o3 model), approximately 65% higher than human participants' 0.84. This tendency arises from LLMs' optimization for rapid generalization. While instructing models to be "fair" was ineffective, offering a bonus for diverse hiring or providing relevant personal information about candidates substantially reduced bias. This has serious implications as companies deploy LLMs for resume screening.
Key takeaway
For Directors of AI/ML or HR professionals deploying LLMs for hiring, be aware that these models can develop novel biases from experience, stereotyping applicants more than humans. Your systems, optimized for generalization, may quickly segregate candidates based on limited feedback. To mitigate this, you must design explicit diversity bonuses into your LLM objectives and ensure the models receive relevant, individualized candidate information, not just group affiliations. This proactive approach is crucial to prevent unintended discriminatory outcomes.
Key insights
LLMs develop novel biases from experience, stereotyping more than humans due to their generalization optimization.
Principles
- LLMs prioritize generalization from limited data.
- Goal design influences LLM social behavior.
- Relevant personal data reduces group-based bias.
Method
Researchers simulated a hiring game where LLMs (ChatGPT, Claude, Gemini) hired for 20 jobs over 40 rounds, learning success outcomes. Candidates from four fictional ethnic groups were equally qualified.
In practice
- Incorporate diversity bonuses into LLM objectives.
- Provide specific, relevant candidate data to LLMs.
- Monitor LLM hiring decisions for emergent segregation.
Topics
- LLM Bias
- AI Hiring
- Algorithmic Fairness
- Stereotyping
- Machine Learning Ethics
- Exploration-Exploitation
Best for: CTO, VP of Engineering/Data, Executive, AI Scientist, AI Ethicist, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial intelligence – MIT Technology Review.