LLMs showed stronger hiring bias than humans

· Source: Dataconomy · Field: Technology & Digital — Artificial Intelligence & Machine Learning · Depth: Intermediate, quick

Summary

Researchers at Princeton University and the University of Chicago found that large language models (LLMs), including tools like ChatGPT, exhibit stronger hiring biases than human participants, developing stereotypes not only from training data but also from simulated experiences. In a simulated hiring game across 20 job roles and four fictional ethnic groups (Tufa, Aima, Reku, and Weki), models tasked with maximizing successful hires over 40 rounds frequently segregated candidates based on demographic information after interpreting early outcomes. LLMs averaged a segregation score of approximately 1.83, significantly higher than humans' 0.84. This tendency to generalize quickly from limited data poses significant implications as AI chatbots gain advanced memory. While instructing models to prioritize fairness had limited impact, incentivizing diverse hiring and providing more relevant candidate information (like age and education) reduced biases. The study highlights risks of LLMs developing novel biases in real-world applications like résumé screening without immediate feedback.

Key takeaway

For Directors of AI/ML deploying LLMs in hiring, you must actively design systems to counteract inherent bias. Your models can develop novel stereotypes from limited data, even without direct human input. Prioritize integrating explicit social goals and diversity incentives into AI objectives, rather than solely relying on fairness instructions. Additionally, ensure your models receive comprehensive, relevant candidate information to mitigate bias risks in critical evaluations like résumé screening.

Key insights

LLMs develop stronger hiring biases than humans, generalizing stereotypes from limited simulated experience.

Principles

Method

Researchers used a simulated hiring game with 20 job roles and four fictional ethnic groups, tasking LLMs to maximize successful hires over 40 rounds without knowing all candidates had equal success chances.

In practice

Topics

Best for: Research Scientist, CTO, VP of Engineering/Data, AI Scientist, AI Ethicist, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Dataconomy.