The Famous Name-Bias Experiment, Rerun on ChatGPT-Era Models

Are Emily and Greg Still More Employable than Lakisha and Jamal? Investigating Algorithmic Hiring Bias in the Era of ChatGPT

Akshaj Kumar Veldanda, Fabian Grob, Shailja Thakur, Hammond Pearce, Benjamin Tan, Ramesh Karri, Siddharth Garg · arXiv:2310.05135 (PDF) · submitted October 8, 2023

Summary by Andrew Johnson, OneTwo Resume editorial team · updated · We are not affiliated with the authors. arXiv papers are preprints and may not be peer-reviewed.

In 2003, Bertrand and Mullainathan mailed identical resumes that differed only in the name at the top — Emily and Greg versus Lakisha and Jamal — and measured callback rates. It became the gold-standard method for detecting hiring discrimination. This paper reruns that exact logic on the LLMs that now read resumes: GPT-3.5, Bard, and Claude.

The models are tested on matching resumes to job categories while the researchers vary protected attributes: racially suggestive names, gender, and maternity status. The question is whether the model’s outputs shift when nothing about the qualifications changes.

What the paper reports

What this means for your resume

Our editorial interpretation — the paper does not give job-seeker advice.

Read it with these caveats

arXiv preprint from late 2023 — the specific models tested (GPT-3.5, Bard) are now several generations old, and newer models show different (not necessarily absent) bias profiles; see the 2025 job-matching audit in this library for a more recent read.

Primary source: Are Emily and Greg Still More Employable than Lakisha and Jamal? Investigating Algorithmic Hiring Bias in the Era of ChatGPT — always read the paper before citing it. Spotted an error in our summary? Tell us and we'll fix it with a visible correction.

More in bias & fairness in ai screening