The Famous Name-Bias Experiment, Rerun on ChatGPT-Era Models
Are Emily and Greg Still More Employable than Lakisha and Jamal? Investigating Algorithmic Hiring Bias in the Era of ChatGPT
Akshaj Kumar Veldanda, Fabian Grob, Shailja Thakur, Hammond Pearce, Benjamin Tan, Ramesh Karri, Siddharth Garg · arXiv:2310.05135 (PDF) · submitted October 8, 2023
Summary by Andrew Johnson, OneTwo Resume editorial team · updated · We are not affiliated with the authors. arXiv papers are preprints and may not be peer-reviewed.
In 2003, Bertrand and Mullainathan mailed identical resumes that differed only in the name at the top — Emily and Greg versus Lakisha and Jamal — and measured callback rates. It became the gold-standard method for detecting hiring discrimination. This paper reruns that exact logic on the LLMs that now read resumes: GPT-3.5, Bard, and Claude.
The models are tested on matching resumes to job categories while the researchers vary protected attributes: racially suggestive names, gender, and maternity status. The question is whether the model’s outputs shift when nothing about the qualifications changes.
What the paper reports
- The study adapts the classic identical-resume audit methodology to LLM-based resume-to-job-category matching.
- Protected attributes tested include race-associated names, gender, and maternity/paternity signals.
- It provides one of the earliest systematic bias audits of consumer-grade LLMs (GPT-3.5, Bard, Claude) in a hiring task.
What this means for your resume
Our editorial interpretation — the paper does not give job-seeker advice.
- You cannot control which screener reads your application — but you control what identity signals your resume volunteers. City-level location, no photo, no birth date, and no affiliations that only signal demographics are standard advice this literature reinforces.
- If an employer’s AI screening harmed you and you can document it, US EEOC guidance treats algorithmic discrimination like any other discrimination — the audit methodology in papers like this is how such cases get measured.
Read it with these caveats
arXiv preprint from late 2023 — the specific models tested (GPT-3.5, Bard) are now several generations old, and newer models show different (not necessarily absent) bias profiles; see the 2025 job-matching audit in this library for a more recent read.
Primary source: Are Emily and Greg Still More Employable than Lakisha and Jamal? Investigating Algorithmic Hiring Bias in the Era of ChatGPT — always read the paper before citing it. Spotted an error in our summary? Tell us and we'll fix it with a visible correction.
More in bias & fairness in ai screening
Resume Search Engines Built on Embeddings Show Race and Gender Bias
Not the chatbots — the embedding models that power "find me candidates like this" search. Audited for gender, race, and intersectional bias.
Newer LLMs Show Less Name Bias — But Judge Your University Instead
A 2025 audit finds explicit gender/race bias has receded in recent LLMs — while implicit bias around educational background remains significant.
Overcorrection Is Real: Benchmarking Reverse Gender Bias in LLM Resume Scoring
A benchmark grounded in labor economics finds significant reverse gender bias and over-debiasing in LLM resume scoring.