Overcorrection Is Real: Benchmarking Reverse Gender Bias in LLM Resume Scoring
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
Ze Wang, Zekun Wu, Xin Guan, Michael Thaler, Adriano Koshiyama, Skylar Lu · arXiv:2406.15484 (PDF) · submitted June 17, 2024
Summary by Andrew Johnson, OneTwo Resume editorial team · updated · We are not affiliated with the authors. arXiv papers are preprints and may not be peer-reviewed.
Most bias audits ask one question: are protected groups scored lower? JobFair builds a more careful measurement framework, grounded in labor economics and legal principles, distinguishing level bias (average score differences between demographic counterfactuals) from spread bias (variance differences), and statistical from taste-based bias.
Applying the framework to LLM resume scoring, the authors report significant issues of reverse gender bias and over-debiasing — models that have been tuned so hard against historical bias that they now systematically favor the historically disadvantaged group, which is itself a fairness and legal problem.
What the paper reports
- Introduces a two-axis construct: level bias vs spread bias, with level bias further split into statistical and taste-based components.
- Reports significant reverse gender hiring bias and over-debiasing in the tested LLMs’ resume scoring.
- Grounds bias categories in labor economics and legal principle rather than ad-hoc metrics.
What this means for your resume
Our editorial interpretation — the paper does not give job-seeker advice.
- The practical takeaway is about unpredictability: different vendors tune models in different directions, so the same resume can be advantaged or penalized for the same attribute depending on whose model reads it. Optimizing your resume for any assumed demographic tilt is a coin flip — optimizing its substance is not.
Read it with these caveats
arXiv preprint. Findings are model- and prompt-specific; the paper’s value is as much the measurement framework as the specific results. "Reverse bias" findings in benchmarks do not tell you any particular employer’s configuration.
Primary source: JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models — always read the paper before citing it. Spotted an error in our summary? Tell us and we'll fix it with a visible correction.
More in bias & fairness in ai screening
The Famous Name-Bias Experiment, Rerun on ChatGPT-Era Models
The landmark 2003 field experiment — identical resumes, racially suggestive names — replicated against GPT-3.5, Bard, and Claude.
Resume Search Engines Built on Embeddings Show Race and Gender Bias
Not the chatbots — the embedding models that power "find me candidates like this" search. Audited for gender, race, and intersectional bias.
Newer LLMs Show Less Name Bias — But Judge Your University Instead
A 2025 audit finds explicit gender/race bias has receded in recent LLMs — while implicit bias around educational background remains significant.