Overcorrection Is Real: Benchmarking Reverse Gender Bias in LLM Resume Scoring

JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models

Ze Wang, Zekun Wu, Xin Guan, Michael Thaler, Adriano Koshiyama, Skylar Lu · arXiv:2406.15484 (PDF) · submitted June 17, 2024

Summary by Andrew Johnson, OneTwo Resume editorial team · updated · We are not affiliated with the authors. arXiv papers are preprints and may not be peer-reviewed.

Most bias audits ask one question: are protected groups scored lower? JobFair builds a more careful measurement framework, grounded in labor economics and legal principles, distinguishing level bias (average score differences between demographic counterfactuals) from spread bias (variance differences), and statistical from taste-based bias.

Applying the framework to LLM resume scoring, the authors report significant issues of reverse gender bias and over-debiasing — models that have been tuned so hard against historical bias that they now systematically favor the historically disadvantaged group, which is itself a fairness and legal problem.

What the paper reports

What this means for your resume

Our editorial interpretation — the paper does not give job-seeker advice.

Read it with these caveats

arXiv preprint. Findings are model- and prompt-specific; the paper’s value is as much the measurement framework as the specific results. "Reverse bias" findings in benchmarks do not tell you any particular employer’s configuration.

Primary source: JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models — always read the paper before citing it. Spotted an error in our summary? Tell us and we'll fix it with a visible correction.

More in bias & fairness in ai screening