Researchers publish rigorous studies about AI resume screening — bias audits, hidden-text attacks, validity tests — and almost none of it reaches the people being screened. This library translates that research into plain English, one paper per page, always linked to the primary source, with the practical implications kept strictly separate from what the papers actually claim.
Curated by Andrew Johnson · OneTwo Resume editorial team · Last updated · 16 papers
What happens when the resume is written by an AI, screened by an AI, or contains text designed to manipulate the screener.
Do AI Screeners Prefer AI-Written Resumes?
When LLMs sit on both sides of hiring — writing resumes and screening them — do they favor their own kind of text? This study tests it empirically.
Jiannan Xu et al. · arXiv:2509.00462 · 2025-08
Hidden Text in Resumes Can Hijack AI Screeners — With >80% Success
The academic verdict on the viral "white text" trick: adversarial instructions hidden in resumes hijacked LLM screeners with attack success rates above 80%.
Honglin Mu et al. · arXiv:2512.20164 · 2025-12
Audit studies measuring whether names, gender, race, national origin, or education change how AI systems rank identical resumes.
The Famous Name-Bias Experiment, Rerun on ChatGPT-Era Models
The landmark 2003 field experiment — identical resumes, racially suggestive names — replicated against GPT-3.5, Bard, and Claude.
Akshaj Kumar Veldanda et al. · arXiv:2310.05135 · 2023-10
Resume Search Engines Built on Embeddings Show Race and Gender Bias
Not the chatbots — the embedding models that power "find me candidates like this" search. Audited for gender, race, and intersectional bias.
Kyra Wilson et al. · arXiv:2407.20371 · 2024-07
Newer LLMs Show Less Name Bias — But Judge Your University Instead
A 2025 audit finds explicit gender/race bias has receded in recent LLMs — while implicit bias around educational background remains significant.
Hayate Iso et al. · arXiv:2503.19182 · 2025-03
Overcorrection Is Real: Benchmarking Reverse Gender Bias in LLM Resume Scoring
A benchmark grounded in labor economics finds significant reverse gender bias and over-debiasing in LLM resume scoring.
Ze Wang et al. · arXiv:2406.15484 · 2024-06
Deep-Learning Screeners Can Infer — and Penalize — National Origin
Before the LLM wave, deep-learning resume screeners were already encoding national-origin signals. This study examines how.
Sihang Li et al. · arXiv:2307.08624 · 2023-07
They Tried to Remove Gender from 709,000 Resumes. It Mostly Didn’t Work.
Using 709k real IT resumes: gendered information pervades resume text, and scrubbing it only works up to a point.
Prasanna Parasurama et al. · arXiv:2112.08910 · 2021-12
Whether automated screening actually ranks candidates validly, and how qualified people get filtered out by mechanics rather than merit.
Forget Bias for a Moment — Do LLM Screeners Even Rank Correctly?
Bias audits ask if screeners are fair. This Princeton-line study asks the prior question: are their rankings valid at all?
Jane Castleman et al. · arXiv:2602.18550 · 2026-02
General-Purpose LLMs vs a Purpose-Built Hiring Model, on 10,000 Real Pairs
OpenAI, Anthropic, Google, Meta, and DeepSeek models benchmarked on ~10,000 real candidate-job pairs — against a domain-specific model that beat them all.
Eitan Anzenberg et al. · arXiv:2507.02087 · 2025-07
A Name for the "200 Applications, Zero Replies" Problem
High vacancies AND prolonged unemployment, together — this framework blames deterministic screening rejecting qualified candidates via semantic misinterpretation.
Ibrahim Denis Fofanah · arXiv:2601.14534 · 2026-01
Measuring How Many Qualified Candidates Keyword Filters Wrongly Reject
Friction defined as excess false-negative rejection: keyword screening shows high friction in controlled simulation; semantic matching much less.
Ibrahim Denis Fofanah · arXiv:2602.04087 · 2026-02
Can AI Tell Real Seniority from Inflated Titles? Researchers Built Traps to Find Out
A benchmark seeded with deliberately exaggerated resumes tests whether LLMs detect overstated experience — the machine version of the sniff test.
Matan Cohen et al. · arXiv:2509.09229 · 2025-09
Inside an LLM-Agent Screening Pipeline That Reads Resumes 11× Faster Than Humans
A research prototype of exactly what vendors sell: LLM agents that summarize, grade, and decide on resumes at scale — 11× faster than manual screening.
Chengguang Gan et al. · arXiv:2401.08315 · 2024-01
How AI recommendations change what human recruiters do, and what job seekers actually want from automated hiring.
When the AI Is Biased, How Long Do Humans Look Before Agreeing?
The last line of defense against biased AI screening is a human reviewer. This study measures what actually happens to their attention.
Kyra Wilson et al. · arXiv:2606.22213 · 2026-06
What Job Seekers Actually Want From AI Hiring: Explanations
Built with and evaluated by 24 active job seekers: an AI recruitment system designed to explain decisions instead of hiding them.
Aditya Bhattacharya et al. · arXiv:2505.20312 · 2025-05
Experiments we run ourselves, with raw data published for re-annotation.
Do AI Cover Letter Generators Invent Facts? We Tested Ours
10 letters, 5 fictional candidates with closed fact inventories, every claim annotated — including our own product's failure modes. Raw data published.
Suggest a paper
Working on resume screening, algorithmic hiring, or candidate-side AI? We add papers that are on-topic and public — send us the arXiv link. We don't charge for inclusion and we don't accept payment for it.