AI Hiring Research

Researchers publish rigorous studies about AI resume screening — bias audits, hidden-text attacks, validity tests — and almost none of it reaches the people being screened. This library translates that research into plain English, one paper per page, always linked to the primary source, with the practical implications kept strictly separate from what the papers actually claim.

Curated by Andrew Johnson · OneTwo Resume editorial team · Last updated · 16 papers

How to read this library: most arXiv papers are preprints — early research that may not have completed peer review. Every summary here states what the paper reports, flags its limitations, and labels our job-seeker interpretation as editorial. When a summary and the paper disagree, the paper wins; corrections welcome.

Gaming, security, and AI-vs-AI

What happens when the resume is written by an AI, screened by an AI, or contains text designed to manipulate the screener.

Bias & fairness in AI screening

Audit studies measuring whether names, gender, race, national origin, or education change how AI systems rank identical resumes.

How AI screeners work — and fail

Whether automated screening actually ranks candidates validly, and how qualified people get filtered out by mechanics rather than merit.

Forget Bias for a Moment — Do LLM Screeners Even Rank Correctly?

Bias audits ask if screeners are fair. This Princeton-line study asks the prior question: are their rankings valid at all?

Jane Castleman et al. · arXiv:2602.18550 · 2026-02

General-Purpose LLMs vs a Purpose-Built Hiring Model, on 10,000 Real Pairs

OpenAI, Anthropic, Google, Meta, and DeepSeek models benchmarked on ~10,000 real candidate-job pairs — against a domain-specific model that beat them all.

Eitan Anzenberg et al. · arXiv:2507.02087 · 2025-07

A Name for the "200 Applications, Zero Replies" Problem

High vacancies AND prolonged unemployment, together — this framework blames deterministic screening rejecting qualified candidates via semantic misinterpretation.

Ibrahim Denis Fofanah · arXiv:2601.14534 · 2026-01

Measuring How Many Qualified Candidates Keyword Filters Wrongly Reject

Friction defined as excess false-negative rejection: keyword screening shows high friction in controlled simulation; semantic matching much less.

Ibrahim Denis Fofanah · arXiv:2602.04087 · 2026-02

Can AI Tell Real Seniority from Inflated Titles? Researchers Built Traps to Find Out

A benchmark seeded with deliberately exaggerated resumes tests whether LLMs detect overstated experience — the machine version of the sniff test.

Matan Cohen et al. · arXiv:2509.09229 · 2025-09

Inside an LLM-Agent Screening Pipeline That Reads Resumes 11× Faster Than Humans

A research prototype of exactly what vendors sell: LLM agents that summarize, grade, and decide on resumes at scale — 11× faster than manual screening.

Chengguang Gan et al. · arXiv:2401.08315 · 2024-01

Humans + AI deciding together

How AI recommendations change what human recruiters do, and what job seekers actually want from automated hiring.

Our own research

Experiments we run ourselves, with raw data published for re-annotation.

Do AI Cover Letter Generators Invent Facts? We Tested Ours

10 letters, 5 fictional candidates with closed fact inventories, every claim annotated — including our own product's failure modes. Raw data published.

Suggest a paper

Working on resume screening, algorithmic hiring, or candidate-side AI? We add papers that are on-topic and public — send us the arXiv link. We don't charge for inclusion and we don't accept payment for it.