Forget Bias for a Moment — Do LLM Screeners Even Rank Correctly?
Measuring Validity in LLM-based Resume Screening
Jane Castleman, Zeyu Shen, Blossom Metevier, Max Springer, Aleksandra Korolova · arXiv:2602.18550 (PDF) · submitted February 20, 2026
Summary by Andrew Johnson, OneTwo Resume editorial team · updated · We are not affiliated with the authors. arXiv papers are preprints and may not be peer-reviewed.
Almost all scrutiny of LLM resume screening measures bias — differences in how demographic groups are treated. This study targets something more basic: validity. When an off-the-shelf LLM ranks candidates, is the ranking connected to actual qualification, in a way that would survive ground truth?
Measuring that is hard precisely because researchers lack large resume corpora with known correct rankings that models haven’t already trained on. The authors’ contribution is a systematic construction that overcomes this — enabling validity measurement, not just fairness measurement, for general-purpose LLMs that many organizations deploy without task-specific adaptation.
What the paper reports
- Many organizations use general-purpose LLMs for screening without adapting them to the task.
- The study contributes a method for measuring ranking validity despite the ground-truth corpus problem.
- It reframes the evaluation question from "is it fair?" to "is it measuring anything real?"
What this means for your resume
Our editorial interpretation — the paper does not give job-seeker advice.
- A rejection from an automated screen is weak evidence about your qualifications — the tool itself may not be validly ranking anyone. Treat volume rejection as a systems problem: keep applying, vary your materials, and route around the machine with referrals where possible.
Read it with these caveats
arXiv preprint (Feb 2026, v1) — recent enough that peer review and replication are still ahead of it. We summarize the framing and method; specific validity numbers should be read in the paper.
Primary source: Measuring Validity in LLM-based Resume Screening — always read the paper before citing it. Spotted an error in our summary? Tell us and we'll fix it with a visible correction.
More in how ai screeners work — and fail
General-Purpose LLMs vs a Purpose-Built Hiring Model, on 10,000 Real Pairs
OpenAI, Anthropic, Google, Meta, and DeepSeek models benchmarked on ~10,000 real candidate-job pairs — against a domain-specific model that beat them all.
A Name for the "200 Applications, Zero Replies" Problem
High vacancies AND prolonged unemployment, together — this framework blames deterministic screening rejecting qualified candidates via semantic misinterpretation.
Measuring How Many Qualified Candidates Keyword Filters Wrongly Reject
Friction defined as excess false-negative rejection: keyword screening shows high friction in controlled simulation; semantic matching much less.