Forget Bias for a Moment — Do LLM Screeners Even Rank Correctly?

Measuring Validity in LLM-based Resume Screening

Jane Castleman, Zeyu Shen, Blossom Metevier, Max Springer, Aleksandra Korolova · arXiv:2602.18550 (PDF) · submitted February 20, 2026

Summary by Andrew Johnson, OneTwo Resume editorial team · updated · We are not affiliated with the authors. arXiv papers are preprints and may not be peer-reviewed.

Almost all scrutiny of LLM resume screening measures bias — differences in how demographic groups are treated. This study targets something more basic: validity. When an off-the-shelf LLM ranks candidates, is the ranking connected to actual qualification, in a way that would survive ground truth?

Measuring that is hard precisely because researchers lack large resume corpora with known correct rankings that models haven’t already trained on. The authors’ contribution is a systematic construction that overcomes this — enabling validity measurement, not just fairness measurement, for general-purpose LLMs that many organizations deploy without task-specific adaptation.

What the paper reports

What this means for your resume

Our editorial interpretation — the paper does not give job-seeker advice.

Read it with these caveats

arXiv preprint (Feb 2026, v1) — recent enough that peer review and replication are still ahead of it. We summarize the framing and method; specific validity numbers should be read in the paper.

Primary source: Measuring Validity in LLM-based Resume Screening — always read the paper before citing it. Spotted an error in our summary? Tell us and we'll fix it with a visible correction.

More in how ai screeners work — and fail