General-Purpose LLMs vs a Purpose-Built Hiring Model, on 10,000 Real Pairs
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
Eitan Anzenberg, Arunava Samajpati, Sivasankaran Chandrasekar, Varun Kacholia · arXiv:2507.02087 (PDF) · submitted July 2, 2025
Summary by Andrew Johnson, OneTwo Resume editorial team · updated · We are not affiliated with the authors. arXiv papers are preprints and may not be peer-reviewed.
An industry study (from the team behind a commercial matching model, which matters — see limitations) benchmarking frontier LLMs on real candidate-job matching: roughly 10,000 recent real-world candidate-job pairs, scored for predictive accuracy and for fairness via impact-ratio analysis across gender, race, and intersections.
The reported result: the purpose-built domain model outperformed the general-purpose LLMs on accuracy (ROC AUC 0.85 vs 0.77) while also achieving more equitable outcomes across demographic subgroups. The interesting part for outsiders isn’t who won — it’s the documented gap between what generic chatbots and tuned systems do on the same hiring data.
What the paper reports
- Frontier LLMs from five major providers were benchmarked on ~10,000 real candidate-job pairs.
- The domain-specific model reported ROC AUC 0.85 vs 0.77 for the best general-purpose models, with better subgroup impact ratios.
- Fairness was evaluated via impact-ratio cutoff analysis across declared gender, race, and intersectional subgroups.
What this means for your resume
Our editorial interpretation — the paper does not give job-seeker advice.
- The screening quality you face varies wildly by employer — a company using a tuned, audited system and one pasting resumes into a chatbot are different lotteries. This is more reason not to read any single automated rejection as a verdict on you.
Read it with these caveats
arXiv preprint authored by the vendor of the winning model (Eightfold’s Match Score) — a textbook conflict-of-interest case where the methodology matters more than the ranking. We include it because the LLM-vs-tuned-model gap is independently interesting; read the numbers with the authorship in mind.
Primary source: Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions — always read the paper before citing it. Spotted an error in our summary? Tell us and we'll fix it with a visible correction.
More in how ai screeners work — and fail
Forget Bias for a Moment — Do LLM Screeners Even Rank Correctly?
Bias audits ask if screeners are fair. This Princeton-line study asks the prior question: are their rankings valid at all?
A Name for the "200 Applications, Zero Replies" Problem
High vacancies AND prolonged unemployment, together — this framework blames deterministic screening rejecting qualified candidates via semantic misinterpretation.
Measuring How Many Qualified Candidates Keyword Filters Wrongly Reject
Friction defined as excess false-negative rejection: keyword screening shows high friction in controlled simulation; semantic matching much less.