General-Purpose LLMs vs a Purpose-Built Hiring Model, on 10,000 Real Pairs

Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions

Eitan Anzenberg, Arunava Samajpati, Sivasankaran Chandrasekar, Varun Kacholia · arXiv:2507.02087 (PDF) · submitted July 2, 2025

Summary by Andrew Johnson, OneTwo Resume editorial team · updated · We are not affiliated with the authors. arXiv papers are preprints and may not be peer-reviewed.

An industry study (from the team behind a commercial matching model, which matters — see limitations) benchmarking frontier LLMs on real candidate-job matching: roughly 10,000 recent real-world candidate-job pairs, scored for predictive accuracy and for fairness via impact-ratio analysis across gender, race, and intersections.

The reported result: the purpose-built domain model outperformed the general-purpose LLMs on accuracy (ROC AUC 0.85 vs 0.77) while also achieving more equitable outcomes across demographic subgroups. The interesting part for outsiders isn’t who won — it’s the documented gap between what generic chatbots and tuned systems do on the same hiring data.

What the paper reports

What this means for your resume

Our editorial interpretation — the paper does not give job-seeker advice.

Read it with these caveats

arXiv preprint authored by the vendor of the winning model (Eightfold’s Match Score) — a textbook conflict-of-interest case where the methodology matters more than the ranking. We include it because the LLM-vs-tuned-model gap is independently interesting; read the numbers with the authorship in mind.

Primary source: Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions — always read the paper before citing it. Spotted an error in our summary? Tell us and we'll fix it with a visible correction.

More in how ai screeners work — and fail