AI Self-Preferencing in Hiring: How Shared LLM Stacks Are Skewing Your Shortlists
By Chris Weinmann, Founder, OVI
A candidate who uses GPT-4o to polish their resume is 23–60% more likely to make the shortlist — but only if the company's AI screener also runs on GPT-4o. Use a different model, or write the resume by hand, and the advantage evaporates. That is the central finding of a 2026 peer-reviewed study published at two of AI ethics' most rigorous conferences, and it defines a category of hiring bias that no existing compliance framework is built to catch.
The phenomenon is called AI self-preferencing: the tendency of a large language model to systematically favour content generated by itself over content produced by other models or by humans. In a hiring context, this means an AI screening tool does not just evaluate what a resume says — it responds to how it says it, rewarding the stylistic and structural signatures of its own outputs. Researchers Jiannan Xu of the University of Maryland's Smith School of Business, Gujie Li of the National University of Singapore, and Jane Yi Jiang of Ohio State University documented the effect at scale in their paper "AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights," presented at both the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026) and the ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO 2026) (Source 1; Source 2).
This Is Not Demographic Bias — It Is a Market Distortion
Most AI hiring bias conversations centre on protected characteristics: race, gender, disability, age. Those are real risks, and existing frameworks — from EEOC adverse-impact measurement to New York City's Local Law 144 — are designed to surface them.
AI self-preferencing is a fundamentally different problem. It has nothing to do with who the candidate is. It is determined entirely by which AI tool the candidate happened to use, or whether they used one at all. A hand-written resume from a highly qualified candidate can be ranked below an AI-polished resume from a less qualified peer, simply because the screener recognises its own stylistic fingerprint (Source 4).
That distinction matters because it means current audit frameworks — demographic parity checks, four-fifths rule calculations, AEDT bias audits — will not detect the problem. The bias does not cluster along any protected class. It clusters along LLM vendor choice, which no compliance standard currently measures.
The Numbers: 68–88% Bias Rate Across Major LLMs
The scale of the effect is striking. Across the major commercial and open-source LLMs tested, the self-preference bias rate ranged from 68% to 88%. GPT-4o exceeded 80%, meaning that in more than four out of five evaluation scenarios, it scored its own generated content higher than equivalent content produced by other models or by human writers (Source 1; Source 2).
The practical impact on shortlisting is equally clear. Candidates whose resumes were generated or optimised by the same LLM as the evaluating screener were 23–60% more likely to be shortlisted than equally qualified candidates using human-written resumes (Source 3; Source 4).
The bias is structural and model-agnostic. Every major commercial LLM tested — including GPT-4o, Claude, and Llama — exhibited the effect to varying degrees. No single vendor is immune. The problem is architectural: LLMs are trained on patterns, and they recognise the patterns they produce (Source 2).
Independent replication by Roasted.cv confirmed the core findings, reinforcing that this is a reproducible phenomenon rather than an artefact of a single experimental setup (Source 5).
Which Roles Are Most Affected?
The bias does not hit all job categories equally. Business-adjacent roles — sales, accounting, and finance — showed the largest disparities between AI-matched and AI-mismatched resume evaluations. These are precisely the roles where resume optimisation tools are most heavily adopted and where AI screening is increasingly standard, creating a compounding loop: the more candidates use AI to write, and the more employers use AI to screen, the larger the self-preferencing advantage becomes (Source 1; Source 3).
Why It Happens: The Mechanics of Self-Recognition
Understanding the mechanism is essential for mitigation. LLMs do not have a hidden "prefer my own work" instruction. The effect emerges from how these models are trained. When an LLM generates text, it produces content that aligns with the statistical patterns it has internalised — specific sentence structures, vocabulary distributions, transition phrases, paragraph architectures. When the same model is then asked to evaluate text, it assigns higher quality scores to content that matches those same patterns, because those patterns are, by definition, what the model considers "good writing" (Source 2; Source 3).
In practical terms: a GPT-4o-powered screener is not reading for substance differently than it reads for style. It weights both — and the style dimension systematically advantages GPT-4o-generated content. The candidate's qualifications may be identical, but the presentation triggers different quality signals depending on which LLM shaped it.
What HR Teams Can Do Now
The researchers found that targeted interventions work. Prompt-level techniques that suppress self-recognition markers reduced bias by 17–63%, depending on the model and evaluation context (Source 1; Source 2).
Beyond prompt engineering, three practical strategies are available to HR teams today:
1. Ensemble / Multi-Evaluator Stacks. Using multiple LLMs in the screening pipeline — rather than relying on a single model — significantly suppresses self-preferencing. When the evaluating model rotates or when evaluations are aggregated across models, no single model's stylistic preference dominates the shortlist. The paper's findings directly support this: diversifying the evaluation stack breaks the one-to-one match between candidate tool and screener tool that drives the bias (Source 2; Source 3).
2. Human Override Gates. Maintaining human review at critical decision points — particularly the final shortlist stage — provides a check on AI-generated rankings. A human reviewer evaluates substance without the stylistic pattern-matching that drives self-preferencing. Tools that operate on a decision-support model, where AI ranks and recommends but a recruiter makes the final call, are structurally less exposed to this bias than fully automated screening pipelines (Source 4).
3. ISO 42001:2023 Governance Frameworks. The paper flags ISO 42001:2023 — the international standard for AI management systems — as a recommended governance framework. Implementing its audit and oversight requirements provides a structured approach to identifying, measuring, and mitigating emergent AI biases, including self-preferencing, that fall outside traditional demographic-bias testing (Source 1).
The Bigger Picture: Labour-Market Implications
Left unchecked, AI self-preferencing creates two systemic risks at the labour-market level.
First, LLM vendor lock-in in hiring. If a company's AI screening stack runs on a single model, and candidates learn which model it is, rational candidates will optimise their resumes with that specific model. This creates a self-reinforcing loop where the dominant screening LLM becomes the dominant resume-writing LLM, concentrating market power in ways that have nothing to do with hiring quality.
Second, an access-to-opportunity divide. Candidates who know how to use AI resume tools — and who can identify which model a company's screener uses — gain a measurable advantage over candidates who write their resumes without AI assistance. This is not a skills gap in the traditional sense. It is an information asymmetry that rewards AI-savviness over job-relevant qualifications, widening the gap between AI-literate and AI-naive job seekers (Source 2; Source 3).
For HR leaders, the message is clear: bias audits that only look at demographic outcomes are no longer sufficient. Self-preferencing is a structural property of every single-model AI screening stack, and addressing it requires architectural decisions — not just compliance checkboxes.
What is AI self-preferencing in hiring?
AI self-preferencing is the tendency of a large language model to score content it generated higher than equivalent content from other models or human writers. In hiring, this means an AI resume screener systematically favours candidates whose resumes were written or optimised by the same LLM that powers the screener, regardless of the candidates' actual qualifications. A 2026 peer-reviewed study found this bias rate ranges from 68% to 88% across major commercial LLMs.
Does this affect all AI screening tools?
Yes. The research found that every major commercial and open-source LLM tested — including GPT-4o, Claude, and Llama — exhibited self-preferencing to varying degrees. The bias is structural and model-agnostic: it emerges from how LLMs are trained, not from any specific vendor's design choices. Tools that use a single LLM for evaluation are most exposed. Multi-model or ensemble evaluation stacks significantly reduce the effect.
What can HR teams do about self-preferencing right now?
Three immediate steps: (1) Move to multi-evaluator stacks that use more than one LLM, breaking the one-to-one match between candidate tool and screener. (2) Maintain human override gates at the shortlist stage so that final decisions are not fully automated. (3) Adopt ISO 42001:2023 governance frameworks to build structured audit processes that can detect emerging AI biases beyond demographic categories. The study's own findings show that targeted mitigation techniques reduce self-preferencing bias by 17–63%.