AI Interview Assessment Accuracy: What the Research Actually Shows in 2026
By Chris Weinmann, Founder, OVI
AI Interview Assessment Accuracy: What the Research Actually Shows in 2026
Current date (UTC): 2026-07-28
Current time (UTC): 06:30
HireVue processes more than 40 million interviews per year. Roughly 30 percent of large enterprises now use some form of AI-powered interview scoring, according to Gartner's 2025 Market Guide for AI-Augmented HR Services. Adoption is accelerating — but what does the independent research actually say about whether these tools predict job performance, treat candidates fairly, and hold up under regulatory scrutiny?
The answer is more complicated than most vendor pitch decks suggest. A growing body of peer-reviewed research, federal guidance, and candidate sentiment data reveals a significant gap between what AI interview assessment tools promise and what they deliver. Here is what HR leaders need to know.
The Predictive Validity Gap
The central promise of AI interview assessment is predictive validity — the tool's ability to forecast actual on-the-job performance. This is where the evidence gets uncomfortable for vendors.
HireVue's own 2022 Technical Validation Report cites a criterion validity coefficient of .35, positioning this figure as competitive with traditional interview methods. On the surface, .35 sounds reasonable.
The problem is context. Schmidt and Hunter's foundational 1998 meta-analysis in the Journal of Applied Psychology — the most widely cited validity benchmark in industrial-organizational psychology — established that structured behavioral interviews achieve validity coefficients of .48 to .51. Independent studies of AI interview scoring tools consistently cluster between .20 and .35, placing them in the range of unstructured interviews, not the structured methods they are marketed as replacing.
This distinction matters enormously. Many organisations adopt AI interview tools specifically to improve on human judgment and reduce inconsistency. If the independent evidence places these tools at unstructured-interview levels of accuracy, employers may be paying for technology that performs no better than a well-run phone screen — and potentially worse than a structured behavioral interview conducted by a trained interviewer.
The validity gap does not mean AI interview tools are useless. It means the accuracy claims need to be evaluated against independent benchmarks rather than vendor-supplied studies, and that structured approaches to AI assessment consistently outperform unstructured ones — just as they do in human interviewing.
The Bias Evidence Is Mounting
Accuracy is only half the equation. Fairness is equally critical, and the research raises flags that HR leaders cannot afford to ignore.
A 2023 University of Melbourne study examined six commercially available AI hiring tools and found that all six exhibited statistically significant bias against non-native English speakers. The bias was consistent across different tools and different candidate demographics, pointing to a systemic issue with how natural language processing models evaluate spoken and written responses rather than a problem limited to any single vendor.
Earlier research by Raghavan et al., published in the 2020 ACM Conference on Fairness, Accountability, and Transparency (FAccT), documented algorithmic bias in video-based hiring platforms more broadly. The study found that AI scoring systems can encode and amplify existing biases present in training data, producing adverse impact along racial and gender lines without transparent mechanisms for detection or correction.
Perhaps the most concerning finding is not about the tools themselves but about how employers use them. Only 22 percent of employers using AI hiring tools actively test those tools for bias, according to SHRM's 2024 AI in HR survey. Nearly four out of five organisations are deploying AI screening without ongoing monitoring for the adverse impact that federal law requires them to prevent.
Candidates Notice — and They Walk Away
Even a perfectly accurate, perfectly fair AI interview tool would still face a trust problem with the people it screens.
LinkedIn's 2024 Global Talent Trends report found that 56 percent of candidates hold a negative perception of AI-powered video interview screening. More consequentially, 38 percent of candidates said they would decline a job offer from a company that used AI screening without disclosing it.
For HR leaders, this is not an abstract perception problem — it is a pipeline and conversion issue. In a competitive talent market, every candidate who self-selects out of a process represents a measurable cost to sourcing effectiveness and time-to-fill. The question of whether and how to disclose AI use in hiring is not merely ethical. It is a talent acquisition strategy decision with direct impact on offer acceptance rates.
Regulators Are Closing In
The regulatory environment around AI hiring tools has tightened substantially, with three developments that HR leaders must track:
EEOC Technical Assistance (May 2023). The U.S. Equal Employment Opportunity Commission issued guidance clarifying that employers — not vendors — bear responsibility for monitoring AI hiring tools for adverse impact under Title VII and the ADA. The guidance explicitly states that if an employer uses an AI tool that screens out individuals with disabilities or creates adverse impact by race or gender, the employer is liable regardless of whether it designed the tool.
FTC Enforcement Signals (2024). The Federal Trade Commission flagged AI hiring tools as an area of active enforcement concern, focusing on deceptive accuracy claims and insufficient disclosure to candidates about how their data is being analysed.
NYC Local Law 144. New York City's automated employment decision tool (AEDT) law, effective since July 2023, requires annual independent bias audits for AI tools used in hiring or promotion decisions, with results publicly available. The law has set a regulatory precedent that other U.S. jurisdictions and international regulators are now following.
EU AI Act. The European Union's AI Act classifies employment-related AI systems as "high-risk," requiring conformity assessments, transparency obligations, and human oversight before deployment. Full enforcement timelines extend to August 2026, giving employers a narrowing compliance window.
The common thread across all four developments: regulators are placing the compliance burden on the employer, not the vendor. Buying a tool does not transfer legal responsibility.
What Good Looks Like: Five Questions to Ask Before You Buy
The research points to clear criteria that separate defensible AI interview assessment from risky deployment. Before adopting or renewing a tool, HR leaders should require answers to five questions:
"Can you share independent validity evidence — not just internal studies?" Vendor-commissioned research is a starting point, not proof. Ask for third-party validation from industrial-organizational psychologists with no financial relationship to the vendor. Compare any reported validity coefficients to the .48–.51 structured interview benchmark.
"What are your adverse impact ratios by race, gender, age, and language background?" A valid bias audit calculates selection rates by protected group and applies the four-fifths rule. Given the University of Melbourne findings on accent and language bias, this question is critical for any employer with a multilingual candidate pool.
"Does your tool analyse facial expressions, vocal tone, or other biometric signals?" The strongest bias findings in the research involve tools that score non-verbal cues rather than the substance of candidate answers. Transcript-content-only analysis avoids the most documented bias vectors.
"What is your candidate disclosure policy?" With 38 percent of candidates willing to walk away over undisclosed AI screening, transparency is a retention issue, not just a compliance checkbox.
"How does your architecture handle the employer-liability standard?" Under EEOC guidance, employers are legally responsible for their tools' outcomes. Tools that provide explainable, auditable decision-support — rather than opaque automated scoring — give employers a stronger position if outcomes are challenged.
A Structured Approach to AI Screening
The research consistently points in one direction: structured, content-focused AI assessment with human oversight produces the best validity and fairness outcomes. The documented problems — bias, weak validity, candidate distrust — concentrate around opaque scoring, biometric analysis, and automated decision-making with limited human review.
Among AI-native platforms taking a different approach, OVI's Milo screening agent evaluates candidates through structured AI audio chats, analysing transcript content only — no facial recognition, no emotion detection, no vocal-tone scoring. Final hiring decisions remain with the recruiter, with Milo providing decision-support through configurable rubrics with transparent weighting. This human-in-the-loop architecture reduces exposure under AEDT frameworks like NYC Local Law 144, since the tool does not fit the "automated employment decision" definition. OVI plans start at $29 per month, and the platform's practices align with SOC 2 Type II, ISO 27001, GDPR, and EU AI Act readiness standards. Full details are available at the OVI Trust & Compliance Center (ovi-me.com/standards).
Frequently Asked Questions
How accurate are AI interview assessment tools compared to human interviewers?
Independent research places AI interview scoring tools at criterion validity coefficients of .20 to .35, comparable to unstructured human interviews. This falls below the .48 to .51 range established for structured behavioral interviews by Schmidt and Hunter's 1998 meta-analysis. The accuracy depends heavily on design: structured, transcript-based approaches show stronger validity evidence than those relying on facial or vocal analysis.
Are AI interview tools biased?
Multiple studies document bias risks. A 2023 University of Melbourne study found statistically significant bias against non-native English speakers across all six AI hiring tools tested. Raghavan et al.'s 2020 FAccT research documented racial and gender bias in video-based scoring platforms. With only 22 percent of employers actively testing their AI hiring tools for bias (SHRM, 2024), the majority of deployments lack ongoing fairness monitoring.
What laws regulate AI interview tools in the United States?
The EEOC's May 2023 technical assistance clarifies that employers bear liability for adverse impact from AI hiring tools under Title VII and the ADA. NYC Local Law 144 requires annual independent bias audits for automated employment decision tools. The FTC has flagged AI hiring tools as an enforcement priority regarding accuracy claims and candidate disclosure.
What should a valid bias audit include?
A meaningful bias audit calculates selection rates by protected group — including race, gender, age, and language background — and applies the four-fifths rule. It should be conducted by an independent third party, repeated annually at minimum, and test the tool under real-world conditions using the employer's actual candidate population rather than synthetic datasets.
How can companies reduce legal risk when using AI screening?
Choose tools that operate as decision-support rather than automated decision-makers — this reduces exposure under AEDT laws. Ensure human reviewers make final hiring decisions. Require vendors to provide independent validity and bias audit data. Disclose AI use to candidates. Monitor outcomes for adverse impact on an ongoing basis. Document your compliance posture before regulators ask for it.
Sources: HireVue Technical Validation Report (2022); Schmidt & Hunter, "The Validity and Utility of Selection Methods in Personnel Psychology," Journal of Applied Psychology 83(2): 262–274 (1998); Raghavan et al., "Mitigating Bias in Algorithmic Hiring," ACM FAccT Proceedings (2020); EEOC Technical Assistance on AI and ADA (May 2023); FTC AI Hiring Enforcement Guidance (2024); LinkedIn Global Talent Trends (2024); SHRM AI in HR Survey (2024); Gartner Market Guide for AI-Augmented HR Services (2025); University of Melbourne AI Hiring Tools Study (2023).
How accurate are AI interview assessment tools compared to human interviewers?
Independent research places AI interview scoring tools at criterion validity coefficients of .20 to .35, comparable to unstructured human interviews. This falls below the .48 to .51 range established for structured behavioral interviews by Schmidt and Hunter's 1998 meta-analysis. The accuracy depends heavily on design: structured, transcript-based approaches show stronger validity evidence than those relying on facial or vocal analysis.
Are AI interview tools biased?
Multiple studies document bias risks. A 2023 University of Melbourne study found statistically significant bias against non-native English speakers across all six AI hiring tools tested. Raghavan et al.'s 2020 FAccT research documented racial and gender bias in video-based scoring platforms. With only 22 percent of employers actively testing their AI hiring tools for bias (SHRM, 2024), the majority of deployments lack ongoing fairness monitoring.
What laws regulate AI interview tools in the United States?
The EEOC's May 2023 technical assistance clarifies that employers bear liability for adverse impact from AI hiring tools under Title VII and the ADA. NYC Local Law 144 requires annual independent bias audits for automated employment decision tools. The FTC has flagged AI hiring tools as an enforcement priority regarding accuracy claims and candidate disclosure.
What should a valid bias audit include?
A meaningful bias audit calculates selection rates by protected group — including race, gender, age, and language background — and applies the four-fifths rule. It should be conducted by an independent third party, repeated annually at minimum, and test the tool under real-world conditions using the employer's actual candidate population rather than synthetic datasets.
How can companies reduce legal risk when using AI screening?
Choose tools that operate as decision-support rather than automated decision-makers — this reduces exposure under AEDT laws. Ensure human reviewers make final hiring decisions. Require vendors to provide independent validity and bias audit data. Disclose AI use to candidates. Monitor outcomes for adverse impact on an ongoing basis. Document your compliance posture before regulators ask for it.