Most AI recruiting tools we come across are good at measuring the basics. They report time-to-hire, completion rates, and application volume. What none of them measure is whether the people hired through an AI-screened pipeline perform better than the people hired the old way. We dug deeper, and it’s a question literally nobody can answer as of this writing.
The few quality-of-hire metrics touted typically rely on vendor-led efforts. For instance, the most-cited evidence we could find that AI screening works best (as of Summer 2026) is a field experiment on roughly 37,000 applicants for a junior developer role conducted a few months ago. In the experiment, candidates screened by the AI passed the final human interview at a 54 percent rate, against 34 percent for those screened from resumes alone. It is a genuinely well-built study, but it was co-authored by the founder and CEO of micro1, the company whose interview tool it evaluates. It also focuses solely on one type of role, and one type of industry. It’s a start, but far from what we’d like to see to make more generalized claims about a technology type.
Then there are the issues with AI hiring. A Stanford-led audit of the hiring platform Pymetrics, forewent all measures of quality, and tried to focus on measuring whether the system treated candidates fairly. The result? It found that “these tools increase racial bias and shut the same people out of jobs everywhere they apply”.
None of this means AI recruiting doesn't work. It means that if you are buying with quality of hire in mind and not only speed and volume, the outcome evidence is thinner than the marketing suggests, and there are genuine concerns you should be aware of before you sign. Below we’ll explore what the evidence does and does not show thus far.
Beware of the Headline
Granted, as an HR Tech buyer, you don’t always have time to be reading academic papers, but a similar, healthy skepticism should be applied to every metric put out by vendors. For example, while researching for this piece, I was reading a recruitment vendor's 2026 guide on measuring quality-of-hire. Making the case for AI in hiring, it lists as a key fact that AI-screened candidates pass final interviews at 54 percent versus 34 percent (yes, same figure from the study I mentioned above). The figure is hyperlinked, so I click it. It does not go to a study. It goes to another page on the same vendor's site ranking the best AI recruiting software, where that vendor has ranked themselves as the first option. Big surprise there. Within, the figure links again, this time to the real paper on arXiv. Open the paper and you find the authors measured whether candidates passed the next interview, and state directly that they did not measure job performance or retention. Three hops from "quality of hire" to a study that says it did not measure quality of hire.
That is just one example, but it goes to show that you should always take these bold claims with a grain of salt. Hiring AI is sold on headline numbers, and sometimes the metric measures the wrong thing.
Another problem we’ve found with headlines from data is that the aggregate hides what the disaggregated data shows. In other words, they can claim the data says one thing, at the average level, but if somebody drills into it, the story can change substantially. Another two of the most serious studies on algorithmic hiring published in 2026 echo this point from different directions.
That Stanford-led study drew on Pymetrics data covering 4.2 million applications across 156 employers. In aggregate, every racial group cleared the legal fairness threshold (the four-fifths rule required by law), as the vendor's own prior research had reported. But that regulation was written to be applied per position, not to a pool of every job posted by an employer, since pooling can hide a violation in one role behind high selection rates in another. When the researchers applied the test the way the law intends, position by position, 10.62 percent of the 1,746 positions showed adverse impact against Black applicants. Adverse impact is what is supposed to trigger further scrutiny, it's not proof of illegality.
So the study found out there was more to the headline than what was touted, and that these tools require that level of analysis in order to properly understand the outcomes. Still, the study cannot say the rejected candidates would have been worse hires; it has no measure of quality, but did point out a potential fairness issue.
The same pattern recently turned up the other side of the Atlantic. An independent audit of Barcelona Activa, the city’s public employment agency, examined roughly 497,000 candidate records from a hiring system built on a different vendor, in a different country, under a different labor market. Different everything, except the technology, which is the same principle: algorithms filtering and ranking applicants. And it had the same problem. In aggregate, outcomes across genders looked equitable, and a compliance check reading that top-line number would have passed the system. Disaggregated by salary band, sector, and age, the same records showed women shortlisted less for mid-salary and full-time roles and workers over 55 years of age almost entirely absent from the pipeline.
That two independent 2026 audits, on different continents, found a similar gap between the headline and the disaggregated reality is something I didn’t expect to find when I began researching the impact of AI in quality of hire. It is also why the four-fifths rule is written to apply job by job rather than company-wide: with real people and real jobs, the average is exactly where a disparity goes to hide. But there is a larger implication for anyone buying this technology for quality of hire. Reducing bias was one of algorithmic hiring's original selling points, a claim few thought to challenge. Years in, some of these tools still produce biased outcomes their own aggregate numbers conceal. If the technology has not delivered on one of the main things it was sold to fix and that we know how to measure, perhaps it is a long shot to expect that it will deliver on quality of hire. But, should we expect it to?
Where AI Leaves Quality of Hire
Quality of hire is a challenge to measure as it is, but fortunately, many companies are meeting it head-on.
Metaview, for instance, proposes measuring the quality of hire at the interview layer, reading signals off the transcript within a day instead of waiting a quarter for a manager's review. Most tools take the older route: optimize for speed, ship the hire, and evaluate quality years later, once no one remembers the why behind the hiring decision.
I’d say they are on to something. You can gauge quality at every step, and you should. Hiring is less like a lab result that arrives at a fixed moment in time and more like dating: you form a read from the first message, sharpen it through each exchange, and keep choosing whether to go on. A hire is the same, sort of. There are signals at the screen, at the interview, at 90 days, at the one-year mark when you learn whether they stayed and whether their manager would do it again, up to the first promotion or even 10 years in. The mistake is treating quality as a thing you check once.
But assuming you start trying to get a read of hiring quality at the onset of the relationship, AI is wreaking havoc there as well. A screening tool hands you a shortlist faster than any human process could, and tells you “these are the strongest candidates.” How do you know they are? How do you know the people it surfaced are the ones you should have been assessing, and not simply the ones its scoring happened to favor? To answer that you would need to see the candidates it ranked below the line, interview them too, and compare how both groups performed a year on. Almost nobody runs that experiment, because it is slow and expensive and cuts against the entire value proposition, which is not interviewing the people the tool filtered out.
Sourcing tools are amazing technology, but they don't do that. Then there’s what AI has done to the candidate side. People are now using artificial intelligence to mass-apply to jobs, spruce up their resumes, leading to a true ‘slopification’ in hiring.The recent answer from the recruiter side, which might also impact quality of hire, are ‘Interview Cheat Detection Tools’, which are supposed to flag candidates using AI. Aside from the outright hypocrisy we could call when recruiters use AI to sort or even interview candidates, but cut off those candidates who use AI to polish their resume or interview answers, it’s also a question whether AI use can be truly systematically detected in the first place. This is a question of the broader tech world, not just recruiting, which you likely know if you’ve ever used an ‘AI-detection’ tool for a text, like the one Substack recently added to all posts.
Still, the idea is interesting, and we’d argue that interview intelligence tools which already use some form of AI to help you better interview candidates would be doing their customers a favor by incorporating some form of this tech, at least as a ‘FYI’ within candidate’s profiles. Fabric, for example, has built much of their value proposition around this. The interesting thing will be whether candidate assessment tools start to 1) incorporate more forms of AI (which we haven’t seen as much as in other recruiting categories) and 2) whether they also start to build ‘AI-detection’ mechanisms into their assessments.
Alas, while AI's effect on hire quality sits mostly unmeasured, a combination of these tools, along with human-centric teams devoted to truly understanding the people on the other side of the job application, is bound to deliver some good candidates.
What AI Recruiting Actually Does Well
To be fair, if we step back from the quality-of-hire question, AI earns its keep in a specific place: the top of the funnel. Sourcing and screening at volume is an actual, measurable win. When a role draws 2,000 applications, no human team is able to invest the time and manpower necessary to read them all. Tools that surface, rank, and route candidates cut time-to-hire and clear backlogs that used to swallow weeks. As Josh Bersin put it on his podcast in August 2026, this is a $100 billion-plus market full of systems that genuinely work, sitting inside an industry that is still immature and where nobody is happy. Both halves of that are true. AI is very good at moving high volumes of applicants through a pipeline quickly and lets recruiters focalize their efforts where they might be more effective, but it is not a panacea, and it has not shown it produces better hires by itself.
The distinction worth underscoring is between sourcing and evaluation, which is a combination, typically unique to each business, of interviews and assessments. Sourcing is finding and filtering candidates, deciding who reaches the shortlist. That is where most AI recruiting tools operate, and where the speed gains have a more translatable benefit in time savings, especially for teams working against the widespread challenge of candidate volume. There is even decent evidence they catch things humans miss: in the same 2026 field experiment behind that 54-percent headline, the AI highlighted candidates who listed skills they could not actually demonstrate, roughly one in five, and a human expert agreed with the AI's assessment 96 percent of the time.
The evaluation stage, whether heavier on assessments or interviews, is the more important one for quality because it’s supposed to evaluate how good the shortlisted candidates are, measuring the profile against the job rather than a resume against a keyword salad. The reason this stage matters most for quality is not opinion: a 2022 reanalysis in the Journal of Applied Psychology found the structured interview to be the single strongest predictor of job performance, ahead of cognitive-ability tests and well ahead of resume screening. This echoes the conversation we recently had with Rob Devlin, a Global TA Program Manager for various companies. After delving into the question of candidates using AI to cheat in interviews, he posited that there’s a difference between finding better words to share what one has done, and making up stuff. The way to find out which is true for each candidate? An interview. "How did you do that? What did you do first? Who was in the room? Which stakeholders did you engage?" he said. "If you ask me that, you're going to know within 10 seconds whether I've made it up."
So the practical split for a buyer is: if your problem is volume, too many applicants, too slow a funnel, roles sitting open, AI sourcing tools may have a strong, evidenced case. If your problem is quality and retention— whether the people you hire actually work out, the honest answer is that no tool has proven it fixes that, but validated assessments, if done with enough training data, and with a human at the end of every decision, might be the best bet as of now.
And if you would rather not wade through the task of vetting tools alone, that is what our advisor service is for. Tell us the problem you are actually trying to solve, volume or quality, and we will point you to the two or three tools that fit, free. The one thing we would not do is let a headline number make the decision for you.





















