The Ceiling on AI Evidence
8 August 2026
An examination of the claim that AI can improve hiring by 12%

AI-generated summary
A field experiment at PSG Global Solutions randomised 53,660 applications between an AI voice interviewer and the company's own recruiters, and the AI route produced offers for 9.73 per cent of applications against 8.70 per cent. Foster-Fletcher reads what that result can support. Assigning an application to a route changed four parts of the process at once: who conducted the interview, who scored it, how closely the company's interview guide was covered, and how recruiters weighed the material once they knew an agent had produced it.
"Does AI Beat Humans at Recruiting?" said the headline in Chicago Booth Review, about a new field experiment, and the paper's own abstract is only a little more careful: applicants interviewed by an AI voice agent were, it reports, 12 per cent more likely to receive a job offer. The results are being pitched; PSG Global Solutions, the recruitment company whose hiring was studied, now cites the research in marketing the agent as Anna AI.
What follows is not exactly a debunking, the trial is close to the best evidence anyone has produced on AI’s effects on hiring, but more a consideration of the limits of the evidence.
What we have are numbers that show that the hiring process improved. What we don’t have is proof that an AI agent improved it.
Firstly, let’s break down the study. PSG randomised 53,660 applications for entry-level customer service work in the Philippines between two interview routes, in a split of roughly three to one, and left a further 13,396 applicants to choose their own route, a group that chose for itself, whose results the trial can report but cannot compare with the randomised routes. In one route the company's recruiters interviewed as before; in the other, a voice agent conducted the conversation. Human recruiters made every offer decision in both conditions. The company registered the design and the outcome measures before any applications were assigned, and followed the hires for four months.
And the results are interesting. The route through the agent produced offers for 9.73 per cent of the applications sent to it, against 8.70 per cent for the recruiter route. And more of those offers turned into starts, 6.71 per cent against 5.65 per cent.
Among the 4,300 hires the company could match to its employee records, 27.86 per cent of those who started after an agent interview were still in the job at four months, against 26.09 per cent of those who started after a recruiter interview, a difference well inside what chance would produce (p = 0.25).
Four things changed at once
While the agent route produced more offers, four separate things changed between the two routes. Randomisation assigned each application to an entire route rather than to an individual interviewer, so all four variables shifted at once. Two of those changes were assigned directly: who conducted the interview, and who scored it. In the agent route, an AI voice agent conducted the conversation, and scoring passed to a second recruiter, chosen by rotation, who had not spoken to the applicant. Because recruiters in both routes worked from recordings and transcripts, that second change reduces to a single fact: the scorer knew where the material had come from.
The remaining two differences were never assigned at all; they emerged as the routes ran. The agent covered more of the company's own interview guide than the human recruiters typically managed, and the recruiters scoring the agent interviews weighed the material differently on the authors' own later analysis.
A trial that assigns a bundle of changes can measure the effect of the bundle as a whole. It cannot divide the gain between the parts that were assigned, and it cannot divide it at all between the parts that emerged along the way, because those lie directly on the path between the assignment and the outcome.
The authors do try to separate those parts. Their analysis of 34,109 interview transcripts shows the agent covering 45 per cent of the guide's topics against 38 per cent for human interviewers, though the recruiters varied widely among themselves, with some covering 60 per cent and others under 40 per cent. But because the company changed the interviewer and the consistency of the interview in the same move, the authors can show that the two shifted together without showing how much either contributed. What's more, the transcript analysis, performance data, and study of recruiter behaviour were all added after pre-registration, making them exploratory, and the paper itself was posted on 30 July 2026 without peer review.
Standardisation and recruiter judgement impact the results
The most plausible alternative explanation is older than generative AI. A 1995 meta-analysis of 111 reliability coefficients from selection interviews found that standardising an interview raises the agreement between the people scoring it, and that a highly structured interview has roughly double the validity ceiling of an unstructured one. The company already had the guide and had written it for its own recruiters. If enforcement of that guide produced much of the gain, the company bought consistency in a method it already owned, and the voice agent was the means of imposing it.
A second explanation runs through the recruiters who made the offer decisions. Among walk-in applications, where recruiters had to log a score and a written justification for each interview they reviewed, they rated the agent's interviews above their own and wrote more positive justifications for them. At the offer decision the recruiters read the two scores differently: the interview score predicted the outcome less strongly when the agent had conducted the interview, and the language test score predicted it more strongly. The design does not appear to be able to separate an improvement in the interviews from a change in how the recruiters read them once they knew an agent had done the work.
The limits of the trial and transferability of the result
PSG is trying to attribute a financial improvement directly to an AI agent in its PR and marketing materials. This might stand up in its promotional content, but if you examine companies' earnings calls, you get something quite different. Earlier in 2026, I coded every AI-related statement in the quarterly earnings calls of 22 large listed companies, counting a statement as a measured contribution only where it attached a number to work the company itself does. Of 714 statements made across their 22 calls, 404 described adoption or containment and 308 were product or marketing claims. Just two of the statements measured a contribution to the company's own work. Once the conversation is made accountable, company reporting states only indirect improvements and will rarely commit to saying that the AI made the improvement.
PSG Global Solutions can bank the result of its AI trial no matter what the explanation is. It has a method of increasing the number of hires, reducing the time involved for staff and increasing the time those hires stay in the company. But even this result, one of the strongest pieces of evidence available about how AI improves the hiring process, cannot say how much of that improvement would be available to the next company that tries. The headline said the agent beat the recruiters. The trial itself says that one hiring route, which changed in four ways at once, beat another. The distance between those two sentences is, unfortunately, where most claims about AI at work currently reside.