JobCraftly.

AI and your job search

What to expect from an AI job interview, and what a 70,000-applicant trial found

In the one large randomised test so far, an AI voice agent worked through the employer's own interview guide over the phone, a human recruiter read the transcript and decided, and the applicants assigned to the software were offered the job slightly more often than the ones assigned to a recruiter, which says less about AI than about how much of a first-round interview was already a checklist.

A photographic still life in the JobCraftly brand colours. A plain deep-navy telephone handset with no dial or buttons lies on a pale warm off-white surface, its coiled navy cord running to a small featureless deep-navy cube at the upper right. In front of them lies a straight row of six identical blank rectangular cards, five pale off-white and, at the right end of the row, one deep rust red.

The email says your first interview will be conducted by an AI. Not scheduled by one, or transcribed by one: conducted. That can mean more than one thing. In the version this article has evidence about, a voice rings you, asks the questions, follows up on your answers and ends the call, and at some later point a person reads what you said. Other versions record you on video with nobody on the line, and some score your answers with nobody reading them at all; what the evidence can and cannot say about those is set out further down. Whichever you have been sent, the natural reading is the same: that you have been handed the cheaper version of the process, the one for applicants the company would rather not spend a recruiter on.

There is now one large randomised test of the half of that reading that can be tested. It cannot tell you whether an employer is sorting applicants into tracks, because in the experiment nobody was: the interviewer was assigned by lot. What it can test is whether the process the firm built around the software, a phone call with an agent and a person deciding afterwards, cost applicants anything compared with the recruiters' own process, and on that it points the other way. Between March and June 2025 a recruitment firm in the Philippines assigned 40,103 applicants for customer-service jobs to an AI voice agent and 13,557 to its own recruiters, at random, and the group assigned to the software was offered the job slightly more often, started slightly more often, and was still employed a month later slightly more often.1 What follows is what that experiment found, what it cannot tell you, and what it suggests about how to answer a machine whose only product is a transcript.

What the call was like

The AI interview was a phone call, whether the applicant had applied online or had walked into one of the firm's recruitment centres, where the call was taken on site. The human recruiters it was compared with interviewed walk-in applicants face to face and online applicants by phone, so for walk-ins the AI arm also swapped a room for a phone line, and the trial compares the two processes as the firm ran them rather than the software on its own. The agent itself was a large language model wired to speech recognition on the way in and a synthetic voice on the way out, and it was given the same written interview guidelines the human recruiters used.1 The paper says that when the call started the agent "immediately discloses its artificial identity" and "explicitly states that a human recruiter will review the interview, evaluate it, and make the hiring decision, not the AI itself".1 That sentence describes the whole division of labour: the software collected, and people decided.

The guidelines allowed up to 14 topics, a mix of verification questions and open-ended ones, starting with the practical screens, such as where the applicant lived, whether they could commute and whether they could work changing shifts, before moving to motivation, work history and education, with details of the job and a chance to ask questions at the end. A full interview took between 10 and 20 minutes. After it, applicants sat a separate standardised test of English and reasoning of about half an hour, and a recruiter then scored the interview and the test and decided whether the applicant met the firm's standard.1 If your invitation describes something like this, a short structured phone screen followed by a human decision, you are looking at the same shape.

What the experiment was

The scale is what makes it worth reading. The firm, PSG Global Solutions, a subsidiary of Teleperformance, received 70,884 applications between 7 March and 7 June 2025 for 48 postings across 41 client accounts, all entry-level customer-service roles in the Philippines serving US and European companies, paying between 16,000 and 25,000 pesos a month, roughly $280 to $435. Of those, 67,056 were eligible and were randomly assigned: 40,103 to the AI interviewer, 13,557 to a human recruiter, and 13,396 to a third group who were offered the choice.1 The study was pre-registered, and the authors are Brian Jabarian of the University of Chicago Booth School of Business and Luca Henkel of Erasmus University Rotterdam; it was first circulated in August 2025 and the version read here is dated January 2026.12

One thing to know about the authors' position. The paper states that the firm supplied the data but "had no role in analyses, manuscript preparation, or the decision to publish", and that two months after data collection ended, Jabarian accepted an unpaid role as the firm's chief economist.1 That is disclosed on the title page, which is where it belongs, and it is a reason to read the numbers rather than the summary.

The applicants assigned to the software did slightly better

Of the applicants assigned to a recruiter, 8.70 per cent were offered a job; of the applicants assigned to the AI, 9.73 per cent were. Both figures count everyone assigned to the group, including the applicants whose AI call failed or who ended it rather than talk to a machine, so they are not success rates among completed interviews. The paper describes the difference as a 12 per cent higher likelihood of an offer; in absolute terms it is one extra offer per hundred applicants. Counting everyone who was randomised, the AI-interviewed group was also 18 per cent more likely to start the job and 18 per cent more likely to still be employed after a month, with the gap holding at 17, 16 and 17 per cent at two, three and four months.1

The paper also looks only at the people who accepted an offer. That is no longer a comparison of the randomised groups, because who received an offer already differed between them, so read it as description rather than as a cleaner test. Among acceptors, 68.84 per cent of the recruiter-assigned and 73.36 per cent of the AI-assigned actually started, and 60.60 against 64.32 per cent were still there a month later. At two months the figures were 55.62 and 59.13 per cent. At three and four months the gap pointed the same way but was no longer statistically distinguishable from zero.1 The causal result is the one in the previous paragraph, counted over everyone assigned. Among the people who were offered a job, acceptance was similar whichever interviewer they had been assigned, 93.64 per cent in the recruiter group and 92.14 per cent in the AI group, a difference the paper cannot distinguish from zero at the conventional level. Among people hired, the split between leaving voluntarily and being let go was the same in both groups, and on the three productivity measures the employer tracked, handle time, customer satisfaction and quality scores, the paper found no meaningful difference.1

Two things are worth holding onto from those numbers. First, the offer rate was under 10 per cent either way: the AI interview did not make the job easy to get, it made it slightly less hard. Second, on the measures the paper could see, the people hired after an AI interview looked like the people hired after a human one. That is a description of the hires it observed, not proof about the extra hires the AI produced, but it is the finding the firm will read, and the reason this kind of interview is likely to spread.

Why: the software ran the interview the firm had written down

The authors analysed the transcripts, and the explanation they arrive at is unglamorous. The AI stuck more closely to the guideline's topic order, covered a more consistent number of topics from one interview to the next, and used more standardised wording.1 Human recruiters were allowed to tailor coverage, order and phrasing to the applicant in front of them, and they did, with the result that two applicants interviewed by two recruiters were not really being asked the same interview.

The paper then asks what the recruiters were rewarding when they read a transcript. Within the recruiter-led interviews, an offer was more likely the more exchanges the interview contained and the richer and more complex the applicant's language, and less likely the more the applicant used short attention cues, the "yes", "okay" and "mhm" that the paper counts as backchannel cues, and the more questions the applicant asked.1 The AI-led interviews contained more of the features that predicted an offer and fewer of the ones that predicted a rejection. The authors' phrase for the mechanism is controlled variance: the agent adapted its follow-ups to each person inside a framework it did not depart from.

The recruiters noticed, and reacted in two directions at once. They gave the AI-led interviews higher scores, mostly by rating fewer of them low rather than by rating more of them high. But when the paper models the offer decisions themselves, recruiters "place less weight on interview scores and more on language scores" for applicants the AI had interviewed than for the ones they had interviewed themselves.1 The conclusion puts it in four words: recruiters discount AI signals. A transcript someone else collected was trusted a little less than a conversation they had had.

What the applicants thought

The firm surveyed applicants afterwards, and the response deserves its caveat: the survey went to 19,200 people and 2,764 answered, a 14 per cent response.1 Among those who did, the willingness to recommend the firm to a friend scored 8.97 out of 10 after an AI interview and 8.84 after a human one, a difference the paper reports as not significant. Stress and comfort were rated about the same. So were the flow of follow-up questions, the amount of feedback and the fairness of the interview. The one clear difference was naturalness: applicants found the AI conversation significantly less natural.1 A smaller share of respondents reported feeling discriminated against on grounds of gender after an AI interview: 62 of the 1,818 respondents in the AI group, about 3 per cent, against 22 of the 346 in the human group, about 6 per cent, a rate the paper describes as nearly halved and which it notes rests on small counts.1

Then there is the group who were given the choice. Of the 13,396 applicants in that arm, 13,079 recorded a choice, and 78 per cent of those chose the AI: 69 per cent of the walk-in applicants, who had come to a recruitment centre to apply, and 82 per cent of those who had applied online.1 The invitation page is reproduced in the paper's appendix, and it offered two things at once: the AI call "can be scheduled at your convenience", while a human interview had to be scheduled "based on the human recruiter's availability".1 So the choice bundled the interviewer with the scheduling, and the paper neither separates the two nor asks people why they chose. What it does report is that applicants' attitudes to AI, which were generally positive in this sample, predicted their choice.1 The paper adds a finding it calls negative sorting: the applicants who chose the AI scored significantly lower on the language and reasoning tests than the ones who chose a recruiter.1 It does not report offer rates by the interviewer people chose, so nothing in it says which choice paid off.

What went wrong

Not everything ran. The paper reports that "5% of applicants ended their interview because they were unwilling to speak to an AI, and in 7% of cases, the AI voice agent faced technical difficulties".1 The technical failures were classified from the transcripts into two kinds, telephony faults such as signal loss or an unstable internet call, and faults in the agent itself, which the classification scheme defines as the agent stalling, crashing or failing to respond.1 Roughly one AI call in fourteen, then, had a fault of one kind or the other, which is a rate worth knowing before you take the call somewhere with one bar of signal.

What the paper does not say is what happened next to those applicants. It does not report whether people whose calls failed were re-interviewed, or whether the five per cent who declined the AI were offered a recruiter instead.1 So this study cannot tell you what refusing costs. It can only tell you that in this firm a refusal ended the interview.

What this does and does not tell you about your own interview

Every number above comes from one firm, in one country, hiring for one kind of job, and the job matters more than the country. These were entry-level customer-service roles in which fluent spoken English is the work itself, the interview topics were mostly practical screens and work history, and the whole thing took a quarter of an hour. That is the setting in which speaking in full sentences to a machine is the closest possible rehearsal of the job.1 The authors say as much in their conclusion: the benefits of automation "appear highest in high-volume, high-turnover environments where tasks are repetitive, outcomes are rapidly observable", while "settings that depend on tacit knowledge, relational inference, or screening of highly specialized skills may benefit more from human screening".1

Four more limits are structural. The agent was a voice on a phone line, not a video interview, and the study says nothing about tools that record your face. For walk-in applicants the AI arm replaced a face-to-face interview with a phone call as well as replacing the interviewer, so what was randomised was a whole process, and the result is about that process rather than about software as such. The software collected and a human decided; where a vendor's software scores you and nobody reads the transcript, this experiment is not evidence about that. And the recruiters knew which interviews the AI had run, which is precisely why they were able to discount them. Whether your interview will be read by a person, and what you are entitled to be told about that, depends on where the job is, and we set out the rules for New York City and for the EU in the article on automated screening. This article does not add to them.

How to answer an interviewer that is producing a transcript

With those limits stated, the experiment does point at how to prepare, because what the AI removed from the interview is as informative as what it kept.

Treat it as a structured interview, because that is the direction it leans. The agent kept closer to the guide's topic order and covered a more consistent number of topics from one interview to the next than the recruiters did, while still adapting its follow-up questions to each applicant, so the topics a recruiter might have skipped for you are more likely to come up.1 The preparation is the one that works for any structured interview, which we described in the article on scripted and unscripted interviews: for each topic the job description implies, have one specific, complete example ready, and let the follow-up questions draw out the detail.

Answer in sentences. In this experiment's human-led interviews, offers went more often to applicants whose interviews had more back-and-forth, richer vocabulary and more complex sentences, and less often to applicants who leaned on "yes" and "okay".1 That is a correlation in one dataset for a job that is spoken English, and other jobs may weigh it differently. But the underlying point travels to any AI interview where a person reads the transcript afterwards, as one did here: a transcript contains nothing a human interviewer would have inferred from your face, your handshake or the room. A one-word answer that would have been fine in person, because you were visibly engaged, is a one-word answer on the page. Where software scores you and nobody reads it, this study has nothing to say.

Ask your questions, and know what the transcripts showed. In the recruiter-led interviews, the more questions an applicant asked, the less likely the offer.1 That is an association inside one set of interviews, not a test: nobody was assigned to ask more or fewer questions, the paper does not say why the pattern exists, and it did not check whether the same holds when the interviewer is the agent. So it is not a reason to stay silent, and if the AI call is your only contact with the employer it may be your only chance to ask about the job. It is a reason not to let your questions crowd out your answers, and to ask them at the point set aside for them, which in this firm's guide was the end of the call.1

Take the call somewhere it will not drop. One in fourteen of these calls failed for a technical reason, and the study does not tell you what a failed call costs, so the cheap move is to give it a charged phone, a quiet room and a strong signal.1

If you are offered the choice, know what the study can and cannot tell you about it. In this experiment the AI option came bundled with flexible scheduling, so the paper cannot say how many chose the machine for its own sake; it does not report whether either choice led to more offers; and the group that chose the machine was, on the tests, the weaker one.1 None of that is a signal about which you should pick. Choose on what you want from the call, and do not read the machine as the lesser option: the randomised comparison of the two processes says it was not.

The single most useful fact in the experiment is that the AI ran the interview the firm had already written down, and the firm's own recruiters, reading the results, preferred what came back. What a voice agent takes out of a first-round interview is the small talk and the reading of the room. What it leaves is your answers, in full, to a list of questions you can largely predict. That is a narrower test than a human interview, and a far more rehearsable one. If you want to rehearse it, JobCraftly can run a practice round against the actual job description and give you specific feedback on your answers, which is the part of this interview that will still be there when the voice is not.

1: Brian Jabarian and Luca Henkel, Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews, working paper dated 26 January 2026, arXiv:2607.28222, posted 30 July 2026. All figures are from the full text; the section for each is given in the source notes. Accessed 13 September 2026. 2: Erasmus University Rotterdam, Study finds AI matches humans in conducting job interviews, news release, 19 August 2025. Accessed 13 September 2026.

References

Sources

  1. Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews
    Brian Jabarian (University of Chicago Booth School of Business) and Luca Henkel (Erasmus University Rotterdam); working paper dated 26 January 2026, posted to arXiv as 2607.28222 [econ.GN] on 30 July 2026; first circulated as an SSRN working paper in August 2025 · accessed 13 September 2026
  2. Study finds AI matches humans in conducting job interviews
    Erasmus University Rotterdam, news release, 19 August 2025 · accessed 13 September 2026