AI adoption in recruitment has moved fast. Most agencies now use some form of automated screening or matching, and it is easy to see why. It promises speed, and speed matters when clients want a shortlist yesterday. But a stat that is doing the rounds in industry conversation right now is worth sitting with before you trust that speed completely: when the same AI screening tool was run twice on identical candidate data, the shortlist overlap was only 14%.

Run that the same way, twice, expect the same answer, and get something close to a coin flip instead. If your agency is leaning on a tool like that to make first-round decisions, it is worth asking what else might be slipping through.

Why the same input produces a different answer

Most “AI screening” tools are not doing what the name suggests. A lot of them are closer to keyword matching with a confident interface, scoring CVs against a job description using pattern recognition rather than genuine understanding of context or transferable skill. That approach is inconsistent by nature. Change the order candidates are processed in, tweak an unrelated setting, or simply run the same batch again, and the output can shift because the scoring was never as deterministic as it looked.

The practical effect is that strong candidates get missed, not because they lack the skills, but because the tool scored them differently the second time round. For an agency, that is not a minor technical quirk. It is candidates who never make it in front of a client, through no fault of their own.

The signal problem is getting worse, not better

Inconsistent screening would be a manageable issue on its own. The bigger concern is what is happening on the other side of the process. Estimates suggest that around 40% of tech candidates have meaningfully inflated their CVs, and reports of deepfake interviews and AI generated applications are becoming more common, undermining the signals recruiters have relied on for decades to judge who is genuinely qualified.

Put those two things together and you get a screening process that is both less reliable at evaluating real candidates and more exposed to candidates who are not being straightforward about their experience in the first place. Neither problem cancels the other out. They compound.

What this means for how agencies use AI

None of this is an argument against using AI in recruitment. It is an argument for being precise about what you are asking it to do. A few practical principles worth applying:

Treat AI output as a first pass, not a final answer. If a tool is producing a shortlist, a person still needs to look at who got filtered out, not just who made it through. The candidates sitting just below the cut line are often where the inconsistency shows up most.

Build verification into the process itself, not just at offer stage. With CV inflation and AI generated applications on the rise, a skills based, evidence first evaluation step earlier in the process protects the quality of your shortlist rather than catching problems after a client has already met someone.

Ask specific questions before buying or renewing any AI tool. What accuracy has it shown on CVs from your specific market, not a generic benchmark. How does it handle skills that are implied rather than explicitly stated. Does it improve over time in a way you can actually see, or does performance just fluctuate. A vendor who cannot answer those clearly is asking you to trust a black box with your shortlists.

Remember where the highest value AI use case actually sits. For most agencies, the biggest return from AI is not in screening new applicants at all. It is in activating and cleaning the database of candidates you already have. A large share of placements come from people already known to the desk, so keeping that data accurate and current often does more for fill rates than any external sourcing tool.

Talking to clients about it

Clients are hearing the same headlines about AI in hiring that everyone else is, and plenty of them now assume that “AI screened” automatically means “thoroughly vetted.” That assumption is worth gently correcting rather than leaving unchallenged. A client who believes the technology has already done the hard work may push back when you present a shortlist that took longer to compile, or ask why you are not simply forwarding a bigger batch of AI matched CVs straight through.

This is a genuine opportunity to explain your value rather than defend it. Being able to say plainly that automated tools are used to speed up the early stages, but that every shortlisted candidate has been reviewed and verified by a person who understands the role, is a stronger pitch than pretending the technology is infallible. Clients generally respond well to that kind of honesty, particularly once they have seen a headline about resume fraud or a deepfake interview story themselves.

The candidate side of the equation

It is worth remembering that inconsistent screening does not only cost agencies placements. It costs candidates fair consideration too. Someone who is genuinely well suited to a role can be filtered out on one run of a tool and would have made the shortlist on another, through no difference in their actual ability. For candidates who already find the process opaque and frustrating, an inconsistent AI layer added on top makes that experience worse, not better.

Agencies that can speak to this honestly, explaining how candidates are actually assessed rather than leaving them to assume a machine made the call, tend to build stronger long-term relationships with the people in their database. That matters more than it might seem, given how much of future business comes from candidates agencies have already placed once.

The bigger shift underneath this

There is a broader point here that goes beyond any single tool. As AI takes on more of the repetitive, high volume parts of recruitment, the judgement calls, the relationship building, and the ability to spot when something does not quite add up become the parts that actually differentiate one agency from another. Automating the admin is worth doing. Automating the decisions that require genuine human scrutiny is where things start to go wrong, and the 14% figure is a useful reminder of why.

Agencies that build in a deliberate human checkpoint, rather than assuming the tool has already done the thinking, are the ones who will keep placing strong candidates while everyone else quietly wonders why good people keep slipping through the net.