NewRecruitly + WhatsApp — message from your CRM
Recruitly LogoRecruitly
Industry

Why the AI in your ATS gets things wrong

Not because the model is stupid. Because almost nobody gives it a way to say it does not know, and a system that must always answer will always answer, including when it should not.

Agencies who have had AI in their CRM for a year tend to describe the same experience. It is impressive at first, useful in places, and then somewhere around month three a quiet distrust sets in. Consultants start double-checking things. Somebody keeps a spreadsheet. Nobody can point at a disaster and everybody has stopped quite believing the screen.

I want to explain exactly what causes that, because it is a design decision rather than a limitation of the technology, and once you can see it you can test for it in twenty minutes.

A model will never refuse to answer

The thing underneath almost all of this is a language model, and a language model produces the most plausible continuation of what it has been given. That is its entire nature. Ask it a question with a clear answer and it gives you the clear answer. Ask it a question with no answer and it gives you the most plausible-looking one instead, in the same tone, with the same apparent certainty.

There is no internal moment where it notices the question was unanswerable. Nothing in it corresponds to hesitating.

Put that inside a CRM and the consequence follows immediately. An email arrives saying please send this to Sara. Your agency has a Sarah and a Sara Collingwood. The software assigns it to one of them. The record it writes looks exactly like every correct record in your database, because it is the same shape, in the same place, with no mark on it.

Why a wrong record is invisible
The three records are indistinguishable on screen and one of them is a guess. This is the mechanism behind a database nobody quite trusts, and no amount of a better model fixes it.

Nothing about that requires a bad model. A better model gets the answerable cases right more often and behaves identically on the unanswerable ones, because the problem is not accuracy, it is that there is no route out other than answering.

The second cause: the software was given the wrong shape of job

The other common failure is asking a model to produce text when what is needed is a choice.

Take setting the seniority on a job. The tidy way to do it is to hand the software the list your agency actually uses, along with the words your consultants use for each one, and ask it to pick one of those or decline. The common way is to ask a model to write down the seniority, get back a word like mid-weight, and then try to match that word to your list in code.

That matching step is where the errors live. The model said something reasonable, your list has slightly different words, and the match either fails or lands on the wrong entry. The failure gets blamed on the AI, when the AI was never asked the right question.

This matters to you because it is testable from outside. Ask a vendor whether their AI picks from your list or writes an answer that then gets matched. It is a question about design, and the answer tells you where a whole category of error is coming from.

The third cause: nothing was tested against the live model

Models change underneath everybody. A provider improves something, adjusts something, retires a version, and the answers shift. Products that pass your text straight out to a model find out when a customer complains.

The defence is a set of cases run against the real model before anything reaches a customer, including cases that assert the software stays unsure about genuinely ambiguous inputs. That second kind is rare and it is the one that catches the dangerous drift, because a model becoming more confident is not an improvement when it is confidently wrong.

We hold a battery of cases like that and run it against the live model before release. Some of them exist purely to prove that an ambiguous input stays below the line where anything gets written.

What a well-built version does instead

Four things, none of them exotic.

It asks for a choice from a real list rather than for a sentence to be matched afterwards. It gets back a confidence figure alongside the answer. It compares that figure to a line, and below the line it writes nothing and hands the case to a recruiter. And it keeps the older, simpler method working underneath, so that when the clever path declines or fails, something still happens.

The effect on your day is that the only cases reaching a person are the ones genuinely worth a person, and everything on your records was either certain or approved by someone.

Symptom in your CRMWhat is actually causing itWhat to ask the vendor
Records that are confidently wrongNo route for the AI to declineWhat does it do when it is unsure?
Fields set to the wrong value from a listText generated, then matched in codeDoes it pick from our list, or write an answer you match?
Behaviour that changed without a releaseA model changed underneath themWhat do you run against the live model before release?
Consultants double-checking everythingAll of the above, compoundingHow do we tell an AI value from a human one?

The twenty-minute test

Take a genuinely ambiguous case into the demo. An email with a first name that matches two of your people, or a CV where the most recent role has no dates. Watch what the software puts on the record.

Then ask how you would tell, six months later, which values on a record were set by a person and which by the software. If there is no answer to that, every AI-set value in your database is permanently indistinguishable from a checked one, and that is the thing that eventually makes people stop trusting the screen.

None of this is an argument against AI in recruitment. It is an argument for the unfashionable half of the work, which is deciding what the software does when it does not know. Getting that right is most of the job and none of the demo.


Lokesh is Founder and Head of Engineering at Recruitly.

ai-recruitmentatsdata-qualityai-accuracy

The product this came out of

Nineteen modules on one record: sourcing, screening, campaigns, calls, e-signature and billing, without a second system to keep in step. Free to start, no card, no call.