Why the AI in your ATS gets things wrong
Not because the model is stupid. Because almost nobody gives it a way to say it does not know, and a system that must always answer will always answer, including when there is nothing to answer with. Here are the three causes, how to tell which one you have, and what a product that got it right does instead.
Agencies who have had AI in their CRM for a year tend to describe the same arc. Impressive in month one. Genuinely useful in places by month two. And then, somewhere around month three, a quiet distrust settles in. Consultants start double-checking things. Somebody keeps a spreadsheet. Nobody can point at a disaster, because there wasn't one, and everybody has stopped quite believing the screen.
That arc is not bad luck and it is not a bad vendor in the ordinary sense. It is the predictable output of three design decisions, and once you can see them you can test for them in twenty minutes on a demo call.
Cause one: a model will never refuse to answer
The thing underneath almost all of this is a language model, and a language model produces the most plausible continuation of what it has been given. That is its entire nature and it is not a defect.
Ask it a question with a clear answer and you get the clear answer. Ask it a question with no answer available and you get the most plausible-looking one instead, in the same tone, at the same apparent confidence, with no hedge. There is no internal moment where it notices the question was unanswerable. Nothing in it corresponds to hesitating, and nothing in it corresponds to saying I would need to check.
Put that inside a CRM and the consequence follows immediately. An email arrives saying please send this one to Sara. Your agency has a Sarah in Leeds and a Sara Collingwood in Manchester. The software assigns the record to one of them. The record it writes looks exactly like every correct record in your database, because it is the same value, in the same field, in the same format, with no mark on it.
Nothing about this requires a poor model. A better model gets the answerable cases right more often and behaves identically on the unanswerable ones, because the problem is not accuracy. The problem is that there is no route out other than answering.
And notice what it does to the cost of the errors. A system that is wrong ten per cent of the time and tells you which ten per cent is useful, because you check those and trust the rest. A system that is wrong two per cent of the time and cannot tell you which two per cent is worse in practice, because the only rational response is to check everything. Consultants work that out within a quarter, which is exactly when the distrust arrives.
Cause two: the software was given the wrong shape of job
The second common failure is asking a model to produce text when what was actually needed was a choice, and it accounts for a whole category of errors that get blamed on the AI unfairly.
Take setting the seniority on a job. The tidy way is to hand the software the list your agency actually uses, plus the words your consultants use for each one, and ask it to pick one of those or decline. The common way is to ask a model to state the seniority, get back a word like mid-weight, and then try to match that word against your list in code.
That matching step is where the errors live. The model said something entirely reasonable, your list uses slightly different words, and the match either fails silently or lands on the wrong entry. Nothing was hallucinated. The software was simply never asked the right question, and the failure surfaces three steps away from its cause, which is why it is so rarely diagnosed correctly.
This matters to you because it is testable from the outside without knowing anything technical. Ask a vendor whether their AI picks from your list, or writes an answer that then gets matched to your list. It is a question about design, and the answer tells you where a whole family of errors in your data is coming from.
It also explains a pattern you may already have noticed: the AI being reliable on free-text work like summarising a call, and unreliable on exactly the structured fields that ought to be easier. Structured fields are the ones being matched.
Cause three: nothing was tested against the live model
Models change underneath everybody. A provider improves something, adjusts something, retires a version, and the answers shift. Products that pass your text straight out find out when a customer complains, because nothing in their own code changed and so nothing alerted them.
The defence is a set of test cases run against the real model before anything reaches a customer. That much is obvious once stated. The part that is not obvious, and that almost nobody builds, is cases that assert the software stays unsure about genuinely ambiguous input.
Without those, a dangerous kind of drift is invisible. A model becoming more confident looks like an improvement on every measure anyone tracks, and it is not an improvement at all when the thing it has become confident about is a coin toss. The accuracy score goes up and your data quality goes down, at the same time, for the same reason.
We learned this the honest way. Two of our cases failed when we first ran them, both returning the right answer at low confidence. The model was being truthful and the gate was doing its job. The test was wrong, so we added a way for a case to assert stay unsure about this as a contract in its own right.
What a well-built version does instead
Four things, none of them exotic, and all four are absent from most products in this market.
It asks for a choice from a real list rather than for a sentence to be matched afterwards. It gets back a confidence figure alongside the answer. It compares that figure against a line, and below the line it writes nothing and hands the case to a recruiter. And it keeps the older, simpler method working underneath, so that when the clever path declines or fails, something still happens.
The effect on your day is worth stating plainly, because it is not more AI, it is less. Fewer things are done automatically. The ones that are done automatically are right. A handful of cases a day arrive on a person's desk with a reason attached, and those are genuinely the cases that needed a person. Everything on your records was either certain or approved by somebody.
Matching the symptom to the cause
| What you notice in your CRM | What is actually causing it | What to ask the vendor |
|---|---|---|
| Records that are confidently wrong | No route for the AI to decline | What does it do when it is unsure? |
| Free text fine, list fields wrong | Text generated, then matched in code | Does it pick from our list, or match an answer to it? |
| Behaviour changed without a release | A model changed underneath them | What runs against the live model before release? |
| Consultants double-checking everything | All of the above, compounding | How do we tell an AI value from a human one? |
The twenty-minute test
Take a genuinely ambiguous case into the demo. Not a tidy one from their sample data. An email with a first name that matches two of your people, or a CV where the most recent role has no dates. Watch what the software puts on the record, and whether anything on that record says it was unsure.
Then ask the question that outlasts the demo: six months from now, how would we tell which values on a record were set by a person and which by the software. If there is no answer to that, every AI-set value in your database is permanently indistinguishable from a checked one. That information was never captured, so it cannot be recovered later by better reporting or a better model.
None of this is an argument against AI in recruitment. We build it, and the version that works gives an agency real hours back. It is an argument for the unfashionable half of the work, which is deciding what the software does when it does not know. That half is most of the job, none of the demo, and the entire difference between a database you trust in three years and one you do not.
Lokesh is Founder and Head of Engineering at Recruitly.



