Can you trust AI to update your CRM records?
Yes, under three conditions, and almost no product in this market meets all three. Here they are, how to check each one before you switch anything on, which fields to let it write and which to hold back, and how to roll it out without spending a quarter finding out you got it wrong.
This is the question agencies ask last and should ask first. Everything else about AI in recruitment software is a matter of convenience and hours saved. This one decides whether your database is still worth something in three years, and your database is the asset the whole business rests on.
The instinct most owners have is one of two extremes. Switch it all on, because it saves obvious time. Or switch it all off, because it cannot be trusted with anything that matters. Both are wrong, and the right position is in between and quite specific.
Condition one: it has to be able to decline
A language model always answers, because producing an answer is the whole of what it does. There is no internal moment where it notices a question was unanswerable and stops. That single fact is the basis of everything else here.
Left alone with your records, that means every ambiguous case gets resolved rather than flagged. An email says pass this to Sara, you have a Sarah and a Sara Collingwood, and the software assigns it. A CV has no dates on the most recent role and a plausible date appears. A reply is genuinely unclear and it gets filed as interested.
Each of those records then sits in your database looking exactly like the ones that are right, with nothing on it saying this one was a guess. And that is the real damage. A database with errors in it is survivable, because errors get found and corrected. A database where the wrong entries are indistinguishable from the right ones is a different problem, because the only rational response is to distrust all of it, and that is what your consultants will quietly do.
The fix is not a better model. A better model gets the answerable cases right more often and behaves identically on the unanswerable ones. What you need is a product that asks for a confidence figure alongside the answer, compares it to a threshold, and below that threshold writes nothing and gives the case to a person. Ours draws the line at three quarters. The number is unremarkable; the existence of a route out of answering is the entire point.
How to check. Ask the vendor to show you their software handling a genuinely ambiguous case. Bring your own, because nothing in their demo data will be ambiguous. Then watch whether anything reaches the record at all.
Condition two: you have to be able to tell afterwards
Six months from now somebody will look at a record and need to know whether a person checked that value or the software inferred it. If your CRM cannot answer that, every inferred value is permanently indistinguishable from a verified one, and the doubt spreads from the field to the whole record to the whole database.
This is the condition most products fail, and it fails silently. The software writes to the same field, in the same way a person does, with no mark. A year later nobody can separate the two, and there is no later fix: the information was never captured, so it cannot be recovered.
What you want is that every AI-set value carries its origin, and ideally how certain the software was when it set it. Then a doubt about one field stays a doubt about one field. It also gives you something to audit: you can pull a hundred AI-set values at random and actually check them, which is impossible if you cannot tell which ones they are.
How to check. Ask to see a record where the software has filled something in, and ask how you would know that in six months. If the answer involves a separate activity log that nobody will ever open, treat it as a no.
Condition three: the old way still has to work
The third condition is about ordinary bad days. Models are slow sometimes, providers have outages, and a well-built product declines rather than guesses, which means on some days the software simply does not do the work.
In a product built properly the previous method is sitting underneath and runs, and most recruiters notice nothing. In a fragile one the feature stops, and somebody discovers it at four o'clock on a Thursday with a client waiting for a shortlist.
This condition gets skipped because it sounds like an edge case. It is not. Once AI is doing work in the normal path, its failure is operational rather than annoying, and a vendor who has moved it into the path without building the fallback has created a risk and handed it to you without mentioning it.
How to check. Ask, feature by feature, what happens when the model is unavailable. Anybody who has thought about it answers immediately and specifically. A pause means nobody has.
Where I would let it write, and where I would not
Even with all three conditions met, I would not switch everything on at once. Some fields are cheap to get wrong and some are expensive, and the difference is not how important the field looks.
| Let it write | Let it suggest only |
|---|---|
| Call and meeting notes on the record | Who owns a client relationship |
| Sorting replies by intent | Fee and commercial terms |
| Structured details pulled from a CV | Pipeline stage on a live process |
| Tags, categories and topics | Anything a client might ask you to evidence |
| Flagging a client who has gone quiet | Right to work, compliance and consent |
The rule behind the two columns is not importance. It is how expensive the mistake is and how likely anyone is to notice it. A wrong tag is discovered the next time somebody searches and takes ten seconds to fix. A wrong relationship owner quietly reroutes commission, nobody says anything for a quarter, and when it surfaces it is a conversation about money between two consultants. The second one is worse despite the field looking less important.
The same logic explains why call notes are a good place to start even though notes feel significant. The error is visible immediately, to the person who was on the call, on the day. Self-correcting mistakes are cheap mistakes.
How to roll it out without a bad quarter
Turn it on for one thing, on one team, for a month. Pick something from the first column, ideally call notes, because the time saving is obvious and the errors are self-correcting.
At the end of the month, read fifty records chosen at random rather than chosen by the software. You are looking for two numbers, and the second one matters more than people expect.
How often was it wrong. The obvious one, and the one everybody measures.
How often did it decline. The one that tells you whether the confidence gate is real. A system that never declines across fifty records is either handling only trivial cases or it is guessing somewhere and not telling you. Zero declines is not a good score; it is an unanswered question.
Then widen it, one column-one field at a time. Every agency I have seen have a genuinely bad experience with AI in their CRM did the opposite: switched everything on at once, could not tell afterwards which values came from where, and ended up with a clean-up project instead of a time saving.
What good looks like after six months
The AI does the sorting and the note-taking without anybody thinking about it. A handful of cases a day come back to a person with a reason attached, and those are genuinely the cases worth a person's judgement. The records are more complete than they were before, not less, because the friction that stopped consultants writing things down has gone.
And when a client asks how you know something about a candidate, you can answer, because the record says where the value came from.
That is an unexciting description and it is the right one. Trust it to write when it can decline, when you can tell afterwards what it did, and when something sensible happens on the days it cannot run. Those three conditions are not difficult to build, and most products do not have them, which is a statement about what vendors have chosen to prioritise rather than about what the technology can do.
We publish how every decision model we use scores against this first condition, including how often each one is sure of a wrong answer, on the Decisions board. Siva explains how that testing works in is AI accurate enough to change your CRM.
Lokesh is Founder and Head of Engineering at Recruitly.



