WhatsApp PersonalBusiness
Recruitly LogoRecruitly
Comparisons

Which recruitment CRM has the best AI?

Every product in this market now claims AI, so the claim has stopped carrying information. Four tests separate the ones who built something from the ones who added a chat box, you can run all four yourself, and none of them requires you to understand how any of it works.

I build the AI side of our product, which means I spend a lot of time looking at what everyone else has shipped. The honest summary is that the gap between the best and the worst in this market is enormous and almost completely invisible from outside, because every vendor describes theirs with the same six words.

Comparison sites do not help, because they compare feature lists and every serious product now lists the same features. What separates them is not which features exist. It is what the software does at the moments the demo never shows you.

These are the four things I would test if I were buying, in the order I would test them.

Test one: does it ever say it does not know?

A language model always produces an answer. That is what it is for, and it is the single most important thing to understand before you let one near your database.

Give it an email that says pass this one to Sara, when your agency has a Sarah and a Sara Collingwood. A model asked to pick will pick. It will not hesitate, and the record it writes will look precisely as trustworthy as a record that is correct. Nothing on the screen will say this one was a guess.

Two ways of handling an uncertain answer
Both of these look identical on a demo, because demo examples are never ambiguous. The difference only appears in your data, about a month in, and by then it is thousands of records deep.

Multiply that across every ambiguous email, CV and call note in a month. The result is not a database with some errors in it, which would be manageable, because errors get found and fixed. The result is a database where the wrong entries are indistinguishable from the right ones, which is the expensive kind of broken because the only rational response is to distrust all of it.

There is a counter-intuitive consequence worth sitting with. A system that is wrong more often but tells you when it is unsure is more valuable than a system that is wrong less often and never flags anything. The first one you check selectively. The second one you check entirely, which costs more than the AI ever saved.

How to test it. Ask the vendor to show you their software handling a genuinely ambiguous case, and bring your own rather than using theirs. If the software writes an answer regardless, you have learned exactly what it will do two thousand times a month in your account.

Ours answers with a number for how certain it is, and there is a line at three quarters. Below it nothing is written and the case goes back to a recruiter. We also keep test cases whose only purpose is to prove it stays unsure about things it should be unsure about, which sounds like an odd use of an afternoon until you have seen the alternative sitting in somebody's database.

Test two: is the AI doing work, or making conversation?

There are two kinds of AI in recruitment software and they are worth quite different amounts to you.

The first is a chat box. You type a question, it answers, and you decide what to do with the answer. It is useful, it took about a fortnight to build in 2024, and everybody has one. It is also the kind that demos well, because you can watch it happen.

The second does a specific job without being asked. It sorts every reply to your outreach so the interested ones sit at the top of your morning rather than in the middle of ninety. It writes the notes from a call and puts them on the record. It reads the tone of a meeting. It fills fields on a job from your own lists rather than from its imagination. It flags the client who has gone quiet.

Nobody has to remember to use any of that, which is the whole point. The thing recruiters reliably stop doing under pressure is the optional step, so an optional tool is most absent in exactly the weeks you most needed it. A product with one chat box and no working agents has AI in the sense that a car has a radio.

How to test it. Ask to see what the software did overnight in an account where nobody was logged in. A list is a good answer. An explanation is not.

Test three: can they name the models, and say what happens when one fails?

Ask which models they use. A vendor who owns their AI properly will name them, tell you what happens when one is slow or down, and find the question slightly dull. A vendor who does not will talk about their partnership with a large technology company, which is an answer about procurement rather than engineering.

The follow-up matters more. Models change underneath everybody, sometimes overnight, and the answers change with them. If a product sends your text straight out to whatever is on the other end, nobody there knows when the behaviour shifted until a customer complains, because nothing in their own code changed.

What ought to sit in between is a layer the vendor controls: one that tests a model against their own cases before it reaches you, keeps a second one in reserve, swaps it when a better one arrives, and measures what each call costs. Building that is unglamorous, demos badly, and is the difference between a product that survives a provider's bad Tuesday and one that does not.

There is a fit question underneath this that almost nobody asks. Some work is about language: summarising a call, drafting an email, pulling details from a CV. Some work is about judgement between real options: which colleague owns this, which stage this belongs in, whether this reply is interested or a bounce. Those want different tools, and using a conversational model for the second kind is where a great deal of the bad data in this industry comes from. A vendor who has noticed the distinction will raise it themselves.

Test four: can they tell you where AI touches your data?

Your clients are going to ask you a version of this, so you need to be able to ask your vendor. It is already turning up in enterprise supplier questionnaires and in PSL renewals.

Most vendors cannot answer it. The honest ones say they will have to check. The revealing answer is a confident number that turns out to be a slide from last year, because a list maintained by hand goes stale the week after it is written: every new feature adds to it and nobody updates the document.

We stopped keeping ours by hand for that exact reason. A program reads every line of code we have, here and across the other systems we run, and refuses to finish if it finds anything touching AI it cannot account for. The first time we pointed it at our own code it found more than forty places the hand-written list had missed, which is not a flattering number and is the whole argument for counting by machine.

What good and bad look like side by side

AskA product with a chat boxA product with an AI estate
What happens when it is unsure?It is very accurateIt reports confidence and hands low ones to a person
What does it do unprompted?Waits to be askedSorts replies, writes notes, fills fields, flags risk
Which model is it?Leading AI technologyNamed, with a reserve and a measured cost
Where does AI touch our data?Across the platformA count, and how it is kept current
How do you know an update has not broken it?We test thoroughlyCases run against the live model before release

What I would actually do

Take one ambiguous email and one messy CV into every demo you attend. Not tidy ones. The kind that arrive on a Friday afternoon with half the information missing and a name that could be two people.

Watch what the software does with them. Everything worth knowing about a vendor's AI is visible in that one moment, and none of it is visible in the part of the demo they prepared for you.

Then ignore the feature comparison entirely. Every serious product will match on features within a year anyway, because features are the easy part to copy. What does not get copied quickly is the confidence gate, the fallback, the measurement and the counting, because none of those show up in a sales conversation and all of them take real engineering to build.

The products worth your money are the ones where somebody has thought hard about being wrong. That is a much rarer thing in this market than a chat box, and it is the only part of it that will still matter to you in three years.


Lokesh is Founder and Head of Engineering at Recruitly.

ai-recruitmentrecruitment-crmcomparisonai-features

The product this came out of

Nineteen modules on one record: sourcing, screening, campaigns, calls, e-signature and billing, without a second system to keep in step. Free to start, no card, no call.