How to tell if a vendor's AI is real
Every recruitment CRM on the market now says it has AI, so the claim has stopped carrying information. Four questions separate the products where somebody built something from the products where somebody wired a chat box to a text field, and you can ask all four on a demo call without knowing anything technical.
Two years ago AI was a differentiator in recruitment software. Today it is a line on every website and a badge on every feature list, which means it tells you nothing at all. The useful question is no longer whether a vendor has AI. It is whether anyone at that company understands what they have bought.
There is a real difference underneath the identical marketing, and it is large. It also does not show up on a demo, because a demo is built from examples where the software works. It shows up in your database somewhere around month three, when your consultants quietly start double-checking things.
These are the four questions I would ask, in the order I would ask them, and what a real answer sounds like.
Question one: what does it do when it does not know?
A language model always answers. That is the single most important thing to understand before you let one anywhere near your records, and almost nobody selling to you will volunteer it.
A model produces the most plausible continuation of whatever it has been given. Ask it something with a clear answer and you get the clear answer. Ask it something with no answer and you get the most plausible-looking one instead, in the same tone, at the same apparent confidence. There is no internal moment where it notices that the question could not be answered. Nothing in it corresponds to hesitating.
Put that inside a CRM and the consequence arrives immediately. An email comes in saying please pass this one to Sara. Your agency has a Sarah in Leeds and a Sara Collingwood in Manchester. The software assigns the record to one of them. It does not flag it, because from the model's point of view nothing unusual happened.
Now multiply that by every ambiguous email, CV and call note your agency handles in a month. What you end up with is not a database with some errors in it, which would be survivable, because errors get found. You end up with a database where the wrong records are indistinguishable from the right ones, and every consultant knows it. That is the expensive kind of broken, because it does not look broken and there is nothing to point at.
How to test it. Take a genuinely ambiguous case into the demo. Not a tidy one from their sample data. A real email with a first name that matches two of your people, or a CV where the most recent role has no dates on it. Hand it over and watch what the software puts on the record.
What a real answer sounds like. The software returns a number alongside its answer saying how certain it is, the product compares that number to a threshold, and below the threshold it writes nothing and gives the case to a recruiter. Ours draws that line at three quarters. The specific number is unremarkable and easy to copy. What matters is that a route out of answering exists at all, because building one means somebody sat down and thought about being wrong, which is the rarest thing in this market.
We go one step further and keep test cases whose entire purpose is to prove the software stays unsure about things it should be unsure about. That sounds like a strange way to spend an afternoon until you notice the failure it catches: a model quietly becoming more confident is not an improvement when the thing it is confident about is a coin toss, and without a test asserting doubt, nothing detects that.
Question two: does it do work, or does it wait to be asked?
There are two kinds of AI in recruitment software and they are worth very different amounts to you.
The first is a chat box. You type a question, it answers, you decide what to do with the answer. It is genuinely useful, it took about a fortnight to build in 2024, and every product has one. It is also the kind that gets demoed, because it is the kind you can see.
The second does a specific job without being asked. Every reply to your outreach is sorted the moment it lands, so the interested ones sit at the top of your morning instead of buried in the middle of ninety. The notes from yesterday's calls are already on the records. Something has noticed the client who has gone quiet. A job that came in as an email arrives with its fields filled from your own lists rather than from the software's imagination.
The difference in value is not subtle, and it comes down to one thing. The chat box only helps the consultants who remember it is there, which in most agencies is about a third of them, and usually the third who needed the least help. Work that happens in the normal path helps everybody, including the person having a bad week, which is exactly where your time is actually going.
How to test it. Ask them to show you what the software has done in an account overnight, with nobody logged in. A product with real AI in it has a list. A product with a chat box has an empty screen and an explanation.
Question three: can they name the models, and say what happens when one fails?
Ask which models they use. This sounds like a technical question and it is not; it is a question about whether anyone there owns the decision.
A vendor who owns their AI properly will name them, tell you what happens when one is slow or unavailable, and be faintly bored by the question, which is the real tell. People are bored by questions they have answered internally a hundred times. A vendor who does not own it will talk about their partnership with a large technology company, which is an answer about procurement rather than about engineering.
The follow-up matters more than the first answer. Models change underneath everybody. A provider improves one, adjusts one, retires a version, and the answers shift, sometimes noticeably. If a product passes your text straight out to whatever is on the other end, nobody at that company finds out the behaviour changed until a customer complains, and the customer is you.
What ought to sit between the product and the model is a layer the vendor controls. That layer chooses which model answers, runs a change against their own test cases before it reaches a customer, keeps a second model configured behind the first, and measures what every call costs. None of that is glamorous and none of it demos well, which is exactly why its presence tells you something about the company.
There is also a question of fit that almost nobody asks. Some work is about language: summarising a call, drafting an email, pulling details out of a CV. Some work is about judgement between real options: which colleague owns this, which stage this belongs in, whether this reply is interested or a bounce. Those want different tools, and using a conversational model for the second kind is where a great deal of the bad data in this industry comes from. A vendor who has noticed that distinction will say so unprompted.
Question four: can they tell you where AI touches your data?
Your clients are going to ask you a version of this, so you need to be able to ask your vendor. It is already appearing in enterprise supplier questionnaires, and the answer is now yours to provide rather than theirs.
Most vendors cannot answer it. The honest ones say they will have to check, which is fine. The revealing answer is a confident number that turns out to come from a slide somebody made last year, because a list of AI touchpoints maintained by hand is stale the week after it is written. Every new feature adds to it and nobody updates the document.
We stopped keeping ours by hand for that exact reason. A program reads every line of code we have, here and across the other systems we run, and refuses to finish if it finds anything touching AI that it cannot account for. The first time we pointed it at our own code it found more than forty places our hand-written list had missed. That is not a flattering thing to publish and it is the entire argument for counting by machine.
The reason this matters to you is practical rather than philosophical. A vendor who cannot say where their AI runs cannot tell you what happens to a candidate's CV, cannot tell you which of their suppliers sees it, and cannot tell you what quietly starts behaving differently the next time a model is updated underneath them. Three answers you will eventually need, all resting on one count nobody took.
What the answers sound like side by side
| Ask | A product with a chat box | A product with an AI estate |
|---|---|---|
| What happens when it is unsure? | It is very accurate | It reports confidence, and below a line a person gets it |
| What does it do unprompted? | It waits to be asked | Sorts replies, writes notes, fills fields, flags risk |
| Which models, and what if one fails? | Leading AI technology | Named, with a reserve behind each one |
| Where does AI touch our data? | Across the platform | A count, and how it is kept current |
| How do you know an update has not broken it? | We test thoroughly | Cases run against the live model before release |
None of the right-hand answers require you to evaluate anything technically. You are listening for whether the person can answer at all, and whether the answer is about their software or about their marketing. Twenty minutes of this separates a shortlist more reliably than a fortnight of demos.
The two things to take into every demo
One genuinely ambiguous email, and one messy CV. Not tidy ones. The kind that arrive on a Friday afternoon with half the information missing and a first name that could be two different people.
Hand them over and watch what the software does. Everything worth knowing about a vendor's AI is visible in that one moment, and none of it is visible in the part of the demo they prepared. A product that writes a confident answer to an unanswerable question has told you exactly what it will do two thousand times a month in your account.
Then ask the question that outlasts the demo: six months from now, how would we tell which values on a record were set by a person and which were inferred by the software. If there is no answer, every inferred value in your database is permanently indistinguishable from a checked one. That is the thing that eventually stops people trusting the screen, and no amount of a better model repairs it.
What this looks like a year in
Agencies who chose on the demo describe the same arc. Impressive in month one. Useful in places by month two. A quiet distrust settling in around month three, consultants double-checking, someone keeping a spreadsheet, and nobody able to point at a disaster because there wasn't one. The software was never broken. It was just never allowed to say it did not know.
Agencies who asked these four questions first describe something duller, which is the point. The AI does the sorting and the note-taking, a handful of cases a day come back to a person with a reason attached, and nobody thinks about it much. That is what working software looks like, and it is almost impossible to tell apart from the other kind at the point where you sign.
The products worth your money are the ones where somebody thought hard about being wrong. That is a much rarer thing in this market than a chat box, and it is the only part of any of this that will still matter to you in three years.
Gowri is the CEO of Recruitly. The figures here come from our own code records and were last counted on 20 September 2026.



