WhatsApp PersonalBusiness
Recruitly LogoRecruitly
Engineering

Which AI models does your recruitment CRM use?

It is a fair question, your clients are starting to ask you a version of it, and the way a vendor answers tells you more about their engineering in four minutes than an hour of demo does. Here is why it matters commercially, what a real answer sounds like, and what we would say if you asked us.

Almost nobody asks this during a sales process. The few who do are usually told something about leading AI technology and a partnership with a large technology company, which contains no information at all. The fact that this is the standard answer, given fluently, by most vendors, is itself the finding.

You do not need to understand models to ask about them. You need to notice whether anyone at the company owns the decision.

Why the answer matters to an agency

Three reasons, none of them technical, and all of them arrive eventually whether or not you asked.

Your clients are starting to ask. Where candidate data goes, and which third parties see it, is now appearing in enterprise supplier questionnaires and in the security reviews that come with a PSL. When a client asks you, the answer has to come from your CRM vendor, and if they will not give it to you, that becomes your problem to explain rather than theirs.

Models change underneath everybody. Providers improve them, adjust them and retire versions, and the answers shift when they do. If a product passes your text straight out to whatever is on the other end, the behaviour of your CRM can change on a Tuesday for reasons nobody at that company controls or even notices. You find out when something that worked for a year stops working, and nobody can tell you what changed because nothing in their code did.

AI is a running cost, not a fixed one. Every call has a price. A vendor who cannot tell you which model answered also cannot tell you what their own product costs them to run. That shows up eventually in your pricing, usually as a surprise, and usually framed as a new AI add-on rather than as a cost they failed to measure.

What a real answer sounds like

A vendor who owns their AI properly tells you three things without preparation: the names of the models they use, what happens when one of them is slow or unavailable, and how they know an update has not changed the behaviour.

They will also be slightly bored by the question. That is the tell worth watching for. People are bored by questions they have already answered internally a hundred times, and animated by questions they are hearing for the first time.

The weak answers are distinctive too, and each one points at the same underlying gap.

AskWeak answerWhat it actually means
Which models do you use?Leading AI technologyNobody there owns the decision
What happens when one is slow?That has not happened to usThere is no reserve, and they have not looked
How do you know an update has not broken it?We test thoroughlyNothing runs against the live model before release
Where does AI touch our data?Throughout the platformIt has never been counted
What does it cost you to run?It is includedUnmeasured, and repriced later

Why we use several models rather than one

We use more than one because the work is not one kind of work, and using a single model for all of it is where a great deal of the bad data in this industry comes from.

Some of what a CRM needs is about language. Summarising a call, drafting an email, pulling structured details out of a CV, writing a job description from a brief. That is generative work, the output is prose, and a conversational model is the right tool.

Some of it is about judgement between real options. Which of your colleagues owns this job. Which stage this candidate belongs in. Whether this reply is interested, not interested, out of office or a bounce. Whether this message is a request to stop contacting someone. The output there is not prose, it is one option from a list you already hold, and what the software most needs to know is how confident it should be.

Asking a conversational model to do the second kind works, right up until it does not. It will return a word, that word gets matched against your list in code, and the failures live in that matching step. The model said something reasonable, your list uses slightly different words, and the match either fails or lands on the wrong entry. The error then gets blamed on the AI, when the AI was never asked the right question.

For the judgement work we use Jev, from TypeSafe. It takes a situation and a typed question, returns one of the options you offered, and reports how sure it is. That last part is the reason we use it: a number we can compare against a threshold is what lets the software decline instead of guessing. Credit to the TypeSafe team, because that shape changed how we build these features.

For the language work we use established general models, and every one of them has a second model configured behind it, so a slow or failing provider does not stop a recruiter's afternoon.

Passing straight through, against a layer the vendor controls
Both of these are described as AI on a website, in the same words. Only the lower one can answer a client questionnaire honestly, because only the lower one knows the full list of places data leaves from.

Everything goes through a layer we built and control. That layer picks the model, records what each call cost, holds the reserve, and lets us run a change against our own cases before a customer ever sees it. None of that is visible in the product and none of it demos well, which is exactly why its presence or absence tells you something about a company's priorities.

How we know an update has not broken something

We keep a battery of test cases drawn from real recruitment situations, and we run it against the live models rather than a copy. A release does not go out on a failing battery.

The less obvious half is that some of those cases exist to prove the software stays unsure. A genuinely ambiguous input is supposed to come back below the confidence line, so the work goes to a recruiter rather than onto the record. Without a case asserting that, this particular drift is completely invisible: a model quietly becoming more confident looks like an improvement on every dashboard, and is not an improvement at all when the thing it has become confident about is a coin toss.

We found this the honest way. Two cases failed when we first ran them, both returning the right answer at low confidence. The model was being truthful and the gate was working correctly. The test was wrong, so we added a way for a case to assert "stay unsure about this" as a contract in its own right.

Where AI touches your data, and how we count it

We do not keep this list by hand. A hand-kept list of where models are used goes stale the week after it is written, because every new feature adds to it and nobody updates the document. A stale list is worse than no list, because people believe it.

Instead a program reads every line of code we have, across this product and the other systems we run, and refuses to finish if it finds anything touching AI that it cannot account for. When we first pointed it at our own code it found more than forty places the hand-written list had missed. That is not a flattering number to publish, and it is the whole argument for counting by machine rather than by memory.

The four minutes

Ask which models. Ask what happens when one is slow. Ask what gets run against the live model before a release. Ask where AI touches your data.

You are not evaluating the answers technically, and you do not need to. You are finding out whether the person in front of you can answer at all, because a vendor who cannot name their own models has outsourced a decision that your clients will eventually hold you responsible for, and a vendor who has never counted their AI touchpoints cannot tell you what happens to a candidate's CV.

Four minutes, four questions, and they are almost impossible to answer well without having done the work.


Lokesh is Founder and Head of Engineering at Recruitly.

ai-recruitmentai-modelsvendor-evaluationdata-protection

The product this came out of

Nineteen modules on one record: sourcing, screening, campaigns, calls, e-signature and billing, without a second system to keep in step. Free to start, no card, no call.