NewRecruitly + WhatsApp — message from your CRM
Recruitly LogoRecruitly
Engineering

Which AI models does your recruitment CRM use?

It is a fair question, your clients will eventually ask you a version of it, and the way a vendor answers tells you more about their engineering than an hour of demo does.

Ask AI about this

ChatGPT
Perplexity
Grok
Claude
Google AI

Almost nobody asks this during a sales process, and the few who do are usually told something about leading AI technology and a partnership with a large technology company. That answer contains no information, and the fact that it is the standard answer is itself informative.

Here is why the question matters commercially, what a real answer sounds like, and what we would say if you asked us.

Why the answer matters to an agency

Three reasons, none of them technical.

The first is that your clients are increasingly asking you where their candidates' data goes. Enterprise clients have started putting it in supplier questionnaires. If you cannot answer because your CRM vendor will not, that is your problem rather than theirs.

The second is that models change underneath everybody. Providers improve them, adjust them, and retire versions, and when that happens the answers shift. If a product passes your text straight out to whatever is on the other end, the behaviour of your CRM can change on a Tuesday for reasons nobody there controls or noticed.

The third is cost. AI is a running cost, not a fixed one, and a vendor who cannot tell you which model is answering also cannot tell you what it costs them. That shows up eventually in your pricing, usually as a surprise.

What a real answer sounds like

A vendor who owns their AI properly will tell you three things without preparation: the names of the models they use, what happens when one of them is slow or unavailable, and how they know an update has not changed the behaviour.

They will also be slightly bored by the question, which is the tell. People are bored by questions they have answered internally many times.

Passing straight through, against a layer you control
Both of these are described as AI on a website. Only the lower one can tell you what changed when the answers start looking different.

What we use, and why more than one

We use several models rather than one, because they are not interchangeable and the differences matter.

Some work is about language: summarising a call, drafting an email, pulling structured details out of a CV. Some work is about judgement between real options: which of your colleagues owns this, which stage this belongs in, whether this reply is interested or a bounce. Those two kinds of work want different tools, and using a conversational model for the second kind is where a great deal of bad data in this industry comes from.

For the judgement work we use Jev, from TypeSafe. It takes a situation and a typed question, returns a typed answer, and reports how sure it is. That last part is the reason we use it, because a number we can compare to a threshold is what lets the software decline instead of guessing.

For the language work we use established general models, and we keep a second one configured behind each use so that a slow or failing provider does not stop a recruiter's afternoon.

Everything goes through a layer we built and control. That layer picks the model, records what each call cost, keeps the reserve, and lets us run a change against our own test cases before any customer sees it.

How we know an update has not broken something

We keep a battery of test cases, drawn from real recruitment situations, and we run them against the live models rather than against a copy. A release does not go out on a red battery.

The less obvious half is that some of those cases exist to prove the software stays unsure. A genuinely ambiguous input is supposed to come back below the confidence line so the work goes to a recruiter. A model quietly becoming more confident is not an improvement when the thing it is confident about is a coin toss, and without cases asserting that, this particular drift is invisible.

Where AI touches your data, and how we count it

We do not keep this list by hand. A hand-kept list of where models are used is stale the week after it is written, and a stale list is worse than none because people believe it.

Instead a program reads every line of code we have, across this product and the other systems we run, and refuses to finish if it finds anything touching AI that it cannot account for. When we first ran it against our own code it found more than forty places the hand-written list had missed, which is exactly why we stopped keeping one.

The four questions, and what the answers tell you

AskWeak answerWhat it means
Which models do you use?Leading AI technologyNobody there owns the decision
What happens when one is slow?That has not happenedThere is no reserve
How do you know an update has not broken it?We test thoroughlyNo cases run against the live model
Where does AI touch our data?Throughout the platformIt has never been counted

Four questions, four minutes. You do not need to evaluate the answers technically. You need to notice whether the person can answer at all, because a vendor who cannot name their own models has outsourced a decision that your clients will eventually hold you responsible for.


Lokesh is Founder and Head of Engineering at Recruitly.

ai-recruitmentai-modelsvendor-evaluationdata-protection

The product this came out of

Nineteen modules on one record: sourcing, screening, campaigns, calls, e-signature and billing, without a second system to keep in step. Free to start, no card, no call.