NewRecruitly + WhatsApp — message from your CRM
Recruitly LogoRecruitly
Product

Does AI candidate matching actually work

It works at one job and fails at a different one, and most vendors sell the second as if it were the first. A match score measures how alike two documents are. Placing someone needs to know things that are in neither document. Here is exactly what the number is made of, from someone who built one.

Ask AI about this

ChatGPT
Perplexity
Grok
Claude
Google AI

Agency owners ask me this in two forms. The optimistic form is whether the match score is good enough to trust, so a consultant can send the top five without reading the rest. The sceptical form is whether it is doing anything at all, since the top of the list keeps producing people the client turns down. Both questions come from the same misunderstanding about what the score is, and once that is clear, the answer to both is the same.

I built the matching that runs inside Recruitly, so I can say precisely what it does. It is useful, and it is useful for a narrower job than the one it gets credited with. The gap between those two is where agencies lose money.

What a match score is actually made of

Under every AI matching product I have looked at, including ours, the mechanism is the same. The text of the job and the text of the candidate are each turned into a long list of numbers, an embedding, that captures what the text is about. The score is how close those two lists sit to each other. Close means the documents talk about the same things in the same way. That is the whole calculation. There is no model of the person in there, only a model of the words about the person.

How a match score is computed
Two documents in, one distance out. The score is a statement about how alike the writing is, and everything the score is later asked to do rests on that.

That has a consequence worth stating plainly. A match score can only be as good as the two documents are honest, specific and different from each other. A vague job spec matches everything a little. A polished CV matches a lot of things well. And a pile of CVs that were all written by the same AI to the same job ad will sit close together in the same corner of the space, with scores that cluster so tightly the ranking between them is noise.

The job it does well

Ordering. Given four hundred applications for one role, the score puts the ones that are obviously about the right thing near the top and the ones that are obviously about the wrong thing near the bottom, in a second, and it does that better and faster than a tired person on a Thursday afternoon. On Recruitly this shows up in three places: the AI review on each application, which puts a skills score, a language score and an overall assessment on it; the AI Matches tab on a job, which runs the job against your own database; and the per-job AI search, which runs it against the wider index.

The second thing it does well is recall. It finds the person you would have missed because their title was odd or their CV led with the wrong job, because it reads the whole document rather than the first line. A consultant scanning four hundred CVs in six seconds each is matching on titles. The score is matching on everything, and that is a genuine edge at the top of a large pile.

What the score does to a large pile
The score moves the good ones towards the front without knowing which ones they are, which is why it is a reading order.

The job it cannot do

Everything a placement depends on that is not written down. Whether the person is actually open to moving, or updated their profile out of boredom. Their notice period, their real salary floor, the thing they will not relocate for. How they talk, which is the first thing the client will judge. What the client actually wants, which is frequently different from what the client wrote in the spec, and which the consultant learned on a call the model was not on. None of that is in either document, so none of it is in the score, and a score that looks precise to two decimal places is precise about the wrong thing.

Then there is the problem that arrived in the last two years. Every CV is now written to the job ad by an AI that read the job ad. That means the CV and the spec are close by construction, and the score, which measures closeness, rises for everyone. In practice the scores on a pile of AI-written applications bunch up near the top, and the difference between an 88 and an 84 stops meaning anything. The document has lost the signal the score was built to read, and the score cannot tell you that has happened. It just keeps producing numbers.

Score distributions before and after AI-written CVs
Illustrative shapes, and the second one is what applicant scores look like on most desks now. The documents the model reads have converged, and the score reports that convergence as excellence.

Where the judgement lives, and why we kept it there

This is the design decision that separates a matching tool you can live with from one that quietly costs you clients. On Recruitly the relevancy score and the AI review order the pile, and advancing or rejecting anyone is a separate action a person takes. The thing that records a judgement about the person is the scorecard, and the scorecard is a different object: a set of criteria your agency chose, answered yes, partial or no, with a comment and a name against each answer.

The AI drafts that scorecard from the CV and the spec, because the reading is the slow part and a model is good at reading. Then the consultant changes whatever is wrong, after the call, which is where the things the document does not contain get found out. So the score orders, the model drafts, and a person decides, and the person's name is on the decision. We wrote about why our AI asks before it acts in June, and this is the same rule applied to matching: the model is allowed to be fast at reading and never allowed to be final about a person.

How to use a match score without being used by it

Matching works if you hold it to the job it can do. Four habits keep it there.

  1. Read it as a reading order. The top of the list is where to start, and position five is still worth opening, because the good one with the odd CV is sitting there.
  2. Cap the pile before you score it. A score on four hundred is a sort. A score on forty, sourced to what the client can decide this week, is a sort you will actually finish reading.
  3. Write the spec the model can read. Vague specs produce flat scores. A spec with the three things the client will actually reject on, written as plainly as the consultant would say them on the phone, produces a ranking that means something.
  4. Judge it by interview-to-offer. If the client's yes rate on the people you send goes up after matching went in, it is working. If the scores are high and the client keeps saying no, the score is measuring the polish, and the scorecard is where the real reading has to happen.

So it works, for ordering a pile and finding the one you would have missed, and we would not have spent the engineering on it otherwise. It stops working the moment anyone treats a distance between two documents as an opinion about a person. Keep those two things apart and a match score is one of the more useful numbers on a desk. Let them blur and you will send the client a beautifully ranked list of people who all read the same job ad.

Written for agency owners who were shown a 94 and wondered why the client said no.

ai-matchingcandidate-matchingai-recruitmentscorecardrecruitment-crm

The product this came out of

Nineteen modules on one record: sourcing, screening, campaigns, calls, e-signature and billing, without a second system to keep in step. Free to start, no card, no call.