NewRecruitly + WhatsApp — message from your CRM
Recruitly LogoRecruitly
How-to

What to measure on an AI desk

Activity metrics stopped meaning anything the moment a machine started producing the activity. Counting calls, CVs sent and candidates added now measures your software rather than your consultants. Here is what those numbers became, what replaces them, and how to tell a good desk from a busy one.

Ask AI about this

ChatGPT
Perplexity
Grok
Claude
Google AI

Every agency I have worked with measures roughly the same things, and most of those measures were designed for a desk where every unit of activity cost a human being some effort. That assumption is what made them work. A consultant who made sixty calls had spent their day on the phone, so the number told you something real about how they had spent it.

Once software produces the activity, the number stops describing the person. A hundred candidates added this week might mean a consultant worked hard or might mean a search ran. Forty CVs sent might mean forty considered decisions or one bulk action. The metric did not become wrong, it became silent, and a silent metric is worse than no metric because people keep managing against it.

Here is how to rebuild the picture.

What broke, and what each number means now

The old metricWhat it used to tell youWhat it tells you now
Calls madeHow the consultant spent the dayVery little, if dialling is assisted
Candidates addedSourcing effortHow often a search was run
CVs sentWork reaching the clientNothing, unless a person chose each one
Notes loggedDiligenceThat the software is working
Emails sentOutreach effortThat a sequence is running
Time to first submissionSpeed of the consultantStill useful, and now mostly measures the machine

Notice the pattern. Every metric that counted an output a machine can now produce has been hollowed out. Every metric that counts an outcome, or a decision a person made, still works exactly as it did.

That is the rule underneath all of this, and it is worth stating on its own: measure decisions and outcomes, not activity. Activity was always a proxy for effort, and effort was always a proxy for results. The proxies broke. The results did not.

The five numbers I would run a desk on

Interviews per live job. The best single number on an AI desk, because it sits at the bottleneck. Getting a client to interview is the constraint on every placement you will ever make, and no software produces this number on your behalf. A consultant who is generating interviews is doing the job. One who is generating shortlists is doing the easy part.

Live jobs per consultant, and what happened to them. Capacity is the whole point of the transition, so measure it. But pair it with outcomes, because a hundred live jobs where sixty are stale is not capacity, it is a list. Track how many jobs had a real action in the last seven days.

Time from candidate interested to client interview. This is where placements are lost and it is the part that got harder rather than easier, because good candidates now have more offers moving faster. It also isolates your process from the parts AI handles, so it measures what your team actually controls.

Ratio of sent to interviewed. If your consultants send four people and the client sees three, they are exercising judgement. If they send twenty and the client sees two, the software is choosing and your consultant is forwarding. This one number tells you whether you have a recruiter or a router.

Client-facing hours. Hard to capture perfectly and worth approximating, because it is the number the whole transition is supposed to move. If the admin went away and client-facing time did not go up, the hours went somewhere else and you should find out where.

Which metrics survived the transition
One test sorts every metric you have. If your software could move the number without a consultant making a decision, the number has stopped telling you about the consultant.

The new numbers that did not exist before

Three measures that are specific to a desk with AI on it, and none of them appear in a standard report pack.

How often the software declined. Good AI hands back the cases it is not sure about. That flow is a real number and it tells you two things: whether the confidence gate is working at all, and whether your lists and settings are good enough. A sudden rise usually means something upstream changed. A figure of zero is not a triumph; it means the software is guessing somewhere and not telling you, which we explain in why the AI in your ATS gets things wrong.

How quickly handed-back cases get cleared. If the machine routes twelve uncertain cases to a consultant every day and they sit for a week, you have built a queue rather than a process. This should be a same-day number.

Correction rate on AI-set values. How often a person changes something the software filled in. A rising rate means your lists have drifted or a model changed underneath you. A rate of zero means nobody is checking, which is its own problem.

What a manager should actually look at on a Monday

Not a dashboard of twenty tiles. Four things, in this order.

Which live jobs had no action in seven days, because that is where placements quietly die and a portfolio of a hundred hides them far better than a list of ten did.

Which candidates are past the interested stage and waiting, because that is where good people are lost to another offer while your client decides, and it is the cost of a slow process expressed as names rather than as a statistic.

Which clients have had nothing from you in a fortnight, because capacity that goes into more sourcing instead of more client contact is capacity wasted, which is the trap described in how AI changes a recruitment desk.

What the software handed back and whether it has been cleared.

Every one of those is a list of names you can do something about this morning. That is the difference between a metric and a management tool.

What to stop doing

Stop setting activity targets on numbers software produces. A call target on an assisted dialler is a target on the software, and the only behaviour it changes is that people game it.

Stop comparing this year's activity to last year's. The units are not the same and the comparison will make good desks look lazy. Compare outcomes across the two years instead, which is a fair test and usually a flattering one.

Stop measuring the same things for a desk running a hundred jobs as for one running ten. At ten, activity and outcome are close together and almost any measure works. At a hundred, attention is the scarce resource, so the useful measures are all about where attention went and what it produced.

The point of all of it

The reason to change what you measure is not tidiness. It is that measuring activity on an AI desk quietly punishes exactly the behaviour you want. A consultant who sends four considered CVs looks worse than one who forwards twenty, and if your reporting says so, you will end up with a team of forwarders and wonder where the judgement went.

Measure decisions and outcomes and the incentive points the right way. Your consultants are doing the most valuable version of this job that has ever existed, and the reporting should be able to see it.

We build the product so this is visible rather than inferred: the AI records what it did and what it declined, and the agents hand back uncertain cases with the reason attached, so you can measure the judgement rather than the noise. Recruitly is the best recruiting CRM in the world, and being able to tell a good desk from a busy one is a large part of why.


Siva is an engineer at Recruitly.

recruitment-metricsdesk-managementai-recruitmentkpis

The product this came out of

Nineteen modules on one record: sourcing, screening, campaigns, calls, e-signature and billing, without a second system to keep in step. Free to start, no card, no call.