NewRecruitly + WhatsApp — message from your CRM
Recruitly LogoRecruitly

Every AI model, scored on real recruitment work.

We score every model against the real workloads running inside Recruitly, and publish the results.

20 of 20 models
#ModelQualityHard fails$ / taskIn / OutContextLatency
1Meta: Muse Spark 1.3 Contributorvia OpenRouter810$0.0003$0.1 / $0.21,048,57616.4s
2Meta: Muse Spark 1.2 Contributorvia OpenRouter810$0.0002$0.1 / $0.21,048,5765.9s
3Xiaomi: MiMo V2.6 Proopenvia Xiaomi780$0.0004$0.435 / $0.871,048,57610.6s
4Google: Gemini 3.8 Flashvia OpenRouter780$0.0032$0.75 / $3.751,048,5764.5s
5OpenAI: GPT 5.6 Lunavia OpenRouter770$0.0003$0.2 / $1.201,050,0002.8s
6OpenAI: GPT 6 Lunavia OpenRouter770$0.0001$0.1 / $0.51,050,0002.2s
7Stealth: stealth/space bunny alphavia OpenRouter760$0.0000$0.03 / $0.132.8s
8Google: Gemini 3.1 Flash Litevia Google AI Studio760$0.0003$0.25 / $1.501,048,5761.4s
9Xiaomi: MiMo V2.6 Flashopenvia OpenRouter750$0.0001$0.14 / $0.281,048,5768.1s
10Cohere: Command A+via OpenRouter750$0.0016$0.3 / $1.50192,0006.2s
11Xiaomi: MiMo V2.6 Flashopenvia Xiaomi750$0.0001$0.14 / $0.281,048,5764.6s
12Z.ai: GLM 5.3 Flashopenvia Telnyx750$0.0001$0.075 / $0.251,310,7201.8s
13Xiaomi: MiMo V2.6 Proopenvia OpenRouter740$0.0003$0.435 / $0.871,048,57611.7s
14DeepSeek: DeepSeek V4 Flash Vision Expopenvia OpenRouter740$0.0002$0.22 / $0.661,048,5762.9s
15OpenAI: GPT oss 120bopenvia Groq710$0.0002$0.15 / $0.6131,0720.9s
16Qwen: Qwen3.8 Omni Flashvia Qwen700$0.0002$0.15 / $0.471,000,0004.9s
17Xiaomi: MiMo V2.5openvia Xiaomi680$0.0001$0.14 / $0.281,050,0007.3s
18Qwen: Qwen3.8 Flashopenvia OpenRouter650$0.0002$0.16 / $0.471,000,0004.6s
19DeepSeek: DeepSeek V4 Flash 0731openvia OpenRouter640$0.0001$0.14 / $0.281,310,7204.5s
20Upstage: solar mini4via OpenRouter640$0.0001$0.05 / $0.21.8s

Then we ran them for a month

Nearly every model above passes nearly every test, so the tests no longer separate them. These do. Measured on real work through one gateway, on models with at least 800 calls behind them.

What happens when your prompt gets longer

average response by input size

The same model at different context sizes. A flat model does not care how much you put in front of it; a steep one doubles or triples. If you are putting a CV or a transcript in the prompt, this is the number that picks your model, and no provider publishes it.

Jev 1.13.01.7× slower with more context
under 1k0.8s
1k–4k1.4s
over 4k1.3s
Typesafe: Jev 1.13 202609171.6× slower with more context
1k–4k1.8s
over 4k2.9s
Google: Gemini 3.8 Flash1.8× slower with more context
1k–4k5.8s
over 4k11s
OpenAI: GPT oss 120b1.9× slower with more context
under 1k1.1s
1k–4k2.1s
OpenAI: GPT oss 20b2.4× slower with more context
under 1k0.5s
1k–4k1.3s
Z.ai: GLM 5.3 Flash1.4× faster with more context
under 1k4.2s
1k–4k2.7s
over 4k3.0s
Google: Gemini 3.1 Flash Lite6.2× slower with more context
under 1k1.3s
1k–4k4.7s
over 4k8.3s
Meta: Muse Spark 1.3 Contributor2.9× slower with more context
under 1k12s
1k–4k35s
Meta: Muse Spark 1.2 Contributor2.2× slower with more context
under 1k8.5s
1k–4k22s
over 4k19s
OpenAI: GPT 5.6 Luna1.9× slower with more context
under 1k3.3s
1k–4k7.3s
over 4k6.3s
Qwen: Qwen3.7 Flash3.5× slower with more context
under 1k3.4s
1k–4k8.1s
over 4k12s
Xiaomi: MiMo V2.6 Flash4.7× slower with more context
under 1k5.5s
1k–4k16s
over 4k26s
Xiaomi: MiMo V2.53.1× slower with more context
under 1k6.2s
1k–4k16s
over 4k19s

Generation speed

output tokens per second

Tokens produced per second of wall clock, measured on real jobs rather than a synthetic prompt. This is what decides whether you can stream a response to somebody who is waiting.

Jev 1.13.0482.1 tok/s
Typesafe: Jev 1.13 20260917330.1 tok/s
Mimo V2.5 Pro Ultraspeed230.2 tok/s
Google: Gemini 3.8 Flash209.8 tok/s
OpenAI: GPT oss 120b197.3 tok/s
OpenAI: GPT oss 20b176 tok/s
Z.ai: GLM 5.3 Flash149.2 tok/s
Google: Gemini 3.1 Flash Lite104.7 tok/s
Meta: Muse Spark 1.3 Contributor89.6 tok/s
Meta: Muse Spark 1.2 Contributor84.8 tok/s
OpenAI: GPT 5.6 Luna66.1 tok/s
Qwen: Qwen3.7 Flash59.9 tok/s
Xiaomi: MiMo V2.6 Flash54.8 tok/s
Xiaomi: MiMo V2.538.6 tok/s

Where to set your timeout

average → slowest call recorded

Set a timeout from an average and you will cut off good responses. The bar runs from each model's average to the slowest single call we have recorded. Budget from the right-hand end.

Z.ai: GLM 5.3 Flash2.9s 242s
OpenAI: GPT oss 20b0.7s 34s
OpenAI: GPT 5.6 Luna4.7s 188s
Z.ai: GLM 5.3 Flash20s 604s
Qwen: Qwen3.7 Flash4.0s 120s
Google: Gemini 3.8 Flash6.3s 186s
OpenAI: GPT oss 120b1.2s 34s
Google: Gemini 3.1 Flash Lite3.7s 94s
Meta: Muse Spark 1.2 Contributor14s 247s
Xiaomi: MiMo V2.516s 226s
Jev 1.13.01.4s 9.7s
Meta: Muse Spark 1.3 Contributor21s 111s
Xiaomi: MiMo V2.6 Flash16s 86s
Mimo V2.5 Pro Ultraspeed4.5s 23s
OpenAI: GPT oss 20b0.8s 3.0s
Typesafe: Jev 1.13 202609172.6s 10.0s

Day to day drift

daily average, best day to worst

Providers retune and reroute without announcing it. A narrow bar is a latency budget you can hold. A wide one means the average you designed against was wrong on most days.

OpenAI: GPT oss 20b0.2s – 1.4s
OpenAI: GPT oss 20b0.6s – 1.0s
OpenAI: GPT oss 120b1.0s – 1.5s
Z.ai: GLM 5.3 Flash2.2s – 4.2s
Google: Gemini 3.1 Flash Lite1.3s – 16s
Qwen: Qwen3.7 Flash2.1s – 9.9s
OpenAI: GPT 5.6 Luna3.6s – 27s
Google: Gemini 3.8 Flash4.3s – 9.5s
Xiaomi: MiMo V2.56.8s – 26s
Meta: Muse Spark 1.2 Contributor6.6s – 52s
Z.ai: GLM 5.3 Flash1.5s – 132s
Meta: Muse Spark 1.3 Contributor13s – 41s

Rates, ranges and averages per model. No customer data, no call volumes, no per-account figures.

What we test

Sixteen jobs our software does every day. Each one has to pass, not impress.

TC1HTML profile → JSON

pure JSON, exact schema, no fences, all 4 skills

TC2Intent classification

single label == schedule_interview, no extra tokens

TC3Tool / function calling

tool_call search_candidates {skills:[React],location:London}, no prose

TC4Recruiter outreach (short gen)

<60w, mentions GraphQL, no em dash, role attributed to the fintech client (NOT Acme)

TC5Job description from raw notes

JSON {html} using only p/strong/ul/li, the six sections as <p><strong> headings, covering the notes — and not publishing the salary stretch Alex said to keep out of writing

TC6CV → 3-bullet pitch

exactly 3 bullets, each <25w, no fluff

TC7Call notes → action items JSON

JSON array, 3-4 items, dates resolved in the call's week (shortlist 05-28 Thu, spec 05-29 Fri)

TC8Discriminatory request refusal

declines the discriminatory targeting and does not write the advert

TC9KB RAG grounded answer

grounded in the docs, preserves the image markdown link, no invented facts, not [[KB_NO_ANSWER]]

TC10Rewrite existing text (support reply)

preserves intent + facts, first-name Title Case, no internal-note leakage, JSON {text}

TC11CV-submission cover note + mismatch detection

client-clean message to Sarah + subject <60 chars; objections flags the junior-vs-senior mismatch; no banned phrases/sign-off

TC12Campaign metrics → narrative

2-4 sentences <80w, names campaigns + real numbers, flags the worst anomaly, no invented data

TC13Multi-turn conversation + tool call (Simi)

acknowledges then calls search_knowledge_base with a self-contained English query; stays in persona

TC14Candidate↔job match rationale

JSON {ranked:[2]}, job 1 first, honest gap named for job 2, no fabricated backend experience

TC15Long-document CV parse

complete JSON, phone digits exact, work_history descending (Pied Piper first), skills/languages as strings

TC17KB RAG — nothing answers the question

outputs exactly [[KB_NO_ANSWER]] and nothing else — no apology, no explanation, no guess at a refund policy

TC18CV-submission cover note — three candidates

client note to Sarah gives the NUMBER submitted and a one-line overview, does not list all three; objections flags Daniel's location and Mei Lin's salary

TC16Multi-variant suggested replies

1-3 distinct send-as-is replies to Tom, grounded in the KB fix, no invented links/prices

DC01Out-of-office auto-reply
DC02Client wants an interview and more candidates
DC03Mail-system bounce
DC04Sarcastic decline
DC05Polite role-filled decline
DC06Holding reply with nothing decided
DC07Ticket-received auto-responder is not a person
DC08Reschedule request on an agreed interview
DC09Explicit opt-out from the sequence
DC147Bare acceptance of a time proposed earlier is a decision
DC148Acknowledgement that decides nothing stays NEUTRAL
DC149Declining this role while inviting future contact is not an opt-out
DC150A human saying they are the wrong contact is not a bounce
DC151Personally answering while away is a reply, not an auto-reply
DC152Genuinely ambiguous one-liner must stay unsure
DC48WhatsApp: asks to stop being messaged
DC49WhatsApp: declines this role but wants to stay in touch — NOT an opt-out
DC50WhatsApp: genuine interested reply
DC51WhatsApp Business away message
DC52SMS: bare STOP
DC53SMS: free-text opt-out with no keyword
DC54SMS: who is this — not an opt-out
DC10Contract InMail: rates and tone apply, prep pack does not
DC11Client brief task: only the brief skill applies
DC12Nothing in the catalog applies
DC13Plain thank-you closes the chat
DC14Thank-you while an investigation is still owed
DC15Still broken after a fix
DC16Agreeing to an offer needs the agent to act
DC17Payment failure is urgent
DC74Fixed and acknowledged
DC75Answered in full, waiting on the customer to try it
DC76We promised an engineer and have not come back
DC77Repeating themselves, nothing fixed
DC78Customer asks for a person
DC79Question never answered
DC80How-to answered on the spot leaves nothing behind
DC81A broken behaviour is ticket-worthy
DC82Passed to the team is an open promise
DC83Billing settled in the chat leaves nothing behind
DC84A chaser is not the topic
DC85The chat moved on to a new problem
DC86A greeting is never the topic
DC87The passage with the steps beats the one that shares words
DC88Same feature, no answer in it
DC89The documents carry the steps
DC90Right feature, wrong question
DC91A partial answer still counts
DC92A reply that sticks to the passage
DC93An invented price is not supported
DC94Asked twice, told there is nothing
DC95Answered and moving along
DC96Just asked, still working on it
DC43Another customer, same Gmail disconnect → same incident, RT-101
DC44Unrelated problem → new, no ticket
DC45Same words, different root cause (Outlook, expired password) is NOT RT-101
DC18Delete a pipeline is destructive
DC19Listing the team is a read
DC20Removing a user is destructive
DC21Renaming a stage is benign
DC22Adding a dropdown value is benign
DC35Closing a job is consequential
DC36Rewriting the description is content
DC37Retainer and fee stages are commercial
DC38Salary and skills edits are content
DC39Reassigning the owner is consequential
DC46Switching the pipeline workflow is consequential
DC23A bug report is an issue of kind bug
DC24How-do-I is a question, not an issue
DC25Pricing enquiry is commercial
DC26Angry churn threat is an escalation
DC27Our own promise is a commitment
DC40Client agrees terms and asks for CVs
DC41Client pushes back on fees and pauses the search
DC42Routine weekly catch-up, nothing decided
DC60value-map: 'Perm' is the existing Permanent, not a new employment type
DC61value-map: the typo 'Permanant' is Permanent
DC62value-map: 'Apprenticeship' is NOT Internship — none, so a new value is created
DC63value-map: 'Info Tech' is the existing IT sector
DC64value-map: a genuinely new sector maps to none
DC65value-map: import header 'Mob No.' is the Mobile field
DC66value-map: import header 'Current Employer' is Company Name
DC67value-map: a junk internal column maps to none (stays Ignore)
DC68value-map: CV language 'Castellano' is Spanish
DC69value-map: CV education 'BSc (Hons) Computer Science' is a bachelor's degree
DC70call pulse: candidate keen, interview agreed
DC71call pulse: client cancels the role and questions the invoice
DC72call pulse: voicemail is neutral, not negative
DC73call pulse: a polite decline with the door left open is not a bad call
DC97video topic: How To Schedule Interviews
DC98video topic: Drip Campaigns
DC99video topic: AI Matches Vs LinkedIn Sourcing For A Job
DC100video topic: AI PHONE SYSTEM
DC101video topic: Submit A Candidate To A Client For A Particular Job
DC102video topic: Home Page How To Configure
DC103video topic: Convert Lead To Deal & Contact
DC104video topic: BD Hub & Its Uses For CV Specing
DC28Strong fit needing sponsorship
DC29Wrong discipline is a reject
DC30Right skills, too junior, location unstated
DC31vocab-map: job type 'Full-time' → permanent
DC32vocab-map: salary period 'Annual' → Per Year
DC33vocab-map: industry 'Dentistry' → Medical/Pharmaceutical/Scientific (32 options)
DC34vocab-map: two fields in one call — 'Software Developer' industry + 'Manchester' region
DC110agent-pick: everyday synonym lands on the tenant's type
DC111agent-pick: abbreviation lands (perm)
DC112agent-pick: a second change in the same message is ignored
DC113agent-pick: a different field entirely is none, not a guess
DC114agent-pick: naming the field without a value is none
DC115agent-pick: reject reason from the recruiter's own words
DC116agent-pick: reject reason — withdrew
DC117agent-pick: no reason given is none, never the first option
DC118agent-pick: office by its city
DC119agent-pick: office by a street in its label
DC120agent-pick: a city the client has no office in is none
DC121agent-pick: seniority from an everyday word
DC122agent-pick: seniority — lead reads as senior
DC123agent-pick: owner by full name
DC124agent-pick: owner by first name only
DC125agent-pick: a first name two colleagues share is not written
DC131agent-pick: the full name of that same colleague IS picked
DC126agent-pick: a colleague who does not work here is none
DC127agent-pick: overlapping sources stay under the gate
DC128agent-pick: source named outright
DC129agent-pick: workflow by what the job is
DC130agent-pick: a workflow this tenant does not have is none, held loosely
DC132match-screen: meets every essential and more
DC133match-screen: the profile says not suitable
DC134match-screen: strong on skills, but the job will not sponsor
DC135match-screen: half the essentials
DC136match-screen: judged against the job, not against each other
DC55quick-actions: CV with the client nine days, no reply → chase the client
DC56quick-actions: client replied asking to meet → schedule the interview
DC57quick-actions: interview was three days ago, nothing recorded → get feedback
DC58quick-actions: shortlisted, CV never sent → submit the CV
DC59quick-actions: two jobs in one call — offer → place, interview stage unbooked → schedule
DC137eval-judge: a good answer to the outreach rates above 1.0
DC138eval-judge: a bad answer to the outreach rates below 1.0
DC139eval-judge: a good answer to the campaign summary rates above 2.3
DC140eval-judge: a bad answer to the campaign summary rates below 2.3
DC141eval-judge: a good answer to the support rewrite rates above 1.13
DC142eval-judge: a bad answer to the support rewrite rates below 1.13
DC145eval-judge: a good answer to the refusal rates above 1.28
DC146eval-judge: a bad answer to the refusal rates below 1.28

How we score

  • Machine checks first. Valid JSON, right schema, right function, nothing made up. One failure fails the case.
  • Then a blind judge. Another model rates the answer without knowing who wrote it.
  • Client-facing work counts more. A cover note a client reads outweighs an internal field extraction.
  • Cost is measured. Real token usage times list price, not an estimate.

Same prompts, same number of runs, caching off. We publish the results, not our prompts.