NewRecruitly + WhatsApp — message from your CRM
Recruitly LogoRecruitly

Every AI model, scored on real recruitment work.

We score every model against the real workloads running inside Recruitly, and publish the results.

24 of 24 models
ModelQualityHard fails$ / taskIn / OutContextLatency
Meta: Muse Spark 1.2 Contributorvia OpenRouter1000$0.0003$0.1 / $0.21,048,5769.6s
Meta: Muse Spark 1.3 Contributorvia OpenRouter1000$0.0003$0.1 / $0.21,048,5767.9s
Google: Gemini 3.8 Flashvia OpenRouter1000$0.0029$0.75 / $3.751,048,5763.7s
Z.ai: GLM 5.3 Flashopenvia OpenRouter990$0.0004$0.075 / $0.51,310,72016.5s
Xiaomi: MiMo V2.5openvia Xiaomi990$0.0001$0.14 / $0.281,050,0004.1s
OpenAI: GPT 5.6 Lunavia OpenRouter990$0.0001$0.2 / $1.201,050,0002.5s
Xiaomi: MiMo V2.5openvia OpenRouter981$0.0001$0.112 / $0.2241,050,0005.3s
Google: Gemini 3.1 Flash Litevia Google AI Studio980$0.0003$0.25 / $1.501,048,5761.1s
Xiaomi: MiMo V2.5 Proopenvia OpenRouter982$0.0002$0.348 / $0.6961,050,0005.9s
Z.ai: GLM 5.3 Flashopenvia Telnyx982$0.0001$0.075 / $0.251,310,7201.2s
Inception: Mercury 2.5 Previewvia OpenRouter963$0.0001$0.04 / $0.15260,0008.0s
Upstage: Solar Pro 4via OpenRouter942$0.0000$0.03 / $0.12524,2883.3s
OpenAI: GPT oss 120bopenvia Groq914$0.0002$0.15 / $0.6131,0720.7s
Qwen: Qwen3.8 Flashopenvia OpenRouter911$0.0001$0.16 / $0.471,000,0003.5s
Qwen: Qwen3.7 Flashvia OpenRouter914$0.0000$0.03 / $0.131,000,0002.2s
Xiaomi: MiMo V2.5 Proopenvia Xiaomi913$0.0002$0.435 / $0.871,050,0006.8s
DeepSeek: DeepSeek V4 Flash 0731openvia OpenRouter901$0.0001$0.14 / $0.281,310,7203.4s
Qwen: Qwen3.7 Plusvia OpenRouter892$0.0003$0.32 / $1.281,000,0004.1s
Mimo V2.5 Pro Ultraspeedvia Xiaomi891$0.0024$1.30 / $2.613.3s
Meta: Llama 4 Scout 17B 16E Instructvia Nscale872$0.0001$0.09 / $0.293.0s
DeepSeek: DeepSeek V4 Flash Vision Expopenvia OpenRouter862$0.0002$0.22 / $0.661,048,5762.5s
OpenAI: GPT oss 20bopenvia Groq846$0.0001$0.075 / $0.3131,0720.5s
Meta: Llama 4 Maverickopenvia OpenRouter843$0.0001$0.2 / $0.81,048,5763.3s
IBM: Granite 4.1 8Bopenvia OpenRouter818$0.0000$0.05 / $0.1131,0721.6s

Then we ran them for a month

Nearly every model above passes nearly every test, so the tests no longer separate them. These do. Measured on real work through one gateway, on models with at least 800 calls behind them.

What happens when your prompt gets longer

average response by input size

The same model at different context sizes. A flat model does not care how much you put in front of it; a steep one doubles or triples. If you are putting a CV or a transcript in the prompt, this is the number that picks your model, and no provider publishes it.

OpenAI: GPT oss 120b1.1× slower with more context
under 1k1.1s
1k–4k1.3s
Google: Gemini 3.1 Flash Lite1.4× slower with more context
under 1k1.6s
1k–4k2.1s
Z.ai: GLM 5.3 Flash1.8× faster with more context
under 1k5.0s
1k–4k3.6s
over 4k2.8s
Meta: Muse Spark 1.2 Contributor2.6× slower with more context
under 1k7.1s
1k–4k18s
Qwen: Qwen3.7 Flash1.8× slower with more context
under 1k2.4s
1k–4k4.3s
OpenAI: GPT 5.6 Luna1.6× slower with more context
under 1k3.3s
1k–4k7.1s
over 4k5.3s
Xiaomi: MiMo V2.51.3× slower with more context
under 1k6.7s
1k–4k12s
over 4k8.6s

Generation speed

output tokens per second

Tokens produced per second of wall clock, measured on real jobs rather than a synthetic prompt. This is what decides whether you can stream a response to somebody who is waiting.

Xiaomi: Mimo V2.5 Pro Ultraspeed230.3 tok/s
OpenAI: GPT oss 120b200.7 tok/s
Google: Gemini 3.1 Flash Lite170.7 tok/s
OpenAI: GPT oss 20b129.7 tok/s
Z.ai: GLM 5.3 Flash129.5 tok/s
Meta: Muse Spark 1.2 Contributor108.8 tok/s
Qwen: Qwen3.7 Flash72.6 tok/s
OpenAI: GPT 5.6 Luna67.5 tok/s
Xiaomi: MiMo V2.533.2 tok/s

Where to set your timeout

average → slowest call recorded

Set a timeout from an average and you will cut off good responses. The bar runs from each model's average to the slowest single call we have recorded. Budget from the right-hand end.

OpenAI: GPT oss 20b0.8s 34s
OpenAI: GPT 5.6 Luna4.6s 188s
OpenAI: GPT oss 120b1.1s 29s
Google: Gemini 3.1 Flash Lite2.3s 57s
Xiaomi: MiMo V2.511s 226s
Qwen: Qwen3.7 Flash2.7s 45s
Z.ai: GLM 5.3 Flash3.4s 30s
Xiaomi: Mimo V2.5 Pro Ultraspeed4.5s 23s
Meta: Muse Spark 1.2 Contributor7.5s 38s

Day to day drift

daily average, best day to worst

Providers retune and reroute without announcing it. A narrow bar is a latency budget you can hold. A wide one means the average you designed against was wrong on most days.

OpenAI: GPT oss 120b1.0s – 1.5s
Google: Gemini 3.1 Flash Lite1.4s – 4.8s
Qwen: Qwen3.7 Flash2.1s – 9.9s
OpenAI: GPT 5.6 Luna3.6s – 12s
Xiaomi: MiMo V2.56.8s – 16s

Rates, ranges and averages per model. No customer data, no call volumes, no per-account figures.

What we test

Sixteen jobs our software does every day. Each one has to pass, not impress.

TC1HTML profile → JSON

pure JSON, exact schema, no fences, all 4 skills

TC2Intent classification

single label == schedule_interview, no extra tokens

TC3Tool / function calling

tool_call search_candidates {skills:[React],location:London}, no prose

TC4Recruiter outreach (short gen)

<60w, mentions GraphQL, no em dash, role attributed to the fintech client (NOT Acme)

TC5Long-form job description

4 H2 in order, 250-400w total, no emojis

TC6CV → 3-bullet pitch

exactly 3 bullets, each <25w, no fluff

TC7Call notes → action items JSON

JSON array, 3-4 items, dates resolved in the call's week (shortlist 05-28 Thu, spec 05-29 Fri)

TC8Discriminatory request refusal

declines discriminatory targeting + offers a skills-based advert, brief, no lecture

TC9KB RAG grounded answer

grounded in the docs, preserves the image markdown link, no invented facts, not [[KB_NO_ANSWER]]

TC10Rewrite existing text (support reply)

preserves intent + facts, first-name Title Case, no internal-note leakage, JSON {text}

TC11CV-submission cover note + mismatch detection

client-clean message to Sarah + subject <60 chars; objections flags the junior-vs-senior mismatch; no banned phrases/sign-off

TC12Campaign metrics → narrative

2-4 sentences <80w, names campaigns + real numbers, flags the worst anomaly, no invented data

TC13Multi-turn conversation + tool call (Simi)

acknowledges then calls search_knowledge_base with a self-contained English query; stays in persona

TC14Candidate↔job match rationale

JSON {rationales:[2 strings]}, specific fit for job 1, no fabricated backend experience for job 2, same order

TC15Long-document CV parse

complete JSON, phone digits exact, work_history descending (Pied Piper first), skills/languages as strings

TC16Multi-variant suggested replies

1-3 distinct send-as-is replies to Tom, grounded in the KB fix, no invented links/prices

How we score

  • Machine checks first. Valid JSON, right schema, right function, nothing made up. One failure fails the case.
  • Then a blind judge. Another model rates the answer without knowing who wrote it.
  • Client-facing work counts more. A cover note a client reads outweighs an internal field extraction.
  • Cost is measured. Real token usage times list price, not an estimate.

Same prompts, same number of runs, caching off. We publish the results, not our prompts.