Every AI model, scored on real recruitment work.
We score every model against the real workloads running inside Recruitly, and publish the results.
| Model | Quality↓ | Hard fails | $ / task | In / Out | Context | Latency |
|---|---|---|---|---|---|---|
| 100 | 0 | $0.0003 | $0.1 / $0.2 | 1,048,576 | 9.6s | |
| 100 | 0 | $0.0003 | $0.1 / $0.2 | 1,048,576 | 7.9s | |
| 100 | 0 | $0.0029 | $0.75 / $3.75 | 1,048,576 | 3.7s | |
| 99 | 0 | $0.0004 | $0.075 / $0.5 | 1,310,720 | 16.5s | |
| 99 | 0 | $0.0001 | $0.14 / $0.28 | 1,050,000 | 4.1s | |
| 99 | 0 | $0.0001 | $0.2 / $1.20 | 1,050,000 | 2.5s | |
| 98 | 1 | $0.0001 | $0.112 / $0.224 | 1,050,000 | 5.3s | |
| 98 | 0 | $0.0003 | $0.25 / $1.50 | 1,048,576 | 1.1s | |
| 98 | 2 | $0.0002 | $0.348 / $0.696 | 1,050,000 | 5.9s | |
| 98 | 2 | $0.0001 | $0.075 / $0.25 | 1,310,720 | 1.2s | |
| 96 | 3 | $0.0001 | $0.04 / $0.15 | 260,000 | 8.0s | |
| 94 | 2 | $0.0000 | $0.03 / $0.12 | 524,288 | 3.3s | |
| 91 | 4 | $0.0002 | $0.15 / $0.6 | 131,072 | 0.7s | |
| 91 | 1 | $0.0001 | $0.16 / $0.47 | 1,000,000 | 3.5s | |
| 91 | 4 | $0.0000 | $0.03 / $0.13 | 1,000,000 | 2.2s | |
| 91 | 3 | $0.0002 | $0.435 / $0.87 | 1,050,000 | 6.8s | |
| 90 | 1 | $0.0001 | $0.14 / $0.28 | 1,310,720 | 3.4s | |
| 89 | 2 | $0.0003 | $0.32 / $1.28 | 1,000,000 | 4.1s | |
| 89 | 1 | $0.0024 | $1.30 / $2.61 | — | 3.3s | |
| 87 | 2 | $0.0001 | $0.09 / $0.29 | — | 3.0s | |
| 86 | 2 | $0.0002 | $0.22 / $0.66 | 1,048,576 | 2.5s | |
| 84 | 6 | $0.0001 | $0.075 / $0.3 | 131,072 | 0.5s | |
| 84 | 3 | $0.0001 | $0.2 / $0.8 | 1,048,576 | 3.3s | |
| 81 | 8 | $0.0000 | $0.05 / $0.1 | 131,072 | 1.6s |
Then we ran them for a month
Nearly every model above passes nearly every test, so the tests no longer separate them. These do. Measured on real work through one gateway, on models with at least 800 calls behind them.
What happens when your prompt gets longer
average response by input sizeThe same model at different context sizes. A flat model does not care how much you put in front of it; a steep one doubles or triples. If you are putting a CV or a transcript in the prompt, this is the number that picks your model, and no provider publishes it.
Generation speed
output tokens per secondTokens produced per second of wall clock, measured on real jobs rather than a synthetic prompt. This is what decides whether you can stream a response to somebody who is waiting.
Where to set your timeout
average → slowest call recordedSet a timeout from an average and you will cut off good responses. The bar runs from each model's average to the slowest single call we have recorded. Budget from the right-hand end.
Day to day drift
daily average, best day to worstProviders retune and reroute without announcing it. A narrow bar is a latency budget you can hold. A wide one means the average you designed against was wrong on most days.
Rates, ranges and averages per model. No customer data, no call volumes, no per-account figures.
What we test
Sixteen jobs our software does every day. Each one has to pass, not impress.
pure JSON, exact schema, no fences, all 4 skills
single label == schedule_interview, no extra tokens
tool_call search_candidates {skills:[React],location:London}, no prose
<60w, mentions GraphQL, no em dash, role attributed to the fintech client (NOT Acme)
4 H2 in order, 250-400w total, no emojis
exactly 3 bullets, each <25w, no fluff
JSON array, 3-4 items, dates resolved in the call's week (shortlist 05-28 Thu, spec 05-29 Fri)
declines discriminatory targeting + offers a skills-based advert, brief, no lecture
grounded in the docs, preserves the image markdown link, no invented facts, not [[KB_NO_ANSWER]]
preserves intent + facts, first-name Title Case, no internal-note leakage, JSON {text}
client-clean message to Sarah + subject <60 chars; objections flags the junior-vs-senior mismatch; no banned phrases/sign-off
2-4 sentences <80w, names campaigns + real numbers, flags the worst anomaly, no invented data
acknowledges then calls search_knowledge_base with a self-contained English query; stays in persona
JSON {rationales:[2 strings]}, specific fit for job 1, no fabricated backend experience for job 2, same order
complete JSON, phone digits exact, work_history descending (Pied Piper first), skills/languages as strings
1-3 distinct send-as-is replies to Tom, grounded in the KB fix, no invented links/prices
How we score
- Machine checks first. Valid JSON, right schema, right function, nothing made up. One failure fails the case.
- Then a blind judge. Another model rates the answer without knowing who wrote it.
- Client-facing work counts more. A cover note a client reads outweighs an internal field extraction.
- Cost is measured. Real token usage times list price, not an estimate.
Same prompts, same number of runs, caching off. We publish the results, not our prompts.