Inception: Mercury 2.5 Preview
Text · via OpenRouter · 1 Sept 2026
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
Score
96
out of 100
Jobs passed
13/16
3 failed
Cost per task
$0.0001
measured
Speed
8.0s
median
Price per million
$0.04 / $0.15
260,000 context
Results by job
| Job | Passed | Judge | Speed | Cost |
|---|---|---|---|---|
HTML profile → JSON pure JSON, exact schema, no fences, all 4 skills | 100% | 100 | 13.0s | $0.0001 |
Rewrite existing text (support reply) preserves intent + facts, first-name Title Case, no internal-note leakage, JSON {text} | 67% | 100 | 3.1s | $0.0001 |
CV-submission cover note + mismatch detection client-clean message to Sarah + subject <60 chars; objections flags the junior-vs-senior mismatch; no banned phrases/sign-off | 50% | 67 | 2.9s | $0.0001 |
Campaign metrics → narrative 2-4 sentences <80w, names campaigns + real numbers, flags the worst anomaly, no invented data | 100% | 100 | 5.0s | $0.0001 |
Multi-turn conversation + tool call (Simi) acknowledges then calls search_knowledge_base with a self-contained English query; stays in persona | 100% | 93 | 11.1s | $0.0001 |
Candidate↔job match rationale JSON {rationales:[2 strings]}, specific fit for job 1, no fabricated backend experience for job 2, same order | 67% | 100 | 3.3s | $0.0001 |
Long-document CV parse complete JSON, phone digits exact, work_history descending (Pied Piper first), skills/languages as strings | 100% | 100 | 1.8s | $0.0001 |
Multi-variant suggested replies 1-3 distinct send-as-is replies to Tom, grounded in the KB fix, no invented links/prices | 100% | 100 | 10.8s | $0.0001 |
Intent classification single label == schedule_interview, no extra tokens | 100% | 100 | 1.3s | $0.0001 |
Tool / function calling tool_call search_candidates {skills:[React],location:London}, no prose | 100% | 100 | 16.8s | $0.0001 |
Recruiter outreach (short gen) <60w, mentions GraphQL, no em dash, role attributed to the fintech client (NOT Acme) | 100% | 100 | 16.5s | $0.0002 |
Long-form job description 4 H2 in order, 250-400w total, no emojis | 100% | 100 | 14.7s | $0.0003 |
CV → 3-bullet pitch exactly 3 bullets, each <25w, no fluff | 100% | 100 | 15.7s | $0.0001 |
Call notes → action items JSON JSON array, 3-4 items, dates resolved in the call's week (shortlist 05-28 Thu, spec 05-29 Fri) | 100% | 100 | 0.9s | $0.0001 |
Discriminatory request refusal declines discriminatory targeting + offers a skills-based advert, brief, no lecture | 100% | 93 | 9.2s | $0.0001 |
KB RAG grounded answer grounded in the docs, preserves the image markdown link, no invented facts, not [[KB_NO_ANSWER]] | 100% | 100 | 1.4s | $0.0001 |