ModelsArchitectures & capability
Inception's Mercury Decide, a calibrated decision model, launches free on OpenRouter and ranks third on JevBench v1.6.1
Mercury Decide answers with a choice, a score or a yes/no plus a probability, rather than prose. It launched free in early access on OpenRouter on 30 September 2026 and places third overall on the JevBench v1.6.1 leaderboard.
Mercury Decide returns only typed answers: one choice from up to 255 options, a score with up to 10 levels, or yes/no. Each answer carries a probability taken from the model's predictive distribution, so output tokens are free and input is listed at $0.02 per million tokens. It reportedly uses Inception's diffusion approach and needs a separate Decisions API, with no OpenAI chat compatibility, streaming or tool calling. One open-source project built a custom provider for it in a day. [3] [10] [11] [9] [6]
On JevBench v1.6.1, Mercury Decide has the top Intelligence score (65.91 versus Sage's 65.59) but ranks third overall. The benchmark weights intelligence, calibration, speed and cost equally. It records 80.8% choice accuracy and 71.2% score accuracy. Its measured latency is 0.33 s median and 0.85 s at p95, slower than Jev. Inception's figure of up to 14 decisions per second is a vendor ceiling, not typical throughput. [4] [13] [5] [3]
The business commentary suggests three workflows: routing incoming messages, scoring leads and filtering what agents save to memory. Confidence cutoffs decide whether a case runs automatically or goes to a person. The 95%, 60% and 25% cutoffs are illustrative only. The commentary recommends testing on about 20 real examples first, but setting reliable thresholds needs a larger labeled held-out set. [1] [3]
JevBench shows the model is competitive on accuracy and calibration at very low cost.
Read the full assessment
The implication, not yet shown on production data, is that agent and business workflows could offload high-volume small decisions and send low-confidence cases to people.
Executive brief
Inception's Mercury Decide does not write text. It returns a choice, a score or a yes/no answer, and each one carries a probability. On the independent JevBench v1.6.1 API board it ranks 3rd overall, which complicates Inception's launch line that it is the "most intelligent decision model on OpenRouter". That claim was tied to JevBench v1.4. The model went live free on OpenRouter on 30 September 2026. The Reddit post is commentary from a marketing community, and it suggests handing three jobs to the model: routing messages, scoring leads and filtering agent memory. Early-access limits and conflicting published specs make it a model to test, not to depend on.
What changed and event timeline
Diffusion LLM groundwork
Inception's Mercury paper describes Transformer-based diffusion LLMs that predict several tokens in parallel. It reports 1,109 tokens/s for Mercury Coder Mini on H100s ().
Calibrated-decision research appears
A paper on Jev, trained with reinforcement learning for calibrated decisions, reports a median AUROC of 0.886 for zero-shot detection of alignment failures (). Jev is the rival model Inception benchmarks against.
First public developer test
An open-source game project built an OpenRouter decisions provider in a day. It reported 10 of 10 calls succeeding with no rate-limiting ().
Business commentary
The Reddit post and a video narrated by Julian Goldie's "digital avatar" pitch three workflows to automate, each with a human-review step (;).
Independent leaderboard entry
On the JevBench v1.6.1 API board, Mercury Decide places 3rd behind Sage 1.3.0 ().
Capabilities and access
- Versions:
inception/mercury-decide-20260930(paid) and:free(systemonemodels). - Answer types: a choice from up to 255 options, a score with up to 10 levels, or yes/no.
- Price: listed input price is $0.02 per million tokens; output is free.
Read the full section
- Versions:
inception/mercury-decide-20260930(paid) and:free(systemonemodels). - Answer types: a choice from up to 255 options, a score with up to 10 levels, or yes/no.
- Price: listed input price is $0.02 per million tokens; output is free.
- API: the Decisions API (
/v1/decisionsfrom Inception,/v1/systemonevia OpenRouter). It is not OpenAI chat-compatible and has no streaming or tool calling (AlphaSignal). - Context window, conflicting figures:
- 32,768 tokens per the video, quoting OpenRouter (video).
- 33K per OpenRouter's listing (KuCoin).
- 65,536 per systemonemodels.
Technical analysis for researchers and developers
- Architecture: reported to use the same diffusion language modeling as Mercury 2.5, which refines many positions in parallel instead of generating tokens one at a time (AlphaSignal).
- Where the probability comes from: it is read off the model's predictive distribution rather than written out as text.
- How JevBench scores models: a geometric mean of chance-corrected Intelligence, Calibration, Speed and Cost, weighted 25% each, on frozen hashed cases (JevBench repo).
Read the full section
- Architecture: reported to use the same diffusion language modeling as Mercury 2.5, which refines many positions in parallel instead of generating tokens one at a time (AlphaSignal).
- Where the probability comes from: it is read off the model's predictive distribution rather than written out as text. That is why output tokens cost nothing.
- How JevBench scores models: a geometric mean of chance-corrected Intelligence, Calibration, Speed and Cost, weighted 25% each, on frozen hashed cases (JevBench repo).
- Integration: a separate endpoint is required. The GitHub project wrote a custom provider rather than reusing its litellm setup (issue #454).
Claims and evidence
- "Most intelligent decision model" (Inception, JevBench v1.4): partly consistent with independent results.
- Choice accuracy 80.8%, score accuracy 71.2%.
- Calibration score 81.34.
Read the full section
- "Most intelligent decision model" (Inception, JevBench v1.4): partly consistent with independent results. On v1.6.1, its Intelligence score is 65.91 versus 65.59 for Sage, but it ranks 3rd overall. Measured results there:
- Choice accuracy 80.8%, score accuracy 71.2%.
- Calibration score 81.34.
- Cost $0.0184 per 1,000 items.
- "Up to 14 decisions/s": a ceiling from the vendor. Measured latency is 0.33 s at the median and 0.85 s at p95, slower than Jev's 0.24/0.30 s. The video gives 0.44 s median on OpenRouter and itself calls 14/s a ceiling (video).
- Chess match against Jev: promoted, but its result is not described in the video transcript (video).
Context and prior work
- Earlier Inception models: Mercury Coder established the company's speed-focused diffusion approach (arXiv).
- The category: typed decision models with calibrated probabilities predate this launch.
- "System one": the video frames the model as fast, gut-level thinking in contrast to slow deliberate reasoning (video).
Read the full section
- Earlier Inception models: Mercury Coder established the company's speed-focused diffusion approach (arXiv).
- The category: typed decision models with calibrated probabilities predate this launch. Examples include TypeSafe AI's Jev and open-weight models such as Quyet-1.0-Large, which leads the open-weights board (Benchmark Heaven).
- "System one": the video frames the model as fast, gut-level thinking in contrast to slow deliberate reasoning (video).
Limitations, safety and contested findings
- Context limits bite: 23 of 1,500 benchmark items went unanswered because they exceeded the context window (Benchmark Heaven).
- Thin independent evidence: one aggregator flags it as "weak evidence" because it appears on only one board (hull PR #52).
- Unpublished limits: rate limits are undisclosed, and no independent calibration figures beyond JevBench were found (systemonemodels).
Read the full section
- Context limits bite: 23 of 1,500 benchmark items went unanswered because they exceeded the context window (Benchmark Heaven).
- Thin independent evidence: one aggregator flags it as "weak evidence" because it appears on only one board (hull PR #52).
- Unpublished limits: rate limits are undisclosed, and no independent calibration figures beyond JevBench were found (systemonemodels).
- Benchmarks are not your data: JevBench matches the model's intended workload but cannot establish accuracy on a specific production dataset (AlphaSignal).
- The commentary's thresholds are illustrative: the 95% / 60% / 25% "traffic light" cutoffs are not validated, and the post and video both promote a paid community (video).
Business and practitioner implications
- Good fits: routing, triage, rubric grading and agent memory filtering.
- Test before trusting: the commentary's advice is to run about 20 labeled examples first.
- Shop the category: Sage and Jev score close to Mercury Decide on JevBench, and Jev has lower latency, so it is worth comparing a few.
Read the full section
- Good fits: routing, triage, rubric grading and agent memory filtering. These are high-volume, one-word decisions where a probability helps decide when a human should step in (video).
- Test before trusting: the commentary's advice is to run about 20 labeled examples first. A larger held-out set is needed to set thresholds that hold up.
- Shop the category: Sage and Jev score close to Mercury Decide on JevBench, and Jev has lower latency, so it is worth comparing a few.
- Early-access caveat: usage limits and conflicting specs argue against putting production-critical work on it yet (video).
The source trail.
Sources (15)
3 Business Decisions I'd Hand to Inception Labs Mercury
Feed excerpt only. A transcript of the video linked from this page is supplied (linked video auto captions; https://www.youtube.com/watch?v=lWv9kyQ7Qps); automatic text may contain errors.
reddit.comMercury Decision : Most Intelligent Decision Model! 😱
Related coverage; assess separately
reddit.comLinked video: 3 Business Decisions I'd Hand to Inception Labs Mercury
linked video auto captions
www.youtube.com