Oct 9 edition/Reporting & analysis
ModelsCodingBusinessSafetyMultimodal

ModelsArchitectures & capability

Mistral releases Large 4 'Le Chonk' preview, a 1-trillion-parameter multimodal model, with open weights planned for late October

Mistral has released an API preview of Mistral Large 4, a roughly 1-trillion-parameter multimodal mixture-of-experts model, ahead of a planned late-October open-weights release; Artificial Analysis ranks it 64th of 226 models, and its license remains unannounced.

THE CORE IDEAS4 TAKEAWAYS
01

Large 4 is a granular mixture-of-experts model with about 1 trillion total parameters. Roughly 5% of them are active per token, with stated counts of 49B and 52B. It accepts text and images, outputs only text, covers more than 160 languages, and was trained on 3,800 Grace Blackwell GPUs. Its predecessor had 675B total parameters. [6] [7] [8] [9]

02

Artificial Analysis gave the preview 38 on its Intelligence Index, about level with GPT-6 Luna, and called it the most capable model from outside the US and China. The model is very verbose, using about 200M output tokens in testing. Cost per task (about $1.13 on average) is therefore a better comparison than its low list prices of $1.36 and $4.18 per million tokens. [4] [5] [10]

03

The coding and security headline numbers come from Mistral's own announcement: 61.7% on DeepSWE, 3.74/5 in a Surge AI blind rating where Claude Opus 5 scored 4.22, and 93.3% attack resistance on B3. On the public DeepSWE leaderboard, GLM-5.3 and Kimi K3 reach 69%. Its CyberGym lead also partly reflects closed models refusing the tasks. [5] [6] [8]

04

For now, access is limited to the API through Mistral Studio. Weights are expected around October 27, after about three weeks of red-teaming. During that window, vetted partners and state authorities get a variant with lighter moderation. Large 3 shipped under Apache 2.0, but the license for Large 4 has not been announced. [5] [6] [8]

WHY IT MATTERS

The evidence so far is an API preview, one third-party index score and benchmark figures from the company.

Read the full assessment

The implication is that a permissive release would give EU and regulated buyers a self-hostable option, but the weights would still need about 1T parameters' worth of memory.

Executive brief

Mistral Large 4, nicknamed "Le Chonk", is a mixture-of-experts model with about 1 trillion parameters. Mistral says it is the best open-weight model from the US or Europe. The first independent measurement is more modest. Artificial Analysis scores the preview at 38 on its Intelligence Index, ranked #64 of 226 and roughly level with GPT-6 Luna. Weights are promised for the end of October, but the license hasn't been announced. Today it is an API-only preview. The 2nd-of-5 human rating and the "nine in ten attacks resisted" figure that the commentary repeats both come from Mistral's own announcement.

What changed and event timeline

  1. Mistral raises €3B

    The Series D valued Mistral at €21B or more, a month before the launch. Its enterprise customers include Airbus, ASML and HSBC ().

  2. Large 4 preview launches

    API model mistral-large-4-0 is billed as a hybrid instruct-and-reasoning multimodal MoE, trained on 3,800 Grace Blackwell GPUs (;).

  3. First independent score published

    Artificial Analysis scores the preview 38 and notes it is very verbose, using 200M output tokens during evaluation ().

  4. Creator commentary amplifies the launch

    A YouTube explainer presented by an AI avatar of Julian Goldie, posted alongside the Reddit item, repeats Mistral's claims and promotes a paid community ().

  5. (planned): Weights release

    The weights follow about three weeks of red-teaming with developers, security leaders and government authorities ().

Capabilities and access

  • Model: Mistral Large 4 preview, API ID mistral-large-4-0. It takes text and image input and returns text only, and covers 160+ languages (Mistral).
  • Price: $1.36 per million input tokens and $4.18 per million output tokens (Artificial Analysis). Cached input costs $0.14 per million (Let's Data Science).
  • Context window, reported two ways: 1M tokens in press coverage (MarkTechPost), 524k on Artificial Analysis.
Read the full section
  • Model: Mistral Large 4 preview, API ID mistral-large-4-0. It takes text and image input and returns text only, and covers 160+ languages (Mistral).
  • Price: $1.36 per million input tokens and $4.18 per million output tokens (Artificial Analysis). Cached input costs $0.14 per million (Let's Data Science).
  • Context window, reported two ways: 1M tokens in press coverage (MarkTechPost), 524k on Artificial Analysis.
  • Access now: API preview through Mistral Studio. Weights are expected around October 27.

Technical analysis for researchers and developers

  • Architecture: a granular MoE with a 1.6B-parameter vision encoder (MarkTechPost). Active parameters are stated as 49B on X but 52B on Mistral's announcement page.
  • Post-training: large-scale RL that generates about 33B tokens a day, of which about 16B are trainable.
  • Self-hosting cost: with only about 5% of parameters active per token, compute stays lower, but the full ~1T weights still have to be loaded into memory (MarkTechPost).
Read the full section
  • Architecture: a granular MoE with a 1.6B-parameter vision encoder (MarkTechPost). Active parameters are stated as 49B on X but 52B on Mistral's announcement page. Mistral says full architecture details will come with the weights.
  • Post-training: large-scale RL that generates about 33B tokens a day, of which about 16B are trainable. Rewards come from reward models, unit tests, LLM judges and static checks (Mistral).
  • Self-hosting cost: with only about 5% of parameters active per token, compute stays lower, but the full ~1T weights still have to be loaded into memory (MarkTechPost).
  • Reproducibility: the agentic-coding results came from an evaluation harness that is not yet public (Let's Data Science).

Claims and evidence

  • Coding: DeepSWE v1.1 61.7%, SWE-Atlas-QnA 59.4%, AutomationBench 59.9% (Mistral). Under other configurations, the public DeepSWE leaderboard shows GLM-5.3 and Kimi K3 at 69% (VentureBeat).
  • Blind human rating: Surge AI scored coding quality across five models. Large 4 got 3.74/5, behind Claude Opus 5 at 4.22 (Mistral).
  • Security: Mistral reports 93% on Cybench, 93.3% attack resistance on the B3 agent-security benchmark, and 82% on vulnerability patching.
Read the full section
  • Coding: DeepSWE v1.1 61.7%, SWE-Atlas-QnA 59.4%, AutomationBench 59.9% (Mistral). Under other configurations, the public DeepSWE leaderboard shows GLM-5.3 and Kimi K3 at 69% (VentureBeat).
  • Blind human rating: Surge AI scored coding quality across five models. Large 4 got 3.74/5, behind Claude Opus 5 at 4.22 (Mistral). This is the test the video describes (video).
  • Security: Mistral reports 93% on Cybench, 93.3% attack resistance on the B3 agent-security benchmark, and 82% on vulnerability patching. The video's "more than nine in ten" figure matches the B3 number (video).
  • Vision: 42% on Dense200 versus 41% for GPT-6 Astra. VentureBeat could not find the competitor scores from Mistral's slides in public sources.
  • "657 workflows": this figure from the video (video) does not appear in any of the reviewed sources.

Context and prior work

  • Predecessor: Mistral Large 3 had 675B total and 41B active parameters, was trained on 3,000 H200s (VentureBeat), and shipped under Apache 2.0 (Let's Data Science).
  • Competitive field: the main open-weight rivals are Chinese: DeepSeek V4, Qwen 3.8, Kimi K3 and GLM-5.3.
  • Independent ranking: Artificial Analysis calls it the most intelligent model from outside the US and China (Trending Topics).
Read the full section
  • Predecessor: Mistral Large 3 had 675B total and 41B active parameters, was trained on 3,000 H200s (VentureBeat), and shipped under Apache 2.0 (Let's Data Science).
  • Competitive field: the main open-weight rivals are Chinese: DeepSeek V4, Qwen 3.8, Kimi K3 and GLM-5.3. The video frames Large 4 as Europe's answer to them (video).
  • Independent ranking: Artificial Analysis calls it the most intelligent model from outside the US and China (Trending Topics).

Limitations, safety and contested findings

  • Spec conflicts: sources disagree on active parameters (49B vs 52B) and context length (1M vs 524k).
  • Cyber benchmark caveat: closed models score near zero on CyberGym-E2E because their refusal filters block the tasks, so Mistral's lead there overstates any capability gap (Let's Data Science).
  • Coding gap: on the public DeepSWE leaderboard, the model trails Chinese open models by about 7 points and closed frontier models by about 12 (same source).
Read the full section
  • Spec conflicts: sources disagree on active parameters (49B vs 52B) and context length (1M vs 524k).
  • Cyber benchmark caveat: closed models score near zero on CyberGym-E2E because their refusal filters block the tasks, so Mistral's lead there overstates any capability gap (Let's Data Science).
  • Coding gap: on the public DeepSWE leaderboard, the model trails Chinese open models by about 7 points and closed frontier models by about 12 (same source).
  • Unlocked variant: during the red-team window, vetted partners and state authorities get a version with reduced moderation and expanded cyber capabilities (Mistral).
  • Video caveat: the presenter says he has not tested the model himself (video).

Business and practitioner implications

  • Sovereignty: a European-hosted, self-hostable frontier-class model suits regulated buyers. Mistral offers end-to-end EU operation under European law.
  • License risk: don't commit roadmaps until the license is published.
  • Cost check: list prices are low, but the model's verbosity means comparing cost per task (Artificial Analysis measured $1.13 on average) is more useful than comparing token prices.
Read the full section
  • Sovereignty: a European-hosted, self-hostable frontier-class model suits regulated buyers. Mistral offers end-to-end EU operation under European law.
  • License risk: don't commit roadmaps until the license is published. A custom license, rather than Apache 2.0, would change how the weights can be used commercially.
  • Cost check: list prices are low, but the model's verbosity means comparing cost per task (Artificial Analysis measured $1.13 on average) is more useful than comparing token prices.
  • Before relying on it: wait for third-party runs on the released weights, especially for coding, where the vendor and leaderboard numbers diverge.
FOLLOW THE EVIDENCE

The source trail.

Sources (10)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief
Connect with us

Find us where you already read.