Mistral Large 4 "le Chonk": what is verified, and which hardware can serve it
On October 6, 2026, Mistral AI opened the preview of its largest model, with open weights promised for the end of the month, and we cross-checked 382 claims from official sources, press, videos and independent rankings: what holds, what conflicts, and which machines can host it.
Methodology
This methodology tags every figure with its origin, because in twenty-four hours the preview produced more conflicting values than official ones, and a number repeated ten times by the press is no more reliable for it.
We read Mistral's own pages (the English and French posts, the documentation card, the Hugging Face card, the governance page), operators that publish their own measurements (OpenRouter, Vercel, Artificial Analysis, Vals), French and international press, and fourteen videos whose full transcripts we read.
Each value carries a level: primary when it comes from Mistral or from the operator measuring it, concordant secondary when several outlets repeat it without a primary source, and calculation when it comes from our arithmetic. The research, run on October 7, 2026 with save-agent-rs, our local agent swarm, was then checked by hand against the source pages.
What Mistral published on October 6, 2026
Mistral Large 4 shipped as a "Public Preview", version v26.10, under the API identifiers mistral-large-4 and mistral-large-4-0, and was presented by Arthur Mensch at AI Everything in Abu Dhabi. The post nicknames it itself: "Unofficially ML4, very officially: le Chonk", echoing June's "fat kitten" meme.
The documentation card lists most published characteristics: a "granular" mixture of experts, a hybrid model that answers directly or reasons depending on reasoning_effort (none or high), a 1.6-billion-parameter vision encoder, and more than 160 languages including every official EU language. Output is text only.
There is no technical paper yet: Mistral promises the architecture, more benchmarks and the post-training method together with the weights. The number of experts, the number of layers, the attention design and the vocabulary size therefore remain unknown.
What do 1.05 trillion parameters and 49 billion active mean?
In a mixture of experts, each token only goes through a small part of the network: a router picks a few experts among hundreds and the rest waits in memory. That is how Mistral Large 4 computes each token with only 49 billion parameters, under 5 % of the total, while still keeping all 1,050 billion loaded.
Two practical consequences follow for anyone hosting it: memory is sized on the total, since every expert must be resident, while generation speed mostly depends on the active parameters, since each token only reads their weights, so such a model is heavy to load but relatively fast once loaded.
The gap between 49 and 52 billion is resolved by the Hugging Face card, which states 49 billion per token in the compute layers, 52 counting embeddings and output layer, hence the repository name Mistral-Large-4.0-1T05-A52B. The English post says 49, and the documentation card went from 49 to 52 on October 6 at 17:08 UTC.
Which values conflict?
Eight points differ from one source to another, and none can be settled before the weights and the technical report are out, so the table gives the value we keep and why.
| Topic | What circulates | Value kept |
|---|---|---|
| Total parameters | "1 T" (post, X); "675 B with 41 B active" on a third-party site | 1.05 T (card). The 675 B / 41 B figures are Mistral Large 3's |
| Active parameters | 49 B (English post); 52 B (French post, card) | 49 B per token, 52 B with embeddings and output |
| Context | 1 M (card); 524,288 (API, OpenRouter, Vercel) | 1 M announced, 524,288 served in preview, 262,144 output |
| Training GPUs | 3,800 (post); "close to 4,000" (Mistral's LinkedIn); 4,000 (press) | 3,800 Grace Blackwell GPUs, generation not stated |
| Location | "Europe" (post); France or Sweden (press, videos) | Mistral's own European data centers, country not published |
| Weights date | "end of the month" (post); October 26, 27 or 31 | October 31, the date shown on the Hugging Face card |
| Licence | "custom Mistral license" (one outlet); assumed modified MIT | Not published |
| Inference hardware | "4 to 8 B200 or B300 in FP8 and FP4" (press) | Not found in any Mistral document |
The last row deserves a calculation: four B200 offer 720 GB, which cannot hold the 1,050 GB of FP8 weights and leaves barely 41 GB of cache in NVFP4, which means "4 to 8 GPUs" only holds if 8 means FP8 and 4 means 4-bit, with a short context.
config.json is not public. Source: huggingface.co, October 7, 2026.Training and sovereignty: what is European and what is not
According to Mistral, training happened in Europe, from scratch, on 3,800 NVIDIA Grace Blackwell GPUs in its own data centers, and the preview is served from the same infrastructure, operated under European law.
The reinforcement learning phase uses 3,000 GPUs and processes about 33 billion tokens a day, 16 billion of them trainable, and it is still running without saturation according to Guillaume Lample, which means the preview tested today is not an end point.
Sovereignty still has limits that press releases leave out, and a public or regulated buyer needs to know them: compute remains NVIDIA silicon, and the €3 billion Series D funding this model is led by Samsung Electronics.
Above all, the licence that will define what you may do with the weights is not published, and self-hosting the model from the end of October will settle where data lives, not the hardware dependency.
The often-quoted 10 MW comes from an oral statement by Pierre Stock reported by the press, and the two-month duration from a single family of secondary sources, so we only repeat them as such.
Benchmarks: what Mistral declares, what third parties measure
Every score in the post is self-reported, and three of them are contradicted by third-party measurements or by Mistral's own charts, and the two charts below are reproduced as published, scales included.
| Benchmark | Declared by Mistral | Independent measurement |
|---|---|---|
| DeepSWE v1.1 | 61.7 % | Absent from the public leaderboard cited by VentureBeat, where Kimi K3 and GLM-5.3 near 69 % |
| Terminal-Bench 4 | 28.3 % | 26.8 % (Artificial Analysis), 22.73 % (Vals) |
| CyberGym-E2E | 82 %, best score | 82 % confirmed by Artificial Analysis, ahead of MiMo-V2.6-Pro at 79 %; several closed models refuse the task |
| Finance Agent v2 | 54.7 | 54.68 % (Vals), consistent |
| Harvey Legal Agent | 15.8 | 15.83 %, 6th of 75 (Vals), consistent |
| SciCode-Verified | 91.8, "open state of the art" | Mistral's own chart puts GLM-5.3 at 92.5 and Qwen3.8 2.4T at 93.8 |
With more distance, Artificial Analysis gives it 38 on its Intelligence Index v4.3.2, 64th of 225 and 8th among open-weight models, behind seven Chinese models from MiMo-V2.6-Pro to DeepSeek V4.1 Flash. The same organisation calls it the most capable model built outside the United States and China.
The measured weak point is verbosity: 200 million output tokens to complete the index, against a median of 81 million, so the index cost per task reaches $1.13 at list price, despite a moderate per-token price, while Vals ranks it 32nd of 44 at 48.05 %.
API, pricing and usage on OpenRouter
Mistral's API, also offered through OpenRouter and Vercel, lists $1.36 per million input tokens, $4.18 for output and $0.14 for cache reads, with a 50 % launch discount for two weeks. Vercel sets the end of that discount on October 20, 2026, a date Mistral has not confirmed on its pricing page.
OpenRouter has one provider, Mistral itself, and measured on October 7 a median throughput of 67 tokens per second with 1.68 seconds of latency and 97.24 % uptime over three days, whereas Artificial Analysis measures 116.1 tokens per second on the direct API, Vercel 72.
These values depend on the time of day, the load and the reasoning mode: they measure the preview service, not the model. On its weekly ranking, OpenRouter places it third among new models with 12.7 billion tokens processed, but no major US cloud, Azure, AWS or Google Cloud, lists it yet.
Which hardware can host Mistral Large 4?
No official checkpoint size exists yet, and Mistral leaves the minimum memory fields of its card empty. The figures below are therefore QDNA calculations: parameters times bytes per format, plus a 15 % runtime margin, without the key-value cache, whose size depends on an unpublished architecture.
The diagram reads directly: in BF16 no eight-GPU server is enough and a GB300 NVL72 rack is needed, and in FP8 the line falls between the 8 × H200 (1,128 GB, too small) and the DGX B200 (1,440 GB). In NVFP4, a DGX Station GB300 or an eight-card RTX PRO 6000 server suffices, with under 90 GB left for cache.
| Machine | Memory | Dense FP8 PFLOPS | Dense FP4 PFLOPS | Format that fits | Left for cache |
|---|---|---|---|---|---|
| DGX Spark (GB10) | 128 GB | not published | 0.5 NVFP4 | None | n/a |
| Mac Studio M5 Ultra | 512 GB | not published | no FP4 hardware | None, even at 4-bit (604 GB) | n/a |
| DGX Station GB300 | 748 GB | 5 | 15 NVFP4 | NVFP4 | 69 GB |
| 8 × RTX PRO 6000 Blackwell | 768 GB | 8 | 16 NVFP4 | NVFP4 | 89 GB |
| 8 × H200 | 1,128 GB | 15.8 | no FP4 hardware | Weight-only 4-bit (604 GB) | 524 GB |
| DGX B200 | 1,440 GB | 36 | 72 NVFP4 | FP8 | 232 GB |
| DGX B300 | 2,304 GB | 36 | 108 NVFP4 | FP8 | 1,096 GB |
| 8 × MI355X | 2,304 GB | 40 | 80.8 MXFP4 | FP8 | 1,096 GB |
| GB300 NVL72 | 20,736 GB | 360 | 1,080 NVFP4 | BF16 | 18,321 GB |
PFLOPS are the dense values from vendor datasheets; NVIDIA headlines values with structured sparsity, twice as high, that should not be compared with the others. Since the H200 and the Mac Studio have no FP4 unit, a 4-bit model runs there with compressed weights and 8- or 16-bit compute.
This table says what loads, not what serves well, because a 69 GB cache strongly limits the number of users and the context length, and the real answer depends on the cache size per token, which only config.json will give.
Mistral Large 3 used compressed latent attention over 61 layers, but nothing guarantees Large 4 inherits it. The QDNA memory calculator sizes weights and cache for the open models it covers; Mistral Large 4 is in the French version, with a warning while the cache cannot be computed.
What does the AI Act change for this model?
The European regulation presumes systemic risk above 1025 floating-point operations of cumulative training compute, which requires notifying the Commission within two weeks, adversarial testing, incident reporting and stronger cybersecurity. Penalties on general-purpose models have been possible since August 2, 2026.
On its governance page, Mistral classifies Large 4 as a general-purpose model and states that it provides no systemic-risk model, yet an order of magnitude is enough to raise the question, starting from the published GPU count and the duration reported by the press.
At about 2.5 dense BF16 PFLOPS per GPU, 3,800 GPUs over sixty days yield 4.9 × 1025 operations at full load, and still 1.5 × 1025 at 30 % utilisation, which is our calculation but clears the threshold by a wide margin.
Open weights do not exempt a systemic-risk model from these obligations, since the Article 53 exemption only covers models below the threshold. For a buyer the consequence is simple: ask for the technical documentation and the training data summary that the regulation requires in every case.
Where does it stand among open-weight models?
If the weights ship as announced, Mistral Large 4 will become the largest European open model, without being the most capable or the largest on the market, since the Chinese references are bigger or better rated, and their licences are already published.
| Model | Size | Licence | AA index v4.3.2 |
|---|---|---|---|
| Mistral Large 4 (preview) | 1.05 T / 49 B active | not published | 38 |
| Kimi K3 | 2.8 T / 104 B active | Kimi K3 License | 43.6 |
| GLM-5.3 | 744 B / 40 B active | GLM-5.3 licence | 44.8 |
| GLM-5.3 Flash | 320 B / 18 B active | MIT | 41.8 |
| Qwen3.8 2.4T | 2.4 T / 95 B active | Qwen licence, agreement above $50 M | 39.9 |
| DeepSeek V4.1 Flash | 552 B | MIT | 39.5 |
| Mistral Large 3 | 675 B / 41 B active | Apache 2.0 | 9 |
The jump from Large 3, rated 9 on the same index, is considerable. The remaining gap to the top of the open ranking, about seven points, matches what Guillaume Lample describes himself: the level of the best Chinese models from one or two months ago.
What the videos say
No official Mistral video existed on October 7, but about thirty creators published within twenty-four hours, and we transcribed fourteen, in English and French; these six bring elements that written articles do not.
- Bruno Vega, "Is It Actually A Failure?": the most technical, it computes memory, notes the 524,288-token served context and attributes the cost per task to cache reads.
- Bijan Bowen, first test: commented specifications and long sessions in Vibe CLI, with the explanation of the 49 versus 52 billion gap.
- Matthew Berman, "Mistral is BACK!": the most viewed, focused on benchmarks and per-token price.
- BFM Tech, Tech&Co of October 6: French debate, source of the Sweden training hypothesis and of the "4 or 8 B200 B300 chips" phrase.
- IA Technologie, "is France really back?": critical summary of the open-model rank and the cost per task.
- Arena AI, first impressions: WebDev Arena rank and the open question of the GPU generation, GB200 or GB300.
Automatic transcripts mangle proper names, "Lechon" for the Chonk or "deep 1.1" for DeepSWE, so we checked each quote in context before keeping it, and set aside very low-audience channels whose content looked machine-generated.
What to watch before the end of October
Six points remain to watch until the weights ship at the end of October, and each can change server sizing, running cost or the right to use the model, and we will update this article as each one is published:
- The
config.jsonfile and the technical report, which will give layers, experts, attention and therefore the real cache size. - The licence, which will decide commercial use and deployment at an end customer.
- Whether official FP8 and NVFP4 checkpoints ship, and their exact size.
- The context window of the published weights, 1 M or 524,288.
- A possible notification to the European Commission for systemic risk.
- The first throughput measurements on local hardware, which we will publish in our measurements when they exist.
Frequently asked questions
Is Mistral Large 4 open weight?
Not yet. As of October 7, 2026, it is only available through the API as a preview. Mistral announces the weights for the end of October and the Hugging Face repository shows October 31. The licence is not published.
How many parameters does Mistral Large 4 have?
1.05 trillion parameters in total, in a mixture of experts. 49 billion are active per token, 52 billion counting embeddings and output layer. The 675 billion sometimes quoted belong to Mistral Large 3.
What hardware is needed to run Mistral Large 4 locally?
By our calculation, the weights alone need about 1,210 GB in FP8 and 680 GB in NVFP4, 15 % margin included. A DGX B200 or DGX B300 host it in FP8, a DGX Station GB300 or eight RTX PRO 6000 in NVFP4 with little room for cache. A DGX Spark or a Mac Studio is not enough.
What is its context window?
The card announces one million tokens, but the preview API serves 524,288, with at most 262,144 output tokens. The window of the published weights remains to be checked.
How much does the Mistral Large 4 API cost?
$1.36 per million input tokens, $4.18 for output and $0.14 for cache reads, with a 50 % discount during the two launch weeks. Its verbosity brings the cost per task measured by Artificial Analysis to $1.13.
Official sources
- Mistral AI, Mistral Large 4 announcement, October 6, 2026
- Mistral AI, Mistral Large 4 documentation card
- Hugging Face, mistralai/Mistral-Large-4.0-1T05-A52B repository
- Mistral AI, model governance page
- OpenRouter, mistralai/mistral-large-4-0 page
- Artificial Analysis, Mistral Large 4 analysis
- Vals AI, Mistral Large 4 results
- Reuters, Arthur Mensch's statements in Abu Dhabi
- Simon Willison, first tests of the Chonk
- NVIDIA, DGX B300 datasheet and DGX B200 datasheet
- European Commission, AI Act, Article 51