QDNA Book a call

Mistral Large 4 "le Chonk": what is verified, and which hardware can serve it

On October 6, 2026, Mistral AI opened the preview of its largest model, with open weights promised for the end of the month, and we cross-checked 382 claims from official sources, press, videos and independent rankings: what holds, what conflicts, and which machines can host it.

Screenshot of the Mistral Large 4 announcement page, titled Le chonk, with a large pixel-art number 4 in orange.
Screenshot of the announcement page, mistral.ai, October 6, 2026. Image: Mistral AI.
Short answer. Mistral Large 4 is a 1.05-trillion-parameter mixture of experts with 49 billion parameters active per token (52 billion with embeddings and output layer), reading text and images. The preview is served by Mistral's API with a 524,288-token context, while the model card announces one million. Weights, licence and detailed architecture are not published yet, and Mistral targets the end of October. Artificial Analysis scores it 38, eighth among open-weight models behind seven Chinese ones. To host it, count about 1,210 GB in FP8 and 680 GB in NVFP4 for the weights alone: a DGX B200 or DGX B300 in FP8, a DGX Station GB300 in NVFP4.
Mistral Large 4 identity card with the Mistral logo: 1.05 T parameters, 49 B active, 1 M context announced and 524,288 served, text and image, 1.6 B vision encoder, list price, weights expected October 31, 2026, licence not published.
The published characteristics in one view. Logo and values: Mistral AI.

Methodology

This methodology tags every figure with its origin, because in twenty-four hours the preview produced more conflicting values than official ones, and a number repeated ten times by the press is no more reliable for it.

We read Mistral's own pages (the English and French posts, the documentation card, the Hugging Face card, the governance page), operators that publish their own measurements (OpenRouter, Vercel, Artificial Analysis, Vals), French and international press, and fourteen videos whose full transcripts we read.

Each value carries a level: primary when it comes from Mistral or from the operator measuring it, concordant secondary when several outlets repeat it without a primary source, and calculation when it comes from our arithmetic. The research, run on October 7, 2026 with save-agent-rs, our local agent swarm, was then checked by hand against the source pages.

What Mistral published on October 6, 2026

Mistral Large 4 shipped as a "Public Preview", version v26.10, under the API identifiers mistral-large-4 and mistral-large-4-0, and was presented by Arthur Mensch at AI Everything in Abu Dhabi. The post nicknames it itself: "Unofficially ML4, very officially: le Chonk", echoing June's "fat kitten" meme.

Screenshot of the Mistral Large 4 announcement text on mistral.ai, introduction paragraph of the post.
The announcement post, screenshot taken October 7, 2026 (French edition). Source: mistral.ai.

The documentation card lists most published characteristics: a "granular" mixture of experts, a hybrid model that answers directly or reasons depending on reasoning_effort (none or high), a 1.6-billion-parameter vision encoder, and more than 160 languages including every official EU language. Output is text only.

Screenshot of the Mistral Large 4 documentation card: preview status, parameters, context and pricing.
The preview's technical card. Source: docs.mistral.ai.

There is no technical paper yet: Mistral promises the architecture, more benchmarks and the post-training method together with the weights. The number of experts, the number of layers, the attention design and the vocabulary size therefore remain unknown.

Timeline to scale of Mistral models, from Large 2 in July 2024 to Large 4 in October 2026, with weights expected at the end of October.
Ten months separate Large 3 from Large 4, on a timeline drawn to scale. Dates: Mistral AI.

What do 1.05 trillion parameters and 49 billion active mean?

In a mixture of experts, each token only goes through a small part of the network: a router picks a few experts among hundreds and the rest waits in memory. That is how Mistral Large 4 computes each token with only 49 billion parameters, under 5 % of the total, while still keeping all 1,050 billion loaded.

Diagram of 105 squares, each standing for 10 billion parameters, 5 of them colored for the parameters active per token.
Active versus total parameters, to scale. The placement of active squares is illustrative, since routing is unpublished.

Two practical consequences follow for anyone hosting it: memory is sized on the total, since every expert must be resident, while generation speed mostly depends on the active parameters, since each token only reads their weights, so such a model is heavy to load but relatively fast once loaded.

The gap between 49 and 52 billion is resolved by the Hugging Face card, which states 49 billion per token in the compute layers, 52 counting embeddings and output layer, hence the repository name Mistral-Large-4.0-1T05-A52B. The English post says 49, and the documentation card went from 49 to 52 on October 6 at 17:08 UTC.

Which values conflict?

Eight points differ from one source to another, and none can be settled before the weights and the technical report are out, so the table gives the value we keep and why.

TopicWhat circulatesValue kept
Total parameters"1 T" (post, X); "675 B with 41 B active" on a third-party site1.05 T (card). The 675 B / 41 B figures are Mistral Large 3's
Active parameters49 B (English post); 52 B (French post, card)49 B per token, 52 B with embeddings and output
Context1 M (card); 524,288 (API, OpenRouter, Vercel)1 M announced, 524,288 served in preview, 262,144 output
Training GPUs3,800 (post); "close to 4,000" (Mistral's LinkedIn); 4,000 (press)3,800 Grace Blackwell GPUs, generation not stated
Location"Europe" (post); France or Sweden (press, videos)Mistral's own European data centers, country not published
Weights date"end of the month" (post); October 26, 27 or 31October 31, the date shown on the Hugging Face card
Licence"custom Mistral license" (one outlet); assumed modified MITNot published
Inference hardware"4 to 8 B200 or B300 in FP8 and FP4" (press)Not found in any Mistral document

The last row deserves a calculation: four B200 offer 720 GB, which cannot hold the 1,050 GB of FP8 weights and leaves barely 41 GB of cache in NVFP4, which means "4 to 8 GPUs" only holds if 8 means FP8 and 4 means 4-bit, with a short context.

Screenshot of the Hugging Face page of Mistral-Large-4.0-1T05-A52B marked Upcoming release with a countdown.
The Hugging Face repository is reserved but empty: config.json is not public. Source: huggingface.co, October 7, 2026.

Training and sovereignty: what is European and what is not

According to Mistral, training happened in Europe, from scratch, on 3,800 NVIDIA Grace Blackwell GPUs in its own data centers, and the preview is served from the same infrastructure, operated under European law.

Three columns: what is European in Mistral Large 4, what depends on non-European players, and what is still unknown on October 7, 2026.
Large 4 sovereignty, layer by layer. QDNA summary.

The reinforcement learning phase uses 3,000 GPUs and processes about 33 billion tokens a day, 16 billion of them trainable, and it is still running without saturation according to Guillaume Lample, which means the preview tested today is not an end point.

Sovereignty still has limits that press releases leave out, and a public or regulated buyer needs to know them: compute remains NVIDIA silicon, and the €3 billion Series D funding this model is led by Samsung Electronics.

Above all, the licence that will define what you may do with the weights is not published, and self-hosting the model from the end of October will settle where data lives, not the hardware dependency.

The often-quoted 10 MW comes from an oral statement by Pierre Stock reported by the press, and the two-month duration from a single family of secondary sources, so we only repeat them as such.

Benchmarks: what Mistral declares, what third parties measure

Every score in the post is self-reported, and three of them are contradicted by third-party measurements or by Mistral's own charts, and the two charts below are reproduced as published, scales included.

Mistral chart, Terminal-Bench 4.0: Mistral Large 4 Preview 28, DeepSeek V4 Pro 0813 10, Qwen3.8 Max 17, Kimi K3 21, GLM-5.3 40.
Terminal-Bench 4.0 according to Mistral: GLM-5.3 leads at 40. Artificial Analysis measures 26.8 % and Vals 22.73 %. Chart: Mistral AI.
Mistral chart, DeepSWE 1.1 with an axis starting at 40: Mistral Large 4 62, Qwen3.8 Max 51, DeepSeek V4 Pro 57, GLM-5.3 61, Kimi K3 68.
DeepSWE 1.1 according to Mistral. The vertical axis starts at 40, which visually inflates the gaps. Chart: Mistral AI.
BenchmarkDeclared by MistralIndependent measurement
DeepSWE v1.161.7 %Absent from the public leaderboard cited by VentureBeat, where Kimi K3 and GLM-5.3 near 69 %
Terminal-Bench 428.3 %26.8 % (Artificial Analysis), 22.73 % (Vals)
CyberGym-E2E82 %, best score82 % confirmed by Artificial Analysis, ahead of MiMo-V2.6-Pro at 79 %; several closed models refuse the task
Finance Agent v254.754.68 % (Vals), consistent
Harvey Legal Agent15.815.83 %, 6th of 75 (Vals), consistent
SciCode-Verified91.8, "open state of the art"Mistral's own chart puts GLM-5.3 at 92.5 and Qwen3.8 2.4T at 93.8

With more distance, Artificial Analysis gives it 38 on its Intelligence Index v4.3.2, 64th of 225 and 8th among open-weight models, behind seven Chinese models from MiMo-V2.6-Pro to DeepSeek V4.1 Flash. The same organisation calls it the most capable model built outside the United States and China.

Screenshot of the Artificial Analysis article about Mistral Large 4 and French AI.
Artificial Analysis's write-up, screenshot taken October 7, 2026. Source: artificialanalysis.ai.

The measured weak point is verbosity: 200 million output tokens to complete the index, against a median of 81 million, so the index cost per task reaches $1.13 at list price, despite a moderate per-token price, while Vals ranks it 32nd of 44 at 48.05 %.

Screenshot of the Vals AI page for Mistral Large 4: Vals index 48.05%, $13.78 cost per test, 512K context.
The Vals AI page, screenshot taken October 7, 2026. Source: vals.ai.

API, pricing and usage on OpenRouter

Mistral's API, also offered through OpenRouter and Vercel, lists $1.36 per million input tokens, $4.18 for output and $0.14 for cache reads, with a 50 % launch discount for two weeks. Vercel sets the end of that discount on October 20, 2026, a date Mistral has not confirmed on its pricing page.

Screenshot of the OpenRouter page for mistralai/mistral-large-4-0: 524,288-token context, pricing and a single provider, Mistral.
The OpenRouter page: a single provider, Mistral. Source: openrouter.ai, October 7, 2026.

OpenRouter has one provider, Mistral itself, and measured on October 7 a median throughput of 67 tokens per second with 1.68 seconds of latency and 97.24 % uptime over three days, whereas Artificial Analysis measures 116.1 tokens per second on the direct API, Vercel 72.

These values depend on the time of day, the load and the reasoning mode: they measure the preview service, not the model. On its weekly ranking, OpenRouter places it third among new models with 12.7 billion tokens processed, but no major US cloud, Azure, AWS or Google Cloud, lists it yet.

Which hardware can host Mistral Large 4?

No official checkpoint size exists yet, and Mistral leaves the minimum memory fields of its card empty. The figures below are therefore QDNA calculations: parameters times bytes per format, plus a 15 % runtime margin, without the key-value cache, whose size depends on an unpublished architecture.

Bars of Mistral Large 4 weight memory: 2,415 GB in BF16, 1,208 GB in FP8, 679 GB in NVFP4, compared with the capacity of five NVIDIA machines.
Weights only, 15 % margin included, against machine memory. QDNA calculation.

The diagram reads directly: in BF16 no eight-GPU server is enough and a GB300 NVL72 rack is needed, and in FP8 the line falls between the 8 × H200 (1,128 GB, too small) and the DGX B200 (1,440 GB). In NVFP4, a DGX Station GB300 or an eight-card RTX PRO 6000 server suffices, with under 90 GB left for cache.

MachineMemoryDense FP8 PFLOPSDense FP4 PFLOPSFormat that fitsLeft for cache
DGX Spark (GB10)128 GBnot published0.5 NVFP4Nonen/a
Mac Studio M5 Ultra512 GBnot publishedno FP4 hardwareNone, even at 4-bit (604 GB)n/a
DGX Station GB300748 GB515 NVFP4NVFP469 GB
8 × RTX PRO 6000 Blackwell768 GB816 NVFP4NVFP489 GB
8 × H2001,128 GB15.8no FP4 hardwareWeight-only 4-bit (604 GB)524 GB
DGX B2001,440 GB3672 NVFP4FP8232 GB
DGX B3002,304 GB36108 NVFP4FP81,096 GB
8 × MI355X2,304 GB4080.8 MXFP4FP81,096 GB
GB300 NVL7220,736 GB3601,080 NVFP4BF1618,321 GB
Screenshot of the NVIDIA DGX B300 page, an eight-GPU Blackwell Ultra server with 2.1 TB of memory.
The DGX B300, one of the three eight-GPU machines in the table that hold Large 4 in FP8. Source: nvidia.com.

PFLOPS are the dense values from vendor datasheets; NVIDIA headlines values with structured sparsity, twice as high, that should not be compared with the others. Since the H200 and the Mac Studio have no FP4 unit, a 4-bit model runs there with compressed weights and 8- or 16-bit compute.

This table says what loads, not what serves well, because a 69 GB cache strongly limits the number of users and the context length, and the real answer depends on the cache size per token, which only config.json will give.

Mistral Large 3 used compressed latent attention over 61 layers, but nothing guarantees Large 4 inherits it. The QDNA memory calculator sizes weights and cache for the open models it covers; Mistral Large 4 is in the French version, with a warning while the cache cannot be computed.

What does the AI Act change for this model?

The European regulation presumes systemic risk above 1025 floating-point operations of cumulative training compute, which requires notifying the Commission within two weeks, adversarial testing, incident reporting and stronger cybersecurity. Penalties on general-purpose models have been possible since August 2, 2026.

Screenshot of the Mistral legal center: Mistral Large 4, type BASE, classified General Purpose AI Model, released October 6, 2026, status active.
Mistral's governance page lists Large 4 as a general-purpose model. Source: legal.mistral.ai.

On its governance page, Mistral classifies Large 4 as a general-purpose model and states that it provides no systemic-risk model, yet an order of magnitude is enough to raise the question, starting from the published GPU count and the duration reported by the press.

At about 2.5 dense BF16 PFLOPS per GPU, 3,800 GPUs over sixty days yield 4.9 × 1025 operations at full load, and still 1.5 × 1025 at 30 % utilisation, which is our calculation but clears the threshold by a wide margin.

Open weights do not exempt a systemic-risk model from these obligations, since the Article 53 exemption only covers models below the threshold. For a buyer the consequence is simple: ask for the technical documentation and the training data summary that the regulation requires in every case.

Where does it stand among open-weight models?

If the weights ship as announced, Mistral Large 4 will become the largest European open model, without being the most capable or the largest on the market, since the Chinese references are bigger or better rated, and their licences are already published.

Artificial Analysis index bars: MiMo-V2.6-Pro 46.3, GLM-5.3 44.8, Kimi K3 43.6, GLM-5.3 Flash 41.8, Qwen3.8 2.4T 39.9, Qwen3.8-Flash-Next 39.8, DeepSeek V4.1 Flash 39.5, Mistral Large 4 38, Mistral Large 3 9.
Seven points separate Large 4 from the top of the open ranking, twenty-nine separate it from Large 3. Data: Artificial Analysis.
ModelSizeLicenceAA index v4.3.2
Mistral Large 4 (preview)1.05 T / 49 B activenot published38
Kimi K32.8 T / 104 B activeKimi K3 License43.6
GLM-5.3744 B / 40 B activeGLM-5.3 licence44.8
GLM-5.3 Flash320 B / 18 B activeMIT41.8
Qwen3.8 2.4T2.4 T / 95 B activeQwen licence, agreement above $50 M39.9
DeepSeek V4.1 Flash552 BMIT39.5
Mistral Large 3675 B / 41 B activeApache 2.09

The jump from Large 3, rated 9 on the same index, is considerable. The remaining gap to the top of the open ranking, about seven points, matches what Guillaume Lample describes himself: the level of the best Chinese models from one or two months ago.

What the videos say

No official Mistral video existed on October 7, but about thirty creators published within twenty-four hours, and we transcribed fourteen, in English and French; these six bring elements that written articles do not.

Thumbnails of the six videos cited: Bruno Vega, Bijan Bowen, Matthew Berman, BFM Tech, IA Technologie and Arena AI.
The six videos kept. Images: respective YouTube channels.

Automatic transcripts mangle proper names, "Lechon" for the Chonk or "deep 1.1" for DeepSWE, so we checked each quote in context before keeping it, and set aside very low-audience channels whose content looked machine-generated.

What to watch before the end of October

Six points remain to watch until the weights ship at the end of October, and each can change server sizing, running cost or the right to use the model, and we will update this article as each one is published:

  1. The config.json file and the technical report, which will give layers, experts, attention and therefore the real cache size.
  2. The licence, which will decide commercial use and deployment at an end customer.
  3. Whether official FP8 and NVFP4 checkpoints ship, and their exact size.
  4. The context window of the published weights, 1 M or 524,288.
  5. A possible notification to the European Commission for systemic risk.
  6. The first throughput measurements on local hardware, which we will publish in our measurements when they exist.

Frequently asked questions

Is Mistral Large 4 open weight?

Not yet. As of October 7, 2026, it is only available through the API as a preview. Mistral announces the weights for the end of October and the Hugging Face repository shows October 31. The licence is not published.

How many parameters does Mistral Large 4 have?

1.05 trillion parameters in total, in a mixture of experts. 49 billion are active per token, 52 billion counting embeddings and output layer. The 675 billion sometimes quoted belong to Mistral Large 3.

What hardware is needed to run Mistral Large 4 locally?

By our calculation, the weights alone need about 1,210 GB in FP8 and 680 GB in NVFP4, 15 % margin included. A DGX B200 or DGX B300 host it in FP8, a DGX Station GB300 or eight RTX PRO 6000 in NVFP4 with little room for cache. A DGX Spark or a Mac Studio is not enough.

What is its context window?

The card announces one million tokens, but the preview API serves 524,288, with at most 262,144 output tokens. The window of the published weights remains to be checked.

How much does the Mistral Large 4 API cost?

$1.36 per million input tokens, $4.18 for output and $0.14 for cache reads, with a 50 % discount during the two launch weeks. Its verbosity brings the cost per task measured by Artificial Analysis to $1.13.

Official sources

Further reading