K Koda Intelligence
DEEP DIVE DEEP DIVE № 224 · 07 October 2026DOCKODA-20261007-DE93233ACF83sha-256 of date + article + 24 checked + 13 computed

Mistral's trillion-parameter model runs on 49 billion

Mistral opened a preview of Mistral Large 4, nicknamed Le Chonk, on October 6, 2026. It holds 1.05 trillion 25 parameters and activates 49 billion per token, about 4.7% of the model. Weights are reportedly due October 27 under a custom license, after training on 4,000 08 Nvidia Grace Blackwell GPUs by VentureBeat's count. We judge it by active-parameter economics and deployment control.

6 MIN READ · BY THE KODA EDITORIAL TEAM · STRATEGY · OPEN WEIGHTS
TOTAL PARAMS1.05TVERIFIED CLAIM 06MISTRAL DOCS
ACTIVE PER TOKEN49BREPORTED CLAIM 07MISTRAL DOCS
ACTIVE SHARE4.7%COMPUTED CLAIM 34KODA MATH
TOTAL PARAMS1.05TMISTRAL DOCS ACTIVE PER TOKEN49BMISTRAL DOCS ACTIVE SHARE4.7%KODA MATH PREVIEW OPENEDOCT 6Mistral WEIGHTS DUEOCT 27VENTUREBEAT CONTEXT WINDOW1MMISTRAL DOCS TRAINING GPUS4,000VENTUREBEAT SERIES D€3BANADOLU AGENCY

Mistral's documentation lists its newest flagship at 1.05 trillion 25 parameters, with 49 billion active per token. That is about 4.7% of the model, by our math. The other 95% 26 sit in memory and wait for a router to call them.

Mistral announced Mistral Large 4, nicknamed Le Chonk, on October 6, 2026 01. Its launch post on X led with "1T 04 parameters" and put "49B active" second. Anyone paying to run the model should read those numbers in the opposite order. The 49 billion 25 drives compute per token, and the 1.05 trillion drives the memory you must rent to host it.

Nobody outside Mistral knows the real bill yet. Mistral's documentation says the company gave no API pricing in the materials it reviewed. The weights are due by the end of October, reportedly on October 27, under a custom Mistral license. Until then, every cost claim about Le Chonk, including ours, is arithmetic on published specs.

The Capacity Bank Behind 49 Billion

Mistral Large 4 is a mixture-of-experts model. The network is split into many smaller sub-networks called experts. A router sends each token to a few of them. Think of it as a capacity bank: about a trillion parameters on deposit, about 49 billion 07 withdrawn per token. Divide 1,050 billion 27 by 49 billion and the bank is roughly 21 times larger than any single withdrawal.

SELF-HOSTING MATH · OCTOBER 2026MISTRAL DOCS · KODA ARITHMETICBASE: 24 CHECKED + 13 COMPUTED, 4 SHOWN

What it takes to keep 1.05 trillion weights loaded

Weights at 16-bit Koda math · two bytes per weight COMPUTED CLAIM 28
2.1 TB
Weights at 8-bit Koda math · one byte per weight COMPUTED CLAIM 29
1.05 TB
Weights at 4-bit Koda math · half a byte per weight COMPUTED CLAIM 30
525 GB
Context window Mistral docs · KV cache grows with length VERIFIED CLAIM 11
1M tokens

That framing sorts every Le Chonk fact into four accounts. Capacity is the full 1.05 trillion 06, which Mistral's documentation lists alongside a 1.6 10-billion-parameter vision encoder. Compute is the 49 billion 07 active weights, so per-token arithmetic lands near a dense model of that size. Memory is everything that has to stay loaded. Bandwidth is the traffic of tokens moving between experts on different chips.

Memory is where the trillion comes back. At 16 28-bit precision each weight takes two bytes, so 1.05 trillion weights need about 2.1 terabytes. At 8-bit precision that drops to about 1.05 terabytes 29, and at 4-bit to roughly 525 gigabytes 30. And those totals come before the KV cache, the stored attention data for every token already in the prompt.

The KV cache grows with context length, and Mistral's documentation lists a context window of up to 1 million 11 tokens. At full length, the cache can become the binding cost before the active weights do. Bandwidth charges its own tax. When the chosen experts sit on different GPUs, every token triggers network hops, and small batches feel that delay the most.

Deciding Before the October 27 Weights

Mistral is selling control as much as capability. VentureBeat reports Mistral says it trained the model in its own European data centers, with a significant share of its training data spanning more than 160 12 languages. For a regulated buyer inside the EU, jurisdiction can matter more than a benchmark lead.

The best open weights model from US or Europe on aggregated benchmarks.· MISTRAL AI · LAUNCH POST ON X

Geopolitics sharpens the pitch. CNBC reports that open-source AI has become a contentious topic in the US-China technology race, and that the most capable open models have been Chinese. Mistral called Large 4 "the best open weights model from US or Europe on aggregated benchmarks" in its launch post. That is the vendor scoring itself. Independent tests have not landed yet.

The rollout tells you the most. VentureBeat describes a roughly three-week testing window with developers, cybersecurity leaders and government authorities before the weights ship. Anadolu Agency's account of the same window adds vetted partners. Mistral is handing a cyber-capable model to vetted users first, which says a lot about who it expects its early buyers to be.

The sources disagree on the training bill. VentureBeat reports roughly two months on 4,000 08 Nvidia Grace Blackwell GPUs, while Mistral itself, like Anadolu Agency, gives 3,800. Take the larger figure: 4,000 31 GPUs for about 60 days comes to around 240,000 GPU-days. Anadolu Agency also reports Mistral raised €3 billion 14 ($3.37 billion) in a Series D round.

That capital story should shape how a buyer judges the model. Sparsity lowers the cost of using a frontier-sized model. It does nothing for the cost of training one, so the number of labs shipping models like this stays small. For your business, one figure matters: the cash spent per task your system completes correctly.

Some calls can be made today. If you run regulated EU workloads where data residency is a hard rule, start testing the preview now through the mistral-large-4 API identifier in Mistral's documentation. If you run a low-traffic internal tool on one server, keep your dense model and watch. Sparse models pay off only when steady traffic keeps the experts busy. A quiet deployment just pays rent on idle weights.

Three things would change the call. The biggest is the license text on October 27. Published API prices and independent throughput numbers at realistic batch sizes would matter almost as much. A permissive license with small quantized checkpoints would let more teams self-host. A restrictive one would leave Le Chonk an API product with a sovereignty story attached.

It is unclear whether Mistral counts the vision encoder and shared attention layers inside its 49 billion 07 active figure. Labs do not share a standard, so active counts across vendors may not line up. We think the trillion on the box is the least useful number Mistral published this month.

Where Le Chonk's real bill actually hides

MEMORY BILL
2.1 TB 28

The trillion comes back as rented memory.

Only 49 billion 25 weights compute per token, yet all 1.05 trillion must stay loaded. At 16-bit that is about 2.1 terabytes before the KV cache, which can bind first near the 1 million 11 token limit.

LICENSE TEXT
OCT 27

The license decides who can self-host.

Weights reportedly arrive October 27 under a custom Mistral license, and no API pricing appears in the materials VentureBeat reviewed. A permissive license with small quantized checkpoints widens self-hosting, while a restrictive one keeps Le Chonk an API product.

EU RESIDENCY
160 12

European training gives regulated buyers a jurisdiction story.

VentureBeat reports Mistral trained the model in its own European data centers, with training data spanning more than 160 languages. For EU workloads with hard data residency rules, that can outweigh a benchmark lead.

2031: Memory Bills Outgrow Compute Bills

The dated facts are thin, and we should say so. Mistral opened the preview on October 6, 2026 32, and VentureBeat reports weights on October 27, a 21-day gap. That is one release. Any five-year line drawn from it is a guess with stated assumptions.

Suppose the next generations keep the working set near 49 billion 33 and double the expert pool twice by 2031. The pool would reach about 4.2 trillion 34 parameters, and the active share would fall from 4.7% to roughly 1.2%. At 8-bit precision, stored weights would grow from about 1.05 terabytes 35 to about 4.2 terabytes. Per-token arithmetic would barely move.

If that line holds, compute stops being the scarce input for open frontier models, and memory and interconnect take its place. The market then splits into layers. A few labs fund training runs on thousands of GPUs. Serving companies compete on routing and quantization. Application teams compete on cost per task.

The risk is lopsided. A team that builds a cost-per-task test now spends perhaps a week, then scores each later release in an afternoon. A team that sizes hardware to a headline parameter count can lock capital into the wrong memory-to-compute ratio for years. Both choices compound with every release. Only the first compounds in your favor.

Against closed US labs, Mistral is counterpositioning. It offers downloadable weights trained on European soil, something a closed vendor would struggle to match without undercutting its own API business. Against Chinese open models, CNBC's framing suggests the pitch is jurisdiction and trust. By 2031, the evidence suggests, the label on a model will matter less than where its weights can legally run.

Price One Task Before the Weights Land

Step one: pull 25 37 real tasks from last month's work, each with an answer you already know is correct. If your work overlaps the areas Mistral targets, include tasks such as code review or financial document checks. Keep the set small enough to rerun in an hour.

Step two: run all 25 37 through the mistral-large-4 preview and through the model you use now. Log input tokens, output tokens, pass or fail, and time to first token. Time to first token is the wait between sending a request and seeing the first word back, and it is the delay your users feel.

Step three: compute cost per completed task. Add up the token spend and divide by the number of tasks that passed. VentureBeat reports no Mistral API pricing in the materials it reviewed, so leave Mistral's rate as a blank cell and fill it in when the price list appears.

Step four: build a one-row self-hosting sheet. Multiply 1.05 trillion 28 parameters by the bytes per weight at each precision you could run, then add KV cache for your typical prompt length. Compare that total with the GPU memory in the cloud quote you would actually sign.

Step five: on October 27, read the license before anyone downloads anything. VentureBeat says the weights are expected to come under a custom Mistral license, so check commercial use and fine-tuning rights line by line.

Expect breakage. Prompts tuned for your current model will fail on a new one, and VentureBeat reports Mistral is still tuning Le Chonk during the preview. Treat early failures as data about your prompts as much as about the model. Rerun the same 25 37 tasks every week until the weights ship, and keep every result in one sheet.

DOJO · BUILD THIS WEEKEND

Price one task before the weights land

  1. Pull 25 37 graded tasks. Take real work from last month with known correct answers and run each through the mistral-large-4 preview and your current model, logging input tokens, output tokens, pass or fail, and time to first token.
  2. Compute cost per completed task. Sum token spend and divide by tasks that passed. Leave Mistral's rate as a blank cell until a price list appears.
  3. Build a one-row self-hosting sheet. Multiply 1.05 trillion 28 parameters by bytes per weight at each precision, add KV cache for your typical prompt length, and compare the total with GPU memory in a real cloud quote.
Practice: Diagnose and Fix Bad Output
THE BOTTOM LINE

Judge Le Chonk by cost per task and the license on October 27.

Sparsity lowers the price of running a frontier-sized model and leaves the memory bill intact. Teams with steady traffic and EU residency needs have a reason to test the preview now. Teams running quiet internal tools should keep their dense model and watch. A 25 37-task benchmark built this week will score every later release in an afternoon.

LISTEN · AUDIO BRIEFINGThe conversation · ~8 min
WATCH · VISUAL NARRATIVEAnimated breakdown · ~3 min
PLAY · YOUTUBE
EDITORIAL RECEIPTKODA-20261007-DE93233ACF83
As of07 October 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE24 CHECKED + 13 COMPUTED · 11 VERIFIED · 13 REPORTED
11 verified13 reported13 computed
  1. 01Mistral announced Mistral Large 4 on October 6, 2026VERIFIEDTRUEMODELmistral.ai
  2. 02Mistral Large 4 is nicknamed Le ChonkVERIFIEDTRUEMODELmistral.ai
  3. 03The weights are due by the end of October, reportedly on October 27, under a custom Mistral license.REPORTEDMOSTLY TRUEFEATURECORRECTED IN COPYmistral.ai
  4. 04Mistral's launch post for Mistral Large 4 on X led with "1T parameters" and put "49B active" secondVERIFIEDTRUEATTRIBUTIONmistral.ai
  5. 05Mistral's documentation says Mistral gave no API pricing for Mistral Large 4 in the materials VentureBeat reviewedREPORTEDMOSTLY TRUEPRICEdocs.mistral.ai
  6. 06Mistral's documentation lists Mistral Large 4 at 1.05 trillion total parametersVERIFIEDTRUESTATdocs.mistral.ai
  7. 07Mistral Large 4 has 49 billion active parameters per tokenREPORTEDMOSTLY TRUESTATmistral.ai
  8. 08VentureBeat reports roughly two months on 4,000 Nvidia Grace Blackwell GPUs, while Mistral itself, like Anadolu Agency, gives 3,800.REPORTEDMOSTLY TRUESTATCORRECTED IN COPYventurebeat.com
  9. 09Mistral called Mistral Large 4 "the best open weights model from US or Europe on aggregated benchmarks" in its launch postREPORTEDMOSTLY TRUEBENCHMARKmistral.ai
  10. 10Mistral's documentation lists a 1.6-billion-parameter vision encoder for Mistral Large 4VERIFIEDTRUEFEATUREdocs.mistral.ai
  11. 11Mistral's documentation lists a context window of up to 1 million tokens for Mistral Large 4VERIFIEDTRUEFEATUREdocs.mistral.ai
  12. 12VentureBeat also reports Mistral says it trained it in its own European data centers, with a significant share of its training data spanning more than 160 languages.REPORTEDMOSTLY TRUESTATCORRECTED IN COPYmistral.ai
  13. 13Anadolu Agency reports Mistral Large 4 was trained on 3,800 Nvidia Grace Blackwell GPUsVERIFIEDTRUESTATmistral.ai
  14. 14Anadolu Agency reports Mistral raised €3 billion ($3.37 billion) in a Series D roundREPORTEDMOSTLY TRUESTATcnbc.com
  15. 15Mistral Large 4 is a mixture-of-experts modelVERIFIEDTRUEFEATUREmistral.ai
  16. 16Mistral's documentation lists the API identifier mistral-large-4 for the Mistral Large 4 previewVERIFIEDTRUEFEATUREmistral.ai
  17. 17VentureBeat says the weights are expected to come under a custom Mistral license, so check commercial use and fine-tuning rights line by line.REPORTEDMOSTLY TRUEATTRIBUTIONCORRECTED IN COPYventurebeat.com
  18. 18Claim removed during the check; its text is not republished.REPORTEDMIXEDATTRIBUTIONCUT FROM COPYventurebeat.com
  19. 19VentureBeat reports Mistral trained Mistral Large 4 in Mistral's own European data centersVERIFIEDTRUEATTRIBUTIONmistral.ai
  20. 20CNBC reports that open-source AI has become a contentious topic in the US-China technology race, and that the most capable open models have been Chinese.REPORTEDMOSTLY TRUEATTRIBUTIONCORRECTED IN COPYcnbc.com
  21. 21CNBC reports that the most capable open models have been ChineseVERIFIEDTRUEATTRIBUTIONcnbc.com
  22. 22VentureBeat describes a roughly three-week Mistral Large 4 testing window with developers, cybersecurity leaders and government authorities before the weights shipREPORTEDMOSTLY TRUEATTRIBUTIONmistral.ai
  23. 23Anadolu Agency's account of the Mistral Large 4 testing window adds vetted partners as participantsREPORTEDMOSTLY TRUEATTRIBUTIONmobile.aa.com.tr
  24. 24VentureBeat reports Mistral is still tuning Mistral Large 4 (Le Chonk) during the previewREPORTEDMOSTLY TRUEATTRIBUTIONventurebeat.com
  25. 25Mistral Large 4's 49 billion active parameters are about 4.7% of its 1.05 trillion total parametersCOMPUTEDCOMPUTED
  26. 26About 95% of Mistral Large 4's parameters sit idle in memory per token, waiting for a router to call themCOMPUTEDCOMPUTED
  27. 27Mistral Large 4's 1,050 billion total parameters are roughly 21 times its 49 billion active parametersCOMPUTEDCOMPUTED
  28. 28At 16-bit precision, Mistral Large 4's 1.05 trillion weights need about 2.1 terabytes of memoryCOMPUTEDCOMPUTED
  29. 29At 8-bit precision, Mistral Large 4's weights need about 1.05 terabytes of memoryCOMPUTEDCOMPUTED
  30. 30At 4-bit precision, Mistral Large 4's weights need roughly 525 gigabytes of memoryCOMPUTEDCOMPUTED
  31. 314,000 GPUs for about 60 days comes to around 240,000 GPU-days of Mistral Large 4 trainingCOMPUTEDCOMPUTED
  32. 32The gap between Mistral opening the Mistral Large 4 preview on October 6, 2026 and the October 27 weights release is 21 daysCOMPUTEDCOMPUTED
  33. 33If Mistral keeps the working set near 49 billion and doubles the expert pool twice by 2031, the pool would reach about 4.2 trillion parametersCOMPUTEDCOMPUTED
  34. 34In the 2031 scenario with a 4.2 trillion parameter pool, the active share would fall from 4.7% to roughly 1.2%COMPUTEDCOMPUTED
  35. 35In the 2031 scenario, stored weights at 8-bit precision would grow from about 1.05 terabytes to about 4.2 terabytesCOMPUTEDCOMPUTED
  36. 36Building a cost-per-task test takes perhaps a week, after which each later release can be scored in an afternoonCOMPUTEDCOMPUTED
  37. 37The recommended evaluation pulls 25 real tasks from last month's work, each with a known correct answerCOMPUTEDCOMPUTED

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20261007-DE93233ACF83
Filed underStrategyDeep Dive07 October 2026
Browse the Deep Dive archive

Get the morning Signal

194 editions so far, one a day. Unsubscribe anytime.