Mistral's documentation lists its newest flagship at 1.05 trillion 25 parameters, with 49 billion active per token. That is about 4.7% of the model, by our math. The other 95% 26 sit in memory and wait for a router to call them.
Mistral announced Mistral Large 4, nicknamed Le Chonk, on October 6, 2026 01. Its launch post on X led with "1T 04 parameters" and put "49B active" second. Anyone paying to run the model should read those numbers in the opposite order. The 49 billion 25 drives compute per token, and the 1.05 trillion drives the memory you must rent to host it.
Nobody outside Mistral knows the real bill yet. Mistral's documentation says the company gave no API pricing in the materials it reviewed. The weights are due by the end of October, reportedly on October 27, under a custom Mistral license. Until then, every cost claim about Le Chonk, including ours, is arithmetic on published specs.
The Capacity Bank Behind 49 Billion
Mistral Large 4 is a mixture-of-experts model. The network is split into many smaller sub-networks called experts. A router sends each token to a few of them. Think of it as a capacity bank: about a trillion parameters on deposit, about 49 billion 07 withdrawn per token. Divide 1,050 billion 27 by 49 billion and the bank is roughly 21 times larger than any single withdrawal.
What it takes to keep 1.05 trillion weights loaded
That framing sorts every Le Chonk fact into four accounts. Capacity is the full 1.05 trillion 06, which Mistral's documentation lists alongside a 1.6 10-billion-parameter vision encoder. Compute is the 49 billion 07 active weights, so per-token arithmetic lands near a dense model of that size. Memory is everything that has to stay loaded. Bandwidth is the traffic of tokens moving between experts on different chips.
Memory is where the trillion comes back. At 16 28-bit precision each weight takes two bytes, so 1.05 trillion weights need about 2.1 terabytes. At 8-bit precision that drops to about 1.05 terabytes 29, and at 4-bit to roughly 525 gigabytes 30. And those totals come before the KV cache, the stored attention data for every token already in the prompt.
The KV cache grows with context length, and Mistral's documentation lists a context window of up to 1 million 11 tokens. At full length, the cache can become the binding cost before the active weights do. Bandwidth charges its own tax. When the chosen experts sit on different GPUs, every token triggers network hops, and small batches feel that delay the most.
Deciding Before the October 27 Weights
Mistral is selling control as much as capability. VentureBeat reports Mistral says it trained the model in its own European data centers, with a significant share of its training data spanning more than 160 12 languages. For a regulated buyer inside the EU, jurisdiction can matter more than a benchmark lead.
Geopolitics sharpens the pitch. CNBC reports that open-source AI has become a contentious topic in the US-China technology race, and that the most capable open models have been Chinese. Mistral called Large 4 "the best open weights model from US or Europe on aggregated benchmarks" in its launch post. That is the vendor scoring itself. Independent tests have not landed yet.
The rollout tells you the most. VentureBeat describes a roughly three-week testing window with developers, cybersecurity leaders and government authorities before the weights ship. Anadolu Agency's account of the same window adds vetted partners. Mistral is handing a cyber-capable model to vetted users first, which says a lot about who it expects its early buyers to be.
The sources disagree on the training bill. VentureBeat reports roughly two months on 4,000 08 Nvidia Grace Blackwell GPUs, while Mistral itself, like Anadolu Agency, gives 3,800. Take the larger figure: 4,000 31 GPUs for about 60 days comes to around 240,000 GPU-days. Anadolu Agency also reports Mistral raised €3 billion 14 ($3.37 billion) in a Series D round.
That capital story should shape how a buyer judges the model. Sparsity lowers the cost of using a frontier-sized model. It does nothing for the cost of training one, so the number of labs shipping models like this stays small. For your business, one figure matters: the cash spent per task your system completes correctly.
Some calls can be made today. If you run regulated EU workloads where data residency is a hard rule, start testing the preview now through the mistral-large-4 API identifier in Mistral's documentation. If you run a low-traffic internal tool on one server, keep your dense model and watch. Sparse models pay off only when steady traffic keeps the experts busy. A quiet deployment just pays rent on idle weights.
Three things would change the call. The biggest is the license text on October 27. Published API prices and independent throughput numbers at realistic batch sizes would matter almost as much. A permissive license with small quantized checkpoints would let more teams self-host. A restrictive one would leave Le Chonk an API product with a sovereignty story attached.
It is unclear whether Mistral counts the vision encoder and shared attention layers inside its 49 billion 07 active figure. Labs do not share a standard, so active counts across vendors may not line up. We think the trillion on the box is the least useful number Mistral published this month.
Where Le Chonk's real bill actually hides
The trillion comes back as rented memory.
Only 49 billion 25 weights compute per token, yet all 1.05 trillion must stay loaded. At 16-bit that is about 2.1 terabytes before the KV cache, which can bind first near the 1 million 11 token limit.
The license decides who can self-host.
Weights reportedly arrive October 27 under a custom Mistral license, and no API pricing appears in the materials VentureBeat reviewed. A permissive license with small quantized checkpoints widens self-hosting, while a restrictive one keeps Le Chonk an API product.
European training gives regulated buyers a jurisdiction story.
VentureBeat reports Mistral trained the model in its own European data centers, with training data spanning more than 160 languages. For EU workloads with hard data residency rules, that can outweigh a benchmark lead.
2031: Memory Bills Outgrow Compute Bills
The dated facts are thin, and we should say so. Mistral opened the preview on October 6, 2026 32, and VentureBeat reports weights on October 27, a 21-day gap. That is one release. Any five-year line drawn from it is a guess with stated assumptions.
Suppose the next generations keep the working set near 49 billion 33 and double the expert pool twice by 2031. The pool would reach about 4.2 trillion 34 parameters, and the active share would fall from 4.7% to roughly 1.2%. At 8-bit precision, stored weights would grow from about 1.05 terabytes 35 to about 4.2 terabytes. Per-token arithmetic would barely move.
If that line holds, compute stops being the scarce input for open frontier models, and memory and interconnect take its place. The market then splits into layers. A few labs fund training runs on thousands of GPUs. Serving companies compete on routing and quantization. Application teams compete on cost per task.
The risk is lopsided. A team that builds a cost-per-task test now spends perhaps a week, then scores each later release in an afternoon. A team that sizes hardware to a headline parameter count can lock capital into the wrong memory-to-compute ratio for years. Both choices compound with every release. Only the first compounds in your favor.
Against closed US labs, Mistral is counterpositioning. It offers downloadable weights trained on European soil, something a closed vendor would struggle to match without undercutting its own API business. Against Chinese open models, CNBC's framing suggests the pitch is jurisdiction and trust. By 2031, the evidence suggests, the label on a model will matter less than where its weights can legally run.
Price One Task Before the Weights Land
Step one: pull 25 37 real tasks from last month's work, each with an answer you already know is correct. If your work overlaps the areas Mistral targets, include tasks such as code review or financial document checks. Keep the set small enough to rerun in an hour.
Step two: run all 25 37 through the mistral-large-4 preview and through the model you use now. Log input tokens, output tokens, pass or fail, and time to first token. Time to first token is the wait between sending a request and seeing the first word back, and it is the delay your users feel.
Step three: compute cost per completed task. Add up the token spend and divide by the number of tasks that passed. VentureBeat reports no Mistral API pricing in the materials it reviewed, so leave Mistral's rate as a blank cell and fill it in when the price list appears.
Step four: build a one-row self-hosting sheet. Multiply 1.05 trillion 28 parameters by the bytes per weight at each precision you could run, then add KV cache for your typical prompt length. Compare that total with the GPU memory in the cloud quote you would actually sign.
Step five: on October 27, read the license before anyone downloads anything. VentureBeat says the weights are expected to come under a custom Mistral license, so check commercial use and fine-tuning rights line by line.
Expect breakage. Prompts tuned for your current model will fail on a new one, and VentureBeat reports Mistral is still tuning Le Chonk during the preview. Treat early failures as data about your prompts as much as about the model. Rerun the same 25 37 tasks every week until the weights ship, and keep every result in one sheet.
Price one task before the weights land
- Pull 25 37 graded tasks. Take real work from last month with known correct answers and run each through the mistral-large-4 preview and your current model, logging input tokens, output tokens, pass or fail, and time to first token.
- Compute cost per completed task. Sum token spend and divide by tasks that passed. Leave Mistral's rate as a blank cell until a price list appears.
- Build a one-row self-hosting sheet. Multiply 1.05 trillion 28 parameters by bytes per weight at each precision, add KV cache for your typical prompt length, and compare the total with GPU memory in a real cloud quote.
Judge Le Chonk by cost per task and the license on October 27.
Sparsity lowers the price of running a frontier-sized model and leaves the memory bill intact. Teams with steady traffic and EU residency needs have a reason to test the preview now. Teams running quiet internal tools should keep their dense model and watch. A 25 37-task benchmark built this week will score every later release in an afternoon.
