Byte by Mahmud logoByteby Mahmud
The Internet Joked 'Le Chonk' Into Existence. Mistral's Trillion-Parameter Model Just Made It Real.
Technology6 min read4 views

The Internet Joked 'Le Chonk' Into Existence. Mistral's Trillion-Parameter Model Just Made It Real.

Mahmud Hasan

Mahmud Hasan

October 6, 2026

In June, the internet invented a fake French megamodel called "Le Chaton Fat" — thirty trillion parameters, one thousand meows per second, fake benchmark charts and all — and laughed about it for a week. On Tuesday, Mistral released a real trillion-parameter flagship nicknamed "Le Chonk." The joke is dead. The competition with China just got serious.

That's not hype. It's arithmetic. Mistral's models had slipped to 24th place on Artificial Analysis' aggregate intelligence ranking, and the company's last major release was in December 2025. Meanwhile, Chinese open-weight models were eating the open-model world. Mistral Large 4, announced October 6, is the company's answer: public preview through its API today, downloadable weights on October 27. The question that matters isn't whether the nickname is funny. It's whether a European lab with a fraction of the compute can stay in the race.

The meme became a trillion-parameter model

The backstory is worth it, because Mistral is clearly in on the joke. When "Le Chaton Fat" went viral on X and Reddit in June, CEO Arthur Mensch replied, "It's actually le gros chaton." Underneath the shitposting was genuine expectation: by July, TechCrunch was reporting real anticipation of a large open-weight Mistral model, with Mensch and investor Marc Andreessen amplifying the jokes.

Now the model exists, and the fat-cat codename came with it. Internally, ML4 is called "Le Chonk," a deliberate nod to the community that had been rooting for the company to build something enormous. Mistral Large 4 is a granular mixture-of-experts model with 1.05 trillion total parameters — but only 49 billion active during inference. It pairs that with a 1.6-billion-parameter vision encoder, a one-million-token context window, multimodal inputs (text output for now), and fluency in more than 160 languages — every official EU language included. It was trained from scratch in roughly two months on 4,000 Nvidia Grace Blackwell GPUs in Mistral's own European data centers.

Large 3, the previous flagship, was a 675-billion-parameter MoE with 41 billion active parameters, trained on 3,000 Nvidia H200 GPUs. ML4 is a step up in size — but the more interesting claim is about efficiency, and that's where you should read the fine print.

The benchmarks look good. Read the fine print anyway.

Mistral's preliminary numbers: 62–63% on DeepSWE v1.1, a long-horizon software-engineering benchmark. The company's comparison chart lists Reflection's new Beam at 44%, Qwen 3.8 Max at 51%, DeepSeek V4 Pro at 57%, and GLM-5.3 at 61% — which, co-founder Guillaume Lample told Le Monde, puts ML4 "essentially at the same level as the best Chinese models from a month or two ago."

Here's the catch, and credit to VentureBeat for doing the homework: benchmark configuration decides these rankings. The live DeepSWE leaderboard's best-published configurations put GLM-5.3 and Kimi K3 at about 69%, with the top closed models around 74%. So ML4's preview score is competitive against Western open-weight models — it beats Beam convincingly — but it doesn't establish an outright coding lead, and in the best-known configurations it trails the Chinese leaders it's measured against.

The other benchmarks are cleaner. Mistral claims 15% on Harvey's Legal Agent Benchmark, and that one cross-checks: the public Vals.ai leaderboard has Kimi K3 at 12.92%, MiMo V2.6 Pro at 10.83%, and GLM-5.3 at 8.33%. On Finch — 384 enterprise finance-and-accounting tasks from messy real-world-style data, accepted to ACL 2026 — Mistral reports 67%, tied with DeepSeek V4 Pro. Visual grounding scores of 42% on Dense200 and 73% on DIOR-RSVG lead the general-purpose models in its comparison.

The honest summary: strong across coding, legal, finance, and vision; plausibly the best open-weight model built outside China; but every ranking claim is provisional. As of the announcement, ML4 appears on neither Artificial Analysis' public evaluations nor the DeepSWE leaderboard. Mistral gets three weeks of reinforcement-learning tuning and a monitored preview before the weights land. The benchmarks that matter are the ones independent labs run after October 27 — not the vendor's slides.

4,000 GPUs against the world

The most quoted line of the launch is Lample's efficiency claim: 4,000 chips is "two to three times fewer" than what Chinese startups have access to, versus "hundreds of thousands of chips" for the American leaders. Four thousand Grace Blackwell GPUs is roughly 10 megawatts of power; Mistral wants a gigawatt by 2030.

It's an impressive ratio, but the real product was never just the model. Mistral's pitch is a sovereign enterprise stack: the weights, plus the deployment, customization, and engineering around them, sold to governments and enterprises that want to run AI on their own infrastructure with zero-data-retention options. Head of science Pierre Stock put the philosophy bluntly: open-source AI ensures "access to the model can never be cut off," alluding to the White House temporarily restricting access to Anthropic's Mythos and Fable models.

That pitch also explains the cybersecurity positioning. Mistral argues that security teams can't depend entirely on closed providers whose safety systems may refuse dual-use but legitimate defensive requests — code scanning, defensive testing, high-volume security workflows. An open-weight model gives the team control over its own moderation policy. The obvious counter-argument is that the same openness hands the same model to whoever wants to run offense with it — which is exactly why the release is staged: three weeks of preview behind a monitored interface, with security testing completed before the weights go public. "As the sector is rattled by numerous incidents," as Le Monde's framing has it, nobody releases a trillion-parameter model casually anymore.

The business behind it is real, too. Mistral raised €3 billion in September at a post-money valuation above €21 billion — the largest equity fundraising by a European tech company, ever — and says it serves 125+ global enterprises, including Airbus, ASML, and HSBC. Its science team has grown from three researchers to roughly 300. There is one delicious irony: while competing with Chinese labs, Mistral plans to host open Chinese models in its own data centers. Compete with the rival, sell the rival's model, out-run the rival's model. That's not confusion; that's a strategy.

What to actually do with this

First, calibrate your skepticism. "At the same level as the best Chinese models from a month or two ago" is a moving-target claim — the Chinese labs are not pausing for three weeks. And Mistral's models start this race from 24th place on the independent aggregate ranking. The trajectory argument — Large 5 and 6 next year on the same architecture — has to be earned in public benchmarks, not press releases.

Second, if you run an engineering or security team, the preview is worth your time. The API identifier is mistral-large-4, and there's a playground on Mistral's docs page. Benchmark it on your own tasks — coding agents, document workflows, anything you'd never run through a closed provider's moderation filter — because vendor charts have now been shown to be configuration-sensitive. Mistral hasn't published API pricing, so any cost comparison is impossible for now; note that gap.

Third, mark October 27. The weights release — not today's announcement — is the actual event. That's when independent evaluators test the final checkpoint, developers try to reproduce the scores on their own hardware, and we see whether the preview-period reinforcement learning moved the numbers. After that, the "Le Chonk" nickname stops mattering, and what matters is whether a 4,000-GPU European model can keep pace with labs running hundred-thousand-GPU fleets.

Either way, the meme is dead. Long live the cat.

References

Comments

Leave a comment

Your email stays private — only your name is shown.

More in Technology