Mistral Large 4 Launches in Public Preview as a 1-Trillion Parameter MoE
Nicknamed 'Le Chonk,' the European AI lab's flagship mixture-of-experts model pairs 49 billion active parameters with enterprise cybersecurity and multimodal reasoning.

Key takeaways
- Mistral Large 4 features roughly 1 trillion total parameters and activates 49 billion parameters per token via a sparse mixture-of-experts architecture.
- The model is available in public preview through Mistral's API, with weights scheduled for general release on October 27, 2026.
- API pricing is set at $1.36 per 1 million input tokens, $4.18 per 1 million output tokens, and $0.14 per 1 million cached tokens.
- Training took place over two months in European data centers using 3,800 to 4,000 Nvidia Grace Blackwell GPUs.
On October 6, 2026, Paris-based AI company Mistral AI announced the public preview of Mistral Large 4, a trillion-parameter natively multimodal model built to rival leading frontier systems while preparing for an open-weight release. Codenamed unofficially as ML4 and given the official internal nickname "Le Chonk," the model introduces significant architectural scale for the European developer. This official preview launch makes the system directly accessible through Mistral's cloud API ahead of a full open-weight release scheduled for October 27, 2026.
According to Mistral's official announcement, the model was trained from scratch in European data centers on Nvidia Grace Blackwell GPUs. The company is positioning ML4 as a sovereign, self-hostable alternative to closed proprietary systems, emphasizing dual-use cybersecurity workloads, document-level reasoning, and agentic code execution.

Mistral Large 4 Architecture and Training Hardware
Mistral Large 4 is built as a granular mixture-of-experts (MoE) system designed to process multimodal inputs and generate text. The total parameter count stands at approximately 1 trillion parameters—with Marktechpost reporting an exact count of 1.05 trillion—while activating roughly 49 billion parameters per token. In its own blog post, Mistral noted 52 billion active parameters. The architecture incorporates a 1.6-billion-parameter vision encoder and supports a context window of 1 million tokens.
Because of the MoE layout, roughly 4.7 percent of the network's weights activate on any individual token, lowering compute demands during inference even though the full parameter set must reside in memory. Details such as the expert count, top-k routing parameters, and specific layer distribution will be released when weights drop at the end of the month.
Training the model required roughly two months in Mistral's European facilities. Hardware counts vary slightly between reports: Mistral cited 3,800 Nvidia Grace Blackwell GPUs, while VentureBeat reported approximately 4,000 GPUs consuming around 10 megawatts of power, according to The Next Web. The pretraining corpus was heavily multilingual, covering over 160 languages, including all official European Union languages. Reinforcement learning post-training remains underway on an infrastructure fleet processing 33 billion tokens daily.
The model's lighthearted moniker originated from a fictional internet meme in June 2026 dubbed "Le Chaton Fat," which CEO Arthur Mensch humorously referenced on social media before the lab formally embraced "Le Chonk."
Benchmark Performance Across Coding and Workflows
Performance metrics published by Mistral highlight targeted strength across agentic coding, professional enterprise workflows, and visual grounding tasks.
On the DeepSWE v1.1 software engineering evaluation, ML4 scored 61.7 percent (rounded to 62 percent in several disclosures). This places it ahead of DeepSeek V4 Pro 0813 at 57 percent, Qwen 3.8 Max at 51 percent, and Reflection AI's Beam at 44 percent, while running closely alongside GLM-5.3 at 61 percent. Mistral's overall Coding Agent Index score stood at 49.8 percent, supplemented by a 59.4 percent result on SWE-Atlas-QnA and 28.3 percent on Terminal-Bench 4.0.
In blind human evaluations conducted by Surge AI, annotators rated model outputs on a scale from 1 to 5. ML4 achieved an average rating of 3.74, ranking second among five evaluated systems, behind Claude Opus 5 at 4.22, but ahead of GLM-5.3 at 3.60 and Kimi K3 at 3.59.

For enterprise tasks, ML4 scored 67 percent on the FinWorkBench (Finch) financial benchmark, matching DeepSeek V4 Pro and exceeding GLM-5.3 (65 percent). On Harvey's Legal Agent benchmark, ML4 registered a 15 percent task-pass rate, surpassing open-weight peers such as Kimi K3 at 12.92 percent and GLM-5.3 at 8.33 percent.
Visual grounding benchmarks show the model locating specific objects in complex visual inputs. ML4 reached 42 percent on the Dense200 benchmark—a score shared with OpenAI's GPT-6 Astra—and posted 73 percent on the DIOR-RSVG remote-sensing test, outperforming GPT-6 Astra's reported 68 percent.
Cybersecurity Defenses and Self-Hosted Operations
A primary focus of the ML4 rollout centers on defensive cybersecurity operations. In evaluations on the Artificial Analysis Cyber Index, ML4 placed in the global top five. On CyberGym-E2E, which requires models to reproduce and patch real-world software flaws, ML4 scored 82 percent. It resolved 93 percent of tasks on Cybench, an evaluation of 40 cybersecurity challenges.
Mistral noted that several closed models score near zero on specific reproduction tasks due to safety guardrails refusing vulnerability generation. Defending internal infrastructure requires validating whether flaws exist, a workflow that rigid refusal policies can obstruct. Because ML4 weights can be deployed on private servers, security teams can conduct code auditing and patch creation without external intervention.
At the same time, the lab reported safety protections against misuse. On Lakera's B3 benchmark, ML4 resisted 93.3 percent of indirect prompt injection attacks, and achieved a 1.691 score out of 2.0 on KORABench.
Cybersecurity resilience remains urgent for enterprise platforms. In May 2026, Mistral encountered the Shai-Hulud supply-chain attack via the TanStack library compromise, which temporarily forced the invalidation of affected npm and PyPi packages, according to Inc.. Threat researchers at Cisco Talos have separately monitored automated toolkits such as CLOSEDQUORUM that attempt to query multi-model consensus networks for malicious decisions.

Availability, API Pricing, and European Sovereignty
Mistral Large 4 is currently accessible via Mistral Studio API endpoints. Standard API rates are $1.36 per 1 million input tokens and $4.18 per 1 million output tokens, with cached input queries billed at $0.14 per 1 million tokens. Supported capabilities include structured outputs, function calling, batch processing, document Q&A, and agent endpoints.
Full model weights are scheduled to be published on October 27, 2026. Prior to release, vetted government authorities, cybersecurity partners, and enterprise developers are conducting red-teaming with reduced safety filtering to assess security workflows.
Development was backed by Mistral's €3 billion Series D funding round completed in September 2026, which valued the company at over €21 billion (approximately $24 billion). Pierre Stock, Mistral's vice president of science, stated at a recent press briefing that the startup is "definitely closing the gap" with global rivals. With over 125 enterprise clients including ASML, Airbus, and HSBC, Mistral plans to expand its data center compute through the first half of 2027 and use ML4 as the base for specialized and optimized Mistral models.
Frequently asked questions
When will Mistral Large 4 weights be available for download?
Mistral AI plans to release the model weights on October 27, 2026, following a three-week red-teaming and testing phase.
What are the parameter counts for Mistral Large 4?
The model has approximately 1 trillion total parameters (reported as 1.05 trillion) and activates 49 billion parameters per token (with Mistral citing 52 billion active parameters).
How much does Mistral Large 4 cost on the API?
The preview API is priced at $1.36 per 1 million input tokens, $4.18 per 1 million output tokens, and $0.14 per 1 million cached input tokens.
What hardware was used to train the model?
Mistral trained the model over two months using between 3,800 and 4,000 Nvidia Grace Blackwell GPUs in its European data centers.
Sources
- Introducing Mistral Large 4Mistral · Oct 6, 2026 · Official
- Mistral debuts Large 4 'Le Chonk', a 1-trillion parameter text output model with high benchmarks planned for open weights releaseventurebeat.com · Oct 6, 2026
- Europe’s Mistral launches Large 4 to challenge China’s lead in open AI modelsTNW | Launch · Oct 6, 2026
- Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Open-Weight Multimodal MoEmarktechpost.com · Oct 6, 2026
- Mistral Just Revealed a 1-Trillion-Parameter AI Model. It Has a Very Funny Nicknameinc.com · Oct 7, 2026
How this story was made: the newsroom picked it up from Google News, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (33 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.
Published October 8, 2026 at 02:09 UTC


