Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Mistral AI launches Mistral Large 4 with 1 trillion parameters - OpenSmartRoute
Mistral AI launches Mistral Large 4 with 1 trillion parameters
Mistral AI released a public preview of its largest model, Mistral Large 4. The model features over one trillion parameters and excels in cybersecurity and coding tasks.
Key points
The model has 1 trillion total parameters and 49 billion active parameters.
It scores 82% on a real vulnerability reproduction test, the highest globally.
Mistral Large 4 is trained entirely on NVIDIA Grace Blackwell GPUs in Europe.
It outperforms most US and European open-weight models in cybersecurity.
Why it matters: Organizations gain sovereign control over security tools without relying on external providers.
By OpenSmartRoute editorial · written through the router by writer-small
Mistral AI launched Mistral Large 4 today. The company released a public preview API for developers. You can try the model immediately on Mistral Studio. Weights will drop by the end of this month. This marks the largest model in the Mistral family.
Mistral AI announces Mistral Large 4 with a public preview API
Mistral AI officially launched Mistral Large 4 today. The company calls the model le Chonk unofficially. It is a major step forward for open-weight performance. You can access the preview API right now on Mistral Studio. Weights will be available by the end of October. This allows engineers to test capabilities before full release.
The public preview lets users explore new features safely. Engineers can run the model in their own environments. They do not need to wait for official weight releases yet. Mistral is collecting feedback from early adopters globally. The team plans to share more details soon on architecture. Further benchmarks and post-training methods will follow later.
Users should try the preview API to see real-world performance. Share your findings with Mistral on social media channels. This helps improve the model before it goes public. Engineers can also compare it against other open-source options. The team invites you to test its limits in coding tasks.
The model architecture includes 1 trillion parameters and native multimodal support
Mistral Large 4 has over one trillion total parameters. It features 49 billion active parameters for inference. This makes it the largest and most capable model yet. The architecture supports native multimodal input and output. It handles text, images, charts, and documents seamlessly.
The model processes complex documents with high precision. It understands dense natural scenes and technical drawings. Visual grounding helps verify mechanical parts in engineering plans. It can inspect gigapixel satellite imagery for disaster response teams. This capability extends to analyzing massive geospatial images efficiently.
Native multimodal support means the model sees and reasons like humans. It combines vision with agentic workflows effectively. The system grounds evidence from PDFs automatically. It scans objects in hard-to-find locations within large datasets. Engineers can leverage this for engineering and earth observation tasks.
Google launched EmbeddingGemma 2, a compact open model that handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types.
The design focuses on reasoning across diverse complex challenges. It handles out-of-distribution investigation tasks well. Malware reverse-engineering remains a key strength area. The model solves problems that require multi-step logical deduction. Its architecture supports long-horizon tool use without losing context.
Cybersecurity benchmarks show top-tier performance against closed models
Mistral Large 4 ranks among the top five globally on the Artificial Analysis Cyber Index. It leads open-weight models outside China by a wide margin. On one specific test, it scored 82% for reproducing and patching vulnerabilities. This is the highest score reported for any model in this category.
The model solved 93% of challenges in the Cybench security competition. These exercises come from real-world security competitions globally. Several closed models like Claude Opus 5.5 refuse similar tasks entirely. They score near zero because safety filters block the work. ML4 does not have these constraints on critical tasks.
Defenders need systems that match offensive capabilities without refusal. Threat actors increasingly jailbreak closed models for cyber attacks. ML4 can analyze malware and write detection rules effectively. It prioritizes vulnerabilities based on real-world impact data. This allows organizations to run advanced security work autonomously.
On the KORA Benchmark, ML4 achieved a score of 1.691. This is the highest measured score among open-source models. The benchmark measures responsible user engagement and safety alignment. ML4 refuses malicious cybersecurity requests at higher rates than competitors. It engages more responsibly than any previous Mistral model tested.
Training infrastructure relies on European NVIDIA Grace Blackwell GPUs
The model was trained from scratch in Europe exclusively. Mistral used 3,800 NVIDIA Grace Blackwell GPUs for training. These chips reside in Mistral's own datacenters across the continent. The public preview runs on this same European infrastructure. This ensures data sovereignty and compliance with local laws.
A significant share of training data spans over 160 languages. It includes every official language of the European Union. The model benefits from multilingual context during training phases. This supports global deployment while maintaining regional legal standards. Mistral operates these deployments end-to-end independently.
Forged in Europe supports AI sovereignty goals for customers. Organizations gain control over their AI infrastructure choices. They avoid reliance on foreign cloud providers or digital services. European law governs the data handling and model operations fully. This matters for companies with strict compliance requirements globally.
The infrastructure investment spans research, product development, and performance. Mistral aims to deliver state-of-the-art results through open weights. Customers get autonomy to deploy models on private clouds. On-premise deployment options remain available for sensitive workloads.
Why it matters for enterprise AI sovereignty and safety
Enterprise customers face risks when provider-level refusals block research. Losing access to capabilities mid-incident creates critical security gaps. ML4 pairs top-tier cyber performance with open weights directly. This gives organizations the autonomy to run under their own policies.
Safety filters in closed models can block legitimate vulnerability research. Defenders need systems that match offensive cyber capabilities. They require tools that do not refuse malicious requests blindly. ML4 provides this balance between safety and capability effectively.
AI sovereignty means keeping data and models within jurisdictional boundaries. European law protects the deployment of Mistral Large 4 fully. Customers can audit their AI operations without third-party interference. This is crucial for finance, law, and cybersecurity sectors.
The model excels in verticals where precision and trust matter most. Legal and financial benchmarks show it exceeds GPT-6-Astra significantly. Third-party evaluators found it superior in representative tasks globally. Internal evaluations confirmed its strength in CAD and STEM domains too.
How it compares - what existed before, what this changes and what stays the same
Mistral Large 3 was the previous flagship model for Mistral AI. It served as a strong open-weight alternative to major closed models. ML4 replaces that version with significantly more parameters. The new model has one trillion total parameters instead of fewer billions. Only forty-nine billion of those parameters remain active during inference. This reduction improves speed while keeping raw power high.
ML4 adds native multimodal understanding to its core capabilities. Earlier versions handled text and images separately in some workflows. Now it processes visual data directly alongside code and documents. It retains the same training methodology used by Mistral Forge. Customers can still use custom fine-tuning pipelines for specific needs. The model keeps its focus on European sovereignty and privacy standards.
The main difference lies in cybersecurity performance metrics. ML4 scores much higher than previous iterations on vulnerability tasks. It achieves an 82% score on a real-world patching test. Older models often refused to perform such explicit security work. This shift allows defenders to analyze malware without hitting safety blocks. The change enables organizations to run advanced security tools internally.
What stays the same includes the general-purpose agent architecture. ML4 continues to use tools and gather information autonomously. It remains compatible with existing orchestration layers like LangChain or AutoGen. The API structure for developers follows similar patterns to prior versions. Users can integrate it into their current workflows without major rewrites.
The model also keeps its multilingual training foundation. It supports over 160 languages from its initial data set. This feature helps non-English speakers use the system effectively. Previous models had fewer language options in their training sets. ML4 expands this reach to include every official EU language. The consistency ensures reliable performance across diverse global regions.
Questions this leaves open - what the source does not say and how a reader can check it
The article mentions specific benchmark scores but omits raw test data details. Readers cannot see the exact prompts used for the 82% vulnerability score. Independent labs might need to replicate these conditions to verify results. Open-source researchers could request access to the evaluation datasets directly.
The text states ML4 leads open models outside China by a wide margin. It does not define what "wide margin" means numerically. Competitors like DeepSeek V4 Pro or Qwen3.8 Max have their own public scores. Comparing these numbers requires checking official benchmark leaderboards. Third-party evaluators may publish different rankings based on varied criteria.
Readers do not know the exact distribution of training data sources. The article says a significant share was multilingual but gives no percentages. Data provenance audits remain necessary for full transparency checks. Organizations might want to verify if sensitive information leaked into training sets. Mistral AI could release more detailed data lineage reports later.
The preview API lacks specific latency numbers or throughput figures. Engineers need concrete metrics on token generation speed per second. High-volume users might face queue times during peak usage periods. Performance testing under load conditions remains an open area for investigation.
The article hints at future specialized models but offers no roadmap details. No release dates exist for these upcoming optimized variants. Customers cannot predict when vertical-specific versions will become available. Industry analysts might speculate on timelines based on current development pace.
Safety red-teaming results are mentioned as ongoing but lack specific failure rates. Partners and state authorities access the model with reduced moderation. The exact thresholds for triggering refusals remain unclear to the public. Security researchers could attempt to probe these boundaries themselves.
The text claims ML4 solves 93% of Cybench challenges yet does not list them. Specific problem categories from the forty exercises stay unnamed. Users cannot see which tasks the model handles best or worst. Detailed breakdowns might appear in future technical documentation releases.
Mistral Forge customization options are referenced but not described in depth. Customers do not know what parameter ranges they can adjust safely. Fine-tuning limits and cost implications remain unspoken details. Documentation for these tools could provide clearer guidance soon.
The article stops before finishing the AA-Briefcase benchmark description. It mentions spreadsheets, slides, and PDFs without completing the evaluation scope. Long-form document generation capabilities need further testing evidence. Full report cards on professional deliverables are still missing from public data.
Independent verification remains essential for enterprise adoption decisions. Buyers should consult multiple sources before committing to a purchase. Community feedback loops help fill gaps left by official announcements. Open-source audits can provide additional layers of trust assessment.
What to do with the preview API and upcoming weights
Start by trying the Mistral Large 4 preview API today. Test it on coding, cybersecurity, and multimodal tasks. Compare its outputs against your current open-source models. Look for differences in refusal rates and reasoning depth.
Watch for the weight release at the end of this month. Prepare your deployment environment for private cloud or on-premise use. Ensure you have NVIDIA Grace Blackwell GPUs ready for inference. Check compatibility with your existing toolchains and orchestration layers.
Share your feedback with Mistral on social media platforms. The team values insights from early adopters worldwide. Your input helps refine the model before the official launch. Consider joining their community to stay updated on future features.
Compare ML4 against competitors like DeepSeek V4 Pro or Qwen3.8 Max. Use benchmarks like AutomationBench or AA-Briefcase for objective metrics. Blind human evaluations can also reveal nuanced quality differences. Professional annotators rated ML4 Preview second of five models tested.
Monitor the model's performance on visual grounding and scientific tasks. These areas show particular strength against closed models. Test its ability to generate Hartree–Fock simulations in one shot. Evaluate its math reasoning on formal and applied mathematics problems.