Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Mistral Large 4: Europe's Trillion-Parameter Security Model - OpenSmartRoute
Mistral Large 4 is Europe's trillion-parameter model. Mistral released a preview of this new system. The company trained it in European data centers. Weights are expected at the end of October. This marks a major shift for the firm. It aims to compete with US giants.
The model has one trillion parameters total. Ninety-nine percent of these remain inactive. Only 49 billion parameters are active during use. Mistral calls this a fine-grained mixture-of-experts architecture. The vision encoder itself holds 1.6 billion parameters. This allows the system to see and understand images.
Engineers can access the preview API now. They must use Mistral Studio for this task. Input costs $0.68 per million tokens. Output costs $2.09 per million tokens. Cached inputs are cheaper at $0.07 per million. These rates apply during the current preview phase.
Mistral plans to release official weights soon. The company expects a full launch by late October. They will share architecture details then. The license and post-training methods will follow too. Engineers should check the documentation for updates.
Benchmark Scores - Performance in security, coding, and visual tasks
An independent index tracks AI performance across many fields. The Artificial Analysis Intelligence Index aggregates ten different benchmarks. Mistral Large 4 scores 38 points on this list. This is a significant jump from its predecessor. Mistral Large 3 only managed 9 points previously.
The gap to the top remains large currently. Claude Opus 5.5 leads with 58 points. It beats ML4 by twenty points in total. Other closed models also lead the rankings. Open-weight models from China compete closely too. GLM-5.2 scores near ML4 at 38 points.
In scientific coding, Mistral performs very well. It scores 49.8 percent in the Coding Agent Index. GPT-6 Astra and Qwen3.8 score higher here. Human evaluators prefer ML4 for STEM tasks. They rate code quality without knowing the model source.
Mistral claims state-of-the-art performance in finance too. Vals.ai evaluation shows it beats GPT-6 Astra. It handles legal and financial documents better. Automated business workflows also show strong results. Mistral trails GLM-5.3 slightly in this area.
Reflection released Beam, an open-weight model that matches GLM 5.2 and Qwen 3.8 on benchmarks while using three to four times less compute.
Visual grounding remains a key strength for ML4. It analyzes gigapixel satellite imagery effectively. Technical drawings require zooming in on details. The Dense 200 benchmark shows 42 percent accuracy. GPT-6 Astra hits 41 percent on the same test.
Security Capabilities - Why rivals refuse to touch vulnerability work
Mistral's main selling point is IT security. Competitors like Claude and GPT-6 refuse certain tasks. Their safety filters block vulnerability reproduction work. Mistral Large 4 reproduces software vulnerabilities instead. It then patches them automatically.
One test asks models to find real bugs. ML4 hits an 82 percent success rate. This is the highest score of any model tested. Claude Opus 5.5 scores near zero on this task. GPT-6 Astra also refuses the vulnerability work entirely.
This highlights a policy difference between providers. Safety filters block legitimate security research. Attackers jailbreak these same models anyway. Losing access mid-incident creates new risks. ML4 is designed to run privately. It supports private cloud and on-premise deployment.
Mistral touts high refusal rates for attacks too. On malicious cyber prompts, it refuses often. Benchmarks include JailbreakBench and AgentHarm. StrongREJECT measures its ability to block harm. ML4 blocks 93.3 percent of attacks in Lakera's benchmark.
The company does not explain how it distinguishes intent. It separates legitimate research from attack prep. Mistral red-teams the model with security firms. Vetted partners and government agencies test it. They get the same version with reduced moderation.
Training Details - Infrastructure, data, and the Series D funding
Mistral trained ML4 on 3,800 Nvidia Grace Blackwell GPUs. These servers sit in European data centers. The company built this infrastructure from scratch. It uses the same training environment as customers. Mistral Forge offers similar tools to partners.
Training involved collaboration with many industries. Finance and manufacturing companies contributed data. Logistics, pharma, shipping sectors participated too. The public sector also provided training inputs. Training covered more than 160 languages total. All official EU languages are included in the dataset.
Reinforcement learning post-training used about 3,000 GPUs. A single training run generates roughly 33 billion tokens daily. The RL run is still ongoing currently. Mistral says it shows no signs of plateauing. Significant improvements are expected over coming weeks.
Compute expansion is funded by a Series D round. This round raised 3 billion euros total. It is the largest equity round ever for a European tech company. Mistral has been building infrastructure in Europe for months. An $830 million loan started in March.
A data center near Paris uses this funding. It plans 200 megawatts of compute capacity by end of 2027. Mistral is shifting focus toward enterprise customers. The chatbot Le Chat renamed to Vibe in May. It now functions as a work tool instead.
Pricing and Access - Costs for the preview API and future weights
Engineers can access ML4 through Mistral Studio today. The public preview runs on existing infrastructure. Mistral charges $0.68 per million input tokens. Output costs $2.09 per million output tokens. Cached inputs cost just $0.07 per million tokens.
Documentation lists prices at double those rates too. Input pricing reaches $1.36 per million tokens. Output pricing hits $4.18 per million tokens. Cached inputs cost $0.14 per million tokens in this tier. These higher rates may apply to future weights.
Mistral plans to release official weights soon. The company expects a full launch by late October. Engineers should check the documentation for updates. Pricing changes might occur before the official release.
ML4 serves as a foundation for specialized models too. Mistral intends to build new generations of tools. These could include domain-specific versions for industries. Finance and legal sectors benefit from this approach.
Why it Matters - The risk of relying on US models for defense
Relying on US models creates cybersecurity risks currently. Europe risks becoming dependent on foreign systems. Mistral CEO Arthur Mensch warned a French parliamentary commission. He stated military codebases should not be scanned by Anthropic's Mythos.
French military vulnerabilities could remain hidden from US tools. ML4 can find the same vulnerabilities that Mythos missed. This capability matters for national defense security. Private clouds and on-premise systems offer more control.
Closed models block legitimate vulnerability research work. Safety filters prevent finding real bugs in software. Attackers jailbreak these models to bypass restrictions. Losing access mid-incident becomes a security risk itself.
US models dominate the current benchmark landscape. Claude Opus 5.5 leads with 58 points on the index. Mistral Large 4 trails by twenty points significantly. European sovereignty remains a business model priority for Mistral.
The Announcement - Mistral Large 4 is Europe's trillion-parameter model
Mistral announced a preview of its new large language model. This model is named Mistral Large 4. Engineers call it ML4 for short. It has one trillion parameters in total. About forty-nine billion of those are active during use. The company trained this system in European data centers. Weights are expected to arrive at the end of October.
Mistral released a public preview today. The API is available now through Mistral Studio. Users can access it immediately for testing. This allows engineers to try the model before full release. Mistral calls this version "le Chonk." It is the largest model the company has built yet.
The model is natively multimodal. This means it handles text and images together. It analyzes documents, charts, and satellite imagery. It can zoom in on gigapixel images automatically. The vision encoder processes one billion parameters for images. This helps with visual grounding tasks specifically.
Mistral claims this is the best open-weight model from the US or Europe. They say it leads across aggregated benchmarks. However, independent indexes show a different picture. Mistral Large 4 scores thirty-eight points on the Artificial Analysis Intelligence Index. Its predecessor, Mistral Large 3, scored only nine points there. The jump represents a major step forward for the company.
The model still trails leading closed models significantly. Claude Opus 5.5 leads with fifty-eight points in that index. GLM-5.2 from Z.ai scores thirty-six points and comes close. Mistral offers GLM-5.2 on its own platform sometimes. The gap to the top remains large despite the progress.
Mistral lists ML4 as proprietary for now. Its weights have not been released to the public yet. This means only authorized users can access it currently. The company plans to release them soon. Mistral expects full weight availability by late October.
Security is the main selling point of this new model. Mistral argues competitors refuse security work entirely. They say rivals block vulnerability research due to safety filters. ML4 handles tasks that others avoid completely. This creates a unique market position for the company.
The company trained ML4 from scratch on specific hardware. It used three thousand eight hundred Nvidia Grace Blackwell GPUs. Training happened in European data centers exclusively. The data covers more than one hundred sixty languages. This includes all official EU languages specifically.
Mistral collaborated with companies across several industries. These partners include finance, manufacturing, logistics, and pharma. They also worked with shipping and public sector firms. The training environment matches what they use for RL today. Mistral plans a European variant entirely under European law.
The Announcement - Mistral Large 4 is Europe's trillion-parameter model (continued)
Mistral says the company operates this infrastructure independently. It does not rely on other service providers. This ensures compliance with European regulations always. Sovereignty remains a key business model priority for them.
Until weights ship, Mistral red-teams the model actively. Security firms, vetted partners, and government agencies get access. They receive the same version with reduced moderation. Expanded cyber capabilities are included in this preview phase.
Mistral turns security into its primary pitch. Defending software starts with proving vulnerabilities exist. Closed models block exactly that kind of work often. Attackers jailbreak those same models to bypass rules. Losing access mid-incident becomes a risk itself.
ML4 is designed to run in private clouds. It can also operate on-premise systems easily. This gives users more control over their data. European sovereignty supports this distributed deployment strategy.
In the Artificial Analysis Cyber Index, ML4 ranks among the top five models worldwide. Among open-weight models outside China, it leads by a wide margin. One test asks a model to reproduce real vulnerabilities. Then it must patch those bugs successfully.
ML4 hits eighty-two percent on that specific test. This is the highest score of any model tested. Claude Opus 5.5 scores near zero on the same task. GPT-6 Astra also scores near zero there. They refuse the task entirely due to safety filters.
The test measures provider policies as much as model capabilities. Mistral argues this proves its unique value proposition. Legitimate vulnerability research needs access to these tools. Safety filters prevent finding real bugs in software often.
At the same time, ML4 has high refusal rates for attacks. It refuses more often than any other open model. Benchmarks include JailbreakBench, StrongREJECT, and AgentHarm. Mistral does not explain how it distinguishes legitimate research from attack prep.
In Lakera's B3 AI Security Benchmark, ML4 blocks ninety-three point three percent of attacks. This demonstrates its ability to handle malicious prompts well. The company claims this makes it safer for enterprise use cases.
ML4 takes second place in coding benchmarks overall. It scores forty-nine point eight percent in the Artificial Analysis Coding Agent Index. Deepseek V4 Pro and Qwen3.8 Max score lower than ML4 here. Mistral ran a blind evaluation with Surge AI too. Professional annotators rated code quality without knowing which model produced it.
ML4 landed second out of five models in that evaluation. It scored three point seven four out of five points total. Claude Opus 5 led with four point two two points. GLM-5.3 and Kimi K3 trailed behind ML4 significantly.
Human evaluators prefer ML4 in STEM and CAD tasks specifically. Both models perform roughly the same in finance and code generally. On image understanding, Mistral claims a major improvement over previous versions. It analyzes documents and technical drawings on its own.
ML4 zooms in to check details automatically. The lead over closed models is slim in visual grounding. On the Dense 200 benchmark, ML4 scores forty-two percent. GPT-6 Astra hits forty-one percent there closely. An evaluation by Vals.ai shows ML4 beats GPT-6 Astra on legal tasks.
It also beats GPT-6 Astra on financial tasks according to Vals.ai. Automated business workflows show mixed results for ML4. It trails GLM-5.3 slightly in those specific areas. But it beats Kimi K3 and Deepseek V4 Pro clearly.
European sovereignty remains a priority for the company's strategy. Mistral plans to release architecture details soon. Engineers should follow official announcements for updates regularly. Weights are expected at the end of October. The company will share the license then too.
Pricing changes might occur before the official release date. Input costs $0.68 per million tokens during this preview phase. Output costs $2.09 per million output tokens currently. Cached inputs cost $0.07 per million tokens today. This helps reduce inference costs for large contexts significantly.
ML4 supports a one million token context window. Engineers should verify this fits their workflow needs exactly. The fine-grained mixture-of-experts architecture allows specialized tasks easily. Vision encoder handles one billion parameters for image analysis specifically.
Mistral intends to build new generations of tools based on ML4. These could include domain-specific versions for industries. Finance and legal sectors benefit from this approach greatly. ML4 serves as a foundation for specialized models too.
Engineers should test ML4 in their private environments now. The preview API is available through Mistral Studio immediately. Comparing ML4 against GPT-6 Astra and Claude Opus 5.5 shows clear differences. Benchmark scores show ML4 leads in security tasks consistently. It trails slightly in scientific coding compared to Qwen3.8.
Visual grounding performance is near the top tier for ML4. Mistral claims state-of-the-art performance in cyber defense, manufacturing, and finance. The company says it beats even closed frontier models in visual grounding.
What to do - How engineers can use Mistral Large 4 today
Engineers should test ML4 in their private environments now. The preview API is available through Mistral Studio. Input costs $0.68 per million tokens during this phase. Output costs $2.09 per million output tokens.
Check the documentation for updated pricing soon. Double rates may apply to future weights. Cached inputs cost $0.07 per million tokens currently. This helps reduce inference costs for large contexts.
ML4 supports a one million token context window. Engineers should verify this fits their workflow needs. The fine-grained mixture-of-experts architecture allows specialized tasks. Vision encoder handles 1.6 billion parameters for image analysis.
Compare ML4 against GPT-6 Astra and Claude Opus 5.5. Benchmark scores show ML4 leads in security tasks. It trails slightly in scientific coding compared to Qwen3.8. Visual grounding performance is near the top tier.
Mistral plans to release architecture details soon. Engineers should follow the official announcements for updates. Weights are expected at the end of October. The company will share the license then too.