🔍 Read the full analysis: Mistral Large 4 And The Ongoing Race To The AI Frontier on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Mistral introduced Mistral Large 4 in API preview on October 6, 2026, describing it as a one-trillion-parameter mixture-of-experts model. Artificial Analysis gives the preview an Intelligence Index score of 38, below several leading US and Chinese models in a comparison published the next day. Its weights are not yet publicly downloadable, and reported performance does not settle how it will fare on specific workloads.
Mistral AI released Mistral Large 4 in public API preview on October 6, introducing a one-trillion-parameter mixture-of-experts model as it competes for a place among leading AI providers. A benchmark snapshot from Artificial Analysis dated October 7 gives the model an Intelligence Index score of 38, behind several US and Chinese models, while its weights are still scheduled for a later release. For more on its potential agent performance, see our analysis.
Mistral says Large 4 has one trillion total parameters, with 49 billion active parameters, and accepts text and images. The company says it trained the model on its own infrastructure in Europe and is continuing to improve it. The release announced so far is an API preview; the weights are scheduled for release later in October, but were not publicly downloadable as of October 7.
In the Artificial Analysis comparison cited in the source material, Large 4 Preview scored 38 on the Intelligence Index. The same snapshot lists Anthropic’s Claude Opus 5.5 at 58, Google’s Gemini 4 Argon at 53, OpenAI’s GPT-6.1 Sol at 52, China’s GLM-5.3 at 45 and Kimi K3 at 44. DeepSeek V4.1 Flash scored 39. GPT-6 Luna also scored 38; Cohere Command A+ scored 13.
Those scores are a dated snapshot, not a direct forecast of success on a particular task. The comparison uses named reasoning settings that are not evaluations at identical compute budgets. Developer locations identify the companies, not where a user’s API request is handled. Artificial Analysis reports a context capacity of about 512,000 tokens, but context capacity measures how much material can be included, not whether a model reasons reliably over it.
Frontier AI • Release snapshot / 07 Oct 2026
Mistral Large 4 And The Ongoing Race To The AI Frontier
Mistral’s one trillion parameter model arrives in API preview with 49 billion active parameters. An early benchmark snapshot puts it in a crowded field, while weights, workload reliability and full pricing remain open questions.
Artificial Analysis snapshot, October 7
Weights scheduled for later in October
Public API preview
Mixture-of-experts design
Per Mistral’s description
Capacity is not reliability
01 / What’s in the release
A preview with a larger promise
Mistral says Large 4 accepts text and images, was trained on its own infrastructure in Europe, and is still being improved.
Architecture
Scale through experts
The model is described as a one trillion parameter mixture-of-experts system, with 49 billion active parameters.
Access today
API preview only
As of the October 7 report, developers could test through the API, but could not publicly download the weights.
Geography
European development
Mistral says training used infrastructure in Europe. Developer location alone does not show where API requests are processed.
02 / Benchmark snapshot
The index puts it mid-pack
Scores are points on an aggregate index, not percentages of intelligence or a forecast for an individual task.
Snapshot dated October 7, 2026. Named reasoning settings differ and are not evaluations at identical compute budgets. Results may change as models and benchmarks evolve.
03 / What the score can and can’t say
Benchmarks meet real workflows
Aggregate comparisons help orient a choice. They do not settle how a model will perform on a developer’s own workload.
Agentic work
Reliability compounds
Agent systems plan, call tools, interpret results and carry decisions across steps. An early mistake can shape everything that follows, so dependable execution matters alongside headline capability.
Evidence limits
Test the work itself
The source author reported hallucinations in personal use and advised against long, demanding agentic tasks. That is an individual assessment, not a controlled comparison of error rates.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
Thorsten Meyer · Author of the October 7 source report04 / What comes next
From preview to evidence
The next milestone may broaden access and make independent evaluation easier. Its timing and terms need confirmation.
Preview API
Available to test as reported on October 7.
Planned weights
Scheduled by Mistral for later in October 2026.
Workload testing
Measure coding, research and agent tasks directly.
Reassess fit
Use updated reliability, benchmark and pricing data.
05 / Key questions
What developers should know
A concise guide to the launch status and the limits of the available evidence.
What did Mistral announce?
A public API preview on October 6, 2026, described as a one trillion parameter MoE model with 49 billion active parameters.
Are the weights available?
Not in the October 7 account. Mistral had scheduled a later October release; the supplied report offers no later confirmation.
Does a score of 38 predict failure?
No. It is an aggregate benchmark result and does not establish success or failure on a specific workflow.
Is there a cost advantage?
The supplied material ends before giving complete figures, so it does not establish a price or per-task cost comparison.
The Benchmark Gap for Developers
The launch puts a new European-developed model into the market, but the early comparison suggests that Large 4 is not yet at the same aggregate benchmark level as several prominent rivals. For developers choosing a model for complex coding, research or agentic workflows, that is relevant evidence—but not a verdict on every application. Index points are not percentages of intelligence, and they do not predict the probability of success on an individual task.
Agentic systems plan, call tools, interpret results and carry decisions across multiple steps. Errors early in a workflow can shape later actions, so dependable execution matters alongside headline capability. The source author says their own use of the preview produced hallucinations and argues that this reduced their confidence in assigning it long tasks. That is an individual experience, not a controlled comparison of hallucination rates. Mistral’s stated strengths in agentic coding and specialized professional tasks likewise need testing on the workloads developers actually use.
The release also matters beyond model rankings. Mistral says it trained Large 4 on infrastructure in Europe, a development relevant to the region’s AI capacity. That fact does not, by itself, establish that the model is the best choice for a given job or that an API request is processed in Europe.
As an affiliate, we earn on qualifying purchases.
Preview Now, Weights Later
The current announcement is narrower than a completed open-weight release: Mistral has made a preview API available, with downloadable weights scheduled later in October. The distinction affects who can test the model and how. At the time of the cited October 7 account, developers could access the preview through the API, but could not independently download the weights.
Artificial Analysis’s scores give readers a way to compare models, but the source cautions that reasoning settings differ and that results may change. The comparison places Mistral below the listed leading US models and several Chinese alternatives, while its score is level with GPT-6 Luna and above Cohere Command A+. That counterexample means the evidence does not support saying every competing lab scores higher.
The source author’s recommendation against choosing Large 4 for demanding agentic work is an assessment of the current preview, informed by benchmark results and personal use. It is not a finding that the model cannot complete such work. The source also raises a cost comparison with DeepSeek, but the supplied material ends before giving figures or a complete cost analysis, so no price conclusion can be established here.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, author of the October 7 source report
machine learning model evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Weights, Reliability and Costs
Several questions remain open. The model weights had not been released as of October 7, and the source provides no later confirmation that the planned release occurred. It is also unclear how much the preview will change as Mistral continues training and refinement, or whether the final model will score differently.
The supplied benchmark comparison does not establish performance on particular coding, research or business tasks. The author’s hallucination observations are personal and not based on a controlled study; no comparative hallucination rates are provided. The source material also cuts off during its cost discussion, so it does not supply enough information to verify the claim that DeepSeek is cheaper per task or to quantify any price difference.
As an affiliate, we earn on qualifying purchases.
The October Weight Release
The next stated milestone is Mistral’s planned release of Large 4’s weights later in October 2026. If released, they would give developers a way to assess the model beyond the current API preview, subject to the terms Mistral sets. The date and availability should be confirmed when the company provides an update.
For now, developers evaluating Large 4 can compare its API performance on their own tasks rather than treating the aggregate index as a substitute for testing. Further benchmark results, independent workload evaluations, reliability data and complete pricing information would help clarify whether the model’s preview limitations persist and where it fits among competing systems.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral announce?
Mistral introduced Mistral Large 4 in public API preview on October 6, 2026. The company describes it as a mixture-of-experts model with one trillion total parameters and 49 billion active parameters.
Are Large 4’s weights publicly available?
Not according to the October 7 source report. Mistral had scheduled the weights for release later in October, but the report said they were not yet publicly downloadable.
How does Large 4 compare in the cited benchmark?
Artificial Analysis gave the preview an Intelligence Index score of 38 in its October 7 snapshot. That was below several listed US and Chinese models, level with GPT-6 Luna and above Cohere Command A+. The comparison uses differing reasoning settings and does not predict success on every task.
Does the score show that Large 4 will fail at agentic work?
No. The score is an aggregate benchmark result, not proof that the model will fail a specific workflow. The source author advises against choosing it for demanding, long agentic tasks, but presents that as a judgment based on benchmarks and personal experience rather than a controlled evaluation.
What is expected next?
Mistral said the model weights were scheduled for release later in October 2026. Further testing, updated benchmark results and complete pricing details could clarify the model’s capabilities and trade-offs.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
