Mistral Large 4 May Be Best Outside The US And China, But Agents Are Another Matter
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 May Be Best Outside The US And China, But Agents Are Another Matter on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, a sharp improvement over Mistral’s previous models and a high mark for a model from outside the US and China. The source’s analysis says it still trails leading US and Chinese systems, costs more per benchmark task than two Chinese models that score higher, and raises concerns about verbosity and hallucinations in agent workflows.

Mistral has released Large 4 as a research public preview, and the model scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2. The score makes it a notable European release, but the supplied analysis says it remains behind leading US and Chinese systems and may be a poor fit for long-running agent tasks because of its benchmark performance, token use and reported hallucinations.

Large 4 is described as a 1 trillion-parameter model with 49 billion active parameters, supporting text and image input with text output and a 512,000-token context window. It is available through Mistral’s API as a research public preview. The company has promised to release model weights at the end of October; the supplied source does not state the year. Until weights are released, the model is proprietary, and the source says its licence has not been published.

On Artificial Analysis Intelligence Index v4.3.2, Large 4 received 38.4 points. For comparison, Mistral Large 3 scored 9 and Medium 3.5 scored 14 on the same index version, according to the source. That is a marked improvement, though the reported scores for leading US models range from 51.8 to 57.6. Several Chinese models also score above Large 4, including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5.

Mistral lists API pricing at $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14 per million tokens. The source says Mistral is offering a 50% discount for the first two weeks. It also reports that Mistral says reinforcement learning is still underway, meaning performance scores could change. Artificial Analysis’s task-cost estimates put Large 4 at $1.13 per Intelligence Index task, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash; those two models scored 41.8 and 39.5, respectively.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral has released Large 4 as a research preview, with an Artificial Analysis score of 38.4 that places it among leading models from outside the US and China but behind current US and Chinese flagships.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

A European Model With a Cost Gap

The result matters because it shows Mistral has made a substantial capability jump while also exposing the limits of the “best outside the US and China” framing. That description is based on the geographic scope of the comparison, not evidence that Large 4 matches the leading systems overall. The reported index score places it below major US models and several Chinese competitors.

For buyers, the issue is not just ranking. The source’s task-cost figures indicate that two Chinese models score higher for less than one-quarter of Large 4’s estimated cost per task. Those figures are specific to the benchmark’s tasks and should not be treated as a complete measure of production costs, which can vary with workload, prompt length, caching and provider terms. Still, they raise a practical procurement question for organisations weighing a European supplier against alternatives.

The distinction may matter for companies seeking a model provider based in Europe or outside the US and China. Large 4 offers a potential option, but its current preview status, unpublished licence and performance trade-offs mean that regional origin alone does not establish suitability for a particular deployment. Teams considering agents face an added concern: errors can carry through multiple steps, while lengthy outputs can add latency and expense.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Benchmark Shows

Artificial Analysis Intelligence Index v4.3.2 combines results from several evaluations, including AA-Briefcase for agentic knowledge work, GDPval-AA for real-world work tasks, AutomationBench for software-as-a-service workflows and Terminal-Bench 4.0 for agentic coding. The source characterises the index as heavily shaped by agentic work. Its score is therefore relevant to multi-step tasks, but it is not a guarantee of how a model will perform in every company’s workflow.

The source’s analysis says Large 4 generated 200 million output tokens across the index evaluation, compared with a median of 81 million for comparable models. That is a benchmark observation, not a universal measure of token use in every application. More output can affect cost and response time in multi-step systems, especially when a workflow calls a model repeatedly.

The source also records a jump from Mistral Large 3’s score of 9 to Large 4’s 38.4 on the same index version. It describes this as the biggest step taken by a European lab, but that characterisation is the source author’s assessment. Mistral says reinforcement learning is ongoing, so both the model and its measured results may change.

“Reinforcement learning is still running.”

— Mistral, as reported in the supplied source

Amazon

large language model API subscription

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Limits and Open Questions

Large 4 is a research public preview, and the supplied material does not provide a final release date beyond Mistral’s promise to ship weights at the end of October. The year is not specified. The source also says the model’s licence is unpublished, leaving its future use and redistribution terms unclear until Mistral releases them.

The 38.4 score and task-cost comparisons are tied to Artificial Analysis Index v4.3.2 and its evaluation setup. They do not show how the model will perform on every real-world task. The source author reports seeing confident hallucinations in hands-on use, but does not provide a test protocol, sample size or a comparable Large 4 hallucination rate. That observation should not be mistaken for a published benchmark result.

It is also not yet clear how much ongoing reinforcement learning will alter Large 4’s scores, output length, reliability or pricing. The supplied source gives no full model card, independent replication of the hands-on findings, or detailed information about deployment safeguards. Those gaps limit conclusions about its readiness for business-critical agents.

Amazon

AI model token usage calculator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Testing Ahead

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. Their availability, accompanying licence and any final release documentation will clarify whether organisations can use or adapt the model beyond the current API preview. The source does not specify a calendar year for that target.

Mistral says reinforcement learning is still in progress, so updated evaluations may change the current picture. Buyers and developers will need to compare later benchmark results with the current v4.3.2 scores and test the model against their own workloads, including multi-step tasks where reliability, output length and total cost all matter. Further evidence about hallucination rates and agent performance has not been supplied.

Amazon

AI model performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is Mistral’s research-preview model. The source describes it as a 1 trillion-parameter, natively multimodal system with 49 billion active parameters and a 512,000-token context window.

How did Large 4 score against leading models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The supplied comparison lists leading US models at 51.8 to 57.6, while several Chinese models also scored above Large 4.

Is Large 4 available to run with its weights?

Not yet, according to the supplied source. Large 4 is available through Mistral’s API as a research public preview, and Mistral has promised to release the weights at the end of October. The source does not give the year, and says the licence has not been published.

Why does the source question Large 4 for agent tasks?

The source points to its index score, high output-token use and the author’s reported observations of confident hallucinations. The hallucination observation is not presented as a formal benchmark result, and actual performance will depend on the workflow and testing conditions.

How much does Mistral Large 4 cost?

The listed API rates are $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. The source reports a 50% discount for the first two weeks, but does not specify the dates of that offer.

Source: ThorstenMeyerAI.com

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Google Will Expand Age Checks On Android Worldwide Till The End Of The Year

Google to extend age verification measures on Android devices globally by the end of 2023, aiming to enhance user safety and compliance.

The Safari MCP Server For Web Developers

Apple introduces Safari MCP server to aid web developers in testing and optimizing their websites, marking a new tool for browser compatibility.

Buried Apple feature turns an iPhone into the perfect kids’ dumb phone

A secret Apple feature enables iPhones to function as simple, kid-friendly devices, offering limited communication and control. Here’s what is known.

AI and Studio Condenser Microphones: The Future of Sound in 2026

In 2026, AI integration with studio condenser microphones is revolutionizing sound quality and recording workflows, promising more natural and precise audio capture.