AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Makes Mistral Large 4 Stand Out Outside The US And China? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4, released as a research preview, scored 38.4 on the Artificial Analysis Intelligence Index. That makes it a leading model from outside the United States and China, but the supplied benchmark data places several US and Chinese models ahead; its weights and license are not yet available.

French AI company Mistral released Large 4 as a research preview, and the model scored 38.4 on the Artificial Analysis Intelligence Index. The result makes it a standout among models developed outside the United States and China, but the same benchmark places multiple US and Chinese models above it, a distinction that matters to buyers comparing capability, price and access to model weights.

Artificial Analysis Index version 4.3.2 gives Large 4 a score of 38.4. The source report describes it as the highest-scoring model from outside the US and China, while noting that the leading US model in its table scores 57.6 and Chinese models also rank above Mistral. These are benchmark results, not a measure of every use case, and Mistral says reinforcement learning is ongoing, so the score may change.

Mistral describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, image and text input, text output and a 512,000-token context window. It is currently available through Mistral’s API as a research preview. The company has promised to release the weights at the end of October, but the source material does not provide a published license.

The listed API price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14. The source says Mistral offered a 50% discount for the first two weeks. Artificial Analysis estimates a cost of $1.13 per Intelligence Index task at standard pricing; the report compares that with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash, both of which it says scored higher. These task-cost figures are benchmark-specific, not universal estimates for customer workloads.

At a glance
reportWhen: Released yesterday relative to the sour…
The developmentMistral released Large 4 as an API research preview, with benchmark results positioning it as a strong European model while leaving it behind leading US and Chinese systems.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

A European Option With Trade-Offs

Large 4 gives companies and developers another high-profile model from a European lab at a time when leading AI systems are concentrated among US and Chinese providers. For organizations seeking a model provider based outside those countries, Mistral’s result is a meaningful benchmark showing progress. It does not, on its own, establish that Large 4 is the best choice for a particular task or that it matches the leading systems overall.

The reported ranking also sets a clear limit on the headline. The source’s table places several US frontier models and Chinese systems ahead of Large 4. It says the model is particularly relevant as a European alternative, rather than evidence that Mistral has caught the overall frontier. Buyers should weigh the benchmark alongside their own testing, including reliability on the steps their workflows require.

Cost and output length may affect those decisions. The source report says Large 4 used 200 million output tokens across the Intelligence Index evaluation, against a median of 81 million for comparable models. That observation suggests it can produce more output during those benchmark tasks, which may raise expense and latency in some agent workflows. It does not establish that every user or application will see the same difference.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Mistral’s Earlier Scores

The source compares Large 4 with earlier Mistral models on the same Intelligence Index version: Large 3 scored 9 and Medium 3.5 scored 14, while Large 4 scored 38.4. Those figures indicate a substantial improvement within Mistral’s own lineup. They do not erase the gap between Large 4 and higher-scoring competitors in the same table.

Artificial Analysis’s index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. The source therefore interprets the result as relevant to multi-step work and software tasks, not simply a general knowledge test. Benchmark performance remains one input among several: real-world results depend on task design, tools, prompts and the model’s reliability under deployment conditions.

At launch, the weights are not yet available, so developers cannot independently host the released model based on the information in the source. Until Mistral publishes them and clarifies the license, access is through its API. The absence of a license also means potential users do not yet have enough information to assess the terms for future use of the weights.

Weights, License and Reliability

Several details remain open. Mistral has promised weights by the end of October, but the source gives no confirmation that they have shipped or states the license under which they would be released. Until those details are public, users cannot determine from this material whether self-hosting will be available on terms that fit their needs.

The source report also offers a personal observation that Large 4 produced confident false statements during hands-on testing. That is not an Artificial Analysis benchmark result, and the report does not provide a test set, sample size or methodology for the observation. The benchmark scores and the reported cost estimates likewise do not predict performance or spending in every customer deployment. Mistral’s ongoing reinforcement learning could also alter results.

Awaiting the Weight Release

The next stated milestone is Mistral’s promised release of Large 4’s weights at the end of October. Users will then be able to assess the availability and license terms if the release proceeds as announced. Mistral’s reinforcement-learning work may also lead to updated performance figures. For now, the model remains a research preview accessible through the API, and organizations considering it will need to test it on their own workloads before drawing procurement conclusions.

Key Questions

What is Mistral Large 4?

It is Mistral’s new model, described as having one trillion total parameters, 49 billion active parameters, image and text input, text output and a 512,000-token context window. It is currently offered as an API research preview.

How did Large 4 score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source report. Several US and Chinese models in the report’s comparison scored higher.

Can developers download its weights now?

Not according to the source material. Mistral has promised the weights for the end of October, but the report says they are not yet available and the license has not been published.

How much does the API cost?

The listed standard rates are $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. The source also reports a 50% discount for the first two weeks; current discount availability is not established.

Does the benchmark show Large 4 is the best model for agents?

No. The index includes agent-oriented tests, but a benchmark ranking cannot determine which model will perform best for every workflow. The source places Large 4 below several US and Chinese systems and reports concerns about output volume and observed hallucinations; those observations need to be checked against a buyer’s own tasks.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AirPods Pro 3 Or AirPods 5? I Compared Them For Weeks, And It’s Surprisingly Close

Search and coverage interest is rising around AirPods Pro 3 and AirPods 5, but the reason for the spike has not been confirmed.

I Tried Google’s New Image Editor, And It Could Replace Canva And Photoshop

Google’s new prompt-driven image editor Google Pics blends Canva-style design and AI photo editing for AI Pro and Ultra subscribers.

The State Of The Tech Industry In 2026

A Pragmatic Engineer report describes AI agents changing software development, while warning that code quality and review practices are under strain.

Sony Threatens To Cut Off Customer Support And Pursue Legal Action Over Harassment As Backlash To PlayStation’s Physical-disc Phaseout Intensifies

Sony warns it may cut support and pursue legal action amid rising harassment linked to PlayStation’s disc phaseout backlash.