AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4’S Appeal Outside The US And China Doesn’t Extend To Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, making it a leading model from outside the US and China but placing it below major US and Chinese competitors. The source analysis argues that its benchmark performance, token use and reported pricing make it a poor fit for some agent workloads; hands-on claims about hallucinations have not been independently verified here.

Mistral released Large 4 as a research preview, with its score of 38.4 on Artificial Analysis’s Intelligence Index placing it among the strongest models from outside the United States and China, but below leading US and Chinese systems. The comparison matters for buyers considering the model for AI agents: the source analysis also reports higher task costs than some better-scoring alternatives and unusually high output-token use.

Artificial Analysis’s Intelligence Index v4.3.2 gives Large 4 a score of 38.4. In the table supplied by the source, the highest-ranked model is Anthropic’s Claude Opus 5.5 at 57.6. Several Chinese models also score above Large 4, including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. The source characterizes Large 4 as the most intelligent model outside the US and China, while noting that this distinction does not put it level with the leading models overall.

The model has one trillion total parameters, with 49 billion active, and accepts text and images while generating text. Mistral lists a 512,000-token context window. It is available through the company’s API as a research preview; Mistral has said weights are expected at the end of October. The source says a licence had not been published at the time of writing. Its listed prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. A 50% discount applies for the first two weeks, according to the source.

On the source’s account of Artificial Analysis data, Large 4 used 200 million output tokens to complete the Intelligence Index tasks, compared with a median of 81 million for comparable models. The analysis estimates a cost of $1.13 per index task at standard pricing. It compares that with $0.25 for GLM-5.3-Flash, which scored 41.8, and $0.27 for DeepSeek V4.1 Flash, which scored 39.5. These figures and comparisons are reported by the supplied source; pricing and benchmark results can change.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral released Large 4 as a research preview, and an independent benchmark comparison shows it remains behind leading US and Chinese models, including on tasks relevant to AI agents.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Agent Work Raises Cost and Reliability Questions

The benchmark is relevant to agent use because the supplied analysis says the index includes tasks involving multi-step work, software workflows and coding, not only short-form question answering. A system that makes errors in one step can pass them into later steps, so performance on longer tasks may differ from success on a single prompt. The benchmark score alone does not establish how often Large 4 will fail in a particular customer’s workflow, but it gives buyers a reason to test it against their own tasks before deployment.

Output volume may also affect the economics of those workflows. If an agent makes repeated model calls, more generated tokens can add cost and latency, even when the per-token price appears competitive. The source’s cost-per-task estimate suggests some alternatives completed the benchmark tasks at lower expense while scoring higher. That is a useful procurement comparison, though actual spending will depend on prompts, task length, caching, usage discounts and the way an application is built.

The source author also reports observing confident but incorrect answers during hands-on testing. That is an attributed observation, not a result established by the Artificial Analysis index. If replicated in a buyer’s tests, such errors could be more consequential in an agent workflow than in ordinary chat, because later actions may rely on an inaccurate claim.

Amazon

AI model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Large Jump From Mistral’s Prior Scores

The source reports that Mistral Large 3 scored 9 on the same version of the index and Medium 3.5 scored 14. Large 4’s 38.4 score marks a substantial reported improvement over those earlier Mistral results. That progress is compatible with the separate finding that it remains behind the current highest-scoring models; improvement within a company’s lineup does not by itself establish parity with competitors.

Mistral says reinforcement learning on Large 4 is still underway, according to the supplied material, so its scores may change. The model is not yet presented there as a fully released open-weights system: the current access described is a proprietary API research preview, with weights expected later and licence terms not yet available. Those release details matter to organizations weighing deployment, data handling and the ability to run a model independently.

“Reinforcement learning is still running.”

— Mistral

Release Terms and Agent Results Remain Open

The supplied material does not include the underlying benchmark report or testing protocol, so readers cannot assess from it alone how each score was produced or how closely those tasks resemble a specific business workflow. The source also does not establish that Large 4 will perform worse on every agent task; the index is a broad comparison, not a guarantee of outcomes in individual applications.

Weights and licence terms remain pending in the account provided, and the timing of the promised end-of-October release is not independently confirmed here. Mistral’s statement that training is continuing means scores and model behavior could change. The hallucination concern is based on one author’s hands-on report; the supplied text provides no sample size, testing method or independent replication. The detailed price comparison also depends on the listed rates and task-cost methodology, and could change with discounts or revised prices.

Watch for Weights and Updated Benchmarks

The next milestones are Mistral’s promised release of Large 4’s weights at the end of October and publication of the associated licence, as described in the source. Buyers can also watch for updated Artificial Analysis results as reinforcement learning continues and for independent testing of the model on agent tasks. Until those details are available, organizations evaluating Large 4 should compare it with alternatives on representative workflows, tracking accuracy, recovery from errors, output tokens and total task cost rather than relying on a single headline score.

Key Questions

What did Mistral announce?

Mistral released Large 4 as a research preview through its API. The supplied source says the model accepts text and images, has a 512,000-token context window, and is expected to have weights released at the end of October.

How did Large 4 score against competitors?

It scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, according to the source. That is below the leading US models listed and below several Chinese models, including GLM-5.3 and DeepSeek V4.1 Flash.

Why does the source question Large 4 for agent use?

The source points to its index score, high reported output-token use and estimated cost per task. The author also reports seeing confident incorrect answers in hands-on testing, but that observation is not an independently verified benchmark result.

Is Large 4 open source now?

Not according to the supplied material. It is available through Mistral’s API as a research preview; weights are promised for the end of October, and the source says the licence had not yet been published.

Could its benchmark score change?

Yes. Mistral reportedly said reinforcement learning is still underway and that scores may move. The source does not provide an updated result or confirm when revised benchmark data will appear.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

ChannelHelm: One Video, Every Platform

ChannelHelm automates the creation of multiple platform-specific assets from a single video, streamlining content distribution and reducing manual effort.

13 Best Guides to AI-Powered Marketing Automation Tools for Smarter Campaigns in 2026

Explore the best books and guides on AI marketing automation, including strategy, tools, and implementation for smarter campaigns.

Revolutionize Your Mobile Workflow With 2026’S AI Laptops

New AI-powered laptops in 2026 promise to revolutionize mobile professional workflows with advanced hardware and software integration.

Client asset intake portal for accountants

A new client asset intake portal for small accounting firms is being tested to streamline document collection and reduce administrative loops.