🔍 Read the full analysis: Mistral Large 4’S Appeal Outside The US And China Doesn’t Extend To Agents on ThorstenMeyerAI.com
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral Large 4 scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, making it a leading model from outside the US and China but placing it below major US and Chinese competitors. The source analysis argues that its benchmark performance, token use and reported pricing make it a poor fit for some agent workloads; hands-on claims about hallucinations have not been independently verified here.
Mistral released Large 4 as a research preview, with its score of 38.4 on Artificial Analysis’s Intelligence Index placing it among the strongest models from outside the United States and China, but below leading US and Chinese systems. The comparison matters for buyers considering the model for AI agents: the source analysis also reports higher task costs than some better-scoring alternatives and unusually high output-token use.
Artificial Analysis’s Intelligence Index v4.3.2 gives Large 4 a score of 38.4. In the table supplied by the source, the highest-ranked model is Anthropic’s Claude Opus 5.5 at 57.6. Several Chinese models also score above Large 4, including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. The source characterizes Large 4 as the most intelligent model outside the US and China, while noting that this distinction does not put it level with the leading models overall.
The model has one trillion total parameters, with 49 billion active, and accepts text and images while generating text. Mistral lists a 512,000-token context window. It is available through the company’s API as a research preview; Mistral has said weights are expected at the end of October. The source says a licence had not been published at the time of writing. Its listed prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. A 50% discount applies for the first two weeks, according to the source.
On the source’s account of Artificial Analysis data, Large 4 used 200 million output tokens to complete the Intelligence Index tasks, compared with a median of 81 million for comparable models. The analysis estimates a cost of $1.13 per index task at standard pricing. It compares that with $0.25 for GLM-5.3-Flash, which scored 41.8, and $0.27 for DeepSeek V4.1 Flash, which scored 39.5. These figures and comparisons are reported by the supplied source; pricing and benchmark results can change.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
Agent Work Raises Cost and Reliability Questions
The benchmark is relevant to agent use because the supplied analysis says the index includes tasks involving multi-step work, software workflows and coding, not only short-form question answering. A system that makes errors in one step can pass them into later steps, so performance on longer tasks may differ from success on a single prompt. The benchmark score alone does not establish how often Large 4 will fail in a particular customer’s workflow, but it gives buyers a reason to test it against their own tasks before deployment.
Output volume may also affect the economics of those workflows. If an agent makes repeated model calls, more generated tokens can add cost and latency, even when the per-token price appears competitive. The source’s cost-per-task estimate suggests some alternatives completed the benchmark tasks at lower expense while scoring higher. That is a useful procurement comparison, though actual spending will depend on prompts, task length, caching, usage discounts and the way an application is built.
The source author also reports observing confident but incorrect answers during hands-on testing. That is an attributed observation, not a result established by the Artificial Analysis index. If replicated in a buyer’s tests, such errors could be more consequential in an agent workflow than in ordinary chat, because later actions may rely on an inaccurate claim.
As an affiliate, we earn on qualifying purchases.
A Large Jump From Mistral’s Prior Scores
The source reports that Mistral Large 3 scored 9 on the same version of the index and Medium 3.5 scored 14. Large 4’s 38.4 score marks a substantial reported improvement over those earlier Mistral results. That progress is compatible with the separate finding that it remains behind the current highest-scoring models; improvement within a company’s lineup does not by itself establish parity with competitors.
Mistral says reinforcement learning on Large 4 is still underway, according to the supplied material, so its scores may change. The model is not yet presented there as a fully released open-weights system: the current access described is a proprietary API research preview, with weights expected later and licence terms not yet available. Those release details matter to organizations weighing deployment, data handling and the ability to run a model independently.
“Reinforcement learning is still running.”
— Mistral
Release Terms and Agent Results Remain Open
The supplied material does not include the underlying benchmark report or testing protocol, so readers cannot assess from it alone how each score was produced or how closely those tasks resemble a specific business workflow. The source also does not establish that Large 4 will perform worse on every agent task; the index is a broad comparison, not a guarantee of outcomes in individual applications.
Weights and licence terms remain pending in the account provided, and the timing of the promised end-of-October release is not independently confirmed here. Mistral’s statement that training is continuing means scores and model behavior could change. The hallucination concern is based on one author’s hands-on report; the supplied text provides no sample size, testing method or independent replication. The detailed price comparison also depends on the listed rates and task-cost methodology, and could change with discounts or revised prices.
Watch for Weights and Updated Benchmarks
The next milestones are Mistral’s promised release of Large 4’s weights at the end of October and publication of the associated licence, as described in the source. Buyers can also watch for updated Artificial Analysis results as reinforcement learning continues and for independent testing of the model on agent tasks. Until those details are available, organizations evaluating Large 4 should compare it with alternatives on representative workflows, tracking accuracy, recovery from errors, output tokens and total task cost rather than relying on a single headline score.
Key Questions
What did Mistral announce?
Mistral released Large 4 as a research preview through its API. The supplied source says the model accepts text and images, has a 512,000-token context window, and is expected to have weights released at the end of October.
How did Large 4 score against competitors?
It scored 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, according to the source. That is below the leading US models listed and below several Chinese models, including GLM-5.3 and DeepSeek V4.1 Flash.
Why does the source question Large 4 for agent use?
The source points to its index score, high reported output-token use and estimated cost per task. The author also reports seeing confident incorrect answers in hands-on testing, but that observation is not an independently verified benchmark result.
Is Large 4 open source now?
Not according to the supplied material. It is available through Mistral’s API as a research preview; weights are promised for the end of October, and the source says the licence had not yet been published.
Could its benchmark score change?
Yes. Mistral reportedly said reinforcement learning is still underway and that scores may move. The source does not provide an updated result or confirm when revised benchmark data will appear.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
