Why Mistral Large 4 Isn’t A Model To Run Your Agents On
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why Mistral Large 4 Isn’t A Model To Run Your Agents On on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 has made a large benchmark jump, scoring 38.4 on Artificial Analysis Intelligence Index v4.3.2. But the source analysis argues it is a poor choice for long-running agents because several rivals score higher or cost less per benchmark task; hands-on hallucination concerns are reported but are not an independent benchmark result.

Mistral released Large 4 as a research preview, and a new analysis of Artificial Analysis results argues that its benchmark score and operating costs make it a weak choice for many agent workflows. The model scored 38.4 on Intelligence Index v4.3.2, below several current US and Chinese models, while the source analysis reports that its benchmark tasks cost more than those of two Chinese models with higher scores.

Large 4 has 1 trillion total parameters, with 49 billion active, and accepts text and images while producing text. Mistral lists a 512,000-token context window. The model is available through Mistral’s API as a research preview; the company says its weights are expected at the end of October. Until then, the weights are not public, and the source material says the licence has not been published.

At standard API rates, the listed price is $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14. Mistral is offering a 50% discount for the first two weeks. The company says reinforcement learning is continuing, meaning the model’s performance may change. The source does not specify the exact calendar date of release or the end-October year.

Artificial Analysis Index v4.3.2 gives Large 4 a score of 38.4. The source analysis compares that with 57.6 for Claude Opus 5.5, 52.6 for Gemini 4 Argon, 44.8 for GLM-5.3 and 39.5 for DeepSeek V4.1 Flash. It also reports that Large 4 used 200 million output tokens across the Index tasks, versus a median of 81 million for comparable models. These figures describe the benchmark run, not token use for every customer workload.

At a glance
analysisWhen: Released yesterday relative to the sour…
The developmentMistral released Large 4 in research preview, prompting scrutiny of its benchmark standing, pricing and suitability for agent workflows.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Benchmark Gaps Affect Agent Costs

The benchmark matters to teams deciding which model should run software agents that perform several steps, use tools or work for extended periods. The Index includes agent-focused evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0, according to the source analysis. A lower score does not translate directly into a specific failure rate for an individual workflow, but it gives buyers a reason to test carefully before putting the model in charge of consequential tasks.

Cost may make that decision harder. The analysis estimates $1.13 per Intelligence Index task for Large 4, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models score 41.8 and 39.5 respectively in the cited comparison. The task-cost figures are benchmark estimates, not a universal quote for an agent deployment; actual spending will depend on prompt length, output and workload.

The analysis also reports that Large 4 produced substantially more output than the comparison-model median. For an agent, extra tokens can add both latency and expense, since the system may call a model repeatedly. That makes per-token prices alone an incomplete basis for procurement: buyers need to measure cost and completion quality over their own tasks.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Jump From Large 3

The source analysis says Mistral Large 3 scored 9 and Medium 3.5 scored 14 on the same version of the Artificial Analysis Index. Large 4’s 38.4 therefore marks a substantial reported improvement within Mistral’s lineup. The analysis describes it as a major step for a European AI lab, while also stressing that a large improvement over earlier Mistral models does not put Large 4 at the top of the current field.

Artificial Analysis describes the model as the most intelligent one from outside the United States and China. That framing depends on the comparison group: the source argues that few labs elsewhere compete directly at this frontier tier. It says Large 4 is behind a number of US and Chinese models in the index, including newer Chinese releases than those Mistral selected for comparison at launch.

The source’s concerns about confident hallucinations come from the author’s hands-on testing, not from the Artificial Analysis score itself. The same analysis cites AA-Omniscience figures for other models, but those measure a separate benchmark and should not be treated as proof of how Large 4 will perform in a particular deployment.

“Reinforcement learning is still running.”

— Mistral

Amazon

large language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Leaves Key Questions Open

Large 4 remains in research preview, and Mistral says training work is continuing. Its eventual benchmark performance, availability and behaviour may differ from the current release. The source says the weights are promised for the end of October but does not establish the year, publication date or final licence terms.

The benchmark comparisons do not establish how Large 4 will perform across every agent task, nor do they show a specific error rate for long-running workflows. The source provides no measured AA-Omniscience hallucination rate for Large 4. Its reported hallucination concerns are from hands-on testing and should be weighed separately from the index scores and task-cost estimates.

It is also unclear how the temporary discount affects practical costs for buyers after the first two weeks, or what pricing and access terms will apply when weights are released. Teams will need workload-specific tests to compare total cost, reliability and completion rates.

Amazon

AI model cost analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch for Weights and Retesting

The next stated milestone is Mistral’s planned release of model weights at the end of October. The licence and final release details will matter to organisations considering self-hosting or adapting the model. Mistral’s ongoing reinforcement learning may also lead to a changed model or revised results.

For buyers, the practical next step is to run representative evaluations rather than infer agent performance from a single leaderboard score. Those tests should include multi-step completion, tool use, factual checks, token consumption and total costs under the intended setup. Updated independent benchmarks could clarify whether Large 4’s current ranking changes as development continues.

Amazon

text and image AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4’s Artificial Analysis score?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source analysis. That is one benchmark result, not a guarantee of performance on a specific task.

Why does the analysis advise caution about using it for agents?

The analysis points to lower scores than several rivals, higher estimated cost per benchmark task than two models that scored higher, and unusually high output-token use in the benchmark. It also reports the author’s own hallucination observations, which are not a formal Large 4 hallucination measurement.

Is Mistral Large 4 open-weight?

Not yet, according to the source material. It is currently offered as a research preview through Mistral’s API, with weights promised for the end of October. The licence is described as unpublished.

How much does Large 4 cost?

The listed standard API prices are $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. The source also reports a 50% discount for the first two weeks.

Could its benchmark standing change?

Yes. Mistral says reinforcement learning is still running, and the model remains a preview. The source does not give a timeline or a predicted size for any performance change.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Looking For AI Workflow Automation Tools? 14 To Explore In 2027

A comparison of 14 books and guides for learning AI workflow automation, from no-code introductions to developer and industry-specific instruction.

Markets Are Competitive If And Only If P != NP

New theoretical findings confirm markets are competitive if and only if P ≠ NP, highlighting a fundamental link between computational complexity and economic theory.

4,400-Year-Old Tomb Of Egyptian Judge Found At Saqqara With Colors On Walls

A 4,400-year-old tomb of an Egyptian judge has been uncovered at Saqqara, featuring well-preserved wall colors, marking a significant archaeological find.

Mac vs GPU Tower for Local LLMs: The Heat-and-Noise Tradeoff

A detailed comparison of Mac Studio with Apple Silicon and GPU towers for running local large language models, focusing on heat, noise, and performance tradeoffs.