AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: A Better Fit For Some AI Uses Than Running Agents on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get office and shipping supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 is available in research preview through Mistral’s API, with a 512K context window and weights promised for the end of October. Artificial Analysis rates it at 38.4, below major US and several Chinese models; its cost per benchmark task and output volume also raise questions about using it for long-running agents. The model’s performance may change while reinforcement learning continues, and the source reports hands-on hallucinations that are not part of the benchmark score.

Mistral has released Large 4 as a research preview on its API, presenting a new flagship model that the supplied source says scores 38.4 on the Artificial Analysis Intelligence Index. That result makes it a significant advance over Mistral’s prior models, but the same comparison puts it below leading US models and several Chinese competitors, while its reported cost and output volume raise questions about using it for long-running AI agents.

Large 4 has 1 trillion total parameters, with 49 billion active, and accepts text and images while producing text. Mistral lists a 512K-token context window. The current release is a proprietary research preview; the company has promised to publish the model weights at the end of October, but the source says the licence has not been published. Standard API pricing is listed as $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. The source reports a 50% discount for the first two weeks.

In Artificial Analysis Intelligence Index version 4.3.2, Large 4 scores 38.4. The source compares that with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version, describing the increase as a major step for Mistral. But the table places Large 4 behind the listed US frontier models, whose scores range from 52.6 to 57.6, and behind several Chinese models, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5.

The source says Artificial Analysis estimates Large 4 costs $1.13 per benchmark task. By comparison, it lists GLM-5.3-Flash at $0.25 per task and a score of 41.8, and DeepSeek V4.1 Flash at $0.27 and 39.5. It also reports Large 4 generated 200 million output tokens across the index, compared with a median of 81 million for comparable models. Those figures are tied to the benchmark and its reported comparison group; they do not by themselves predict the cost or token use of every customer workload.

At a glance
analysisWhen: Released yesterday, according to the su…
The developmentMistral released Large 4 as a research preview, prompting a comparison of its agentic benchmark results, pricing and practical suitability with competing models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Agent Workloads Face Cost Trade-offs

The result matters to organizations choosing a model for multi-step, autonomous work, not just one-off chat. The Intelligence Index includes tasks involving knowledge work, software workflows and coding, according to the source. On a long-running task, errors can affect later steps, while extra output can increase both latency and usage-based costs. The source’s comparison therefore raises a practical question: whether Large 4’s capability is worth its measured task cost for a particular agent workflow.

The benchmark does not establish that Large 4 will fail in every agent deployment, nor that a higher-scoring model will be more reliable for each company’s tools and data. But the source’s reported cost figures show that some lower-priced models scored higher on this index. Buyers would need to test models on their own tasks, including output length, error recovery and total run cost, before making a procurement decision.

The source also reports seeing confident false assertions in hands-on use. That is an attributed observation, not an Artificial Analysis finding in the supplied material. For agent systems, an unsupported statement can become an input to later actions, so operators should check factual reliability and provide safeguards rather than treating a general benchmark score as a guarantee.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Jump From Mistral 3

The source frames Large 4 as both a substantial improvement and evidence that Mistral remains behind leading labs. On the same Artificial Analysis index version, it gives Large 3 a score of 9 and Medium 3.5 a score of 14, compared with 38.4 for Large 4. These figures support a large change within Mistral’s own lineup; they do not erase the gap between Large 4 and the index’s highest-scoring models.

Mistral’s launch framing, as relayed by the source, emphasizes that France is home to the most intelligent model outside the United States and China. The supplied comparison does place Large 4 ahead of the listed models from other countries, but it also shows that the headline describes a narrow geographic comparison. It should not be read as saying Large 4 leads the global field or the open-weight category.

The release is still a preview. Mistral says reinforcement learning is ongoing, so scores may change, and the promised weights are not yet available. Until publication, customers cannot inspect or deploy those weights under a known licence, based on the information provided.

“Reinforcement learning is still running, so scores may move.”

— Mistral, as described in the supplied source

Amazon

large language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Results and Reliability

Large 4’s final benchmark standing is not settled, because Mistral says reinforcement learning is continuing. The supplied material does not provide a later score or a date for when the training-related changes will be reflected in a new evaluation.

The source does not identify the test setup behind its hands-on hallucination observations in enough detail to reproduce them, and the reported observations are not benchmark metrics. It also does not establish how Large 4 performs across customers’ own agent workflows. The exact costs, reliability and task success rates will vary by workload, model settings and how often an agent needs to retry or correct its work.

The source says model weights are expected at the end of October, but does not provide a year in the supplied text. The licence remains unpublished, so the terms for use, modification and redistribution are not yet clear from the material provided.

Amazon

AI model performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Updated Evaluations

The next milestones are the planned release of Large 4’s weights and any updates to its evaluation after Mistral’s ongoing reinforcement learning. Publication of the weights and licence would let developers assess deployment options more directly, though neither is available in the supplied information.

For prospective users, the immediate next step is a workload-specific trial. Teams considering agents can compare Large 4 with alternatives on task completion, factual errors, output volume, latency and total cost, while setting limits and human review for consequential actions. The available index and pricing figures offer a starting point, not a final answer about which model fits a particular use.

Amazon

AI token usage monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is Mistral’s new multimodal model, available as a research preview through its API. The supplied source lists 1 trillion total parameters, 49 billion active parameters and a 512K context window.

How did it score against other models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, according to the source. That is higher than the cited scores for Mistral Large 3 and Medium 3.5, but below the leading US models and several Chinese models in the source’s comparison.

Why does the source question its use for AI agents?

The source points to its benchmark position, reported task cost and high output volume, as well as the author’s own observation of confident false assertions. The observation is not a benchmark result, and performance should be tested on the specific workflow before deployment.

How much does Mistral Large 4 cost?

The listed standard API rates are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. The source also reports a 50% discount for the first two weeks and a benchmark task cost of $1.13; actual costs depend on usage.

When will the model weights be available?

Mistral has promised weights for the end of October, according to the supplied source. The source does not specify the year, and says the licence has not yet been published.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Door: Why the Interface Is Worth More Than the Model

SpaceX’s $60 billion purchase of a coding interface highlights the growing importance of the user interface as the critical chokepoint in AI distribution and control.

Mckesson Surges In Global Coverage

Media coverage of McKesson has surged significantly, with mentions increasing 34-fold, signaling heightened global interest in the healthcare distributor.

EuroHPC. The compute substrate.

EuroHPC’s compute substrate underpins Europe’s AI projects, confirming operational readiness at mid-sized scale but revealing structural gaps for frontier AI training.

How Global Media Is Covering Aretha Franklin’s Music And Live Performances

An analysis of how international media is covering Aretha Franklin’s music and upcoming live shows, highlighting recent surges in coverage and implications for promoters.