Could Mistral Large 4 Be Your Pick Beyond The US And China?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Could Mistral Large 4 Be Your Pick Beyond The US And China? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get monitors, keyboards and dev gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral has released Large 4 as a research public preview, with a 38.4 score on the Artificial Analysis Intelligence Index v4.3.2. The score puts it ahead of the listed models from outside the United States and China, but below leading US models and several Chinese models; pricing, reliability and the unpublished license remain relevant buyer questions.

Mistral has released Large 4 as a research public preview, and it scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The result makes it the top-scoring model in the source’s stated group of models from outside the United States and China, but it trails major US systems and several Chinese models—an important distinction for buyers weighing the launch as a non-US option.

The source describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, text-and-image input, text output and a 512,000-token context window. It is available through Mistral’s API as a research public preview. Mistral says model weights are due at the end of October; until their release, the model is proprietary, and the source says its license has not been published.

Artificial Analysis’ listed score of 38.4 is below the listed US frontier models, which score as high as 57.6, and below several Chinese models, including GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. It is above the listed GLM-5.2 score of 33.7 and DeepSeek V4 Pro score of 36.0. The source characterizes Large 4 as the highest-scoring model outside the US and China, while noting that this comparison group is narrow.

The source lists standard API prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. It also reports a 50% discount for the first two weeks. Artificial Analysis’ task-cost comparison puts Large 4 at $1.13 per Intelligence Index task, versus $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. The source says Mistral reports reinforcement learning is still underway, so benchmark results could change.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral released Large 4 as a research public preview, prompting a new comparison of its benchmark results, costs and position against US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

A European Option With Trade-Offs

Large 4’s result matters because it gives buyers a high-profile model from a French developer in a market whose top scores are led by US and Chinese labs. For organisations seeking a provider outside those two countries, the launch offers a model worth evaluating, though its performance does not establish parity with the leading systems.

The benchmark and cost figures also complicate the case for deploying it broadly. Artificial Analysis scores Large 4 below several alternatives, while the source’s task-cost comparison places it at more than four times the cost of the two cited Chinese models that score higher. These are benchmark-based comparisons, not proof of how any one customer’s workload will perform. Buyers would need to test their own tasks, including output length, latency and error rates.

The source further reports that Large 4 produced 200 million output tokens on the Intelligence Index, compared with a median of 81 million for comparable models. If that difference carries over to a customer’s use, verbosity could add expense and latency, especially in multi-step agent workflows. The source author also says hands-on testing found confident false statements; that is an attributed observation, not a published Artificial Analysis measurement.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Preview Includes

Artificial Analysis Intelligence Index v4.3.2 is the stated basis for the comparisons, allowing the listed model scores to be compared on the same index version. The source says the index includes agent-oriented evaluations such as AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. Its score is therefore presented as a composite benchmark result, not a direct measure of every task a business might assign to a model.

The source compares Large 4 with Mistral’s prior models on the same index version: Large 3 scored 9 and Medium 3.5 scored 14, while Large 4 scored 38.4. That is a marked improvement in this benchmark, though the source does not provide enough detail here to explain how much of the change reflects model capability, evaluation conditions or other factors.

For open-weight comparisons, there is a timing caveat: Mistral’s weights have not yet been released. The source’s prospective ranking among open models depends on that release and on the benchmark results cited; the model is currently described as a proprietary API preview, with its license unpublished.

Amazon

text and image AI input devices

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

License, Reliability And Final Scores

Several details that would affect procurement remain unsettled. Mistral’s weights are promised for the end of October, but the license terms are not yet published in the source material. It is also unclear whether that release will occur on schedule or whether the weights will be available on terms suitable for commercial use.

The benchmark score may change because Mistral says reinforcement learning is ongoing. The source does not provide a sample size or systematic measurement for its reported hallucination observations, and it does not establish whether those results generalize to other prompts or deployments. Nor does the supplied material include independent comparisons of latency, uptime or performance across customer workloads.

The cost figures are tied to the stated benchmark tasks and listed pricing. Actual bills can vary with input and output volume, caching, task design and the introductory discount. The source provides no detail about how the two-week discount period is calculated or when it began.

Amazon

large parameter AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights And Buyer Testing Ahead

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. Buyers will be able to assess the model’s deployment options and license terms only once those details are available. The source does not provide a precise calendar date for the API preview or the discount’s start.

In the meantime, organisations considering the preview can compare it against alternatives on representative tasks rather than relying on a single index score. Testing should account for factual accuracy, multi-step completion, output volume, latency and total cost. Further Artificial Analysis results could also clarify whether the score changes as Mistral’s reported reinforcement learning continues.

Amazon

AI model cost optimization tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

It is a Mistral model available as a research public preview through the company’s API. The source describes it as a one-trillion-parameter, natively multimodal model with a 512,000-token context window.

How does Large 4 score against other models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. That is below the listed leading US models and several Chinese models, but above the listed GLM-5.2 and DeepSeek V4 Pro scores.

Is Large 4 open-weight now?

No. The source describes the current release as a proprietary API preview. Mistral has promised weights for the end of October, but the license is not yet published in the supplied material.

How much does the API cost?

The source lists standard prices of $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. It also reports a 50% discount for the first two weeks.

Are the benchmark and reliability results final?

Not necessarily. Mistral says reinforcement learning is still running, so the score may move. The source author’s comments about confident false statements are personal testing observations, not a published benchmark rate.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AmenGate: The Moment Before the Scroll

AmenGate introduces a prayer-based phone lock system designed to replace mindless scrolling with meaningful prayer, relying on system-level interruption and trust.

Mapquest Surges In Global Coverage

Mapquest’s media mentions have spiked 13-fold recently, signaling increased global interest in the mapping service, though the reasons remain unconfirmed.

VigilSAR Benchmark: There Is No Best Model

VigilSAR Benchmark reveals no model is universally superior; rankings vary based on buyer needs like deployment, compliance, and robustness.

Samsung Says ‘Tim Cook’ Bought A Galaxy Fold

Samsung states that ‘Tim Cook’ purchased a Galaxy Fold, sparking widespread curiosity amid unconfirmed reports and rising search interest.