🔍 Read the full analysis: How The AI Frontier Continues To Outpace Mistral Large 4 on ThorstenMeyerAI.com
Get monitors, keyboards and dev gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral launched an API preview of Large 4 on October 6, 2026; its weights have not yet been released. In an October 7 snapshot, Artificial Analysis gave it an Intelligence Index score of 38, below several leading U.S. and Chinese models. That benchmark does not establish how it will perform on every task, and the source’s concerns about hallucinations and agentic work are personal observations, not controlled comparisons.
Mistral launched a public API preview of Mistral Large 4 on October 6, but an October 7 benchmark snapshot puts it behind several leading U.S. and Chinese models on aggregate intelligence. The gap matters to developers choosing models for complex work, though the available score does not prove how any model will perform on a particular workflow.
Mistral describes Large 4 as its largest model to date, built as a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. It accepts text and images. The company says it trained the model on its own infrastructure in Europe and is continuing to improve it. The release is currently an API preview; Mistral said model weights were scheduled to be released later in October, meaning they were not publicly downloadable on October 7.
Artificial Analysis listed Mistral Large 4 Preview at 38 on its Intelligence Index, in a comparison published on October 7. The same snapshot scored Anthropic Claude Opus 5.5 at 58, Google Gemini 4 Argon at 53 and OpenAI GPT-6.1 Sol at 52. Chinese models Z.ai GLM-5.3 and Moonshot AI’s Kimi K3 scored 45 and 44, respectively. DeepSeek V4.1 Flash scored 39, while OpenAI GPT-6 Luna also scored 38. The comparison includes different named reasoning settings, not tests under identical compute budgets.
The source author argues that the results do not support choosing the preview for demanding, long-running agentic work over stronger-scoring alternatives. That is a judgment about the current preview, not a finding that Large 4 cannot do such work. Artificial Analysis also reports roughly 512,000 tokens of context capacity. A large context window describes how much text can be supplied; by itself, it does not establish accuracy or reliable reasoning across that material.
AI Frontier Briefing · October 7, 2026
How The AI Frontier Continues To Outpace Mistral Large 4
Mistral’s new API preview brings a trillion-parameter mixture-of-experts model to market. In an October 7 benchmark snapshot, its aggregate score trails several leading U.S. and Chinese models. The score is a useful signal—not a verdict on every task.
The leaderboard gap
Artificial Analysis scores indicate aggregate performance on its index suite. Several named competitors scored higher in this dated comparison.
Scores shown as supplied in the October 7 comparison. Reasoning settings differ across models, so the results do not represent a shared compute budget. Company location labels do not indicate where an individual API request is processed.
Agent work tests the whole chain
Long-running tasks depend on more than a large context window. Early mistakes can steer later steps, and a polished final answer may hide a process that went off track.
Plan
Break a request into useful steps and follow its constraints.
Use tools
Call tools appropriately and interpret the results they return.
Verify
Check evidence, catch errors, and revise mistaken assumptions.
Carry through
Sustain reliable execution across a multi-step task.
Scale and modality
Mistral describes Large 4 as its largest model to date, with one trillion total parameters and 49 billion active parameters. It accepts text and images.
Built in Europe
Mistral says it trained the model on its own infrastructure in Europe and is continuing to improve it.
Room is not reliability
A roughly 512,000-token context window describes how much text can be supplied. It does not, by itself, establish accuracy or sound reasoning across that material.
Benchmark results and personal testing answer different questions
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”Thorsten Meyer · ThorstenMeyerAI.com
The source author reports hallucinations and reduced confidence from personal use of the preview. This is an individual experience, not a controlled comparison of error rates. The score of 38 is relevant evidence, but it does not directly test every coding, research, or business workflow.
What this snapshot cannot show
- How Large 4 performs across independently tested, task-specific workloads such as long coding or research assignments.
- Whether the preview is reliable or unreliable for every agentic workflow; the index is an aggregate score.
- How results compare under a common reasoning setting or compute budget; the cited settings differ.
- A controlled hallucination rate: the author provides no sample size or direct error-rate comparison.
- The underlying price figures or method behind the claim that DeepSeek V4.1 Flash had comparable intelligence at much lower measured cost per task.
- Whether weights arrived on the planned schedule. A target for later October is not confirmation of release.
Mistral scheduled the weights release for later in October. They were not publicly downloadable as of October 7.
Test the work you actually need done
Useful next evidence includes multi-step execution, tool use, verification, error rates, and cost measured under comparable conditions. Mistral’s stated strengths in agentic coding and specialized professional work should be evaluated on those workloads. Until then, the benchmark snapshot and personal account are limited evidence about a developing preview—not a final verdict on the model or its future versions.
Quick answers
Are the weights available?
Not as of October 7, 2026, according to the source material. Mistral launched an API preview and scheduled the weights for later in October.
How did Large 4 score?
Artificial Analysis listed Large 4 Preview at 38 in its October 7 Intelligence Index snapshot. Several named U.S. and Chinese models scored higher; reasoning settings differed.
Does the score prove it fails at agent work?
No. It is an aggregate benchmark result, not a direct measure of reliability on every workflow. Test the preview against the tasks and supervision needs that matter to you.
Benchmark Gaps Matter for Agent Work
Agentic work requires a model to plan, call tools, interpret results and carry decisions across multiple steps. A mistaken assumption early in a task can shape later actions, while a fluent final answer may not reveal that the process went off track. For developers, the practical question is not only whether a model can accept a long request, but whether it can follow constraints, check evidence and sustain reliable execution.
The score of 38 is relevant evidence, but it is an aggregate measure rather than a direct test of every coding, research or business workflow. The source author says personal use of the preview produced hallucinations and reduced confidence in assigning it longer tasks. That experience is not a controlled comparison of hallucination rates, and the author does not claim other models are free of errors. It does illustrate why workload-specific testing and supervision costs matter alongside headline benchmark results.
For Mistral, the launch is also a test of its position among frontier AI providers. Its European training infrastructure and a model of this scale are relevant to regional AI capacity. But those facts do not establish performance parity with higher-scoring competitors. Developers should distinguish the significance of the release from evidence that the preview is the best choice for their needs.
As an affiliate, we earn on qualifying purchases.
A Preview, Not the Weight Release
The comparison is a dated snapshot from October 7, 2026, drawing on Artificial Analysis scores. The source material reports that higher scores indicate stronger aggregate performance on the index suite. It also cautions that reasoning settings differ across models, so the table is not a comparison at a shared compute budget. The developers’ country labels refer to where companies are based, not where an individual API request is processed.
The result is more specific than saying every competitor is ahead. Cohere Command A+, a Canadian model, scored 13 on this index, below Mistral Large 4 Preview. The source also says DeepSeek V4.1 Flash had approximately comparable benchmark intelligence at a much lower measured cost per task, but the supplied material does not include the underlying price figures or methodology. Mistral’s own stated strengths in agentic coding and specialized professional work are claims that need evaluation on those workloads.
Artificial Analysis’s Intelligence Index and the source author’s personal testing answer different questions: one provides an aggregate benchmark score, while the other reports an individual experience. Neither alone establishes how the preview will perform for every user. The distinction matters because the weights were still pending and Mistral said it was continuing to improve the model.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, ThorstenMeyerAI.com
As an affiliate, we earn on qualifying purchases.
What the Preview Scores Cannot Show
The available information does not establish how Large 4 performs across independently tested, task-specific workloads, including long coding or research assignments. The Intelligence Index is an aggregate score, not a guarantee of success or failure on an individual task. The cited comparison also uses different reasoning settings, and no common compute budget is specified.
The source author reports hallucinations from personal use but provides no controlled measurement, sample size or direct error-rate comparison. The supplied material also does not include the cost figures behind the claim about DeepSeek V4.1 Flash, or enough detail to compare total costs for equivalent tasks. Mistral’s preview may change as the company continues development, and the planned timing for weights is not confirmation that they were released.
As an affiliate, we earn on qualifying purchases.
Watch for Weights and Workload Tests
The next stated milestone is the planned release of Large 4’s weights later in October, according to Mistral’s announcement. As of October 7, they were not publicly downloadable. Mistral has also said it is continuing to improve the model, so later versions or evaluations may not match the current preview.
For developers, the useful next evidence will be testing on the work they actually need done: multi-step execution, tool use, verification, error rates and cost under comparable conditions. Until those results are available, the benchmark snapshot and the author’s personal account should be treated as limited evidence about a developing preview, not a final verdict on the model or its future releases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Has Mistral released Large 4’s model weights?
Not as of October 7, 2026, according to the source material. Mistral launched an API preview and scheduled the weights for release later in October.
How did Mistral Large 4 score against other models?
Artificial Analysis gave Large 4 Preview a score of 38 in its October 7 Intelligence Index snapshot. Several named U.S. and Chinese models scored higher, but the comparison used different reasoning settings rather than a shared compute budget.
Does the benchmark prove Large 4 is unreliable for agentic tasks?
No. The score is an aggregate benchmark result, not a direct measure of reliability on every workflow. The source author’s recommendation against using it for demanding long tasks is an attributed judgment, supported by the score and personal experience, not proof that it will fail a particular task.
What does the 512,000-token context figure mean?
Artificial Analysis reports roughly 512,000 tokens of context capacity. This describes how much material can fit into a request; it does not guarantee accurate reasoning across that material.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
