🔍 Read the full analysis: Mistral Large 4’S Challenge: Keeping Pace With The AI Frontier on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Mistral launched Mistral Large 4 in public API preview on October 6, 2026. Artificial Analysis gave it an Intelligence Index score of 38, below several leading US and Chinese models; the model’s weights have not yet been released. The benchmark is one measure, not proof of how it will perform on every task.
Mistral AI introduced Mistral Large 4 in public API preview on October 6, presenting its largest model yet as a new step for the French company. But an Artificial Analysis Intelligence Index score of 38 puts the preview below several leading US and Chinese models in the cited comparison, leaving open whether it can keep pace on demanding tasks. Its weights are not yet publicly downloadable.
Mistral describes Large 4 as a mixture-of-experts model with one trillion total parameters and 49 billion active parameters. The preview API accepts text and images. The company says it trained the model on its own infrastructure in Europe and is continuing to improve it. Mistral has scheduled a release of the weights for later in October, but that release had not taken place as of October 7.
Artificial Analysis rated the preview 38 on its Intelligence Index. In the October 7 snapshot cited by ThorstenMeyerAI.com, that score matched OpenAI’s GPT-6 Luna at maximum reasoning effort and was one point below DeepSeek V4.1 Flash at maximum effort. It trailed Z.ai’s GLM-5.3, at 45; Moonshot AI’s Kimi K3, at 44; and several US models, including Anthropic’s Claude Opus 5.5 at 58.
The comparison is a dated benchmark snapshot, not a universal measure of task success. The models were evaluated with different named reasoning settings, so the results do not represent identical compute budgets. Developer locations in the comparison identify the companies, not where an API request is processed. Index-point gaps are not percentages of intelligence or predicted success rates.
Frontier model watch · October 7, 2026
Mistral Large 4’s Challenge: Keeping Pace With the AI Frontier
Mistral’s largest model has entered public API preview. Its benchmark snapshot trails several leading models, while its weights remain pending—leaving real-world competitiveness an open question.
Artificial Analysis preview score in the cited October 7 snapshot
Total parameters / active parameters · mixture of experts
Weights scheduled by Mistral for later in October
A major release, with a key test still ahead
Mistral AI introduced Mistral Large 4 as its largest model yet, opening public API preview on October 6, 2026. The company describes a mixture-of-experts architecture with one trillion total parameters and 49 billion active parameters. It says training took place on its own infrastructure in Europe, and that work on the model is continuing.
Large total capacity
The model has 1 trillion total parameters, with 49 billion active parameters, according to Mistral.
Text and images
The public preview API accepts both text and image inputs.
Preview first
Developers can access the API preview. Public weights had not been released as of October 7.
How the cited scores compare
Artificial Analysis assigned the preview an Intelligence Index score of 38. The October 7 snapshot places several US and Chinese models above it, while also showing that not every listed competitor scored higher.
Selected Intelligence Index scores · Oct 7
*DeepSeek V4.1 Flash at maximum effort. GPT-6 Luna at maximum reasoning effort also scored 38. Selected models only; bars scale to 58.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
Thorsten Meyer · editorial judgmentA signal for testing, not a verdict on every task
The index offers one aggregate comparison. It cannot determine which model will best handle a particular coding, research, or business workflow, and a score of 38 does not establish that Large 4 will fail at a specific job.
Long agentic workflows
Planning, tool use, and decisions across many steps put a premium on dependable execution. An early error can affect later actions, making verification and supervision part of the practical evaluation.
Evidence still developing
Mistral’s claims about agentic coding and specialized professional work need testing on those workloads. Meyer reports hallucinations in his own use, but describes that as personal experience, not a controlled comparison.
Preview now. Weights later.
Mistral scheduled a release of the model weights for later in October. Until then, the offering described here is an API preview. The source material does not specify the exact timing or form of weight access, nor complete cost figures.
Try the preview
Run representative tasks through the API.
Track outcomes
Measure accuracy, tool use, and unsupported claims.
Count oversight
Include supervision needs and cost in the comparison.
Reassess later
Compare again as weights or benchmark results arrive.
Not every competitor outranks it
The cited snapshot includes a lower score as well as several higher ones. These are dated aggregate results; model versions and scores may change.
| Model | Developer | Score | Context |
|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 58 | US model cited among higher scorers |
| GLM-5.3 | Z.ai | 45 | Above Large 4 in the snapshot |
| Kimi K3 | Moonshot AI | 44 | Above Large 4 in the snapshot |
| DeepSeek V4.1 Flash | DeepSeek | 39 | Maximum effort; one point higher |
| Mistral Large 4 | Mistral AI | 38 | Public API preview |
| GPT-6 Luna | OpenAI | 38 | Maximum reasoning effort |
| Command A+ | Cohere | 13 | Below Large 4 in the same snapshot |
What readers should know
What did Mistral announce?
Mistral Large 4 entered public API preview on October 6, 2026. Mistral describes it as a one trillion parameter mixture-of-experts model with 49 billion active parameters, accepting text and images.
How did it score?
Artificial Analysis gave the preview an Intelligence Index score of 38 in the October 7 snapshot. That matched GPT-6 Luna at maximum reasoning effort and trailed several other models.
Can users download its weights now?
Not according to information available on October 7. Mistral had scheduled the weights for release later in October; the current access described was an API preview.
Does the score prove it is unreliable?
No. The index is an aggregate benchmark, not a reliability measurement for every workflow. Meyer’s reported hallucinations reflect personal use, not a controlled comparative study.
The Benchmark Gap for Developers
For developers deciding which model to use for long agentic workflows, the launch is both an availability update and a test of Mistral’s competitive position. Such workflows involve planning, tool use and carrying decisions through multiple steps. Errors early in a process can affect later actions, making dependable execution and checking important alongside a model’s advertised capabilities.
The cited index gives developers a reason to test the preview against higher-scoring options, but it cannot by itself settle which model works best for a particular coding, research or business task. A score of 38 does not establish that Large 4 will fail at a specific job. Mistral’s claims about agentic coding and specialized professional work likewise require evaluation on those workloads.
The source author, Thorsten Meyer, says he would not select the current preview for demanding agentic work or long tasks when stronger-scoring models are available. That is an editorial judgment, not a finding from a controlled head-to-head test. The practical issue is whether the model can complete a user’s task accurately enough to justify the oversight required.
As an affiliate, we earn on qualifying purchases.
Preview Now, Weights Later
The distinction between an API preview and a public release of model weights matters to users evaluating access and deployment. As of October 7, Large 4 was available through a preview API; the weights were still pending. Mistral said they were scheduled for release later in the month, so developers could not yet assess a downloadable version from the information provided.
Mistral says the model was trained on its own European infrastructure. That is relevant to the company’s effort to build AI capacity in Europe, but it does not establish how the preview compares with competitors on accuracy, reliability or cost for a particular customer. Those questions need evidence beyond the model’s size or its training location.
The benchmark table also shows why a sweeping claim that every competitor outranks Mistral would be inaccurate: Cohere’s Command A+ scored 13 in the same snapshot. But the comparison cited by the source still places several leading US models and two stronger Chinese models above Large 4 on aggregate benchmark performance.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, ThorstenMeyerAI.com
large language model training hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance Beyond the Index
The available information does not establish how Large 4 performs across specific customer workloads, or how reliably it completes long, multi-step tasks compared with alternatives. Artificial Analysis’s aggregate index is relevant evidence, but it is not a direct reliability test for every coding or research workflow.
Meyer reports encountering hallucinations in his own use of the preview, while saying that this was his experience rather than a controlled comparative study. That account does not establish the model’s overall hallucination rate or show that competing models no longer make unsupported claims. The source material also does not provide complete cost figures, so its heading about the cost argument cannot support a conclusion about Large 4’s price competitiveness.
The planned weight release could allow further examination, but its exact timing and the form of access are not specified beyond Mistral’s schedule of later in October. Scores and model versions may also change after the October 7 snapshot.
AI model performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Tests After the Weight Release
The next stated milestone is Mistral’s planned release of Large 4’s weights later in October. Until then, the product remains an API preview, and the current benchmark figures describe the version assessed in the cited October 7 snapshot.
Developers evaluating it can compare the preview with their own tasks, tracking accuracy, tool use, unsupported claims, supervision needs and cost. Further benchmark results or changes to the model could alter the comparison; the source material does not specify a date for additional tests or provide a detailed schedule for updates.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral announce?
Mistral introduced Mistral Large 4 in public API preview on October 6, 2026. It is a mixture-of-experts model with one trillion total parameters and 49 billion active parameters, and it accepts text and images.
How did Mistral Large 4 score?
Artificial Analysis gave the preview an Intelligence Index score of 38 in the October 7 snapshot cited by the source. That score matched GPT-6 Luna at maximum reasoning effort and was below several other models in the comparison. The index is not a guarantee of performance on a specific task.
Can users download the model weights now?
Not according to the information available on October 7. Mistral had scheduled the weights for release later in October; the current offering was an API preview.
Does the benchmark prove Large 4 is unreliable?
No. The score is an aggregate benchmark result, not a direct measurement of reliability on every workflow. Meyer reports hallucinations in his own use, but identifies that as personal experience rather than a controlled comparative study.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
