# Mistral Large 4: 1T Parameter Model Enters Preview

> Published 2026-10-07 · https://www.promptzone.com/seojun_sullivan/mistral-large-4-1t-parameter-model-enters-preview-26k

Mistral AI opened the public preview of **Mistral Large 4** on September 28, 2025. The model is a natively multimodal mixture-of-experts system with 1 trillion total parameters and 49 billion active parameters, per a recent Grok AI News thread.

> **Model:** Mistral Large 4 | **Parameters:** 1T total / 49B active | **Type:** Multimodal MoE | **License:** Expected open weights | **Available:** Public preview now

## What It Is

Mistral Large 4 uses a mixture-of-experts architecture that activates only 49 billion parameters per token while maintaining a 1 trillion parameter pool. The model handles text and image inputs natively in a single forward pass.

Weights are scheduled for release by the end of the month. This timing positions the model to enter the open-weight category alongside current top performers from US and Chinese labs.

## Specs and Architecture Numbers

The 49 billion active parameter count places the inference footprint closer to mid-size dense models while the full 1T parameter count supplies capacity for complex multimodal reasoning. No public benchmark scores were released with the preview announcement.

| Feature              | Mistral Large 4     | Typical Frontier Dense Model |
|----------------------|---------------------|------------------------------|
| Total parameters     | 1T                  | 400B–1.8T                    |
| Active parameters    | 49B                 | Full count                   |
| Modalities           | Text + image        | Text + image                 |
| Weight release       | End of month        | Varies                       |

## How to Try It

The public preview is live now. Users can access the model through Mistral’s standard API endpoints without a waitlist. Full weights will appear on Hugging Face once released.

Early testers report standard chat and vision endpoints are already functional. No additional setup steps beyond an API key are required for the preview.

## Pros and Cons

- 49B active parameters reduce inference cost compared with full 1T dense models.
- Native multimodal training removes the need for separate vision adapters.
- Open-weight release planned within weeks lowers barriers for local fine-tuning.
- No benchmark numbers released yet, limiting direct performance comparisons.
- Active parameter count still exceeds most consumer GPUs without quantization.

## Alternatives and Comparisons

**Mistral Large 4** targets the same capability band as GPT-4o, Claude 3.5 Sonnet, and Qwen2-VL-72B. Its MoE design offers a different cost curve once weights are public.

| Model                | Active Params | Multimodal | Open Weights | Current Access      |
|----------------------|---------------|------------|--------------|---------------------|
| Mistral Large 4      | 49B           | Yes        | Planned      | Public preview      |
| GPT-4o               | ~200B est.    | Yes        | No           | API only            |
| Claude 3.5 Sonnet    | ~200B est.    | Yes        | No           | API only            |
| Qwen2-VL-72B         | 72B           | Yes        | Yes          | Hugging Face        |

## Who Should Use This

Developers who need multimodal capabilities and plan to run the model locally after the weight release will find the 49B active parameter count practical. Teams already committed to closed API models for production may see limited immediate benefit until benchmarks appear.

Researchers focused on efficient inference will want the weights when they drop. Organizations requiring verified benchmark leadership before adoption should wait for independent evaluations.

## Bottom Line

Mistral Large 4 brings a 1T-parameter multimodal MoE model into public preview with a clear path to open weights, giving the open-source community a new high-capacity option at a 49B active parameter inference cost.

The release narrows the gap between closed frontier labs and open-weight providers on multimodal tasks.