PromptZone - AI Prompts, Guides and Tools for Builders

Seojun Sullivan
Seojun Sullivan

Posted on

Mistral Large 4: 1T Parameter Model Enters Preview

Mistral AI opened the public preview of Mistral Large 4 on September 28, 2025. The model is a natively multimodal mixture-of-experts system with 1 trillion total parameters and 49 billion active parameters, per a recent Grok AI News thread.

Model: Mistral Large 4 | Parameters: 1T total / 49B active | Type: Multimodal MoE | License: Expected open weights | Available: Public preview now

What It Is

Mistral Large 4 uses a mixture-of-experts architecture that activates only 49 billion parameters per token while maintaining a 1 trillion parameter pool. The model handles text and image inputs natively in a single forward pass.

Weights are scheduled for release by the end of the month. This timing positions the model to enter the open-weight category alongside current top performers from US and Chinese labs.

Specs and Architecture Numbers

The 49 billion active parameter count places the inference footprint closer to mid-size dense models while the full 1T parameter count supplies capacity for complex multimodal reasoning. No public benchmark scores were released with the preview announcement.

Feature Mistral Large 4 Typical Frontier Dense Model
Total parameters 1T 400B–1.8T
Active parameters 49B Full count
Modalities Text + image Text + image
Weight release End of month Varies

How to Try It

The public preview is live now. Users can access the model through Mistral’s standard API endpoints without a waitlist. Full weights will appear on Hugging Face once released.

Early testers report standard chat and vision endpoints are already functional. No additional setup steps beyond an API key are required for the preview.

Pros and Cons

  • 49B active parameters reduce inference cost compared with full 1T dense models.
  • Native multimodal training removes the need for separate vision adapters.
  • Open-weight release planned within weeks lowers barriers for local fine-tuning.
  • No benchmark numbers released yet, limiting direct performance comparisons.
  • Active parameter count still exceeds most consumer GPUs without quantization.

Alternatives and Comparisons

Mistral Large 4 targets the same capability band as GPT-4o, Claude 3.5 Sonnet, and Qwen2-VL-72B. Its MoE design offers a different cost curve once weights are public.

Model Active Params Multimodal Open Weights Current Access
Mistral Large 4 49B Yes Planned Public preview
GPT-4o ~200B est. Yes No API only
Claude 3.5 Sonnet ~200B est. Yes No API only
Qwen2-VL-72B 72B Yes Yes Hugging Face

Who Should Use This

Developers who need multimodal capabilities and plan to run the model locally after the weight release will find the 49B active parameter count practical. Teams already committed to closed API models for production may see limited immediate benefit until benchmarks appear.

Researchers focused on efficient inference will want the weights when they drop. Organizations requiring verified benchmark leadership before adoption should wait for independent evaluations.

Bottom Line

Mistral Large 4 brings a 1T-parameter multimodal MoE model into public preview with a clear path to open weights, giving the open-source community a new high-capacity option at a 49B active parameter inference cost.

The release narrows the gap between closed frontier labs and open-weight providers on multimodal tasks.

Top comments (0)