MiniMax M3
MiniMax

MiniMax M3 handles long context and multimodal input at low cost. It totals 428 billion parameters, of which 23 billion are active, with a one-million-token context window and sparse attention. It suits large documents and image content.
Key points: Multimodal · 428B mixture-of-experts · 1M context · Sparse attention · Cost-efficient.
Architecture
MiniMax M3 totals 428 billion parameters, of which 23 billion are active, with a one-million-token context window and sparse attention that lowers the cost of long sequences.
Strengths
It handles long context and multimodal input at low cost, covering both text and images.
Use cases
MiniMax M3 suits large documents, image analysis, and workflows where cost per token is the priority.
Deployment
Served by vLLM on a GPU server, or consumed via API. It integrates with the platform through LiteLLM.
Every model on the platform is interchangeable through a single gateway: switching models is a matter of changing one configuration line, with no code to rewrite. Local execution remains the priority.
Deploy this model on your infrastructure
On your own hardware, with your data never leaving.
Book a call