QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

MiniMax M3

MiniMax

MiniMax M3

MiniMax M3 handles long context and multimodal input at low cost. It totals 428 billion parameters, of which 23 billion are active, with a one-million-token context window and sparse attention. It suits large documents and image content.

Key points: Multimodal · 428B mixture-of-experts · 1M context · Sparse attention · Cost-efficient.

Architecture

MiniMax M3 totals 428 billion parameters, of which 23 billion are active, with a one-million-token context window and sparse attention that lowers the cost of long sequences.

Strengths

It handles long context and multimodal input at low cost, covering both text and images.

Use cases

MiniMax M3 suits large documents, image analysis, and workflows where cost per token is the priority.

Deployment

Served by vLLM on a GPU server, or consumed via API. It integrates with the platform through LiteLLM.

Every model on the platform is interchangeable through a single gateway: switching models is a matter of changing one configuration line, with no code to rewrite. Local execution remains the priority.

Deploy this model on your infrastructure

On your own hardware, with your data never leaving.

Book a call