QDNAAdvisory and architecture for LLM inference and training platforms, on-premises or hybrid
Architecture

Open-weight models

The open-weight models selected by QDNA, interchangeable via LiteLLM, ranging from frontier-level to compact formats that run on a single card.

GLM 5.2

GLM 5.2, a 744-billion-parameter mixture of experts under the MIT licence, backbone of the dwarfstar stack.

Kimi K3

Kimi K3 targets enterprise AI and large-scale agents.

Kimi K2.7 Code

Kimi K2.7 Code brings stability in agentic context.

DeepSeek V4

DeepSeek V4 delivers frontier-level performance at reduced cost.

Nemotron 3 Ultra

Nemotron 3 Ultra is NVIDIA’s open-weight model, a Mamba 2 and transformer hybrid in NVFP4.

MiniMax M3

MiniMax M3 handles long context and multimodal workloads at low cost.

Mistral

Mistral is the reference French lab for open-weight models.

Qwen 3.8 and Gemma 4

Qwen 3.8 and Gemma 4 make up the compact models.