QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

DeepSeek V4

DeepSeek

DeepSeek V4

DeepSeek V4 delivers frontier-level performance at reduced cost. Its mixture-of-experts architecture reaches 1.6 trillion parameters. The Pro version handles reasoning, while the Flash version serves everyday tasks with speed. This model forms the foundation of dwarfstar.

Key points: Price-to-performance ratio · 1.6T-parameter mixture of experts · Pro and Flash · Foundation of dwarfstar.

Architecture

DeepSeek V4 is a 1.6 trillion parameter mixture-of-experts model, available in a Pro version for reasoning and a Flash version for throughput.

Strengths

It delivers frontier-level performance at reduced cost, making it a benchmark for price-to-performance among open-weight models.

Use cases

The Pro version handles demanding reasoning tasks, while the Flash version serves everyday work with speed. DeepSeek V4 forms the foundation of the dwarfstar stack.

Deployment

Served by vLLM or llama.cpp depending on quantisation. It runs locally on an RTX PRO 6000 server within dwarfstar.

All models on the platform are interchangeable through a single gateway: switching models is a matter of changing one configuration line, with no code rewrite. Local execution remains the priority.

Deploy this model on your premises

On your own hardware, with your data staying in-house.

Book a call