QDNASales and integration of LLM inference and training platforms, on-premises or hybrid
Architecture

Inference runtime

The engines that serve models, from high-throughput production to edge inference, with fine-tuning and the in-house dwarfstar stack.

vLLM

vLLM delivers high-throughput serving for production.

Triton Inference Server

Triton serves models from multiple frameworks on the same infrastructure.

Dynamo

Dynamo distributes inference for reasoning models across large GPU fleets.

llama.cpp

llama.cpp stays lightweight.

Unsloth

Unsloth speeds up fine-tuning by a factor of two while cutting the memory required, through the LoRA and QLoRA methods.

dwarfstar

dwarfstar is the sovereign stack developed by QDNA.