vLLM
vLLM delivers high-throughput serving for production.
Triton Inference Server
Triton serves models from multiple frameworks on the same infrastructure.
Dynamo
Dynamo distributes inference for reasoning models across large GPU fleets.
llama.cpp
llama.cpp stays lightweight.
Unsloth
Unsloth speeds up fine-tuning by a factor of two while cutting the memory required, through the LoRA and QLoRA methods.
dwarfstar
dwarfstar is the sovereign stack developed by QDNA.