QDNASales and integration of LLM inference and training platforms, on-premises or hybrid
Architecture

Real inference measurements

Figures measured on real hardware, one model, format and GPU pair per page, with the procedure that reproduces them.

Method

Each page in this section publishes the measurements of a single pair: an open model, in a given quantization format, on a given machine. The figures come from a measurement script whose full procedure appears in each page, so any reader can check them on an equivalent machine. No figure is extrapolated, copied from a third party or computed: when a measurement does not exist yet, the pair stays listed as planned, with no value.

These measurements complement the sizing matrix, which is computed: it says whether a model fits in memory, while the pages in this section say what the machine actually delivers in throughput and latency.

Published measurements

Planned measurements

Nine more pairs are planned. They will be published as the measurements happen, as access to the machines becomes available, with no value announced in advance.

ModelFormatMachineStatus
Qwen3-VL-Reranker-2BFP16RTX 5080planned measurement
Mistral Small 3.1 24BFP8RTX PRO 6000 serverplanned measurement
Mistral Small 3.1 24BFP8H200 SXM serverplanned measurement
GLM 5.2FP8B200 SXMplanned measurement
Kimi K3FP8B300 SXMplanned measurement
DeepSeek V4NVFP4GB300 NVL72planned measurement
Qwen3 8BFP8Mac Studio Ultraplanned measurement
Qwen3 4BFP8DGX Sparkplanned measurement
Qwen3-Embedding-8BFP8RTX 5080planned measurement

A pair missing from this list can be measured on request, on your hardware or on a reference machine: the about page describes the approach and the contact form frames the procedure.