Method
Each page in this section publishes the measurements of a single pair: an open model, in a given quantization format, on a given machine. The figures come from a measurement script whose full procedure appears in each page, so any reader can check them on an equivalent machine. No figure is extrapolated, copied from a third party or computed: when a measurement does not exist yet, the pair stays listed as planned, with no value.
These measurements complement the sizing matrix, which is computed: it says whether a model fits in memory, while the pages in this section say what the machine actually delivers in throughput and latency.
Published measurements
- Qwen3-Embedding-4B in FP8 on RTX 5080, measured on 31 July 2026: 14,100 tok/s sustained, 19,075 tok/s on long text, 46 requests per second at saturation.
Planned measurements
Nine more pairs are planned. They will be published as the measurements happen, as access to the machines becomes available, with no value announced in advance.
| Model | Format | Machine | Status |
|---|---|---|---|
| Qwen3-VL-Reranker-2B | FP16 | RTX 5080 | planned measurement |
| Mistral Small 3.1 24B | FP8 | RTX PRO 6000 server | planned measurement |
| Mistral Small 3.1 24B | FP8 | H200 SXM server | planned measurement |
| GLM 5.2 | FP8 | B200 SXM | planned measurement |
| Kimi K3 | FP8 | B300 SXM | planned measurement |
| DeepSeek V4 | NVFP4 | GB300 NVL72 | planned measurement |
| Qwen3 8B | FP8 | Mac Studio Ultra | planned measurement |
| Qwen3 4B | FP8 | DGX Spark | planned measurement |
| Qwen3-Embedding-8B | FP8 | RTX 5080 | planned measurement |
A pair missing from this list can be measured on request, on your hardware or on a reference machine: the about page describes the approach and the contact form frames the procedure.