Cost comparator: pay-per-token API or on-site AI server?
From how many users does an AI server installed on your premises cost less than paying per token? This tool applies the dated assumptions published on this site, shows every formula, and checks that the chosen machine sustains the required throughput.
The formulas, as applied
API cost per user per month = working days × (uncached input × input price + cached input × cache price + output × output price), tokens counted in millions.
On-site cost per month = purchase price ÷ write-down months + power × 730 h ÷ 1,000 × kWh price.
Break-even = on-site cost ÷ API cost per user. Required throughput = users × output tokens per day × working days ÷ office seconds in the month (21 days of 8 hours). Machines = required throughput ÷ measured aggregate throughput, rounded up.
Assumptions and sources, dated
- DeepSeek V4 Flash price list of 2 September 2026, peak hours: $0.44 per million uncached input tokens, $0.014 cached, $1.32 output; off-peak half price. Converted at €1 = $1.159 (ECB, 1 September 2026).
- DGX Spark Founders Edition: €4,875 excl. VAT, best idealo.fr price read on 2 September 2026, written down linearly over 36 months. 240 W is the nominal supply of the NVIDIA sheet, a high assumption: ServeTheHome measures 60 to 200 W under load.
- Electricity: EDF Tarif Bleu professional, Base option, price list of 1 August 2026, €0.1624 excl. VAT per kWh, CRE deliberation 2026-147.
- Aggregate throughput of 59 tokens per second at twelve simultaneous requests for DeepSeek V4 Flash on a DGX Spark: Entrpi reading, DSpark engine, NVIDIA forum, 15 July 2026. A third-party measurement on one precise configuration.
- DGX Station: €95,000 excl. VAT (Exxact, read on 27 August 2026); power and throughput not measured by this site, to be entered.
- Not counted: integration, operations, engineering time, premises, and the opportunity cost of a model that has no open weights.
Frequently asked questions
Why does the break-even change so much between agentic and occasional profiles?
Because the API cost is proportional to tokens and an agentic user consumes twenty times more input and ten times more output than an occasional one. The on-site cost does not depend on volume as long as the machine sustains the throughput, hence a break-even of 55 users in one case and 309 in the other.
Is the result a measurement?
No, it is a calculation from named and dated assumptions: purchase price read, nominal power, regulated tariff, the day's API price list. A single measurement enters the calculation, the 59 tokens per second aggregate throughput read by a third party on a DGX Spark, and it serves only the capacity check.
What is missing for this to become a full cost?
Engineering time, operations, supervision and premises, which exist in both scenarios but not in the same proportions, and the key-value cache, which changes real throughput. Our article on choosing between on-site and API details these reservations and the sizing sheets give the memory required per model and per machine.
Further reading
- On-site or API: the full calculation and its reservations
- Which AI server, at what price
- DGX Spark sheet
- Sizing, model by machine
A costing on your real volumes
A no-commitment conversation to set your assumptions and constraints.
Book a call