QDNAAdvisory and architecture for LLM inference and training platforms, on-premises or hybrid

Cost comparator: pay-per-token API or on-site AI server?

From how many users does an AI server installed on your premises cost less than paying per token? This tool applies the dated assumptions published on this site, shows every formula, and checks that the chosen machine sustains the required throughput.

Short answer. With the site's assumptions, a DGX Spark written down over three years costs about €164 excl. VAT per month and becomes cheaper than DeepSeek V4 Flash at the peak rate from 55 agentic users or 309 occasional ones. Change an assumption below and the threshold recomputes. It is a calculation, not a measurement.
Your parameters
API price, € excl. VAT per million tokens
On-site machine

The formulas, as applied

API cost per user per month = working days × (uncached input × input price + cached input × cache price + output × output price), tokens counted in millions.

On-site cost per month = purchase price ÷ write-down months + power × 730 h ÷ 1,000 × kWh price.

Break-even = on-site cost ÷ API cost per user. Required throughput = users × output tokens per day × working days ÷ office seconds in the month (21 days of 8 hours). Machines = required throughput ÷ measured aggregate throughput, rounded up.

Assumptions and sources, dated

Frequently asked questions

Why does the break-even change so much between agentic and occasional profiles?

Because the API cost is proportional to tokens and an agentic user consumes twenty times more input and ten times more output than an occasional one. The on-site cost does not depend on volume as long as the machine sustains the throughput, hence a break-even of 55 users in one case and 309 in the other.

Is the result a measurement?

No, it is a calculation from named and dated assumptions: purchase price read, nominal power, regulated tariff, the day's API price list. A single measurement enters the calculation, the 59 tokens per second aggregate throughput read by a third party on a DGX Spark, and it serves only the capacity check.

What is missing for this to become a full cost?

Engineering time, operations, supervision and premises, which exist in both scenarios but not in the same proportions, and the key-value cache, which changes real throughput. Our article on choosing between on-site and API details these reservations and the sizing sheets give the memory required per model and per machine.

Further reading

A costing on your real volumes

A no-commitment conversation to set your assumptions and constraints.

Book a call