How much does an AI server cost? Prices by tier in 2026
The price of a server for AI depends on the tier, the number of cards and the memory, and in 2026 the memory shortage is pushing everything upward. Here are indicative ranges by segment.

What sets the price of an AI server?
Three factors dominate: the number and type of graphics cards, the amount of memory, and the server form factor, with memory weighing more than ever in the total. The full calculation, hardware versus pay-per-token, is in our article on-premises versus API cost.
Memory, the driver of 2026 prices
Manufacturers have redirected their capacity toward high-bandwidth memory for AI accelerators, which yields more revenue per wafer than standard memory, and the result is a scarcity of ordinary system memory.
Over one year, server memory prices have risen sharply, and Dell warned its customers of increases of at least 15% to 20% from mid-December 2025, while Lenovo let all of its quotes expire on 1 January 2026 (TrendForce, 5 December 2025). A workstation such as the DGX Spark saw its price rise from around four thousand to nearly four thousand seven hundred dollars, purely because of memory cost, and relief is not expected for several years.
Price ranges by tier
The figures below give order-of-magnitude estimates for 2026, taxes and integration excluded, and they are deliberately wide since configuration and memory can double the total.
| Tier | Who it's for | Indicative range |
|---|---|---|
| DGX Spark workstation (128 GB) | Small business, sovereign workstation | €5,850 incl. VAT (€4,875 excl. VAT), best price listed on idealo.fr on 2 September 2026; €5,850 to €7,511 incl. VAT depending on the seller $6,780 to $8,705 at 1 € = 1.159 $ (ECB rate of 1 September 2026) |
| Mac Studio M5 Ultra (96 GB base, up to 512 GB at 1.2 TB/s) | Small business, silent option | from €6,599 incl. VAT for the M5 Ultra with 96 GB (Apple Store France, 2 September 2026); the 512 GB option, announced for late October, has no published price from $7,650 |
| DGX Station workstation (GB300) | SME without a server room | €93,000 to €110,500 excl. VAT depending on the OEM, i.e. €111,600 to €132,600 incl. VAT (pi3g, page dated 27 August 2026, increase announced); €95,000 excl. VAT (Exxact Valence configuration) is the value used in every calculation on the site; other readings: $99,999 excl. VAT at SHI and £98,000 excl. VAT at Scan UK for the ASUS ET900N G3 $107,800 to $128,100 excl. VAT |
| RTX PRO 6000 server (2 to 8 GPUs) | SME, workhorse | on quote: no vendor publishes a price for an 8-card chassis; reference point, single card at $16,000 excl. VAT on the NVIDIA marketplace on 1 September 2026 (thundercompute reading, secondary source) |
| H200 server (8 SXM GPUs) | Large enterprise | about €453,000 excl. VAT: £387,785 excl. VAT at CTO Servers (8 H200 141 GB GPUs, two Xeon Gold 6538Y+), read on 2 September 2026 at the ECB rate €1 = £0.857; a single vendor, quote to be confirmed about $525,000 |
| GB300 NVL72 rack | Sovereign AI cloud | on quote; analyst estimates of $3.7M to $6.5M (Loop Capital, Tom's Hardware, reported 16 August 2026), which are not prices |
These prices are indicative and non-contractual, estimated for 2026 excluding discounts, hosting and integration, with precise pricing depending on the configuration chosen. Dollar figures are converted at the ECB rate of 1 September 2026 (€1 = $1.159) and move with it, so the euro column is the reference.
Turnkey appliance or identifiable hardware?
Alongside the reference hardware listed above sits a second category, the private AI appliance sold as a sealed box with its own software stack. These vendors publish no list price: Zanus AI, one of the better-known names in the segment, shows three tiers on its own site with a quote request button where the price would be.
Third-party estimate sites fill the gap, and their figures for the same tier vary by nearly a factor of three, which says enough about what they are worth.
The ranges on this page take the opposite approach, since they are tied to identifiable hardware and can therefore be checked line by line against a quote: this many cards, this much memory, this chassis. An appliance price that cannot be decomposed cannot be compared, and that is the only reason we publish wide ranges rather than one reassuring number.
Hardware partners
The calculation relies on NVIDIA graphics cards, and servers are certified by HPE, Dell, Supermicro, Lenovo, ASUS, and Gigabyte, which brings the vendor support and warranty expected in enterprise settings. Dell and Lenovo notably raised their server prices as soon as the shortage began.
QDNA integrates and operates this hardware from workstation to rack, and detailed tiers are in our hardware sheets.
Should you buy, colocate, or rent an AI server?
Three paths coexist: buying ties up capital but amortizes the cost over several years, with a cost per token close to zero, while managed colocation in a sovereign French data center avoids buying premises and keeps legal control, and renting capacity suits peaks or validation, ahead of an investment.
The cost break-even point is detailed in the article on-premises versus API cost.
The market is moving AI on-premises
The market is indeed moving AI on-premises. A survey published by Cloudian in March 2026 among 203 IT decision-makers finds that 93% of enterprises are already repatriating AI workloads from the public cloud, are in the process of doing so, or are actively evaluating it, and that 79% have already moved some.
The same year, Broadcom's Private Cloud outlook, a survey of 1,800 senior IT leaders across eight countries, reports that 83% of leaders are considering repatriation, up from 69% a year earlier. The question of on-site price takes on concrete meaning as those workloads come back home.
The reasons overlap from one study to the next. The Cloudera survey conducted in 2024 among six hundred IT leaders places security and compliance at the top of the barriers (74%), ahead of the lack of skills to manage the tools (38%) and their cost (26%). Control over proprietary data, billing predictability, and legal sovereignty recur as the three drivers of this shift.
Over time, the financial trade-off follows the same slope. For a sustained workload, a purchased and amortized server keeps a cost per token well below usage-based billing, and the three-year total cost of ownership can come in below equivalent cloud billing, which the calculation below lets you check rather than assume. This does not condemn hosting, which stays relevant for peaks and validation phases, in a hybrid approach. It clarifies why the price of a server, measured against real usage, reads as an investment rather than an expense.
Sizing correctly before buying
Model size and number of users set the tier, since a compact model fits on a workstation while a large model requires a server. The practical guide is in deploying a local LLM. In a tight memory market, correct sizing avoids overpaying for an oversized configuration.
How to compare the real cost per token?
The purchase price does not tell the whole story. What matters is the cost per token, which depends on hardware throughput for a given model. The open comparator InferenceX by SemiAnalysis publishes reproducible, auditable measurements: latency, throughput in tokens per second, and implied cost per token.
It pits accelerators against each other in pairs, from the H100 and H200 to the Blackwell B200, B300, GB200 NVL72 and GB300 NVL72 and the RTX PRO 6000, against AMD MI300X, MI325X and MI355X cards. The measurements cover about ten representative open-weight models, including DeepSeek V4 Pro, Kimi K3 and K2.7, GLM 5.3, MiniMax M3, Qwen 3.8 Flash Next, gpt-oss and Llama (list re-read on 2 September 2026).
The useful takeaway for a buyer comes down to a rule the spec sheet never shows on its own: the cost per token depends on the throughput achieved on the target model, so the hardware is chosen from the real workload rather than from the sticker price. A more expensive accelerator that delivers more tokens per second frequently costs less per token, once the pairwise comparison verifies this economics card against card, on the exact model the organisation intends to serve in production.
Cost per million tokens: what can be calculated and what remains to be measured
The cost per token depends on the model served and on the throughput achieved, and it cannot be guessed: it is calculated from four named assumptions, hardware price excluding VAT amortised over 1,095 days, electrical power, price per kilowatt-hour, and aggregate throughput published by a third party for the target model.
The table therefore only keeps the machines for which such a throughput exists for DeepSeek V4 Flash, and the full calculation, formula, assumptions, sources and worked numbers, is in our DeepSeek V4 Flash 0731 versus APIs comparison.
| Solution | Daily cost | Aggregate throughput (third-party source) | 10% load | 30% load | 60% load |
|---|---|---|---|---|---|
| DGX Spark workstation | €5.39 (€4,875 excl. VAT over 3 years, 240 W, €0.1624/kWh) | 30 tokens/s, DeepSeek V4 Flash as a 2-bit GGUF, NVIDIA forum, 15 July 2026 | €20.8/1M | €6.93/1M | €3.47/1M |
| 2× DGX Spark cluster | €10.78 (€9,750 excl. VAT, 480 W) | 210.8 tokens/s, DeepSeek V4 Flash 0731 FP8, vLLM with tensor parallel 2, Classmethod, 10 August 2026 | €5.92/1M | €1.97/1M | €0.99/1M |
| DGX Station workstation | €93.00 (€95,000 excl. VAT, 1,600 W) | [TO BE MEASURED]: 159 tokens/s from a customer screenshot without conditions (pi3g, 4 July 2026) | €67.7/1M | €22.6/1M | €11.3/1M |
| RTX PRO 6000 server, H200 server, GB300 NVL72 rack | prices above (H200 about €453,000 excl. VAT; the other two on quote) | [TO BE MEASURED]: no published throughput for this model with its conditions | not computed | not computed | not computed |
The first line in full: 4,875 / 1,095 = €4.45 of amortisation a day, plus 0.24 kW × 24 h × €0.1624/kWh = €0.94 of electricity, i.e. €5.39 a day; at 30 tokens/s and 30% load, the machine produces 30 × 86,400 × 0.30 = 777,600 tokens a day, hence 5.39 / 0.778 = €6.93 per million. DeepSeek's own price list bills the same model at $0.44 to $0.88 per million (2 September 2026) and Claude Opus 5 costs $15 per million: a local workstation drops below closed APIs from a few hundred thousand tokens a day, and only drops below the vendor's price list as a heavily loaded cluster. Per-card throughput can be checked on the InferenceX comparator by SemiAnalysis.
Frequently asked questions
How much does an AI server cost in 2026?
From a workstation at a few thousand euros to a rack at several million, depending on the tier. An SME server with several cards sits between several tens and several hundreds of thousands of euros depending on memory. Figures are indicative and pushed up by memory.
Why are AI servers so expensive in 2026?
Manufacturers have redirected capacity toward high-bandwidth memory for AI, which has made standard memory scarcer. Server memory prices have risen sharply and several vendors have raised their prices.
What AI server for an SME?
A server fitted with RTX PRO 6000 cards, from two to eight depending on need, covers most SME use cases. No vendor publishes a price for an eight-card chassis: it is priced on quote, the single card being listed at $16,000 excl. VAT on the NVIDIA marketplace on 1 September 2026.
Should you buy or rent an AI server?
Buying amortizes the cost over several years with a low cost per token. Renting suits peaks and validation. Sovereign colocation avoids buying premises while keeping legal control.
What is the total cost of ownership of an AI server?
The purchase price is only part of the cost. You must add electricity, which depends on card power draw and cooling, depreciation over three to five years, maintenance, and hosting. Over time, an amortized server keeps a cost per token well below that of a usage-billed API.
Is the market moving AI out of the cloud?
The trend is measurable. A 2026 Cloudian survey finds that 93% of enterprises are repatriating AI workloads from the public cloud or evaluating it, and Broadcom's Private Cloud outlook rose from 69% to 83% in one year. Control over data, cost predictability, and sovereignty drive this shift, without ruling out some hosting for peaks.
Let's size your AI server
Sizing and an estimate based on your actual usage, whether you buy the hardware or opt for sovereign colocation.
Book a callReferences
- NVIDIA DGX Spark Founders Edition on the idealo.fr price comparator (listing of 2 September 2026)
- NVIDIA DGX Spark Founders Edition on the idealo.fr price comparator (listing of 2 September 2026)
- NVIDIA DGX Spark Founders Edition on the idealo.fr price comparator (listing of 2 September 2026)
- NVIDIA DGX Spark Founders Edition on the idealo.fr price comparator (listing of 2 September 2026)
- NVIDIA DGX Spark Founders Edition on the idealo.fr price comparator (listing of 2 September 2026)
- NVIDIA DGX Spark Founders Edition on the idealo.fr price comparator (listing of 2 September 2026)
- DGX Spark price increase, Tom's Hardware
- RTX PRO 6000 pricing, Tom's Hardware
- HBM3E price hike 2026, TrendForce
- InferenceX inference comparator, SemiAnalysis
- AI server cost breakdown, SemiAnalysis
- Enterprise AI Infrastructure Survey 2026, Cloudian
- Private Cloud Outlook 2026, Broadcom
- State of Enterprise AI 2024, Cloudera
- Dell increases of 15% to 20% and Lenovo quote expiry, TrendForce (5 December 2025)
- DGX Station GB300, prices per OEM excl. VAT, pi3g (page dated 27 August 2026)
- Mac Studio M5 Ultra, Apple Store France (read on 2 September 2026)
- DGX Spark specification sheet, 240 W power supply, NVIDIA
- DGX Station specification sheet, 1,600 W, NVIDIA
- CRE deliberation 2026-147 of 15 July 2026, regulated tariffs from 1 August 2026
- DeepSeek V4 Flash throughput on one DGX Spark, NVIDIA forum (15 July 2026)
- DeepSeek V4 Flash 0731 throughput on two DGX Spark, Classmethod (10 August 2026)