QDNAAdvisory and architecture for LLM inference and training platforms, on-premises or hybrid
Architecture

On-premise AI server architecture

Hardware, open models, inference runtime and multi-agent orchestration, in detailed sheets.

Hardware

From the DGX Spark workstation to the GB300 NVL72 rack, eight power tiers.

Open models

GLM, Kimi, DeepSeek, Nemotron, MiniMax, Mistral, Qwen: interchangeable via LiteLLM.

Inference runtime

vLLM, Triton, Dynamo, llama.cpp, Unsloth and dwarfstar to serve the models.

Multi-agent orchestration

Harness, LiteLLM, semantic router, MCP and three-layer memory.

External APIs

OpenAI, Anthropic, Azure, Bedrock, OpenRouter: hybrid overflow under control.

Sovereign on-premises, hosted or hybrid: which mode should you choose?

The platform is deployed in one of three ways. The choice depends on the data, the target compliance and the available server room. Cost details are covered in on-premises or API and in the price guide.

CriterionSovereign on-premisesSovereign hostedHybrid
Datastays within your walls, on your hardwarestays in France, in a managed datacentre under French and European lawconfidential content stays local, only mundane content may overflow
ComplianceGDPR, HDS and SecNumCloud objectives by constructioncarried by the managed host, under French and European lawthe semantic router classifies every request before the call
Costdepreciated hardware, free key-value cache under vLLMpredictable managed-service fee, no heavy upfront investmentdepreciated local capacity, overflow billed per token
Scalingby adding GPUs or nodes on sitein tiers, on request from the hostimmediate on external APIs
Typical profilesmall to large enterprises with a rack or a server roomorganisations without a server roomusage peaks and public content

A platform designed, deployed and operated with this method already runs in production: gironde.qdna.fr, the Gironde case study.