QDNASales and integration of LLM inference and training platforms, on-premises or hybrid
The QDNA blog

LLM inference and training, on-premises or hybrid, explained

Short, verified analyses on on-premises AI platforms, open-weight models, LLM costs, agentic AI, and French and European compliance.

Strategy

Kill switch and local LLMs: what French parliamentary report No. 3054 changes

Access to two major AI models suspended on US government orders, and a proposed open-source tax credit for SMEs.

Published July 19, 2026 · 8 min read
Hardware

Prefill, decode and KV cache: which hardware for an LLM

Why memory bandwidth decides the speed, and when NVLink, SXM or a cluster change the game.

Published July 18, 2026 · 10 min read
Guides

Running AI locally on your PC

From the free PC to the DGX Spark mini-supercomputer and the GB10 ecosystem (ASUS, Dell, Lenovo, MSI, Acer).

Published July 18, 2026 · 9 min read
Guides

Enterprise RAG: connecting AI to your documents

Chunking, hybrid search, reranking and access control: the pipeline for connecting an LLM to your documents, on-premises.

Published July 18, 2026 · 9 min read
Compliance

Generative AI and GDPR: taking back control

What the CNIL recommends, the extraterritorial risk, and how local AI gives data control back to the business.

Published July 18, 2026 · 8 min read
Hardware

Which server for AI? Guide and prices by range

Price ranges by tier, from workstation to rack, in a tight memory market that is pushing costs upward.

Published July 18, 2026 · 8 min read
Models

Local LLM: which open-source AI model to choose?

GLM, DeepSeek, Kimi, Mistral, Qwen: a 2026 comparison of open-weight models, by use case, hardware and license.

Published July 18, 2026 · 9 min read
Guides

Installing a local LLM in the enterprise: the guide

The five decisions: model, hardware (NVIDIA, HPE, Dell, Supermicro, Lenovo), runtime, gateway and security.

Published July 18, 2026 · 8 min read
Agentic AI

Orchestrating several coding agents in parallel

From a single workstation to a managed fleet: a supervisor agent delegates to sub-agents, and review becomes the only real bottleneck.

Published July 18, 2026 · 9 min read
Open format

Open Knowledge Format: an open format for AI agent memory

The open OKF standard represents agent knowledge in Markdown: portable, versionable and interoperable, with no lock-in.

Published July 17, 2026 · 8 min read
Use case

Enterprise chatbot: the internal ChatGPT use case

An internal ChatGPT connected to your documents via RAG, GDPR-compliant and hosted on-premises: support, HR, legal, customer service.

Published July 16, 2026 · 8 min read
Use case

Coding with AI on a local model via LiteLLM

Code assistants on local open-weight models: the code stays inside the infrastructure, and the KV cache stays free.

Published July 16, 2026 · 9 min read
Use case

AI for business applications: CRM, support, ERP, documents

A single internal AI API with virtual keys, budgets and telemetry per application, instead of one provider per tool.

Published July 16, 2026 · 9 min read
Fundamentals

What is a sovereign AI platform?

Definition, the difference between on-premises and hybrid, hardware, open-weight models and compliance. The reference guide for scoping a project.

Published July 5, 2026 · 7 min read
Costs

On-premises or API: from how many users does a local LLM pay off?

The LLM cost break-even point, the role of the KV cache, and the comparison with pay-per-token pricing.

Published July 5, 2026 · 10 min read
Method

From your GPU server to a production LLM service

The seven-phase method, with one deliverable and one measurable exit criterion per phase.

Published July 5, 2026 · 11 min read
Agentic AI

From a chatbot to an agentic platform

The harness, the three-layer memory, and the skills that turn a model into a useful agent.

Published July 5, 2026 · 10 min read
Operations

The pitfalls of putting an on-premises LLM into production

Eight costly mistakes, and the measure that avoids each one. Lessons from the field.

Published July 5, 2026 · 9 min read