QDNAAdvisory and architecture for LLM inference and training platforms, on-premises or hybrid
The QDNA blog

LLM inference and training, on-premises or hybrid, explained

Short, verified analyses on on-premises AI platforms, open-weight models, LLM costs, agentic AI, and French and European compliance.

Compliance

GDPR-compliant AI: the checklist for generative AI in a company

What the GDPR requires, what the CNIL and the EDPB recommend, what the EU AI Act adds and when. Ten checkpoints with their text and their evidence.

Published on 6 September 2026 · 11 min read
Strategy

Local AI platform and open-source LLM: the analysis of an on-premise choice

Confidentiality, permanence and control of the model are secured; licence and the 55-user threshold are manageable. QLoRA fine-tuning, harness, enterprise memory, hybrid router: benchmarks and sources verified.

Published on 5 September 2026 · Updated on 6 September 2026 · 18 min read
Hardware

LLM throughput in tokens per second: the four conditions that change everything, measured on sixteen machines

One Llama 2 7B Q4_0 file on sixteen machines, 14 to 290 tokens per second; computed ceiling, quantisation, engine version, concurrency and fifteen sources of DGX Spark benchmarks, in eight tables.

Published on 30 August 2026 · Updated on 6 September 2026 · 12 min read
Hardware

Qwen3.8-Flash-Next on a Mac Studio M5 Ultra: 112 GB of weights, one million tokens

180 billion parameters of which 6 active, a measured 111.6 GB in 4-bit MLX, and a 24 KiB cache per token.

Published on 30 August 2026 · 9 min read
Hardware

GLM-5.3-Flash on a Mac Studio M5 Ultra: 320 billion parameters within 512 GB

A measured 334.1 GB in 8-bit MLX, an 11 KiB cache per token, a declared one-million-token window.

Published on 30 August 2026 · 9 min read
Architecture

Advanced enterprise RAG patterns: what separates a demo from a system that holds

Naive RAG never breaks down: it answers, with the wrong chunk. Five patterns, and the order to apply them in.

Published August 28, 2026 · 11 min read
Compliance

Sensitive data and AI: what a zero-retention policy really guarantees

The clause covers retention, never transmission. The four places data settles anyway, and what running on-premises changes.

Published August 28, 2026 · 9 min read
Compliance

Legal translation and AI: translating a contract without publishing it

The one AI use that transmits the whole document. Why a processing agreement only half answers it, and how to translate on-premises.

Published August 28, 2026 · 9 min read
Compliance

AI chatbots and GDPR: the four obligations settled in the architecture

Inform, ground, limit, prove: three of these four are decided in the architecture, and the fourth follows from the other three.

Published August 28, 2026 · 10 min read
Orchestration

AI orchestration engine: routing every request to the right model

Classifier, routing policy, gateway: the three parts that decide cost, latency and what leaves the company.

Published August 28, 2026 · 10 min read
Hardware

Mac Studio M5 Ultra: 512GB at 1.2TB/s, what it changes for a local LLM

+47% bandwidth, 512GB unchanged. Which Flash models actually fit, and what MLX quantisation costs in quality : with measured perplexity.

Published August 28, 2026 · 11 min read
Hardware

ASRock AI BOX-A395 + Qwen3.8-27B: 256K context on Strix Halo 128 GB

ASRock AI BOX-A395 (Ryzen AI Max+ 395, Radeon 8060S, 128 GB shared LPDDR5X-8000) runs Qwen3.8-27B multimodal locally. Measured benchmarks llama.cpp + ROCm, KV cache, native 262k context windows, OS comparison, pricing.

Published August 20, 2026 · 6 min read
Hardware

Qwen3.8-27B on-premises: 262,144 tokens of context on a single machine

Gated DeltaNet hybrid architecture, 64 KiB-per-token KV cache, native 262,144-token context: real sizing, card by card.

Published 18 August 2026 · 8 min read
Models

DeepSeek V4 Flash Vision Exp: DeepSeek's first multimodal model

DeepSeek published its first open-weight multimodal model on 31 August 2026, under the MIT license. What the 32-layer visual encoder adds to the V4 Flash architecture, the two changes that are not about vision, the scores the publisher declares, and what the experimental label commits you to.

Published August 31, 2026 · 9 min read
Models

GLM 5.3 deep on-premise: which hardware, at what speed?

In-depth study of Zhipu AI's GLM 5.3: 744B MoE architecture (40B active), 1M context, published FP8 weights (756 GB), throughput on DGX Spark, DGX Station, H200, B200, B300, GB300 NVL72, Z.ai harness scores against GLM 5.2, Kimi K3, DeepSeek V4 Pro, Qwen3.8-Max and Claude Opus 4.8.

Published August 18, 2026 · 9 min read
Hardware

DeepSeek V4 Flash 0731 on DGX Station and 2× DGX Spark cluster

Run the latest DeepSeek model (284B, 13B active, 1M context, FP4+FP8) on-premises on DGX Spark, a 2× Spark cluster and DGX Station GB300 with vLLM 0.25+. Recipes, pricing, benchmarks.

Published August 2, 2026 · 11 min read
Models

Independent benchmark: DeepSeek V4 Flash 0731 vs Claude Opus 5 vs GPT-5.5

Benchmark 2026: DeepSeek V4 Flash 0731 (284B, 13B active, 1M context) vs Claude Opus 5 vs GPT-5.5. MMLU-Pro, GPQA, AIME, Terminal Bench, cost per token and sovereignty.

Published August 2, 2026 · 10 min read
Models

Published benchmark results of DeepSeek V4 Flash 0731 (open weight, 2026)

Vendor-published benchmark results for DeepSeek V4 Flash 0731 and GLM 5.2 (open weight, 2026): Terminal Bench, NL2Repo, GPQA Diamond, HLE. No estimates.

Published August 2, 2026 · 11 min read
Use case

Costed scenario: DeepSeek V4 on DGX Station for an 80-person SME

A sizing scenario, not a customer story: DeepSeek V4 Flash 0731 on a DGX Station GB300 for an 80-person SME. Estimated throughput, 3-year TCO, break-even against APIs.

Published August 2, 2026 · Updated September 2, 2026 · 8 min read
Models

Open weight models for local deployment 2026: complete comparison

2026 comparison of open weight models for local deployment: DeepSeek V4 Flash 0731 and Pro, Kimi K3, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen3.8, Llama 3.1 70B. Parameters, context, published weights, licence and hardware, re-read in the Hugging Face repositories.

Published August 2, 2026 · Updated September 1, 2026 · 10 min read
Orchestration

Semantic layer and ontology for AI agents

The substrate that makes agents reliable: ontology, knowledge graph, kernel-level eBPF enforcement and sovereign serverless deployment.

Published July 23, 2026 · 9 min read
Models

Qwen3.8-Max-Preview: 2.4 trillion parameters, and open questions

2.4 trillion parameters, 95 billion active, and since August the base model weights on Hugging Face. What it changes, and does not, for local LLMs.

Published July 21, 2026 · Updated September 1, 2026 · 7 min read
Strategy

Kill switch and local LLMs: what French parliamentary report No. 3054 changes

Access to two major AI models suspended on US government orders, and a proposed open-source tax credit for SMEs.

Published July 19, 2026 · 8 min read
Hardware

Prefill, decode and KV cache: which hardware for an LLM

Why memory bandwidth decides the speed, and when NVLink, SXM or a cluster change the game.

Published July 18, 2026 · 10 min read
Guides

Running AI locally on your PC

From the free PC to the DGX Spark mini-supercomputer and the GB10 ecosystem (ASUS, Dell, Lenovo, MSI, Acer).

Published July 18, 2026 · 9 min read
Guides

Enterprise RAG: connecting AI to your documents

Chunking, hybrid search, reranking and access control: the pipeline for connecting an LLM to your documents, on-premises.

Published July 18, 2026 · 9 min read
Compliance

Generative AI and GDPR: taking back control

What the CNIL recommends, the extraterritorial risk, and how local AI gives data control back to the business.

Published July 18, 2026 · 8 min read
Hardware

Which server for AI? Guide and prices by range

Price ranges by tier, from workstation to rack, in a tight memory market that is pushing costs upward.

Published July 18, 2026 · 8 min read
Models

Local LLM: which open-source AI model to choose?

GLM, DeepSeek, Kimi, Mistral, Qwen: a 2026 comparison of open-weight models, by use case, hardware and license.

Published July 18, 2026 · 9 min read
Guides

Installing a local LLM in the enterprise: the guide

The five decisions: model, hardware (NVIDIA, HPE, Dell, Supermicro, Lenovo), runtime, gateway and security.

Published July 18, 2026 · 8 min read
Agentic AI

Orchestrating several coding agents in parallel

From a single workstation to a managed fleet: a supervisor agent delegates to sub-agents, and review becomes the only real bottleneck.

Published July 18, 2026 · 9 min read
Open format

Open Knowledge Format: an open format for AI agent memory

The open OKF standard represents agent knowledge in Markdown: portable, versionable and interoperable, with no lock-in.

Published July 17, 2026 · 8 min read
Use case

Enterprise chatbot: the internal ChatGPT use case

An internal ChatGPT connected to your documents via RAG, GDPR-compliant and hosted on-premises: support, HR, legal, customer service.

Published July 16, 2026 · 8 min read
Use case

Coding with AI on a local model via LiteLLM

Code assistants on local open-weight models: the code stays inside the infrastructure, and the KV cache stays free.

Published July 16, 2026 · 9 min read
Use case

AI for business applications: CRM, support, ERP, documents

A single internal AI API with virtual keys, budgets and telemetry per application, instead of one provider per tool.

Published July 16, 2026 · 9 min read
Fundamentals

What is a sovereign AI platform?

Definition, the difference between on-premises and hybrid, hardware, open-weight models and compliance. The reference guide for scoping a project.

Published July 5, 2026 · 7 min read
Costs

On-premises or API: from how many users does a local LLM pay off?

The LLM cost break-even point, the role of the KV cache, and the comparison with pay-per-token pricing.

Published July 5, 2026 · 10 min read
Method

From your GPU server to a production LLM service

The seven-phase method, with one deliverable and one measurable exit criterion per phase.

Published July 5, 2026 · 11 min read
Agentic AI

From a chatbot to an agentic platform

The harness, the three-layer memory, and the skills that turn a model into a useful agent.

Published July 5, 2026 · 10 min read
Operations

The pitfalls of putting an on-premises LLM into production

Eight costly mistakes, and the measure that avoids each one. Lessons from the field.

Published July 5, 2026 · 9 min read