GDPR-compliant AI: the checklist for generative AI in a company
What the GDPR requires, what the CNIL and the EDPB recommend, what the EU AI Act adds and when. Ten checkpoints with their text and their evidence.
Local AI platform and open-source LLM: the analysis of an on-premise choice
Confidentiality, permanence and control of the model are secured; licence and the 55-user threshold are manageable. QLoRA fine-tuning, harness, enterprise memory, hybrid router: benchmarks and sources verified.
LLM throughput in tokens per second: the four conditions that change everything, measured on sixteen machines
One Llama 2 7B Q4_0 file on sixteen machines, 14 to 290 tokens per second; computed ceiling, quantisation, engine version, concurrency and fifteen sources of DGX Spark benchmarks, in eight tables.
Qwen3.8-Flash-Next on a Mac Studio M5 Ultra: 112 GB of weights, one million tokens
180 billion parameters of which 6 active, a measured 111.6 GB in 4-bit MLX, and a 24 KiB cache per token.
GLM-5.3-Flash on a Mac Studio M5 Ultra: 320 billion parameters within 512 GB
A measured 334.1 GB in 8-bit MLX, an 11 KiB cache per token, a declared one-million-token window.
Advanced enterprise RAG patterns: what separates a demo from a system that holds
Naive RAG never breaks down: it answers, with the wrong chunk. Five patterns, and the order to apply them in.
Sensitive data and AI: what a zero-retention policy really guarantees
The clause covers retention, never transmission. The four places data settles anyway, and what running on-premises changes.
Legal translation and AI: translating a contract without publishing it
The one AI use that transmits the whole document. Why a processing agreement only half answers it, and how to translate on-premises.
AI chatbots and GDPR: the four obligations settled in the architecture
Inform, ground, limit, prove: three of these four are decided in the architecture, and the fourth follows from the other three.
AI orchestration engine: routing every request to the right model
Classifier, routing policy, gateway: the three parts that decide cost, latency and what leaves the company.
Mac Studio M5 Ultra: 512GB at 1.2TB/s, what it changes for a local LLM
+47% bandwidth, 512GB unchanged. Which Flash models actually fit, and what MLX quantisation costs in quality : with measured perplexity.
ASRock AI BOX-A395 + Qwen3.8-27B: 256K context on Strix Halo 128 GB
ASRock AI BOX-A395 (Ryzen AI Max+ 395, Radeon 8060S, 128 GB shared LPDDR5X-8000) runs Qwen3.8-27B multimodal locally. Measured benchmarks llama.cpp + ROCm, KV cache, native 262k context windows, OS comparison, pricing.
Qwen3.8-27B on-premises: 262,144 tokens of context on a single machine
Gated DeltaNet hybrid architecture, 64 KiB-per-token KV cache, native 262,144-token context: real sizing, card by card.
DeepSeek V4 Flash Vision Exp: DeepSeek's first multimodal model
DeepSeek published its first open-weight multimodal model on 31 August 2026, under the MIT license. What the 32-layer visual encoder adds to the V4 Flash architecture, the two changes that are not about vision, the scores the publisher declares, and what the experimental label commits you to.
GLM 5.3 deep on-premise: which hardware, at what speed?
In-depth study of Zhipu AI's GLM 5.3: 744B MoE architecture (40B active), 1M context, published FP8 weights (756 GB), throughput on DGX Spark, DGX Station, H200, B200, B300, GB300 NVL72, Z.ai harness scores against GLM 5.2, Kimi K3, DeepSeek V4 Pro, Qwen3.8-Max and Claude Opus 4.8.
DeepSeek V4 Flash 0731 on DGX Station and 2× DGX Spark cluster
Run the latest DeepSeek model (284B, 13B active, 1M context, FP4+FP8) on-premises on DGX Spark, a 2× Spark cluster and DGX Station GB300 with vLLM 0.25+. Recipes, pricing, benchmarks.
Independent benchmark: DeepSeek V4 Flash 0731 vs Claude Opus 5 vs GPT-5.5
Benchmark 2026: DeepSeek V4 Flash 0731 (284B, 13B active, 1M context) vs Claude Opus 5 vs GPT-5.5. MMLU-Pro, GPQA, AIME, Terminal Bench, cost per token and sovereignty.
Published benchmark results of DeepSeek V4 Flash 0731 (open weight, 2026)
Vendor-published benchmark results for DeepSeek V4 Flash 0731 and GLM 5.2 (open weight, 2026): Terminal Bench, NL2Repo, GPQA Diamond, HLE. No estimates.
Costed scenario: DeepSeek V4 on DGX Station for an 80-person SME
A sizing scenario, not a customer story: DeepSeek V4 Flash 0731 on a DGX Station GB300 for an 80-person SME. Estimated throughput, 3-year TCO, break-even against APIs.
Open weight models for local deployment 2026: complete comparison
2026 comparison of open weight models for local deployment: DeepSeek V4 Flash 0731 and Pro, Kimi K3, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen3.8, Llama 3.1 70B. Parameters, context, published weights, licence and hardware, re-read in the Hugging Face repositories.
Semantic layer and ontology for AI agents
The substrate that makes agents reliable: ontology, knowledge graph, kernel-level eBPF enforcement and sovereign serverless deployment.
Qwen3.8-Max-Preview: 2.4 trillion parameters, and open questions
2.4 trillion parameters, 95 billion active, and since August the base model weights on Hugging Face. What it changes, and does not, for local LLMs.
Kill switch and local LLMs: what French parliamentary report No. 3054 changes
Access to two major AI models suspended on US government orders, and a proposed open-source tax credit for SMEs.
Prefill, decode and KV cache: which hardware for an LLM
Why memory bandwidth decides the speed, and when NVLink, SXM or a cluster change the game.
Running AI locally on your PC
From the free PC to the DGX Spark mini-supercomputer and the GB10 ecosystem (ASUS, Dell, Lenovo, MSI, Acer).
Enterprise RAG: connecting AI to your documents
Chunking, hybrid search, reranking and access control: the pipeline for connecting an LLM to your documents, on-premises.
Generative AI and GDPR: taking back control
What the CNIL recommends, the extraterritorial risk, and how local AI gives data control back to the business.
Which server for AI? Guide and prices by range
Price ranges by tier, from workstation to rack, in a tight memory market that is pushing costs upward.
Local LLM: which open-source AI model to choose?
GLM, DeepSeek, Kimi, Mistral, Qwen: a 2026 comparison of open-weight models, by use case, hardware and license.
Installing a local LLM in the enterprise: the guide
The five decisions: model, hardware (NVIDIA, HPE, Dell, Supermicro, Lenovo), runtime, gateway and security.
Orchestrating several coding agents in parallel
From a single workstation to a managed fleet: a supervisor agent delegates to sub-agents, and review becomes the only real bottleneck.
Open Knowledge Format: an open format for AI agent memory
The open OKF standard represents agent knowledge in Markdown: portable, versionable and interoperable, with no lock-in.
Enterprise chatbot: the internal ChatGPT use case
An internal ChatGPT connected to your documents via RAG, GDPR-compliant and hosted on-premises: support, HR, legal, customer service.
Coding with AI on a local model via LiteLLM
Code assistants on local open-weight models: the code stays inside the infrastructure, and the KV cache stays free.
AI for business applications: CRM, support, ERP, documents
A single internal AI API with virtual keys, budgets and telemetry per application, instead of one provider per tool.
What is a sovereign AI platform?
Definition, the difference between on-premises and hybrid, hardware, open-weight models and compliance. The reference guide for scoping a project.
On-premises or API: from how many users does a local LLM pay off?
The LLM cost break-even point, the role of the KV cache, and the comparison with pay-per-token pricing.
From your GPU server to a production LLM service
The seven-phase method, with one deliverable and one measurable exit criterion per phase.
From a chatbot to an agentic platform
The harness, the three-layer memory, and the skills that turn a model into a useful agent.
The pitfalls of putting an on-premises LLM into production
Eight costly mistakes, and the measure that avoids each one. Lessons from the field.