Kill switch and local LLMs: what French parliamentary report No. 3054 changes
Access to two major AI models suspended on US government orders, and a proposed open-source tax credit for SMEs.
Prefill, decode and KV cache: which hardware for an LLM
Why memory bandwidth decides the speed, and when NVLink, SXM or a cluster change the game.
Running AI locally on your PC
From the free PC to the DGX Spark mini-supercomputer and the GB10 ecosystem (ASUS, Dell, Lenovo, MSI, Acer).
Enterprise RAG: connecting AI to your documents
Chunking, hybrid search, reranking and access control: the pipeline for connecting an LLM to your documents, on-premises.
Generative AI and GDPR: taking back control
What the CNIL recommends, the extraterritorial risk, and how local AI gives data control back to the business.
Which server for AI? Guide and prices by range
Price ranges by tier, from workstation to rack, in a tight memory market that is pushing costs upward.
Local LLM: which open-source AI model to choose?
GLM, DeepSeek, Kimi, Mistral, Qwen: a 2026 comparison of open-weight models, by use case, hardware and license.
Installing a local LLM in the enterprise: the guide
The five decisions: model, hardware (NVIDIA, HPE, Dell, Supermicro, Lenovo), runtime, gateway and security.
Orchestrating several coding agents in parallel
From a single workstation to a managed fleet: a supervisor agent delegates to sub-agents, and review becomes the only real bottleneck.
Open Knowledge Format: an open format for AI agent memory
The open OKF standard represents agent knowledge in Markdown: portable, versionable and interoperable, with no lock-in.
Enterprise chatbot: the internal ChatGPT use case
An internal ChatGPT connected to your documents via RAG, GDPR-compliant and hosted on-premises: support, HR, legal, customer service.
Coding with AI on a local model via LiteLLM
Code assistants on local open-weight models: the code stays inside the infrastructure, and the KV cache stays free.
AI for business applications: CRM, support, ERP, documents
A single internal AI API with virtual keys, budgets and telemetry per application, instead of one provider per tool.
What is a sovereign AI platform?
Definition, the difference between on-premises and hybrid, hardware, open-weight models and compliance. The reference guide for scoping a project.
On-premises or API: from how many users does a local LLM pay off?
The LLM cost break-even point, the role of the KV cache, and the comparison with pay-per-token pricing.
From your GPU server to a production LLM service
The seven-phase method, with one deliverable and one measurable exit criterion per phase.
From a chatbot to an agentic platform
The harness, the three-layer memory, and the skills that turn a model into a useful agent.
The pitfalls of putting an on-premises LLM into production
Eight costly mistakes, and the measure that avoids each one. Lessons from the field.