Site map Every section of qdna.fr, from hardware sheets to blog articles. HardwareDGX SparkMac Studio UltraDGX StationRTX PRO 6000 serverH200 SXM serverB200 SXMB300 SXMGB300 NVL72 Open modelsGLM 5.3Kimi K3Kimi K2.7 CodeDeepSeek V4Nemotron 3 UltraMiniMax M3MistralQwen 3.8 and Gemma 4 Inference runtimevLLMTriton Inference ServerDynamollama.cppUnslothninfer Multi-agent orchestrationHermes AgentOpenCodeAgent harnessLiteLLMSemantic routerPrivate MCP serversThree-layer memory External APIsOpenAI (ChatGPT)Anthropic (Claude)Amazon BedrockOpenRouterAzure AI Foundry BlogTutorial: Operating a Digital Organisational Twin (DOTS)LLM inference engines compared in 2026: vLLM, SGLang, TensorRT-LLM, llama.cpp and DynamoSovereign AI inference: 8 × RTX PRO 6000 Blackwell vs 8 × H200 SXMFull Stack AI Company: Building with On-Premises AI, Hugging Face and NVIDIAOn-Premises LLM Training and RFT: From QLoRA to Production ServingGDPR-compliant AI: the checklist for generative AI in a companyLocal AI platform and open-source LLM: the analysis of an on-premise choiceLLM throughput in tokens per second: the four conditions that change everything, measured on sixteen machinesQwen3.8-Flash-Next on a Mac Studio M5 Ultra: 112 GB of weights, one million tokensGLM-5.3-Flash on a Mac Studio M5 Ultra: 320 billion parameters within 512 GBAdvanced enterprise RAG patterns: what separates a demo from a system that holdsSensitive data and AI: what a zero-retention policy really guaranteesLegal translation and AI: translating a contract without publishing itAI chatbots and GDPR: the four obligations settled in the architectureAI orchestration engine: routing every request to the right modelMac Studio M5 Ultra: 512GB at 1.2TB/s, what it changes for a local LLMASRock AI BOX-A395 + Qwen3.8-27B: 256K context on Strix Halo 128 GBQwen3.8-27B on-premises: 262,144 tokens of context on a single machineDeepSeek V4 Flash Vision Exp: DeepSeek's first multimodal modelGLM 5.3 deep on-premise: which hardware, at what speed?DeepSeek V4 Flash 0731 on DGX Station and 2× DGX Spark clusterIndependent benchmark: DeepSeek V4 Flash 0731 vs Claude Opus 5 vs GPT-5.5Published benchmark results of DeepSeek V4 Flash 0731 (open weight, 2026)Costed scenario: DeepSeek V4 on DGX Station for an 80-person SMEOpen weight models for local deployment 2026: complete comparisonSemantic layer and ontology for AI agentsQwen3.8-Max-Preview: 2.4 trillion parameters, and open questionsKill switch and local LLMs: what French parliamentary report No. 3054 changesPrefill, decode and KV cache: which hardware for an LLMRunning AI locally on your PCEnterprise RAG: connecting AI to your documentsGenerative AI and GDPR: taking back controlWhich server for AI? Guide and prices by rangeLocal LLM: which open-source AI model to choose?Installing a local LLM in the enterprise: the guideOrchestrating several coding agents in parallelOpen Knowledge Format: an open format for AI agent memoryEnterprise chatbot: the internal ChatGPT use caseCoding with AI on a local model via LiteLLMAI for business applications: CRM, support, ERP, documentsWhat is a sovereign AI platform?On-premises or API: from how many users does a local LLM pay off?From your GPU server to a production LLM serviceFrom a chatbot to an agentic platformThe pitfalls of putting an on-premises LLM into production SectorsHealthcareFinanceLegalManufacturingPublic sector ToolsMemory calculatorAPI or on-premises cost comparison Case studiesgironde.qdna.frmmr.qdna.frapprendre-ia.qdna.frDGX Spark RDMA clustermedia.qdna.fr, HLS video ↗ (external site) The siteHomeAboutContactLegal noticePrivacy policyAccessibility: non-compliantPlan du site (français)