QDNASales and integration of LLM inference and training platforms, on-premises or hybrid

Qwen 3.6 and Gemma 4

Alibaba and Google

Qwen 3.6 and Gemma 4

Qwen 3.6 and Gemma 4 are the compact models. Qwen 3.6 27B and Gemma 4, from 4 to 31 billion parameters, run on a single card, a workstation or at the edge. They suit small businesses.

Key points: Compact · Qwen 3.6 27B and Gemma 4 · Single card · Edge · Small businesses.

Architecture

Qwen 3.6 27B and Gemma 4, from 4 to 31 billion parameters, are the platform's compact models.

Strengths

They run on a single card, a workstation or at the edge, with a reduced memory footprint.

Use cases

These models suit small businesses, embedded uses and tasks where local latency and cost matter more than raw power.

Deployment

Served by llama.cpp or vLLM on DGX Spark, Mac Studio or a single GPU card.

All the models on the platform are interchangeable through a single gateway: switching models is just changing one configuration line, with no code rewrite. Local execution remains the priority.

Deploy this model at your site

On your hardware, with your data never leaving.

Book a call