Hardware
From the DGX Spark workstation to the GB300 NVL72 rack, eight power tiers.
Open models
GLM, Kimi, DeepSeek, Nemotron, MiniMax, Mistral, Qwen: interchangeable via LiteLLM.
Inference runtime
vLLM, Triton, Dynamo, llama.cpp, Unsloth and dwarfstar to serve the models.
Multi-agent orchestration
Harness, LiteLLM, semantic router, MCP and three-layer memory.
External APIs
OpenAI, Anthropic, Azure, Bedrock, OpenRouter: hybrid overflow under control.