Moonshot Kimi K3 runs on 16x DGX Spark cluster at 38 tokens/s at 16K depth
NVIDIA developer forum users report running Kimi K3 on a 16-node DGX Spark cluster, reaching 38 tokens/s generation speed at 16K context depth.
On September 5, NVIDIA developer forum users shared test results showing a 16-node DGX Spark cluster successfully running the full Kimi K3 model. At a 16,000 token context depth, the model reached a maximum generation speed of 38 tokens per second.
Kimi K3 was open-sourced by Moonshot AI on July 27 via Hugging Face with 2.8 trillion total parameters and 104 billion activated parameters in a MoE architecture, natively supporting 1M-token context. The model uses the self-developed Kimi Delta Attention hybrid linear attention mechanism, activating 16 out of 896 routed experts per token.
The full K3 model weights total roughly 1.56 TB, making independent deployment a substantial infrastructure commitment. DGX Spark is NVIDIA's compact desktop-class AI workstation. The cluster deployment shows that open-weight models can run on a wider range of hardware than previously expected.