Back to news
launchdeepseek2026-09-10

DeepSeek V4.1 Flash released with native multimodal and KV Cache shrunk to one-fourth

DeepSeek released V4.1 Flash on September 10. The 552B MoE model has 8B activation for input and 16B for output. KV Cache demand for HBM drops to one-fourth and SSD to one-eighth of the prior generation.

On September 10, DeepSeek officially released the DeepSeek V4.1 Flash model. It is the smallest model in DeepSeek's new architecture family, with native multimodal support, and is now available on the DeepSeek API.

DeepSeek V4.1 Flash is a 552B-parameter MoE model using a new Causal-Encoder-Decoder structure, with 8B activation for input and 16B activation for output. Benchmarks show it surpasses the prior DeepSeek V4 Pro across performance, cost, speed, and total processing time.

The new architecture significantly compresses KV Cache: HBM demand drops to one-fourth and SSD demand to one-eighth of the prior generation. In agent use cases, where cache hit cost dominates, this compression materially reduces the operating cost of agent tasks. Compared with the first-generation model, KV Cache has shrunk by a factor of 437.

On the API side, the previous V4 Flash and V4 Flash Vision Exp have been retired. For backward compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily route to V4.1 Flash. After September 14 at 12:00, and until the future V4.1 Pro launches, all requests to deepseek-v4-pro will route to V4.1 Flash and be billed at its unit price.

Pricing has been lowered and continues to use peak-valley pricing, with off-peak pricing at half of peak hours. New pricing took effect on September 10 at 12:00. Official partners Tencent (Hunyuan, CodeBuddy) and OpenCode have full access.

DeepSeekV4.1 FlashMoECausal-Encoder-DecoderKV Cache多模态