On September 10th, DeepSeek officially released the DeepSeek V4.1 Flash model. This is the smallest model in the company's new model structure series, with native multi-modal visual understanding. According to reports, the new generation model drastically reduced the size of the KV Cache cache. Compared with the previous model, the demand for HBM was reduced to 1/4, and the demand for SSD was reduced to 1/8. In agent usage scenarios, cache hits often account for a relatively high cost, and compression of KV Cache greatly reduces the usage cost of agent tasks.

Zhitongcaijing · 2d ago
On September 10th, DeepSeek officially released the DeepSeek V4.1 Flash model. This is the smallest model in the company's new model structure series, with native multi-modal visual understanding. According to reports, the new generation model drastically reduced the size of the KV Cache cache. Compared with the previous model, the demand for HBM was reduced to 1/4, and the demand for SSD was reduced to 1/8. In agent usage scenarios, cache hits often account for a relatively high cost, and compression of KV Cache greatly reduces the usage cost of agent tasks.