MODEL
DeepSeek-V4.1-Flash
- LAUNCHED
- CURRENT
- V4 PRO AND FLASH
DeepSeek-V4.1-Flash launches on the API as the smallest model in a new architecture family: a 552B-parameter causal encoder–decoder with 8B active parameters for input and 16B for output, native vision, and lower API prices. It replaces V4-Flash and the experimental V4-Flash-Vision-Exp.
Why it mattered. DeepSeek reported it ahead of V4-Pro on benchmarks, with a KV cache needing a quarter of the memory, which cuts cache charges, often a large share of agent costs.
- AVAILABLE TO
- Open weights · API
- BENCHMARKS
- DeepSeek V4.1 Flash →