DeepSeek has introduced a new model architecture, V4.1-Flash, designed for enhanced inference efficiency and cost-effectiveness. The model utilizes a combination of 8B and 16B parameter sizes, alongside native visual understanding, and incorporates techniques like Sliding-Window Attention Bounded Replay to optimize the KV cache footprint. This results in a significantly reduced footprint compared to previous versions, potentially leading to faster and cheaper operation for long-running agents.
Independent benchmarks, conducted by Artificial Analysis, show V4.1-Flash achieving a score of 40 on their Intelligence Index, surpassing DeepSeek V4 Pro 0813 despite a lower price point of $0.30 per 1M input tokens. This performance is comparable to GLM-5.3-Flash and exceeds the performance of the original V4 Pro. The model operates with a 1M-token context window and supports text+image input.
Pricing details for V4.1-Flash are $0.30 per 1M input tokens, $1.20 per 1M output tokens with cached input at $0.006 / 1M, and an additional 50% off-peak discount. The model is licensed under the MIT license and is available via DeepSeek’s first-party API, with US and global availability. Several platforms, including Baseten and Ollama, have begun rolling out support for the model.
Notably, DeepSeek’s research team now prioritizes data quality improvements over novel post-training algorithms, aligning with observations from Prof. Jie Tang. This shift reflects a focus on maximizing the impact of existing models through optimized data.
Source: https://www.latent.space/p/ainews-deepseek-v41-flash-763b-p8b



