DeepSeek V4.1 Flash Launches With 1M-Token Context and Ultra-Low API Pricing
DeepSeek released V4.1 Flash on September 10, 2026, an open-weight multimodal model featuring a one-million-token context window and up to 384,000 output tokens. The model uses a 552-billion-parameter Mixture-of-Experts backbone alongside a 196-billion-parameter Engram memory module, with asymmetric activation designed to reduce compute costs during input-heavy agent workloads. Its primary competitive advantage is cost-efficiency for long-context agentic tasks such as coding, automation, and tool use, rather than topping benchmarks against the largest closed models. The model is available on the Modelflare platform under the ID deepseek-v4.1-flash-0910, supporting both Chat Completions and Responses APIs. DeepSeek itself acknowledges the model still trails leading closed systems on the most demanding reasoning tasks, and benchmark figures have not been independently verified.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.



Discussion (0)
Log in to join the discussion and vote.
Log in