DeepSeek Launches V4.1 Flash Beta: Native Multimodal Model Hits 420 Tokens/Second
DeepSeek quietly released a 48-hour beta of its V4.1 Flash model on September 8, 2026, with access set to expire on September 10. The model is DeepSeek's first natively multimodal system, integrating text and image processing from the ground up rather than adding vision as an external component. V4.1 Flash recorded a peak speed of 420 tokens per second in long-text reasoning tasks, with benchmarks showing up to 6x speed improvements across coding and retrieval tasks. Unlike its predecessor V4 Flash Vision-Exp, the new model uses a unified latent space for text and images, which the company says improves cross-modal reasoning and reduces latency. DeepSeek is also hiring 150 senior engineers to scale its backend infrastructure, citing the growing complexity of managing large-scale model deployments and agent workflows.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in