SShortSingh.
Back to feed

Alibaba to Release Qwen3 Flash-Next 125B MoE Model Imminently

0
·2 views

Alibaba's Qwen team is set to release a new AI model called Qwen3.8-Flash-Next as early as tomorrow. The model is a Mixture-of-Experts (MoE) architecture with 125 billion total parameters but only 8 billion active parameters per forward pass. The release was announced via a ModelScope model page, suggesting the weights will be publicly hosted on that platform. The post gained early traction on Hacker News, drawing attention from the AI developer community ahead of the official launch.

Read the full story at Hacker News

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

React Native Developers Are the New Target: A Guide to Securing Your Own Machine

A detailed security guide highlights that React Native developers themselves—not just end users—have become the primary targets of supply chain attacks. The guide warns that routine commands like yarn install execute untrusted code with full user privileges, capable of accessing SSH keys, keychains, and environment files without any sandboxing. Configuration files such as metro.config.js, Podfile, and build.gradle are actually executable programs that run automatically on build or project open, creating silent attack surfaces. Real-world incidents like nx/s1ngularity, Shai-Hulud, and GlassWorm demonstrate that these threats are active, with one attack hiding malicious payloads in invisible Unicode characters that evaded code review entirely. The guide also flags AI agent MCP servers as an emerging risk, with over 30% found to carry exploitable vulnerabilities that have already been used to steal private SSH keys.

0
ProgrammingDEV Community ·

Hermes Agent Builds Persistent Skills to Cut Costs for Long-Running AI Tasks

Nous Research released Hermes Agent in February 2026 as an open-source MIT-licensed runtime designed to address a core weakness in AI agent frameworks: the inability to retain and reuse knowledge across sessions. Unlike conventional frameworks that discard reasoning after each task, Hermes logs decision points and tool calls, then enters a reflective phase to assess what worked and convert successful approaches into structured 'skill documents.' These documents are indexed using SQLite FTS5, allowing future similar tasks to query the skill library before invoking the model, with community benchmarks showing up to 40 percent speed gains after around 50 accumulated skills. The architecture is built on five pillars — memory, skills, a persistent behavioral config called 'Soul', scheduled cron jobs, and a self-improvement meta-layer — all designed to compound efficiency over time. Deployment options range from a $59/month managed service to self-hosted builds costing as little as $6–$9 per month, with local inference on an 8B model reportedly achieving 91 percent tool call accuracy on just 8GB of VRAM.