SShortSingh.
Back to feed

Moonshot's Kimi K3 Freezes New Signups Within 48 Hours Due to GPU Shortage

0
·1 views

Chinese AI firm Moonshot launched its Kimi K3 model, only to suspend new subscriptions within 48 hours after demand overwhelmed its available GPU capacity. The model, built on 2.8 trillion open-weight parameters, reportedly outperformed rivals from Anthropic and OpenAI on front-end coding benchmarks, drawing a surge of developer interest. Moonshot allowed existing users to retain access and announced plans to reopen signups in batches as it expands infrastructure. The company also introduced two separate membership tiers — one for general use and another focused on programming — reflecting the heavier compute demands of coding workloads. The incident highlights a broader industry shift, where the primary AI bottleneck is moving from model training to inference, as long-running agentic tasks consume far more compute per user than traditional chatbot interactions.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingHacker News ·

German battery maker Varta files for insolvency

German battery manufacturer Varta has filed insolvency applications, marking a significant financial setback for the company. Varta is a well-known producer of batteries, including small cells used in consumer electronics and hearing aids. The filing was reported on July 24, 2026, signaling the company's inability to meet its financial obligations. This development raises concerns about the future of the firm and its workforce. The insolvency process will now determine how the company's debts and assets are handled going forward.

0
ProgrammingHacker News ·

Why Text-Based MUD Games Still Have a Place in the Modern Era

A recent opinion piece published on andrewzigler.com argues that MUDs, or Multi-User Dungeons, remain relevant despite the dominance of graphically rich modern games. MUDs are text-based multiplayer online games that originated in the late 1970s and were popular through the 1990s. The article makes a case for their continued value in today's gaming landscape, though the specific arguments are detailed at the source link. The piece was shared on Hacker News, where it received modest early traction with six points. No community comments had been posted at the time of sharing.

0
ProgrammingDEV Community ·

Key AWS S3 and CloudFront Edge Cases for Static Site Subdirectory Hosting

A developer deploying a static documentation site on AWS S3 and CloudFront encountered several non-obvious routing and configuration issues. Browser caching of 301 redirects and HSTS policies caused stale behavior in normal sessions even after infrastructure changes had propagated. CloudFront's Default Root Object only resolves index files at the distribution root, not in subdirectories, requiring explicit URL rewriting via a CloudFront Function on the Viewer Request event. Missing trailing slashes in subdirectory URLs caused browsers to misresolve relative asset paths, breaking CSS and JavaScript loading. A concise CloudFront Function that redirects slash-less directory requests and rewrites directory URIs to index.html was found to resolve most of these routing problems.

0
ProgrammingDEV Community ·

Gemini 3.6 Flash's Thinking Dial Can Cut AI Costs 30x With Accuracy Trade-offs

Google's Gemini 3.6 Flash, which became generally available on July 21, 2026, charges users for hidden 'reasoning tokens' in addition to standard output tokens, billed at $7.50 per million. Independent testing conducted on July 24, 2026, found that the model's reasoning effort setting — ranging from minimal to high — can swing costs by up to 30 times on the same task, dropping a 120-word writing task from $0.03316 to $0.00110. Setting reasoning effort to 'minimal' eliminates reasoning tokens entirely and cuts per-call costs by 91–97%, with no noticeable quality loss on retrieval, formatting, and simple factual tasks. However, the minimal setting caused the model to fail all three attempts at a multi-step arithmetic word problem, highlighting a meaningful accuracy risk for complex reasoning tasks. The model's 1-million-token context window and prompt caching behavior were confirmed to match Google's published specifications in testing.

Moonshot's Kimi K3 Freezes New Signups Within 48 Hours Due to GPU Shortage · ShortSingh