Developer runs 125B-parameter AI model on 16GB Mac using SSD streaming trick
A developer has released an open-source tool called Slotstream that enables running Qwen3-235B-A22B-Flash (a 125-billion-parameter model normally requiring over 100GB of memory) on Macs with as little as 16GB of RAM. The tool achieves this by using expert-offloading and SSD streaming to compensate for limited onboard memory. Built natively for Apple hardware using MLX and Swift, it can run the model at around 12 tokens per second on a 48GB Mac. Slotstream ships with an auto-mode that automatically balances memory usage against inference speed. The developer plans to add support for speculative decoding via the MTP module in a future update.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in