Kimi K3 2.8T model runs at 1 token/s on MacBook Pro using four SSDs
A project called Deltafin by Argonaut Labs AI has demonstrated running the massive Kimi K3 2.8-trillion-parameter model on a standard MacBook Pro. The model is streamed directly from four SSDs to work around the memory limitations of consumer hardware. This allows the large language model to generate output at approximately one token per second. The project is open source and hosted on GitHub, making the approach accessible to others interested in running very large models on local hardware.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in