SShortSingh.
Back to feed

Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

0
·1 views

Tuần này trên Hacker News có một bài hơn 600 điểm: chạy một model MoE 125B tham số trên một con RTX 4090 mà vẫn đạt tốc độ sinh token rất đáng nể. Nghe như chuyện đùa, vì 4090 chỉ có 24GB VRAM, trong khi 125B tham số ở mức quantize 4-bit đã chiếm khoảng 70-75GB. Bí quyết không nằm ở phép màu nào cả, mà ở kiến trúc Mixture of Experts (MoE) và một kỹ thuật gọi là expert offloading. Mình đã thử kỹ thuật này với vài model MoE trên máy cá nhân, và bài này tổng hợp lại những gì thực sự có tác dụng, những gì chỉ tốn thời gian, cùng các lệnh bạn có thể copy về chạy luôn. Với model dense (như Llama 3 7

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

Ask Iche: I built a version of me my colleague can ask about git

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend For six months, my colleague has sent me the same message: "Iche, git is doing something again." She's a Flutter developer learning a Node.js backend. The code isn't where she gets stuck. Git is. Push, pull, fetch, and the classic: "why was my push rejected?" Every time, I stop what I'm doing and explain the same few fixes again. She doesn't need a git course.