AMD Launches Instella-MoE-16B-A3B AI Model Trained Entirely on Its Own GPUs
AMD has released a new open AI model called Instella-MoE-16B-A3B, featuring 16 billion total parameters but activating only 2.8 billion per token using a Mixture-of-Experts architecture. The model was trained entirely on AMD's own Instinct MI300X and MI325X GPUs, without relying on Nvidia's CUDA ecosystem. AMD has made the model fully transparent by releasing weights from every training stage, the training code, and data mixture details for public research use. The model requires approximately 32GB of memory in BF16 format, making it accessible to smaller research teams with a single high-end GPU. Model weights are available under a ResearchRAIL license restricting commercial use, while the training code is released under the more permissive MIT license.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in