Back to Newsroom

AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs

By Modelverse Editorial·August 1, 2026·2 min read
AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs

AMD has unveiled Instella-MoE-16B-A3B, a groundbreaking, fully transparent Mixture-of-Experts (MoE) large language model, developed entirely on its Instinct MI300X and MI325X GPUs. This new LLM boasts a substantial 16 billion total parameters, yet its MoE architecture intelligently activates only 2.8 billion parameters for each token processed, optimizing computational efficiency.

The model's design leverages advanced techniques such as Gated MLA and FarSkip-Collective to achieve this efficient parameter activation. Crucially, AMD has committed to full openness, releasing an unprecedented level of detail for Instella-MoE-16B-A3B. This includes model weights from every training phase, comprehensive data mixtures, configuration files, and the complete inference code.

This comprehensive release is a significant boon for the AI community, offering unparalleled transparency and a robust foundation for innovation. For developers and researchers, access to the full training lineage, data, and code provides invaluable insights into MoE model development and performance on AMD hardware. It not only fosters deeper understanding and reproducibility but also empowers the community to build upon and fine-tune this powerful, openly available model, further accelerating advancements in the LLM landscape.

ai-newsbreakingmarktechpost

Footnotes & Primary References

Related content

Disrupting a Criminal Scam Operation

OpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.

Read article

Accelerating Transformer Training with NVIDIA Transformer Engine, Fused Kernels, BF16, FP8, and GPU Benchmarking

Discover how to optimize transformer workloads using the NVIDIA Transformer Engine. This tutorial guides you through configuring fused GPU kernels, implementing FP8 delayed scaling...

Read article

Judge denies xAI’s request to block Minnesota ban on ‘nudify’ apps

Despite a lawsuit from xAI, a Minnesota ban on apps that allow users to “nudify” images can move forward.

Read article
© 2026 Modelverse®. All rights reserved.Modelverse Newsroom