VIDRAFT Releases Aether-7B-5Attn, an Open-Source MoE Model with Five Attention Types
Seoul-based AI startup VIDRAFT has published Aether-7B-5Attn on Hugging Face under the Apache 2.0 license, making it freely available for research and commercial use. The model is a Mixture-of-Experts architecture with 6.59 billion total parameters, activating roughly 2.98 billion per token to keep inference costs low. A distinctive feature is its use of five different attention mechanisms — including full, differential, and sliding window attention — distributed across 49 layers via a 7×7 Latin square arrangement. The model was trained on 144.2 billion tokens, with mathematics, Korean, and English each representing major portions of the data mix, reflecting a focus on multilingual and mathematical capability. VIDRAFT has released not just model weights but also training code, hyperparameters, logs, intermediate checkpoints, and evaluation code, positioning it as one of the most transparently documented open-source model releases to date.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in