VIDRAFT Releases Fully Open-Source 6.59B MoE LLM with Five Attention Mechanisms
Seoul-based AI startup VIDRAFT has launched Aether-7B-5Attn, a 6.59-billion-parameter Mixture-of-Experts language model, on Hugging Face under the permissive Apache-2.0 license. Unlike many so-called open models, the release includes not only model weights but also training data recipes, full training code, hyperparameters, logs, intermediate checkpoints, and evaluation code for complete reproducibility. The model's core architectural innovation combines five distinct attention mechanisms — including full, differential, sliding window, sparse, and hybrid attention — distributed across 49 transformer layers using a 7×7 Latin Square layout to prevent depth-level imbalances. Trained on roughly 144.2 billion tokens with an intentional bilingual focus on Korean and English, the model activates only about 2.98 billion parameters per token, keeping inference costs relatively low. Both a base model and an instruction-tuned variant are available, along with a live demo hosted on Hugging Face.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in