Developer Releases 150 GB Open-Source Dataset for Central Asian Languages and Code
A developer has published a nearly 150 GB multilingual dataset on Hugging Face to address the severe shortage of open-source training data for Central Asian languages, including Kyrgyz, Kazakh, Uzbek, and Tajik. The dataset also includes 20 GB of source code in Python, C++, Rust, and Go, making it one of the few resources combining low-resource languages with technical corpora. Building the dataset took weeks of collecting, parsing, and cleaning data, while compressing the archive to around 27.7 GB caused repeated out-of-memory crashes and took roughly 10 hours to complete. The project is released under the CC BY 4.0 license, allowing free use for research, benchmarking, and commercial purposes. It is intended to support tasks such as code LLM fine-tuning, technical translation, and continual pre-training of lightweight language models tailored to Central Asian contexts.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in