We built MustaqiLLM — Uzbek LLM
Today is an important day for us. We are introducing MustaqiLLM, an Uzbek LLM built from the ground up for Uzbek.
This is not a simple fine-tuned version of an existing foreign model. We designed and trained the model from scratch.
The model
- 5 billion parameters
- A 40-billion-token multilingual dataset
- A 48,000-token BPE tokenizer
- Specially adapted for the Uzbek language
- Grouped-Query Attention
- QK-Normalization
- An architecture ready to work with long context
Why does an independent LLM matter?
AI is becoming an essential part of today’s digital infrastructure. That makes it strategically important to build models that work with our language and data—and that we can understand, develop, and control ourselves.
This matters not only for the advancement of Uzbek in AI, but also for the local AI ecosystem, data security, and our ability to contribute independently to the AI technologies of the future.
MustaqiLLM is not the final destination. We will continue improving its knowledge, reasoning, and overall language capabilities.
But it is an important beginning on the path to independent AI in Uzbek.
An important first Uzbek LLM milestone
As an Uzbek LLM, MustaqiLLM is built to advance independent AI that works with our language, data, and local needs. We see it as an important milestone on the path to the first Uzbek LLM.
Model weights are available on Hugging Face: NeuronUz/MustaqiLLM
© NeuronAI Team
