Africa's AI moment has arrived — this 1.5B model built from scratch for 12 African languages beats Google, Meta, and Alibaba while being 8x smaller
MORENA beat 26 tested models despite having only 1.5 billion parameters The model from Vambo AI was built specifically around African languages from scratch MORENA uses fewer tokens to represent African text than major rivals Vambo AI has released MORENA, a 1.5B parameter language model covering 12 languages spoken across Africa, plus English, French and code.
The developers say it beats models from Google, Meta, Alibaba and HuggingFace on African language modelling and translation while being up to 8x smaller.
On a key benchmark, this LLM scored 1.408 bpb — the best of 26 models tested — beating its closest rival, which carried over five times as many parameters yet still scored only 1.423 bpb.
Training without a borrowed foundation The model covers Nigerian Pidgin, Igbo, Yoruba, Hausa, Kiswahili, ChiShona, isiZulu, isiXhosa, Kinyarwanda, Setswana, Afrikaans, and isiNdebele.
Most projects serving these languages continue training an existing Llama variant, which forces them to keep a vocabulary meant for English and programming text.
Vambo AI instead fixed the tokenizer, the data mixture, and the language list before training began, after comparing several vocabulary sizes for cost and efficiency.
MORENA's vocabulary encodes African text with 1.39 times fewer tokens than Gemma 3 and 1.53 times fewer than Llama 3.2 on identical passages.
African text still costs 0.249 tokens for each byte against 0.234 for English, roughly 6% more, and the team cannot fully explain the difference.
The nearest competitor, Lugha-Llama-8B, an Africa-adapted Llama 3.1 variant, scores 1.423 bpb and loses to MORENA in eight languages out of 12.
The 12B Gemma system, by comparison, consumes roughly 11 times the compute of MORENA while producing a weaker score on the same text.
The chat-tuned instruct version scores 1.441 bpb, trailing Lugha-Llama-8B overall but leading it in five shared languages while using about 20% of its parameters.
The instruct model, after seeing three sample translations, renders English into five languages from Africa at 45.8 chrF++, statistically level with one dedicated translation system.
It trails a larger translation model by about 1.4 points, while general models of similar size typically land in a range from 9 to 14.
A family built on 22,000 GPU hours MORENA used 251.7B tokens during pretraining, followed by another 63B tokens during its mid-training stage.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.techradar.com — the content belongs to TechRadar.