The South African-built tool helping the world’s under-resourced languages enter the AI age
Prof. Vukosi Marivate, Director of the Absa-UP Chair of Data Science and AI, outside the Aula at the University of Pretoria, Pretoria.
A software tool developed by South African researchers to address one of the greatest barriers facing African languages in the digital age is now being used by researchers around the world.
TextAugment was designed to help developers create language technologies when they do not have the vast quantities of digital text normally required to train artificial intelligence systems. Created by Professor Vukosi Marivate, director of the African Institute for Data Science and AI (AfriDSAI) and holder of the Absa UP Chair of Data Science at the University of Pretoria, and researcher Tshephisho Sefara, the open-source software has recorded more than 286,000 downloads and is being used in work involving languages ranging from Swahili and Arabic to Uzbek.
The software’s growing reach was recognised on July 16 when Marivate and the TextAugment team received the inaugural NSTF-SADiLaR Research Software Award for Human Language Technologies. The award recognises their work in developing open-source software that helps researchers create synthetic training data for languages with limited digital resources.
The recognition adds to a series of recent honours for Marivate. On May 19, the Ga-Rankuwa-born computer scientist was awarded the Order of Mapungubwe in Silver for his contributions to data science, artificial intelligence and natural language processing.
The Presidency commended him for “his excellent contributions to data science, artificial intelligence (AI), and natural language processing (NLP) that have significantly advanced both national and continental technological capabilities”.
But his work is about more than teaching machines to understand language. It is about asking a bigger question: what happens when the technologies shaping the future cannot understand the languages spoken by millions of people?
“AI just existing is not enough for it to actually impact people,” he said. “What you have to do is to have the correct conditions, environments to amplify the good, and then you must also reduce the chances of the negatives that come with any type of technology coming into our world.”
African languages are described in artificial intelligence research as “low-resource” languages. Marivate says the term does not mean that these languages have few speakers. Rather, they lack the digital data and tools that allow machines to process them effectively.
“One of the biggest hurdles in training robust machine learning models for low-resource languages (like isiZulu, Sesotho, or Yoruba) is data scarcity. Traditional AI models require mountains of clean, labelled data to perform well. If that data doesn’t exist, these languages are effectively locked out of the modern AI revolution,” Marivate explained.
Prof. Vukosi Marivate (left) posing with the AfriDSAI research group at the University of Pretoria.
That gap becomes particularly visible when people interact with online systems. Sentiment analysis systems, for example, can determine whether an online review is positive, negative, or neutral.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on iol.co.za — the content belongs to IOL Tech.