The Data Science for Social Impact (DSFSI) multidisciplinary lab based at the University of Pretoria is supported by Artificial Intelligence for Development (AI4D), a partnership between the UK’s Foreign Commonwealth and Development Office (FCDO) and IDRC, with additional funding from Swedish SIDA.
Inclusion by design: making African languages visible to AI
Artificial intelligence (AI) is often framed as a tool for inclusion, yet beneath that promise lies an uncomfortable reality: AI systems can exclude entire populations simply by failing to recognise their languages.
Modern large language models (LLMs) — the technology behind ChatGPT, Gemini, Claude, and other generative AI tools — learn from vast collections of text gathered from across the internet. By identifying patterns in language data, they can generate human-like responses. But the processes used to collect and filter this data embed bias at every stage, long before a model produces its first answer.
One key part of this process is the use of web crawlers — automated programs that move across the internet, finding and organizing online content. This is stored in datasets used to train LLMs. Web crawlers start from a list of known webpages, then follow hyperlinks from page to page, discovering, organizing, and storing more pages, continuously expanding the pool of data they collect.
Why African languages are missing from AI datasets
The challenge lies in where that journey begins. Typical web crawlers are designed to start with a small set of Global North-focused, predominantly English-language websites, bypassing African-language content from the beginning. Even when web crawlers encounter African-language content later in the discovery process, automated language-identification tools often misclassify it. When uncertain, these systems may default to English or confuse one African language with another, thereby discarding valid text.
So, while African-language content is available online, web crawlers often miss it because of where — and in what language — that search starts. Over time, these decisions create a feedback loop: AI systems are trained on incomplete datasets, and their limitations reinforce the perception that those languages are "low-resource", meaning they are not present online — when in reality they are simply underrepresented or misrepresented in the data used to build the models.
So, the problem is not only a lack of African-language data as compared to English and other European languages, but also a lack of visibility. While African-language content represents only a fraction of what is available online, it reflects the wisdom, ideas and experiences of hundreds of millions of Africans who create, share and apply knowledge through indigenous languages every day. This matters because ensuring these languages are represented in research and AI is essential to making knowledge systems more inclusive, and by extension, more complete.
Making African-language content easier to find
For Idris Abdulmumin and his team at Data Science for Social Impact (DSFSI) multidisciplinary lab based at the University of Pretoria, this challenge provided an opportunity. They asked a simple question: How could local language communities help web crawlers find high-quality content in underrepresented languages and build datasets that could be used to train AI tools?
This idea was reinforced by a conversation with a team from Common Crawl, an American non-profit organization that crawls the internet and provides free datasets to the public, making it one of the world’s most widely used sources of training data for large language models. The conversation revealed how poorly African-language content was represented on the platform. Led by Abdulmumin, the DSFSI team at the University of Pretoria set out to make existing African-language data visible and accessible. The Africa Common Crawl (AfriCC) initiative was born.
The Pretoria laboratory sent a call out to the Masakhane research community, a continent-wide, volunteer-driven effort working collaboratively on language technologies. They were asked to complete a simple task: submit one or more website URLs containing content in African languages.
Africa’s AI community more than answered that call. Researchers and volunteers across 19 countries have identified approximately 700 seed URLs in 34 African languages so far. Much of the data is identified through grassroots networks, including volunteer contributors who shared links, annotate text, and help validate language use. Participation is often driven not by financial incentives but by a desire to see one’s language represented accurately in digital systems. As Abdulmumin said, “It's something very nice to be able to explain to someone that your language matters, but then seeing your language represented correctly — not only represented but represented correctly — matters because these languages could die eventually.”
The DSFSI team used these seed URLs to expand Common Crawl's web searches and surface existing content in African language, a process that continues today. So far, the AfriCC project has resulted in a collection of approximately 29 million web pages, including about 6 million from the contributor-provided seed URLs, and contributed to the addition of more than 343,000 African-language pages to the January 2026 Common Crawl dataset, which is an 18% increase from the previous monthly update.
From better data to more inclusive AI
The implications of the AfriCC are significant. A small, community-driven intervention has resulted in a measurable increase in the representation of African languages in one of the world’s most widely used AI training datasets. This demonstrates that many African-language resources are already available online; the challenge lies in identifying, collecting, and integrating them into AI training pipelines.
Ensuring linguistic inclusion requires rethinking how AI models are developed. Much of the current work on underrepresented languages focuses on fine-tuning existing models with small datasets. While valuable, this approach does not address foundational gaps in representation. Efforts like AfriCC, the Africa Next Voices project and the Masakhane African Languages Hub aim to influence the pre-training stage of model development, where systems learn the basic structure of language from large-scale datasets. Improving representation at this earlier stage can significantly enhance downstream performance and reduce the need for costly adaptations later.
Scaling community-driven inclusion
Building inclusive AI requires making languages visible, engaging communities as co-creators, and embedding inclusion at every stage of the pipeline. For Idris and his team, the longer-term hope is to keep scaling this approach affordably and continuously. By making it easy for communities to contribute even a single URL, AfriCC can help grow the volume of African-language data available for training future LLMs and other AI tools, steadily improving availability and performance in African languages. Ultimately, this approach aims to ensure that African languages are embedded from the outset, shaping models that better reflect the multilingual diversity of the societies they serve.
Research Highlights
- Addressing bias in AI training data: The AfriCC initiative highlights how web crawling and language identification systems can systematically exclude African languages from the datasets used to train large language models, reinforcing linguistic inequalities in AI.
- A community-driven approach to data inclusion: Researchers and volunteers across 19 African countries contributed approximately 700 seed URLs in 34 African languages, helping surface underrepresented online content for inclusion in major AI training datasets.
- Demonstrating that data exists but remains invisible: The project shows that many African-language resources are already available online; the challenge lies in identifying, collecting, and integrating them into AI training pipelines.
- Shaping AI at the foundation level: By improving language representation during the pre-training stage of AI development, initiatives like AfriCC can enhance model performance and linguistic inclusion while reducing the need for later adaptations.
Contributors : Janani Balasubramaniam, Program Management Officer, IDRC, Abbey Gandhi, Program Officer, IDRC with Idris Abdulmumin, Associate Research Fellow, University of Pretoria
Share this page