South Africa has 12 official languages. Its misinformation detection tools have one working one: English. That’s not an insult to the engineers who built them. It’s just how AI gets made: train it on the biggest pile of text you can find, and the biggest pile is always English. Everything else gets left to “we’ll fix it later.” Meanwhile, the fake COVID cure, the doctored election poster, the deepfaked politician, none of that waits for later. It moves in whatever language the audience actually speaks.
A PhD study out of North-West University just built something that doesn’t wait either. Dr Seani Rananga, now a lecturer at the University of Pretoria, spent her doctorate building a multilingual AI framework that can detect misinformation across English, isiZulu, and Sepedi, not by translating everything into English first and hoping nothing gets lost, but by handling each language on its own terms. She used the COVID-19 pandemic as her pilot case study, testing whether the framework could catch health misinformation as it spread across the three languages.
It worked. The study found multilingual AI can effectively detect misinformation in English, isiZulu and Sepedi while also confirming that machine translation quality plays a major role in how well the system performs for lower-resource languages. In other words: garbage translation in, garbage detection out. The framework is only as good as the language pipeline feeding it.
The bigger ambition isn’t the pilot. Dr Rananga plans to expand the research to more South African languages, with future work focused on misinformation shared during elections, public emergencies, and other high-stakes moments and to eventually tackle AI-generated deepfakes and hate speech too. The research earned her a Google PhD Fellowship, and the long-term vision, in her own words, is to make sure African languages are fully represented in the next generation of AI technologies, not bolted on as an afterthought.
Why this matters way beyond one PhD study is that, think about what “misinformation” actually looks like on the ground in South Africa, or Cameroon, or Nigeria. It’s a voice note in Yoruba forwarded on WhatsApp. It’s a Facebook post in Twi about a fake vaccine side effect. It’s an isiZulu flyer telling people the wrong date to vote. English-only detection tools are structurally blind to almost all of it.
That’s not a hypothetical gap; it’s the exact gap platform moderation and fact-checking pipelines have been quietly living with for years, because nobody built the isiZulu or Sepedi version of the classifier. Rananga’s framework is one of the first serious, published attempts to close it with real engineering instead of a press release. And the timing isn’t subtle. South Africa’s next major election cycles, ongoing public health messaging, and a fast-growing wave of AI-generated content are all converging at once. A misinformation detection tool that only reads English during that convergence is a tool that’s only doing part of its job.
Here’s the honest part: a PhD pilot study tested on COVID-era data across three languages is a proof of concept, not a deployed product. It’s not sitting inside Meta’s moderation stack or South Africa’s Electoral Commission dashboards yet. The gap between “we demonstrated this works in a controlled study” and “this is catching real misinformation at platform scale, in real time, across a messy mix of code-switched WhatsApp voice notes and slang” is enormous, and it’s the gap where most promising academic AI research quietly dies.
There’s also the machine-translation dependency that the study itself flags. If translation quality is the bottleneck for accuracy in low-resource languages, then this framework’s real-world performance is only as strong as its weakest translation layer, and machine translation for isiZulu and Sepedi still has real limitations. That’s not a knock on the research; it’s the honest ceiling on what “multilingual AI” can promise right now. And three languages, out of eleven official indigenous ones, is a start, not coverage. Ndebele, Tsonga, Venda, and Swati speakers are all still waiting.
This is bigger than South Africa. Every country on this continent with more than one dominant language, which is basically all of them, has the same blind spot in its content moderation and fact-checking infrastructure. Nigeria has Hausa, Yoruba, Igbo, and English all carrying misinformation simultaneously. Cameroon runs French, English, and a dozen local languages through the same WhatsApp groups where rumours spread fastest. If a framework like this scales, it’s not just a South African fix; it’s a template.
Building AI that only understands English and calling it “misinformation detection” was never a technical limitation it was a resourcing choice, and everyone building these tools knew exactly which populations they were leaving exposed. This research doesn’t fix that on its own, but it proves the excuse of “it’s too hard to do in African languages” was always weaker than the industry pretended. The next fair question isn’t whether this can be built. It’s why the platforms with billions in moderation budgets didn’t build it first.

