Refiant, a South African AI startup, just launched Protea, a family of language models that can hold up to 10 million tokens in a single prompt. That’s roughly 7.5 million words, or about 15,000 pages, sitting in the model’s working memory at once. No waitlist, no approval process. You go to refiant.ai and use it today.
This is really incredible, as I remember the times I spent nights going through years of maintenance logs offshore, trying to trace when a specific fault pattern first showed up, flipping between spreadsheets, PDFs, and old email threads because no single tool could hold the whole history at once. That’s the exact problem Protea is aimed at. If a model can actually hold five years of records in one conversation instead of forcing you to summarise, chunk, and re-feed it piece by piece, that’s not a party trick, that’s hours given back to actual analysis instead of document wrangling.
The founders, namely: Viroshan Naicker, Siddharth Gutta, and Mathew Haswell, aren’t chasing bigger models; they’re chasing more efficient ones. Their earlier claim to fame was compressing OpenAI’s 120-billion-parameter GPT-OSS model down small enough to run on a MacBook Pro with under 20GB of RAM. That result was enough to pull in a $5 million seed round led by VoLo Earth Ventures, plus research partnerships with Imperial College London and UCL’s Sargent Centre. Protea is what that compression work turned into: three tiers, 1 million, 5 million, and 10 million tokens built using evolutionary search and swarm-style optimisation, techniques borrowed from how ants and bees solve problems without centralised control.
The technical claim that actually matters here isn’t the token count; it’s that Refiant says Protea handles the “lost in the middle” problem where most long-context models stay sharp at the start and end of a huge document but quietly lose the thread on everything buried in between. If that holds up under real stress-testing, it’s the difference between a model that can technically ingest a haystack and one that can actually find the needle.
Why African Businesses should care is that, for example, legal teams reviewing hundreds of contracts in one pass. Insurers running years of claims history without re-querying. Engineering teams are processing an entire codebase instead of uploading files in fragments. Those aren’t hypothetical use cases for African firms; they’re daily friction. Nigerian insurers sitting on a decade of paper-based claims records, South African law firms doing due diligence across sprawling mining or energy contracts, Kenyan fintechs auditing years of transaction logs for compliance, all of that currently gets chopped into pieces because no affordable tool could hold it whole. A free, no-waitlist long-context model built and priced with African compute realities in mind is a genuinely different proposition than waiting for OpenAI or Google to eventually drop context limits on their premium tiers.
Refiant is marketing Protea’s 10 million tokens as “one of the largest ever made publicly available.” That’s broadly true, but it slightly overstates the field. A competitor called Subquadratic reportedly shipped a 12-million-token window two months before Protea launched. Refiant isn’t first; it’s a strong entry in a race that’s heating up faster than the press release implies.
There’s also the internal 100-million-token prototype Refiant keeps mentioning with no release date attached. That’s a classic startup move: dangle the next milestone to keep the narrative alive, and it’s fine, but it’s not a product yet; it’s a lab demo.
And nobody’s talking about the practical cost of actually running a 10-million-token prompt. Compression efficiency is Refiant’s whole pitch, but processing 15,000 pages in one pass still takes real compute and real time, and “free” launch pricing on a startup platform rarely stays free once usage scales. African businesses running this on inconsistent power and bandwidth in Lagos, Nairobi, or Douala need to know the actual latency and cost before building critical workflows on top of it, not just the marketing number.
Protea is the most interesting AI infrastructure story out of South Africa this year, and it deserves the attention of a homegrown team solving a real technical bottleneck instead of wrapping someone else’s API. But the smart move for African businesses right now is to stress-test it hard before trusting it with anything regulatory or financial, because Refiant explicitly invited that stress-testing themselves. our prediction within twelve months, either Protea proves itself on messy African datasets and becomes the default long-context tool for legal, insurance, and fintech teams across the continent, or the compute economics catch up with the “free, no waitlist” pitch and pricing changes the whole conversation. Either way, this is the one to watch, not Ghana’s Google lab.

