Ocular AI, a US-based startup founded by a Zambian and a Zimbabwean, has raised a $2 million pre-seed round led by Drive Capital. Y Combinator, Alumni Ventures, 1745 Ventures (formerly Bertelsmann Digital Media Investments), Orange Collective, MyAsia VC and a group of angel investors also joined. The company builds datasets and benchmarks that AI labs use to train and test voice models. It says some of the top AI labs and Fortune 100 companies now train and test on its data, and that its revenue is in the seven figures and growing quickly. It has not named any customer or given an exact revenue figure, so those claims cannot be checked. The announcement came on October 5 and 6.
The founders are Michael Moyo, who is Zambian, and Louis Murerwa, who is Zimbabwean, according to African tech coverage from 2024. They are former Microsoft and Google engineers, and the company joined Y Combinator’s Winter 2024 batch. That history matters because Ocular AI did not start in voice. Back then it described itself as a platform that let teams search, visualise and run actions across their software tools from one interface. The shift to training data for voice models is a clear change of direction, and it suggests the founders followed where demand was going rather than defending their first idea. One 2024 report called it the first Zimbabwean startup to be backed by Y Combinator, while another said “one of the first.” That claim has not been independently confirmed, so it is left out of this account.
The company’s argument is that voice carries information text does not, including tone, timing and pauses. Most voice AI today works in three steps: one model turns speech into text, a second writes a reply, and a third reads it aloud, so anything that is not a word, such as an interruption, is lost at the first step. Newer full-duplex models skip the transcript and listen and speak at once. Ocular AI says NVIDIA, Thinking Machines Lab, OpenAI and Google have each released or previewed such models. This newsroom confirmed two of them: NVIDIA released its open PersonaPlex model, and Thinking Machines Lab put out a research preview of its first interaction models on May 11, with a public release slated for later this year. The OpenAI and Google claims were not checked here.
The company’s central claim is that the bottleneck is not a shortage of data but a shortage of what it calls “frontier data,” which shows what models cannot yet do. In a May 2026 research post titled Beyond Fisher, Ocular AI argues that major full-duplex models trace their training data to the Fisher English corpus, which was recorded in 2004 as two-party telephone calls at 8 kHz. It cites examples such as Meta’s dGSLM, trained on about 2,000 hours of Fisher, and SyncLLM, which used 1,927 hours of Fisher alongside mostly synthetic speech. It also says the benchmark Full-Duplex-Bench shows no open model handles natural back-channelling and interruption well at once. Ocular’s own answer is a dataset recorded at 48 kHz with a separate channel for each speaker, covering 14 languages and dialects. That is the company’s characterisation and the comparison tables are its own. The post does not report how many hours the dataset holds. Ocular AI also says today’s models struggle with underrepresented accents, but the research post it published does not cite any accent-specific results, so that point rests on the company’s word.
The first product of its new research lab is Converse, a benchmark family that asks whether AI can hold conversations like humans. It looks at four things: understanding natural conversation, speaking with the right emphasis and pacing, managing interruptions, and turning spoken discussions into accurate results. The first test, Converse-STT, is live. Built with the voice-testing firm Cekura, it measures the word error rate of 15 speech-to-text models on real two-person American English conversations and on synthetic audio. On the company’s leaderboard, Reson8 scored best at 2.93%, Cartesia Ink 2 followed at 3.09% and Smallest Pulse at 3.46%, while OpenAI’s GPT-4o Transcribe scored lowest at 12.45%. These are the vendor’s own tests, and a company that sells data also sets the benchmark, so labs will want to replicate the results. The company also plans to extend its Expert Network of specialists beyond voice into medicine, law, finance and software engineering, and says it is already working with labs on audiovisual datasets.
For African founders, the story is less about geography than about position. Moyo and Murerwa built an African-founded company in the supply chain of frontier AI, selling the data and the tests rather than a model. That is a lighter and more specialised business than building a model, and it depends on keeping the trust of a small set of buyers. The next signs to watch are named customers, independent use of Converse, and whether the Expert Network can deliver quality at scale.