Skip to content

Speech AI · Applied ML

ÒreAyò

Automatic speech recognition for Yorùbá, Igbo and Hausa — the first capability of a wider Nigerian language intelligence platform. I took it from research through to a live public preview: model adaptation, evaluation, an inference service, GPU deployment, and a secure browser-facing gateway.

Status
Public preview · v1.0.2-asr
My role
Sole author, end to end
Core stack
PyTorch · FastAPI · SageMaker
Languages
Yorùbá · Igbo · Hausa

Try the speech recognition preview

The problem

Yorùbá, Igbo and Hausa are the three largest languages of Nigeria, with well over a hundred million speakers between them. Mainstream speech recognition is trained overwhelmingly on English and other high-resource languages, so those speakers are largely expected to operate in a second language to use voice technology at all.

Speech recognition is the input layer everything else needs. ÒreAyò is being built toward broader Nigerian language intelligence — voice interfaces, transcription, accessibility tooling — and none of that is possible until a machine can reliably hear these languages. That is the part I set out to build first.

I am Nigerian and Yorùbá-speaking, which is both why I chose this problem and, as it turned out, useful for evaluating it honestly.

What I built

I own this end to end, as the sole engineer: the research direction and the decision to change it, model adaptation and training, the evaluation methodology and the frozen benchmark behind it, the inference service, its containerisation and GPU deployment, the secure gateway that exposes it to a browser, and the feedback loop that now collects real-world corrections.

The work below is organised around the decisions rather than the chronology, because the decisions are the part worth reviewing.

The engineering journey

The original direction was a custom self-supervised speech encoder trained from scratch. It was the more ambitious approach and, had it worked, the more defensible one — a model shaped around these languages rather than adapted to them.

It did not converge sufficiently. Repeated experiments failed to produce a model that justified continuing, so I stopped the direction on the evidence rather than on how much had already gone into it, and moved to adapting Meta’s MMS instead. That produced the project’s first measurable working ASR across all three languages.

The adaptation keeps the 965M-parameter base entirely frozen and trains a small per-language adapter on top — 2.27M trainable parameters, or 0.24% of the model. One base sits in memory and the ~2 MB adapter is swapped per request, so all three languages are served from a single model rather than three copies of a 1B one. That is what makes a genuine three-language system affordable on one GPU for one person.

Training data is a large public read-speech corpus of Nigerian languages — crowd-recorded, human-transcribed, roughly 2,500 hours. The v1 models deliberately used 5–10% of it, as a cost-control decision that later proved consequential.

Results, and what they are worth

Measured on frozen, hash-pinned evaluation sets of 2,000 utterances per language, against the same base model running Meta’s stock adapters. Scored strictly, with every tone mark and sub-dot preserved — the unflattering convention, for reasons in the decisions below.

Word error rate for the stock base model against the adapted model, by language, with relative change, character error rate, and orthographic complexity.

LanguageBaseline WERAdapted WERChangeCEROrthography
Hausa0.4120.422+2.5%0.104 → 0.104No tone marking
Igbo0.6220.484−22.1%0.191 → 0.150Sub-dots, little tone marking
Yoruba0.6430.588−8.5%0.243 → 0.214Fully tone-marked
Adapted figures measured 22 June 2026; baseline measured 27 August 2026 on the same frozen, hash-pinned sets using the stock MMS adapters. Lower is better. Scored strictly with diacritics preserved — see the normalisation policy above.

The ranking tracks orthographic complexity almost exactly. Hausa is written without tone marks, Igbo uses sub-dots with little tone marking, and Yorùbá is fully tone-marked. The language hardest to write consistently is the hardest to score well on — a statement about the writing system as much as the audio.

On Hausa the adaptation did not help. Stock MMS was already strong there and my adapter scores marginally worse. The likely explanation is that 5–10% of the corpus was enough to displace what the stock adapter already knew without adding enough to compensate — likely, but not yet demonstrated. The full-corpus run that would settle it is in progress, and the deployed Hausa adapter is under review in the meantime.

One limitation materially affects how these should be read: the evaluation is speaker-overlapping, so every evaluation speaker also appears in training. The sentences are unseen, the voices are not, which makes these an optimistic bound.

Every figure is backed by a machine-readable record carrying the SHA-256 of each evaluation set, so the numbers tie to exactly the records that produced them.

From model to product

A trained checkpoint is not a capability. The model is served by a FastAPI application over PyTorch, packaged as a CUDA image and deployed to a SageMaker endpoint on an A10G GPU. Warm latency is around 600 ms on short clips against 25–35 seconds for the same request on CPU, which is the entire justification for the GPU.

That service is now public. The ÒreAyò site carries a speech recognition preview where anyone can record or upload audio in any of the three languages, get a transcript, edit it, and submit the correction. Those corrections are the point: they are the first evaluation signal I have from real speech rather than a curated corpus.

The browser never talks to SageMaker. Requests go through a server-side gateway, which authenticates to AWS using Vercel OIDC against a narrowly scoped IAM role — no static cloud credentials are deployed anywhere in the front end — and is rate limited before anything reaches a GPU.

Three decisions

Stopping the research direction that was not working. The custom encoder was the more interesting project and I had invested real time in it. The experiments did not support continuing, so I stopped and adapted an existing model instead. Ending a direction on evidence rather than on sunk cost is what produced a working system across three languages.

Scoring with diacritics intact, and publishing the worse number. Yorùbá is tonal and written with tone marks. Stripping them before scoring was available, defensible-sounding, and would have made the headline error rate look substantially better. But ọkọ̀ and ọkọ are different words, and a metric that treats them as identical is not measuring transcription. I scored strictly and published what that cost.

Declining to buy more capacity. Yorùbá plateaued, and the reflex at a plateau is to unfreeze more of the model and spend more on compute. The error analysis pointed the other way: the residual error was substantially disagreement about how words are conventionally written rather than failure to hear them. More parameters do not resolve an orthography and labelling problem, so I recorded the decision not to spend and redirected the milestone toward understanding the error instead.

Where it goes next

The current milestone is a measured, reproducible quality ceiling for all three languages, trained on the full corpus rather than the subset v1 used. Yorùbá is complete, Igbo in progress, Hausa queued. It is deliberately framed as a decision to reach rather than a WER target to hit, because choosing a number before understanding the error composition invites optimising the metric instead of the system.

Two things changed how I will run the next iteration. Measure the baseline before the adaptation, not after — knowing Hausa was already well served would have redirected that budget to Igbo, where the same effort was worth three times as much. And freeze evaluation sets to a file and hash them, so an improvement is a measurement rather than a hope.

Beyond ASR, the corrections coming back from the public preview become the first real-world evaluation signal, and speech recognition remains the input layer for the broader Nigerian language intelligence work ÒreAyò is building toward.