Asenda Talk
A cheerful call center agent talking on the phone with a headset.

Photo by Yan Krukau on Pexels

The gap between what an AI voice agent promises and what it actually delivers is nowhere wider than in a real Twi conversation about something that matters. An agent can be praised for "decision-making capability" in a pitch deck, then fall apart three sentences into a caller's question about a payment that has not arrived.

Kwame called the district office on a Thursday morning because his mother's monthly disbursement had not landed. She had been on the scheme since the previous year, and the money had always come by the 15th. This month it did not. He dialed the number printed on the letter, expecting a queue, prepared to wait. What he got instead was a machine that answered in English, asked him to select from a menu in English, and then, when he replied in Twi to the English prompt, said it did not understand him and repeated the menu.

He tried again, this time speaking slower. The same loop. He said the words clearly, the way you would to someone learning. The machine asked him to try once more. He hung up and dialed again, because his mother's question was simple: where is the money? He was not asking anything complicated. He was not asking anything an agent should need a decision tree to handle. He was asking one thing, and the machine could not hear it.

The distance between "decision-making" and understanding

When a platform advertises AI decision-making, it usually means the model can pick an intent, route a call, or decide which scripted branch to follow. That is useful machinery. But it assumes the caller can get to the decision point in the first place, which assumes the agent understands what the caller is actually saying.

In Twi, that assumption fails early and often. Twi is tonal. A phrase that means one thing with a high tone means something else with a low one. Pronunciation that passes as acceptable in one of the dialects can mangle a word in another. A word-for-word English-to-Twi translation does not survive contact with a real speaker on a real phone line with background noise from a lorry park. This is why "Twi-compatible" agents that treat the language as a lookup table keep hearing words but missing intent. What Happens When a Voice Agent Can't Actually Speak Twi? walks through the mechanics of where that breakdown happens.

The mismatch matters more when the stakes are public money. A loan application, a health screening, a benefit status check. The caller is not experimenting with a new gadget. They are trying to resolve something concrete, and when the agent fails, they do not blame the agent. They blame the institution that put it in front of them.

Why a wrapper makes the gap wider

The easy way to ship a voice agent is to rent speech recognition and synthesis from a vendor that built its models for English and bolted on an African language later. That approach works for a demo. A caller says "hello", the agent replies, everyone is impressed the recognizer heard anything at all.

But a public service desk is not a demo. It is hundreds of calls a day, each with a caller who has a real question and no patience for a second attempt. A wrapper model that hears Twi "well enough" will handle the first exchange, then stumble on the follow-up, then mishear a figure, then route the caller to the wrong place. Each stumble costs trust, and trust is the one thing a public institution cannot afford to spend on calls about money.

The honest position is that native Twi speech recognition and synthesis, fine-tuned in-house, is not a marketing advantage. It is the difference between an agent that can hold a conversation and one that can recite a script. That is the entire product question for anyone building a voice agent for a Ghanaian audience, and it is a question no amount of model-selection cleverness downstream can fix.

What a call that works actually looks like

The third time Kwame dialed, a different system answered. It greeted him in Twi, not English. It asked, in Twi, what he was calling about. He said his mother's disbursement had not arrived. The agent asked for the registration number, and he read it out slowly, and the agent repeated it back correctly, which was the moment he decided it might be real. Then it told him, in Twi, that the payment had been processed but the bank had rejected the account on file, and that a letter was on its way, and it offered to walk him through updating the details then and there.

Kwame does not think of himself as someone who researches speech models. He noticed two things. The agent did not ask him to repeat himself. And it said his mother's money was not lost, just held up, which was more than the letter had told him. His mother got her payment the following week, after they fixed the account details on that same call.

The difference did not come from a better decision tree. It came from the agent actually understanding the Twi that was spoken to it, well enough that the conversation could move past recognition and into the part where a decision was needed. Does your AI agent truly understand Twi, or just hear the words? presses further on that distinction.

For any team building a voice agent for public services in Ghana, the question is not which orchestration layer to pick. It is whether the agent you put in front of a caller can understand the language that caller actually speaks before it is asked to make a single decision.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.