Asenda Talk
A young man with glasses making a phone call indoors, exhibiting focus and communication.

Photo by David Awokoya on Pexels

A voice AI system earns the “multilingual” label only when it can understand, respond to and safely handle a real conversation in each claimed language. If a caller switches to Twi and the agent guesses, stalls or forces English, the label describes a language menu rather than the call experience.

At 4:40 on a Friday afternoon, Kwame, an illustrative composite, is standing behind the counter of his small appliance shop in Kumasi. A customer is waiting with a faulty fan while his phone plays a recorded reminder about an overdue account. Kwame answers in English, hears the amount, then switches to Twi to explain that he paid earlier that week.

The agent pauses. It repeats the original English prompt.

Kwame tries again, mixing Twi with the account reference in English. The system catches the digits but misses the dispute. If it records silence as acceptance or routes the call under the wrong reason, another collection attempt may follow. He cannot tell whether the agent understood him, and the person waiting at the counter is growing impatient.

This is the point where a multilingual claim becomes testable.

Language selection is easier than language understanding

A system may offer a Twi greeting, detect a few common phrases or route callers after they press a key. Those capabilities can be useful. They do not establish that the agent can carry a Twi conversation through interruptions, code-switching and a high-stakes correction.

Real calls rarely stay inside one language. A caller may greet the agent in Twi, state a reference number in English, describe the problem in Twi and use an English product name. Background noise, regional pronunciation and phone audio add more pressure. A polished opening can conceal weak recognition once the caller leaves the expected script.

That weakness matters because speech recognition sits upstream of every later decision. If the transcript changes “I have already paid” into an unrelated phrase, a capable language model may still produce a confident response to the wrong input. The reply can sound natural while the workflow moves in the wrong direction.

This resembles the problem explored in what happens when English-only automation misroutes a Twi payment dispute?: relevance cannot repair a mistaken understanding of the caller’s intent.

Test the switch, not the language list

A useful evaluation starts with conversations that change direction. Ask the agent to handle a caller who begins in English, moves into Twi when explaining the problem, interrupts the response, then returns to English for a name or number.

Listen for more than pronunciation. Check whether the transcript preserves meaning, whether the reply follows the caller’s latest intent and whether the system knows when confidence is too low to continue. Test consent and opt-out phrases in the language callers will actually use. A Twi-speaking caller should not need an exact English command to stop a call.

The first message deserves the same scrutiny. It sets expectations about who is calling, why and what the person can do next. Testing Twi and English first messages in context helps expose failures that a studio-quality sample can hide.

For Kwame’s call, the decisive test is simple: after he switches languages, does the agent understand that he is disputing the account status? If confidence drops, can it avoid confirming an action he did not request? Can an operator later see what happened from the call record?

Those questions reveal more than a page of language badges.

Native speech changes what teams can evaluate

Asenda Talk is being built around native Twi speech recognition and synthesis fine-tuned in-house. That gives the team control over the speech layer instead of placing Twi on top of a general third-party voice API and accepting whatever support it provides.

The platform currently lets early-access teams create voice agents, configure their persona, first message and voice, and run the assistant through Vapi orchestration. Its telephony lifecycle pipeline records call events so operators can compare the intended workflow with what happened. Consent, opt-out and audit records are part of that call trail.

The product remains in active early access. It has not reached feature parity with established platforms such as Vapi, Retell AI or Bland AI, and outbound calling remains behind an operator-controlled real-money gate while the live telephony-provider decision is unresolved. Those boundaries matter when evaluating a pilot. Native Twi capability should be tested as a specific technical advantage, not treated as proof that every surrounding workflow is finished.

Build the pilot around failure evidence

Before placing production calls, write a compact test set from situations your team already handles: payment disputes, appointment changes, delivery questions or campaign responses. Include code-switching, corrections, interruptions and several natural ways to withdraw consent.

Then preserve the evidence. Review the audio, transcript, agent response, call status and billing event together. A transcript alone cannot show whether the caller interrupted. A final status alone cannot explain why the agent reached it. Call-truth tracking matters because a smooth voice can otherwise disguise a broken outcome.

Return to Kwame at closing time. The successful version of his call does not require a theatrical voice or flawless imitation of every speaker. It requires the system to recognize his switch into Twi, preserve the dispute, avoid an unsupported confirmation and leave an auditable record for the next step. He puts down the phone knowing the payment issue was recorded, rather than wondering which English phrase might make the machine listen.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.