Asenda Talk
A focused team of customer service representatives working on computers and answering calls in an office

Photo by Pavel Danilyuk on Pexels

An English greeting reveals how a caller starts, not which language they will need when the problem becomes difficult. Voice support systems should treat language as something that can change during a call, especially when customers explain disputed charges, consent boundaries, account details, or unfamiliar terms.

Consider an illustrative composite caller named Efua, a market trader in Kumasi who keeps handwritten stock notes beside her phone. At 4:40 p.m., she answers a support call with, “Hello, good afternoon,” then follows the agent’s English questions without hesitation.

The call changes when the agent asks which transaction she wants to dispute. Efua can read the amount, but explaining why the charge is wrong takes her into Twi. She describes who collected the payment, what was promised, and the detail that separates the disputed charge from another valid one.

The agent keeps responding in English.

Efua tries again, shortening her explanation each time. If the system records the wrong transaction, her valid payment could be challenged while the disputed charge remains untouched. She has both references in front of her. The difficult part is making the distinction in the language where she can state it precisely.

A greeting is weak evidence of language preference

“Hello” is common across Ghanaian conversations. So are English names, dates, menu choices, and short confirmations. A caller may use them comfortably without wanting to discuss the full problem in English.

Support interactions also become harder as they progress. The opening usually asks for simple information: a name, a reason for calling, or permission to continue. Later turns demand explanation. The caller may need to correct an assumption, describe a sequence of events, or reject an action the agent proposes.

That is where language preference becomes visible.

A system that classifies the whole call from its first few seconds can mistake familiarity for fluency. It hears a confident English greeting and locks the conversation into English. When the caller switches, the system may treat the Twi as noise, transcribe it poorly, or push the caller back toward shorter English answers.

The result can look tidy in a transcript while preserving the wrong meaning. A customer’s Twi evidence can lose one decisive detail when the system handles language as a fixed call setting.

Difficult moments trigger the switch

People often reach for the language that gives them the most control when the stakes rise. That switch may happen during a correction, an objection, a refusal, or a detailed account of what happened.

For Efua, the turn comes when the agent recognizes her Twi and answers in Twi. She repeats both references, then explains which payment was valid and which one she disputes. The distinction survives the exchange.

That outcome depends on more than detecting isolated Twi words. The system needs speech recognition and synthesis built to handle natural spoken Twi, including transitions from English within the same conversation. It also needs to preserve what changed and when.

Asenda Talk is being built for that setting. Its Twi speech recognition and synthesis are fine-tuned in-house rather than passed through a generic third-party voice layer. Agents can be configured with a persona, first message, and voice, while Vapi orchestrates the assistant runtime.

The platform remains in active early access. More African languages are in progress, and feature parity with established voice-agent platforms has not been reached. Live outbound calling also remains behind an explicit telephony-provider decision and an operator-controlled real-money gate.

Those limits matter because a language demo and a dependable support call are different tests.

Test the transition, not the greeting

A useful bilingual voice test should force the language change to happen after the easy part. Begin in English. Introduce the consequential detail later in Twi. Then check whether the agent preserves the account reference, the requested action, and any refusal or correction.

Typed prompts cannot establish that spoken understanding works. Background sound, pronunciation, pauses, self-correction, and code-switching all affect what the system receives. Twi voice-agent testing needs spoken evidence, especially at the point where meaning could split.

Reviewers should also inspect the event trail. Which language did the caller use when consent was given? Did an opt-out occur after a language transition? Did the transcript preserve the original statement, or only the agent’s interpretation? Asenda Talk’s telephony lifecycle webhook pipeline tracks call events, with consent, opt-out, and audit records attached to each call.

Call truth includes the transition itself.

Design for the sentence that carries the risk

Before enabling a voice workflow, write test calls around the moment a caller has something to lose. Use two similar transaction references. Add a correction in Twi after an English opening. Include a clear refusal. Ask the agent to repeat the requested action before proceeding.

Then listen to the audio and compare it with the transcript and event sequence. A polished English welcome proves very little if the decisive Twi sentence disappears.

Efua ends the call with the disputed reference identified and the valid payment left alone. Her greeting never predicted that outcome. The system’s response at the hardest sentence did.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.