Asenda Talk
Close-up of a man holding a credit card while speaking on the phone in Johannesburg office.

Photo by Phasha 360 on Pexels

A voice agent must carry the customer’s intent across a language switch, because the switch can reveal a more precise dispute than the original English words expressed. Translating each sentence separately can preserve vocabulary while losing the issue the customer needs resolved.

Consider Kojo, an invented composite customer, standing outside a pharmacy in Kumasi at 5:40 p.m. He is holding his bank card in one hand and a receipt in the other when he calls about a balance that looks wrong.

“I checked my balance. The money is not correct,” he says in English.

The sentence sounds like a general balance inquiry. An agent might respond with the available balance, list recent transactions, or ask Kojo to check again later. Then Kojo switches to Twi. He explains that he paid once, received an error message, tried again, and now believes the same purchase has reduced his available balance twice.

His concern has sharpened. He needs the bank to determine whether one payment was duplicated or whether a pending entry is temporarily affecting the balance. If the agent treats the Twi statement as a new request, it may lose the connection between the balance and the two payment attempts.

Kojo is worried he will leave the pharmacy without enough available funds to complete the purchase. The receipt shows only one successful payment, but the balance appears to tell another story.

A language switch can reveal the real job

Kojo’s English question opened the conversation. His Twi explanation defined the job: investigate a possible duplicate charge connected to two attempts at the same payment.

Customers often begin in the language they expect a system to understand, using short and cautious phrasing. They may switch when the first language cannot carry the detail, emotion, or sequence they need to explain. That change should update the agent’s understanding of the conversation.

A basic translation pipeline may turn Kojo’s Twi into acceptable English and still miss the point. The system needs to preserve the entities and relationships already established: the balance, the payment, the first error, the second attempt, and the fear of being charged twice.

This is also why a readable transcript is insufficient on its own. As the example in Ama’s client complaint shows, words on a screen can offer little help when the system fails to preserve what the caller meant.

Intent must survive across turns

The agent should carry one evolving representation of Kojo’s request through both languages. The English opening supplies the object under dispute, his account balance. The Twi explanation supplies the cause he suspects and the outcome he wants.

That continuity affects the next question. “Would you like to hear your balance again?” sends Kojo backward. “Did both payment attempts appear in your recent activity?” moves the dispute forward.

The difference comes from conversation state, not polished phrasing. A useful system should retain:

  • what the customer first asked about;
  • what changed after the language switch;
  • which transaction or event the customer connected to the problem;
  • what remains uncertain;
  • whether consent, opt-out, or escalation requirements affect the next action.

Language recognition also matters before intent tracking can work. Unsupported systems may corrupt Twi input or reduce it to unusable text, a failure explored in how Twi becomes garbage characters in unsupported systems. Once the words are damaged, preserving meaning becomes much harder.

Native Twi changes what can be evaluated

Asenda Talk is in active early access. It lets teams create voice agents, set a persona, choose a voice, and define the first message. Its Twi speech recognition and synthesis are fine-tuned in-house rather than passed through a third-party Twi voice wrapper.

That foundation makes a specific evaluation possible: can an agent hear the English opening, understand the more detailed Twi dispute, and continue with the correct combined intent?

Teams can test that scenario before treating it as production-ready. They can inspect whether the transcript remains usable, whether the assistant asks the right follow-up question, and whether the call record reflects what happened. Asenda Talk includes a telephony lifecycle webhook pipeline with call-truth tracking, plus consent, opt-out, and audit records for each call.

Vapi orchestrates the assistant runtime. Metered billing includes an operator-controlled real-money gate, and outbound calling remains gated while the live telephony-provider decision is unresolved. Those constraints matter. A banking workflow should be evaluated against what the platform supports today, with any unresolved production dependency stated plainly.

Test the switch where meaning changes

The most useful test scripts should include language switches that alter the apparent request. Begin with a broad English statement, then add the decisive Twi detail several turns later. Check whether the agent updates its understanding or starts a separate conversation in its internal state.

For Kojo, the turning point comes when the agent connects “my balance is wrong” to “I attempted the same payment twice after an error.” It asks about both attempts and keeps the possible duplicate charge at the centre of the call.

Kojo no longer has to restart in English or repeat the whole story. He is still outside the pharmacy, receipt folded against his card, but the conversation is finally about the problem he called to resolve.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.