Asenda Talk
A young man with glasses making a phone call indoors, exhibiting focus and communication.

Photo by David Awokoya on Pexels

The sentence before a caller switches from English to Twi often signals that the next detail carries more weight, discomfort or precision. A voice agent should treat hesitation, self-correction and phrases such as “what I mean is” as cues to preserve context and listen closely through the language change.

At 4:40 on a humid afternoon in Accra, Ama stood behind the counter of her small provisions shop, holding a supplier invoice with one corner darkened by spilled malt. In this illustrative scenario, an automated support call had begun in English. She answered the first questions quickly: her name, the delivery reference and the date printed on the paper.

Then her pace changed.

“Okay, the delivery came, but the thing is…”

She stopped. A customer set two tins of milk beside the till. Ama checked the invoice again and continued in Twi, explaining that the amount on the document did not match what had arrived.

The disputed charge was due to be reviewed that afternoon. If the system treated her Twi explanation as noise, a new topic or a failed turn, the wrong amount could stand. She might have to pay for goods she never received.

That unfinished English sentence was the warning.

The cue arrives before the language change

A switch into Twi can begin several words before the first Twi syllable. The caller may slow down, restart a phrase or use a short bridge that announces unfinished meaning:

“I’m trying to explain…”

“Yes, but what happened was…”

“The amount is correct, only that…”

These phrases do more than fill time. They open a thought that the caller expects to complete in the language that gives them better control over the detail.

The change may happen because the topic has become sensitive. It may involve money, consent, a family relationship or a correction the caller does not want softened by uncertain English. Sometimes the caller knows the English term but chooses Twi because the full explanation comes more naturally there.

A system that evaluates each utterance in isolation can miss this structure. It may classify the English fragment as complete, prepare a response and speak over the Twi continuation. It may also carry the wrong assumption into the next turn.

The important unit is the transition: the opening phrase, the pause and the Twi completion belong to one act of meaning.

Preserve the unfinished thought

A bilingual voice agent needs enough conversational memory to hold the English opening while processing what follows in Twi. That means delaying commitment when the sentence is visibly unfinished and keeping references intact across the switch.

In Ama’s call, “the thing” referred to the mismatch between an invoice and a delivery. Her Twi explanation supplied the quantities and the correction. If the agent dropped the English lead-in, it could hear the Twi segment as a fresh complaint without the transaction context. If it dropped the Twi segment, it could record only that the delivery arrived.

Both records would be incomplete.

This is why language-transition tests should examine more than transcription word by word. Reviewers should ask whether the agent preserved the caller’s intent, entities, negation and correction across the boundary. The related review on why one language transition should stop deployment shows how one failed handoff can expose a broader readiness problem.

The same standard applies when a caller corrects the agent mid-call. A missed correction can change which queue, account or action follows, as explored in what happens when an AI service misses a Twi correction.

Test the sentence, pause and switch together

Teams evaluating Twi and English voice agents should build test calls around complete conversational moments. A clean English prompt followed by a clean Twi answer is useful, but it avoids the difficult boundary where meaning is most likely to break.

A stronger test begins in English, introduces hesitation and finishes the consequential detail in Twi. Vary where the switch occurs. Put it after a conjunction, a correction, a pronoun or an incomplete amount. Then check the transcript, the agent’s spoken response and the downstream call record.

The expected result should be explicit. Did the agent wait? Did it connect both language segments? Did it ask for confirmation before an account change or other consequential action? Did the audit trail retain what the caller consented to, corrected or declined?

Asenda Talk is in active early access. It provides native Twi speech recognition and synthesis fine-tuned in-house, alongside configurable voice agents, Vapi-orchestrated runtime, call lifecycle tracking, consent records and opt-out handling. More African languages are in progress. Outbound calling remains behind an operator-controlled real-money gate while the live telephony-provider decision is unresolved.

That status matters because a convincing bilingual demo does not prove deployment readiness. The difficult evidence sits in interrupted sentences, mixed-language corrections and the records produced after the call.

Listen for the moment the stakes rise

Back at the counter, Ama repeats the disputed detail. This time the agent holds the unfinished English clause, processes the Twi explanation and asks her to confirm the corrected interpretation before the case moves forward.

She looks once more at the stained invoice and gives the confirmation in Twi. The call record now reflects the mismatch she described, rather than the simpler but wrong conclusion that the delivery arrived as billed.

When reviewing a bilingual call, start ten seconds before the switch. The most important evidence may be the sentence the caller could not yet finish.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.