Asenda Talk
Two call center employees working together, wearing headsets, in an office setting.

Photo by MART PRODUCTION on Pexels

A mid-sentence language switch can preserve every spoken word while changing the intent the system assigns to it. For a bilingual voice agent, transcript accuracy therefore cannot prove that the caller’s meaning survived the switch.

Consider an illustrative composite: Abena, a support lead in Accra who keeps one earbud in while reviewing difficult calls. At 4:40 p.m., she replays a recording in which a caller begins an account request in English, shifts into Twi for one sentence, then returns to English. The transcript looks clean. Every word appears where she expects it.

But the agent’s proposed action feels wrong.

Abena plays the switch again. The Twi sentence changes the force of the request. What sounded like permission before the switch becomes a correction, a limit, or a withdrawal after it. If the system follows the surrounding English and treats the Twi phrase as supporting detail, it could make an account change the caller was trying to stop.

The transcript passes. The action should not.

A correct transcript can still carry the wrong instruction

Speech recognition answers one question: what words were spoken? A working bilingual agent must answer another: what did the speaker mean across the full turn?

That distinction becomes critical when a caller switches languages to clarify the part that matters most. A person may begin in English because it matches the support script, then move into Twi to express hesitation, urgency, correction, or refusal more naturally. The switch itself can be evidence. It may mark the point where the caller stops following the agent’s framing and states the real instruction.

A system can recognize each word correctly while attaching the Twi sentence to the wrong clause. It can preserve vocabulary and lose scope. A negative phrase may apply to the whole request rather than the nearest verb. A correction may replace the earlier instruction rather than add another detail.

This is why reviewing bilingual calls sentence by sentence is insufficient. Meaning can sit in the relationship between sentences, especially at the transition. A correct transcript can still trigger the wrong action when the runtime treats transcription as proof of intent.

Treat the switch as a decision point

Abena pauses the review before approving the agent configuration. The campaign queue is waiting, and the team wants to move on. If she marks the call as understood, the same interpretation could reach the next caller without anyone listening closely.

She leaves it blocked.

Her next test does not ask whether the model heard the Twi. It asks what changed when the caller entered Twi, which earlier statement the new sentence modified, and what action remained permitted afterward. She also compares the answer against the recording, the transcript, and the agent’s proposed action. Those three records need to agree.

For higher-impact requests, the safe response may be confirmation in plain language: “You want us to leave the account unchanged. Is that correct?” That short question gives the caller a chance to repair the interpretation before the system acts.

The same principle applies to consent and opt-out language. A caller who withdraws permission in Twi has still withdrawn permission. The call record should preserve that event, and later automation should respect it. The risk is examined more closely in The Twi Withdrawal the Transcript Could Miss, and the Calls It Could Allow.

Build evaluation around meaning, action, and evidence

A useful bilingual test set should include complete conversational turns, not isolated audio clips. Each case should record the expected meaning and the allowed action alongside the expected words.

Include switches that carry different conversational jobs:

  • The Twi phrase cancels an instruction given in English.
  • The caller narrows permission after the switch.
  • The speaker corrects a name, amount, date, or account detail.
  • The caller expresses uncertainty that requires confirmation.
  • Consent is withdrawn before the call ends.

Then inspect the entire path. Did speech recognition capture the utterance? Did the agent interpret its relationship to the earlier English? Did the runtime choose a permitted response? Did the call record retain enough evidence for a reviewer to reconstruct what happened?

Asenda Talk has native Twi speech recognition and synthesis fine-tuned in-house, with Vapi orchestrating the assistant runtime. It also has consent, opt-out, audit-trail, and call-truth infrastructure. Those components make this form of evaluation possible, but they do not replace it. The platform remains in active early access, and live outbound calling remains gated until an explicit telephony-provider decision is made.

Keep the consequential action behind the review

At 5:05 p.m., Abena adds the language transition to the blocked evaluation set. The agent can continue through low-risk test conversations, but it cannot treat this pattern as resolved or perform the disputed account action.

The next morning, the team has a concrete acceptance test: preserve the words, explain the correction, confirm the remaining permission, and record the decision. Until the agent can do all four, the switch stays red.

That red mark is more useful than a polished transcript. It shows exactly where language changed, meaning moved, and automation needed to stop.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.