Asenda Talk
A young man with glasses making a phone call indoors, exhibiting focus and communication.

Photo by David Awokoya on Pexels

A bilingual voice agent can preserve every word in a language switch and still lose the caller’s meaning. The failure often appears one turn later, when a pronoun such as “it” points back to the crucial event in Twi and the agent links it to an earlier problem stated in English.

In 1999, NASA’s Mars Climate Orbiter reached Mars carrying numbers that had survived a handoff between engineering teams. Their meaning had not. One team produced thruster data in pound-force seconds; another expected newton seconds. The values crossed the interface, but the units that made those values intelligible did not.

The spacecraft was lost. NASA’s Mars Climate Orbiter Mishap Investigation Board Phase I Report, led by Arthur Stephenson, documented the mismatch and the failed checks around it.

A pronoun crossing a language boundary creates the same class of risk on a smaller scale. The token arrives. Its reference does not.

The failure appears after the language switch

Consider a Monday support call used as an evaluation scenario.

The caller begins in English: “I called on Friday because the transfer was delayed.”

Then she switches to Twi to explain the crucial event. Perhaps the money later reached the wrong recipient, or the account showed a second problem after the first complaint. The exact wording matters because this Twi segment changes what the call is about.

She returns to English: “I need it reversed.”

What does “it” mean?

A voice agent may attach “it” to the delayed transfer mentioned before the switch. The caller may mean the later event described in Twi. A transcript can look complete while the agent’s working context has already split into two incompatible versions of the call.

This is more than a speech-recognition problem. Recognition asks whether the system captured the Twi words. Reference resolution asks whether the system understood which event those words introduced and carried that event forward when English resumed.

A bilingual agent needs both.

A correct transcript can still produce a wrong action

Teams often evaluate multilingual voice systems by reading the transcript and checking whether each sentence looks plausible. That test can miss the failure.

The dangerous sequence is:

  • The caller establishes one issue in English.
  • The caller corrects, narrows, or replaces it in Twi.
  • The caller uses a pronoun after returning to English.
  • The agent resolves that pronoun against the older English context.
  • The workflow records, routes, or acts on the wrong issue.

The error may surface as an incorrect case summary, a transfer to the wrong queue, or a request to reverse the wrong transaction. If the action changes an account, sends money, or records consent, plausible wording offers little protection.

This is why sensitive changes need explicit confirmation. The agent should name the object it believes the caller means: “You want the transfer to the second recipient reversed. Is that correct?” The confirmation must reflect the event introduced in Twi, rather than repeating the earlier English complaint.

The same principle applies when a caller withdraws permission or corrects a destination. A missed correction can remain hidden behind fluent dialogue, as discussed in What Happens When an AI Service Misses a Twi Correction Mid-Call?.

Test references across turns, not languages in isolation

A useful bilingual evaluation set should include more than isolated Twi prompts and clean English replies. It should test state across the switch.

Start with calls where the speaker introduces two possible referents. Put the decisive correction or event in Twi. Then return to English with “it,” “that one,” “there,” or “the other one.” Score the system on the referent it selects, the summary it produces, and the action it proposes.

Include interruptions and short answers. Real calls contain repairs such as “No, the second one” and “I mean the one from yesterday.” Those small phrases often carry more operational weight than a long opening explanation.

For high-impact actions, require the agent to restate the resolved reference before proceeding. If confidence is low, route the call for human review and preserve the relevant audio, transcript, language transitions, consent state, and proposed action in the audit trail.

Asenda Talk is being built for this problem with native Twi speech recognition and synthesis, fine-tuned in-house, plus call-truth tracking and an audit trail for consent and opt-out events. It remains in active early access. Native speech capability does not by itself prove reliable cross-language reference resolution, so this behavior must be evaluated directly before deployment.

Outbound calling is also gated while the telephony-provider decision remains unresolved. A working agent configuration does not mean live outbound operation is approved.

Keep meaning attached at every interface

The Mars Climate Orbiter investigation did not describe missing numbers. It described a failure to preserve what those numbers meant across an interface, followed by missed opportunities to detect the mismatch.

Bilingual call review needs the same discipline. Track entities and events independently of the language used to introduce them. Record when the caller corrects the active issue. Require confirmation before consequential actions. Treat one wrong pronoun reference as a deployment blocker for that workflow, because the next occurrence may reach a billing, consent, or account-change step.

Build the Monday test call before live traffic does. Put the crucial event in Twi, switch back to English, and ask the agent what “it” means.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.