A bilingual voice agent must preserve intent across the language switch, especially when a caller begins a correction in English and finishes it in Twi. Accurate transcription alone cannot prove that the agent understood which charge the caller disputed or what action should follow.
At 4:52 on a humid Friday afternoon in Accra, Kojo is standing outside a pharmacy with a paper receipt folded into his shirt pocket. He calls about two charges on his account. One is valid. The other could push his balance into arrears if it remains unresolved before the next billing run.
“I paid the first amount,” he tells the voice agent in English. “The one I’m disputing is…”
He pauses as a trotro pulls away nearby, then completes the correction in Twi. The crucial words identify the second charge and reject the first account reference the agent repeats back.
The call now has one job: carry Kojo’s correction across the language boundary without breaking its meaning.
One sentence can contain two languages and one intent
A system may transcribe the English and Twi portions correctly as separate pieces while still connecting them incorrectly. It might attach Kojo’s Twi correction to the first charge, treat the language switch as a new request, or retain the reference it heard before he corrected it.
Each fragment can look plausible in a transcript. The resulting action can still be wrong.
This is why bilingual evaluation needs to follow the full intent, from the opening clause through the final correction. Reviewers should ask what Kojo was trying to change, which object the correction referred to, and whether the agent’s next response reflected that change.
The risk becomes clear when the agent says, “I have recorded your dispute against the first charge.”
Kojo goes quiet. The valid payment could now be challenged while the disputed charge remains untouched. His correction has been heard as words and lost as meaning.
That doubt needs to remain visible in the test. A tidy transcript should never erase the possibility of a wrong billing action.
Test the handover point, not only each language
A useful test case marks the exact moment language changes. For Kojo’s call, that means preserving the words immediately before the switch, the Twi continuation, the account reference under discussion, and the agent’s interpretation afterward.
The evaluation should check four things:
- Does the system recognize that the Twi phrase completes the unfinished English sentence?
- Does the correction remain attached to the second charge?
- Does the agent replace the earlier reference instead of storing both as equally valid?
- Does the confirmation state the corrected intent clearly enough for Kojo to accept or reject it?
This is also where call records matter. The language transition, corrected reference, agent response, consent state, and resulting call event should remain available for review. Bilingual call review explains why those transitions belong in the audit trail rather than being flattened into a single transcript.
A stronger evaluation also includes interruptions, hesitation, background noise, and a correction delivered after the agent begins responding. These details test the conversation Kojo is likely to have, not a clean recital written for a demo.
The safest turn is a specific confirmation
With the billing action still in doubt, the agent takes a safer path. It restates the corrected reference and asks Kojo to confirm that the dispute applies to the second charge while the first payment remains valid.
Kojo answers in Twi. The agent confirms the distinction before any downstream action proceeds.
That turn does not depend on the system sounding polished. It depends on reference tracking, cross-language intent continuity, and a confirmation designed around the consequence of getting the correction wrong.
Asenda Talk is being built for this class of Twi and English conversation. Its Twi speech recognition and synthesis are fine-tuned in-house, and its voice agents can be configured with a persona, first message, and voice. The platform also records telephony lifecycle events, consent, opt-out status, and audit history.
It remains in active early access. Feature parity with established voice-agent platforms is still in progress, and live outbound calling remains gated until an explicit telephony-provider decision is approved. Vapi currently orchestrates the assistant runtime. Those boundaries matter when deciding what can be evaluated today and what should remain paused.
Build the evaluation around the possible harm
Start with a transcript pair that looks correct at the word level. Then test whether the agent preserves the disputed object, rejects the superseded reference, and confirms the intended action in language the caller understands.
Add a failure condition that operations teams can enforce. If the agent cannot resolve which charge Kojo means, it should avoid committing the billing action and route the uncertainty for review. The same principle applies when a correct transcript produces the wrong instruction, as explored in why bilingual voice agents can still trigger the wrong action.
Kojo ends the call only after hearing the valid payment separated from the disputed charge. He unfolds the receipt outside the pharmacy, checks the two references once more, and puts it back in his pocket. The final test record should preserve that distinction just as clearly.
Comments
No comments yet.