Asenda Talk
Customer support agent with headset working on laptop in office.

Photo by MART PRODUCTION on Pexels

A convincing greeting cannot prove that a voice agent can hold a bilingual conversation. The real test begins when the caller answers in Twi, shifts into English, or expects the system to preserve meaning across both.

At 9:20 on a Monday morning in Accra, Ama sat in a campaign office with one earbud in and a paper call script beside her keyboard. She was a composite supervisor in a pre-launch test, responsible for deciding whether a voice agent was ready to speak with callers.

The opening sounded natural. The pace was familiar, the Twi pronunciation held, and the first question invited a simple response.

Then the caller asked for clarification.

The agent switched to English and delivered a sentence that was grammatically correct but stiff, detached from the caller’s wording, and out of step with the confidence created by the greeting. Ama stopped writing. If the caller misunderstood the next question, they could give consent without understanding what followed, or hang up before completing the conversation.

The campaign was still paused. It had to stay that way.

A natural opening sets a higher expectation

The first seconds of a call establish more than pronunciation. They tell the caller what kind of conversation is possible.

When the agent opens naturally in Twi, the caller may reasonably expect to continue in Twi, answer partly in English, or move between both. A strong greeting therefore raises the standard for every sentence that follows. If the system falls back to rigid English at the first unexpected reply, the opening becomes a promise the rest of the call cannot keep.

Ama replayed the transition. The failure was easy to miss if she judged each line separately. The Twi greeting sounded right. The English sentence was understandable. Yet the exchange between them felt wrong because the agent had stopped responding to the person and started reciting the script.

That distinction matters in support, campaign, and service calls. A caller may hesitate, correct a reference, ask what a term means, or change language while explaining something sensitive. Testing only the prepared opening reveals little about how the agent handles those moments.

Typed prompts reveal even less. Spoken language includes pace, hesitation, accent, code-switching, and incomplete sentences. Twi voice agent testing requires spoken understanding, because a clean text exchange removes much of the difficulty the caller will introduce.

Test the transition, not only each language

Asenda Talk provides native Twi speech recognition and synthesis fine-tuned in-house, alongside English conversation through a Vapi-orchestrated assistant runtime. That foundation makes natural Twi interaction possible. It does not remove the need to evaluate how the agent moves between languages, preserves context, and recovers when a reply does not match the expected path.

Ama changed the test. Instead of asking another reviewer to repeat the script, she asked him to interrupt after the greeting, answer the question indirectly in Twi, then request an English explanation.

The second run exposed the same break. The agent recognized enough of the reply to continue, but its English response failed to carry forward the caller’s reason for hesitating. The next scripted question assumed agreement that had not been clearly established.

Now the bad ending was visible: a live campaign could record progress through the script while the caller remained uncertain about what they had accepted.

Ama marked the transition as a failed test rather than smoothing it over as an awkward phrase. She also checked whether the call events would preserve enough evidence for review. Language changes belong beside consent and opt-out events in the audit trail, especially when the meaning of a response depends on what was said immediately before the switch. Bilingual call review needs those transitions recorded.

The transcript needs the call’s truth

A transcript can look orderly after a disorderly conversation. That makes operational evidence essential.

Asenda Talk’s telephony lifecycle webhook pipeline is designed to track what happened across the call, while consent, opt-out, and audit records provide evidence for later review. Admin secrets are write-only, masked, and environment-aware. Metered billing also sits behind an operator-controlled real-money gate.

Those controls matter, but they cannot turn an unproven bilingual path into a safe live campaign. They help a team see and govern the system it has. The team still has to test the conversation itself.

Outbound calling remains gated while the live telephony-provider decision is unresolved. Asenda Talk is in active early access and is still reaching feature parity with established voice-agent platforms. A configured agent, a working greeting, and an evaluated runtime should therefore be treated as test evidence, not permission to dial.

Keep the call paused until meaning survives

Near midday, Ama ran the exchange again with a revised persona instruction and a clearer recovery path. The greeting stayed in Twi. When the reviewer hesitated and requested English, the agent acknowledged the uncertainty before continuing. It did not treat the language switch as consent.

Ama removed one mark from her failure column, but she did not approve the campaign. One successful replay could show that the revised path worked once. It could not establish how the agent would handle other accents, interruptions, corrections, or mixed-language answers.

So she left the call list undialed and added three more spoken scenarios to the review: a caller who changes language mid-sentence, one who asks the agent to repeat the purpose of the call, and one who opts out immediately after switching to English.

The greeting still mattered. By then, Ama knew exactly where to listen next.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.