A Twi-English voter-registration agent is ready for real callers only when it handles meaning, consent, language switching, opt-outs, and call outcomes under realistic testing. A 72-hour deadline can support that decision, but it cannot turn an untested outbound channel into a safe launch.
At 8:40 on Monday morning in Accra, Kwame had a headset pressed to one ear and a voter-information script marked with red pen. He was a composite field lead, the kind who knew which English phrases sounded tidy in a planning meeting and confusing over a noisy phone line.
By Thursday, he had to recommend one of three paths: proceed, restrict the agent to further internal evaluation, or stop. The worst outcome was larger than an awkward call. A caller could misunderstand the purpose, miss an opt-out, or appear in a dashboard as successfully contacted when no useful conversation had taken place.
Outbound telephony was still gated because the provider decision had not been made live. Kwame could evaluate the agent, but he could not honestly approve a real outbound campaign.
The first day tests language under pressure
Kwame began with a short test set covering the moments that carried the most risk. The agent had to introduce itself, explain the purpose of the call, obtain consent, respond to questions, and honour a request to stop.
He configured the persona, first message, and voice in Asenda Talk. The assistant runtime used Vapi for call orchestration, while Twi recognition and synthesis came from speech models fine-tuned in-house. That distinction mattered during evaluation. A fluent call flow could still hide weak language understanding.
At 11:15, a tester answered in English, switched to Twi halfway through a date, paused, and corrected herself. The agent captured the final answer but responded as though the first version still applied.
The record looked usable. The conversation was not.
Kwame replayed the exchange and marked it as a comprehension failure. An account match, extracted date, or completed field could not prove that the agent followed the caller’s correction. That is the core warning in Bilingual Voice Agent Testing: Why an Account Match Cannot Prove Understanding.
By the end of Monday, he had a better test plan. Each case needed a transcript review, an expected response, and a clear pass or fail rule. “Sounds natural” was removed from the approval language.
The second day follows the complete call record
Tuesday began with the operational question: what happened to each call from initiation to final status?
Asenda Talk includes a telephony lifecycle webhook pipeline with call-truth tracking. Kwame used those records to separate attempted calls, connected calls, conversations, failures, and completed outcomes. A green status could not carry the whole decision because transport events and human outcomes answer different questions.
One test ended after the opening sentence. Another reached the consent prompt, then the tester asked in Twi not to be called again. Kwame checked whether the opt-out appeared in the audit trail and whether the record showed where the request occurred.
This was the point where the launch could still fail. If the opt-out existed only in a transcript, a later workflow might call the same person again. If the lifecycle record marked the interaction complete without preserving the caller’s instruction, the campaign team could mistake technical completion for permission.
The audit trail captured the test event, giving Kwame a record he could inspect. He still refused to generalise from one clean result. He added variations with code-switching, interruptions, repeated requests, and early hang-ups. What Must a Call Review Prove After a Caller Says “Stop Calling”? offers a useful review standard for this stage.
The third day turns evidence into a bounded decision
By Wednesday afternoon, Kwame had evidence across four areas: language comprehension, consent handling, opt-out recording, and lifecycle accuracy. He also had unresolved failures.
The right readiness decision was conditional.
The team could continue internal and controlled evaluation of the configured Twi-English agent. They could refine the first message, expand difficult code-switching cases, review recognition and synthesis behaviour, and verify that billing records matched test-call duration. Metered per-minute billing remained behind an operator-controlled real-money gate, which prevented testing from quietly becoming uncontrolled spend.
They could not begin live outbound calling. The telephony-provider decision remained open, so the outbound path had no live approval. No deadline, dashboard, or successful test call could supply that missing operational decision.
Kwame wrote the recommendation in two columns: proven now and blocked now. Native Twi speech, agent configuration, audit records, lifecycle tracking, and gated billing belonged in the first. Live outbound campaign readiness belonged in the second.
At 8:35 on Thursday, he closed the laptop before the campaign meeting. The team did not receive the simple yes it had hoped for. It received something safer and more useful: permission to continue controlled evaluation, a list of failed language cases to fix, and a clear block on real outbound calls until the provider path was decided and validated.
That is what a credible 72-hour decision looks like. It names the evidence, preserves the doubt, and refuses to let the calendar approve what the system has not yet proved.
Comments
No comments yet.