A useful Twi voice-agent benchmark should measure how well a system recognizes Twi, speaks it back clearly, and handles Twi-English switching during one conversation. A language count on a product page cannot show whether a caller will be understood at the moment their request matters.
At 8:42 on a Wednesday morning, Kojo, an illustrative composite of a small Accra support lead, was replaying a call through his laptop speakers while a paper cup of tea cooled beside him. The caller had begun in English, then moved into Twi when explaining a delivery problem. The agent caught the greeting, then returned a confident English reply that missed the customer’s actual question.
Kojo had already listened twice. A callback was due before the customer gave up and took the complaint to a human queue. The bad outcome was not a poor benchmark score. It was a customer being sent in circles because the system treated a mixed-language conversation as a sequence of separate language demos.
A credible benchmark would make that failure visible.
Test recognition with the speech people actually use
Recognition testing should begin with audio that reflects the calls an agent is meant to handle. That includes different speakers, recording conditions, speaking speeds, and everyday phrasing. It should also include the moments that cause trouble in real conversations: a caller correcting themselves, a name or location said once, background noise, and a sentence that changes language halfway through.
Word error rate can be useful, but it is incomplete on its own. A voice agent needs to preserve the meaning of the caller’s request. If it transcribes a key word incorrectly, or drops the English product name inside a Twi sentence, the downstream assistant may choose the wrong action even when most of the transcript looks close.
Benchmarks should therefore publish task-level checks alongside transcription results. Can the system identify the requested callback? Can it capture a reference number? Can it distinguish a confirmation from a refusal? Can it retain the caller’s language choice when the conversation turns?
This matters because Twi speech support is often presented as a capability label. CognariAI advertises Twi speech and voice-agent APIs. That tells a buyer where to begin their evaluation. It does not answer how a particular implementation performs on the calls that buyer needs to make or receive.
For an agent built for support, campaigns, or service follow-up, the test set should look like the work. A polished phrase spoken slowly into a quiet microphone belongs in a demo. It should not carry a benchmark.
Synthesis needs listening tests, not a checkbox
Text-to-speech evaluation needs more than confirmation that a system can generate audio in Twi. The key question is whether a caller can follow the response without asking for it again.
A practical benchmark should test pronunciation, intelligibility, pacing, and how the voice handles names, numbers, English terms, and short Twi responses. It should also test turns that a voice agent uses frequently: greetings, confirmations, consent language, opt-out requests, and recovery after the caller says they were misunderstood.
Human listening panels are important here. Automated measures can catch some differences, but they cannot fully judge whether a voice sounds clear in the context of a real conversation. The panel should use defined prompts and a repeatable scoring method, then separate observed results from subjective preference.
Kojo’s team would not need an abstract winner. They would need to know whether a caller who starts in Twi can hear a response, understand the next step, and continue naturally. If the agent asks for a confirmation, the caller should not have to decode a sentence that sounded acceptable in a studio test but failed over a phone connection.
Asenda Talk is being built around native Twi recognition and synthesis fine-tuned in-house, rather than a wrapper around a third-party voice API. That claim should create a higher standard for evaluation: publish the conditions, run the calls, and show where the system succeeds and where it still needs work.
Code-switching is where the conversation becomes real
Mixed Twi-English speech deserves its own benchmark category. Code-switching can happen within a sentence, especially when callers use English for a product name, date, address, number, or technical term. Treating those turns as an edge case makes the benchmark less useful for the actual call.
The benchmark should test at least three things. First, whether recognition preserves both languages without forcing the caller to repeat themselves. Second, whether the assistant keeps the right context when it responds. Third, whether synthesis can return a natural mixed-language response when that is the appropriate conversational choice.
The goal is not to prescribe one language pattern for every caller. A good agent should follow the conversation it receives, while preserving clear consent and opt-out handling. When a caller says they do not want further contact, that request must block future call attempts across the relevant routes. This practical look at Twi stop requests explains why language quality and call controls cannot be evaluated as separate concerns.
By the end of Kojo’s replay, the useful question had changed. He was no longer asking whether the platform listed Twi. He was writing down the exact call types his team needed to test before they trusted an agent with a live customer conversation.
Publish the test conditions alongside the result
A first benchmark does not need to pretend it is final. It needs to be reproducible enough for buyers, builders, and language researchers to challenge and improve it.
Publish the scenario types, language mix, scoring definitions, recording conditions, and known limitations. Separate laboratory-style audio from telephony audio. Identify whether results cover recognition, synthesis, or the full assistant loop. If a platform is still in early access, say so plainly.
Asenda Talk currently provides agent configuration, native Twi speech work, lifecycle webhook tracking, consent and audit records, metered billing controls, and Vapi-orchestrated assistant runtime. Production outbound calling remains gated behind a telephony-provider decision that has not been made live. The current constraint is documented here.
That boundary should shape the benchmark. Test what exists today. Mark what remains roadmap. Then let Kojo’s next call be the one that decides whether the result holds up.
Comments
No comments yet.