Asenda Talk
A diverse group of co-workers engaged in communication and teamwork in a call center environment.

Photo by Pavel Danilyuk on Pexels

Typed-query success cannot prove that a voice agent understands a naturally spoken Twi payment problem. Text removes pronunciation, timing, interruptions, background noise, number repetition, and mid-sentence language changes, which are often the parts that determine whether a call succeeds.

At Monday’s review, the support manager faces a tempting number: the chatbot handled 100,000 typed queries. That figure may show scale and useful text performance. It cannot answer the question everyone in the room needs answered: when a customer explains a payment problem aloud in Twi, does the system understand the account reference, the correction, and the action the caller wants?

The test passed, but the parts did not fit

In April 1970, Apollo 13’s crew had moved into the lunar module after an oxygen tank exploded and forced NASA to abandon the Moon landing. Carbon dioxide was accumulating. The command module carried square lithium hydroxide canisters, while the lunar module used round ones.

The canisters worked. The filtration system worked. The components had already been designed, built, and tested for their intended environments. Yet the square canisters could not fit the round openings where the astronauts now needed them.

Engineers at Mission Control in Houston had to create an adapter from materials available aboard the spacecraft, including plastic bags, cardboard, a hose, and duct tape. The crew assembled the improvised device from ground instructions. Jim Lovell later documented the mission with Jeffrey Kluger in Lost Moon, the account adapted into the film Apollo 13.

The mechanism matters here. Passing a test inside one interface does not prove that the same capability will work inside another interface, under different constraints. The gap can remain hidden until the moment the parts must connect.

Typed Twi and spoken Twi are connected, but they are not interchangeable test environments.

A phone call contains evidence that text removes

A typed support message arrives already segmented into words. The sender has usually corrected some spelling, removed false starts, and chosen what to include. The system receives a relatively clean sequence.

A phone call arrives as sound. The caller may repeat an account number, pause after a merchant name, switch to English for a banking term, then return to Twi to explain what went wrong. A correction may be short and easy to miss. The customer may say the first number was wrong, then provide another while the agent is preparing its response.

A text benchmark can confirm that a model maps a written phrase to the right intent. A voice evaluation must also test whether speech recognition captured the phrase correctly, whether the language transition remained intact, and whether synthesis produced a response the caller could understand.

The transcript alone may still conceal the failure. A cleaned transcript can look correct after punctuation or normalization, while the runtime acted on an earlier interpretation. That is why a correct transcript can still trigger the wrong action.

Evaluate the call as a chain, not a message

A useful Twi voice evaluation begins before the model response and continues after the call ends.

Use recorded test cases with natural speech, including hesitation, corrections, numbers, names, English terms, and language switching. Check the audio against the raw recognition result. Then compare both with the agent’s interpreted intent and its chosen action.

Review the response as audio too. Correct text does little good if pronunciation changes a name, amount, or instruction. For bilingual calls, preserve where each transition happened and what the system understood at that point. Language transitions belong in the audit trail because they can explain a failure that a final transcript cannot.

The rest of the call lifecycle matters as well. Did the person consent? Did an opt-out stop further contact? Did the call connect, fail, or disconnect before the apparent resolution? Call-truth tracking should preserve those events instead of turning every initiated call into a successful conversation.

Asenda Talk is being built for this evaluation surface. Its Twi recognition and synthesis are fine-tuned in-house, rather than delegated entirely to a third-party voice layer. Voice agents can be configured with a persona, first message, and voice. The platform also records telephony lifecycle events, consent, opt-outs, and audit history.

It remains in active early access. Feature parity with established platforms such as Vapi, Retell AI, and Bland AI is still in progress. Vapi currently orchestrates the assistant runtime, and live outbound calling remains behind an operator-controlled real-money gate while the telephony-provider decision is unresolved.

Those constraints should shape the test plan. Evaluate the speech and runtime that exist today. Do not present an undialed campaign as evidence of live outbound performance.

Replace the impressive number with a harder result

The next Monday meeting needs a smaller number with stronger meaning: how many representative Twi and bilingual calls were reviewed end to end, and how many preserved the caller’s intended payment instruction?

Keep the 100,000 typed queries as evidence of text usage. Put it in the right column. Beside it, report recognition errors, misunderstood corrections, language-switch failures, incorrect spoken responses, opt-out handling, and verified call outcomes.

Apollo 13’s canisters were useful hardware. Their proven performance could not make a square cartridge fit a round opening. Voice evaluation demands the same discipline: test the capability where it must operate, with the interfaces, constraints, and failure modes that will be present on the call.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.