Asenda Talk
Focused call center employee reviewing documents while on a call in an office setting.

Photo by Yan Krukau on Pexels

A voice agent review should score what the system correctly recognizes in Twi, English and code-switched speech, with errors recorded by language, utterance and consequence. A polished voice may improve the experience, but it cannot prove that the agent understood the caller.

In Berlin in 1904, psychologist Oskar Pfungst faced a similar evaluation problem. Clever Hans, a horse owned by Wilhelm von Osten, appeared able to answer arithmetic questions by tapping his hoof. The demonstrations looked convincing. Yet Pfungst found that Hans often succeeded only when the questioner knew the answer and remained visible.

The horse was responding to subtle human cues, not calculating. The performance looked right because the test allowed the wrong signal to produce the expected result. Pfungst’s investigation, later documented in his book Clever Hans (The Horse of Mr. von Osten), became a lasting warning about experiments in which observers accidentally influence what they believe they are measuring.

A polished call can hide a recognition failure

A Ghanaian team reviewing a voice agent can fall into the same trap.

The voice sounds natural. The first message is clear. The agent pauses in the right places and produces a confident response. Around a table in Accra, the first reaction may be that the call went well.

Then someone opens the transcript.

A Twi phrase was captured as an English word. A caller switched languages halfway through a sentence, and the system preserved the opening but lost the account detail at the end. The agent replied fluently because the language model found a plausible response, even though the speech recognizer had supplied the wrong meaning.

That is the voice-agent version of watching Clever Hans tap the expected number. The visible output feels persuasive. The underlying mechanism has not yet been proved.

This matters most when the misunderstood phrase changes an action. A wrong product name is inconvenient. A missed refusal, payment amount, appointment date or consent statement can create a false outcome in the call record.

Replace impressions with an utterance ledger

The review changes when the team stops asking, “Did that sound good?” and starts asking, “What did it recognize?”

Each test utterance should have a written expected meaning and a captured result. Record the language pattern too: Twi, English or code-switched. If the caller repeats the phrase, keep both attempts. Repetition is evidence about reliability, especially when the system makes the same error twice.

A useful review record can stay simple:

  • Write the exact phrase spoken, with the speaker’s intended meaning.
  • Save the transcript produced by the recognizer.
  • Record the action the agent took from that transcript.
  • Mark whether the error changed the call outcome.
  • Test the same meaning with different speakers, pacing and language-switch points.

The goal is not a single accuracy score that compresses every failure into one percentage. Teams need to know which speech patterns fail and what those failures cause.

A code-switched sentence that loses a courtesy phrase may still complete the caller’s task. A short Twi refusal recognized as agreement is a release blocker. Treating those errors as equal makes the evaluation easier to summarize and harder to trust.

The same discipline applies when an apparent match masks weak understanding. [Bilingual voice agent testing](\/blog\/bilingual-voice-agent-testing-why-an-account-match-cannot-prove-understanding-d9dd9e2a\/) should separate a correct final record from the reasoning that produced it.

Test the boundary of what exists today

Asenda Talk provides native Twi speech recognition and synthesis fine-tuned in-house, alongside configurable voice agents and Vapi-orchestrated assistant runtime. That gives teams a concrete system to evaluate. It does not remove the need to document where recognition works, where it fails and which language combinations remain unreliable.

The platform is in active early access. More African languages are in progress, and feature parity with established voice-agent platforms has not been reached. Outbound calling also remains behind an explicit telephony-provider decision and an operator-controlled real-money gate.

Those constraints belong in the review record. A test should distinguish speech capability from telephony availability, assistant behavior from transcription quality, and a simulated workflow from a completed live call. Call-truth tracking, consent records, opt-out handling and audit trails help preserve what occurred, but they cannot turn a misheard utterance into a correct one.

When a detail is misheard twice, the next step is a controlled test, not another polished demo. This [voice-agent testing guide](\/blog\/what-should-you-test-when-an-agent-mishears-the-same-account-detail-twice-fb44e44b\/) shows how to isolate the phrase, speaker variation and downstream action.

Make recognition evidence the release criterion

Before approving an agent for a pilot, define the utterances it must understand and the errors that block release. Include ordinary Twi, English and realistic switching between them. Give extra weight to consent, opt-out requests, names, amounts, dates and any phrase that changes what the system records or does next.

Then review failures in the transcript and call lifecycle record. A pleasant voice can remain part of the score, but only after recognition and outcome checks pass.

Pfungst changed the conditions around Clever Hans and discovered what the demonstration was truly measuring. A Ghanaian voice-agent team can do the same. Hide the polish for one review session, inspect the recognized words and document the resulting action. That record tells you what the agent can genuinely handle today.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.