Asenda Talk
Candid black and white photo capturing a street vendor selling mobile accessories in Accra, Ghana.

Photo by Zeal Creative Studios on Pexels

A Twi verification agent must recognize agreement as it is spoken, including forms such as “aane” and shorter contextual affirmations, then distinguish them from hesitation or simple acknowledgment. Direct translation can miss that signal and turn a valid confirmation into a repeat question, an uncertain record, or the wrong next action.

In Berlin in 1907, psychologist Oskar Pfungst faced a related problem of interpretation. Clever Hans, a horse owned by Wilhelm von Osten, appeared able to answer questions by tapping a hoof. The striking part was that Hans often produced correct answers even when von Osten was not asking them.

The result looked convincing. The signal behind it was not what observers thought.

Clever Hans was reading a signal people could not see

Pfungst tested what changed when Hans answered correctly. As documented in his book Clever Hans (The Horse of Mr. von Osten), the horse’s performance depended on whether the questioner knew the answer and whether Hans could see that person.

Hans was not doing arithmetic. He was responding to small, involuntary changes in human posture and expression as his hoof taps approached the expected number. The humans supplied a stopping signal without realizing it.

That distinction matters because a system can appear to understand while tracking the wrong evidence. It may pass a tidy demonstration and still fail when the speaker, phrasing, or context changes.

A voice agent can make the same class of error in reverse. The caller provides a meaningful signal, but the system has been trained or evaluated around a narrower English pattern. The agent waits for “yes” or a clean translation equivalent, while the person has already agreed in natural Twi.

Agreement lives in the exchange, not a lookup table

“Aane” can express yes in Twi. Real calls, however, do not arrive as isolated vocabulary tests. Agreement may be shortened, repeated, blended with English, delivered after a pause, or expressed through a phrase whose force depends on the question immediately before it.

Consider a verification prompt that asks whether a caller approves a stated detail. A clear Twi affirmative should advance the call. A brief acknowledgment during a longer explanation may only mean “I am following.” Treating both as the same event creates risk in either direction.

If the agent misses agreement, it repeats itself. The caller may become unsure whether the first answer was heard, switch languages, or abandon the call. If the agent mistakes a backchannel for consent, the consequence is more serious: it can record approval the caller did not give.

This is why Twi support has to be evaluated at the turn level. The relevant test is not “Can the model translate this word?” It is “Given this prompt, this response, and this conversational position, what did the caller commit to?”

The same problem appears when English-only automation misroutes a Twi payment dispute. A plausible transcript can still carry the wrong operational meaning.

Verification requires language evidence and call truth

A useful verification test set should contain more than polished recordings of full affirmative sentences. It should include short answers, mixed Twi and English, repetitions, corrections, background noise, and responses that acknowledge the agent without granting agreement.

Each sample needs a label tied to the action the system should take:

  • Advance because agreement is clear.
  • Ask a neutral follow-up because the answer is ambiguous.
  • Do not advance because the caller declined or withdrew agreement.
  • Escalate when the transcript and conversational evidence conflict.

Those decisions should survive the whole call pipeline. Recognition output alone is insufficient. Teams need the prompt that preceded the answer, the agent’s interpretation, the state transition, and the resulting consent or verification record.

Asenda Talk’s native Twi speech recognition and synthesis are fine-tuned in-house. The platform also records telephony lifecycle events, consent, opt-outs, and an audit trail for each call. That creates the structure needed to examine what the agent heard and what it did next.

The boundaries matter. Asenda Talk remains in active early access. Vapi orchestrates the assistant runtime, more African languages are in progress, and outbound calling remains behind an operator-controlled real-money gate while the live telephony-provider decision is unresolved. These capabilities should be evaluated against real, consented call scenarios before anyone treats them as finished production behavior.

Test the smallest answer before scaling the call

Start with the moments where one short response changes a record: identity confirmation, appointment acceptance, payment-plan acknowledgment, survey consent, or permission to continue. Collect consented examples from the Twi-speaking audience the agent will actually serve, then review disagreements between the transcript, the inferred intent, and the recorded call state.

Test the first message in context too. The language and formality of the prompt influence the kind of answer a caller gives, as discussed in why the first message must be tested in context.

Pfungst’s work showed how easy it is to celebrate a correct response while misunderstanding the cue that produced it. For Twi verification calls, the practical safeguard is equally concrete: inspect the small affirmations, preserve their conversational context, and verify the action recorded after them.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.