A plausible transcript can still send a case down the wrong path when it drops the word that determines intent. In a Twi or Twi-English payment call, supervisors need the source audio, contextual review and an auditable correction path before acting on a transcript.
At 4:38 p.m. in Accra, Esi was reviewing the last cases before the support desk closed. She had cold tea beside her keyboard and a caller waiting for an answer. The transcript described a “payment” problem, so the case had been placed in the general payment queue.
The caller had meant a payment reversal.
That distinction decided what happened next. A payment question could wait for routine review. A reversal could require someone to check whether money had left the account, returned, or become stuck between states. If Esi accepted the transcript, the caller might finish the day without the right investigation underway.
This is an invented composite, but the failure pattern is concrete: speech recognition produces a sentence that looks reasonable, while losing the word that controls the workflow.
Plausible text can preserve the sentence and erase the case
Obvious transcription errors attract attention. A broken sentence, a string of unrelated words or an empty field tells a reviewer to listen again.
Plausible errors are harder. “Payment” fits the conversation. It belongs in the vocabulary of a support desk, and nothing about it looks damaged. The transcript can pass a quick visual check while carrying less meaning than the caller supplied.
In Esi’s review, the surrounding turns created the first warning. The caller had shifted between Twi and English while explaining that the money appeared to move and then come back. The voice agent’s summary and the case category pointed toward a generic payment issue, but the sequence of the conversation suggested a reversal.
This is why Twi and English code-switching must preserve context. Meaning may sit across languages, several turns and the caller’s correction of an earlier phrase. A clean transcript line cannot carry that full burden by itself.
The source record must remain available
Esi paused the case instead of approving its queue assignment. She opened the call record, checked the relevant audio segment and compared it with the transcript and lifecycle events.
For one uncomfortable minute, the outcome remained uncertain. If the recording did not support the caller’s intended meaning, Esi would have to choose between delaying the case and acting on a category she no longer trusted. Either choice could leave the caller without a defensible answer.
The audio preserved the missing distinction. Esi corrected the case category, recorded why she changed it and routed the issue for reversal review before leaving her desk.
That turn depends on keeping more than an AI-generated summary. A useful call record connects the source audio, transcript, call state, consent status, opt-out events and later human actions. When those records disagree, the disagreement should remain visible. The audit trail should show what changed, who changed it and why.
The same principle applies beyond payment disputes. A transcript may say “call” when the caller said “do not call,” or preserve the name of a service while losing the phrase that withdraws consent. As discussed in why source records, rather than AI summaries, are proof, generated text is an interpretation of the interaction. The underlying record is what lets a supervisor test that interpretation.
Native Twi speech still requires operational checks
Asenda Talk uses native Twi speech recognition and synthesis fine-tuned in-house. It does not pass Twi through a generic third-party voice wrapper and assume the result is adequate. That technical choice gives the team direct responsibility for evaluating how Twi is heard, rendered and improved.
It does not make every transcript correct.
Evaluation must include the words that change actions, especially negation, reversals, consent, amounts, dates and account states. Accuracy measured across whole sentences can hide a failure on one operationally important term. A test set should therefore ask a harder question: did the system preserve the distinction required to handle the call correctly?
Review rules matter too. Low confidence can trigger human review, but confidence alone cannot detect every plausible mistake. Teams should also flag contradictions between the transcript, summary, selected disposition and later call turns. A caller describing money returning should prompt another look when the assigned category says only “payment.”
Asenda Talk is in active early access. Teams can create voice agents, set their persona, first message and voice, and use a Vapi-orchestrated assistant runtime. The platform also has telephony lifecycle webhooks, call-truth tracking, metered billing controls, consent records, opt-out handling and audit trails. More African languages are in progress.
Live outbound calling remains gated until an explicit telephony-provider decision is made. That boundary matters because a responsible launch requires both language performance and a settled calling path.
Build the escalation rule before the first live call
Before deploying a payment workflow, list the words and distinctions that change where a case goes. Add examples in Twi, English and natural code-switched speech. Test them in full conversations, including corrections, interruptions and references to earlier turns.
Then decide what happens when the evidence conflicts. A supervisor should be able to stop automated handling, inspect the source record, correct the disposition and leave an audit entry. The caller should not bear the cost of a transcript that sounded convincing.
Esi’s final screen showed a reversal case, the supporting call segment and her correction in the audit history. The original transcript remained visible. On the next shift, another supervisor would see both what the system heard and why a person changed the decision.
Comments
No comments yet.