A dashboard showing 100,000 monthly queries proves activity, not conversational quality. It cannot tell a bank operations lead whether a caller can explain a frozen account naturally in Twi, switch into English for a technical term, and leave the call correctly understood.
In 1999, NASA’s Mars Climate Orbiter approached Mars with years of engineering work behind it. The spacecraft carried instruments, telemetry, and calculations that appeared precise. Yet one critical interface was wrong: one team supplied thruster data in pound-force seconds while another expected newton-seconds.
The numbers were present. Their meaning did not match.
NASA lost contact with the spacecraft near Mars. Arthur Stephenson led the investigation, and the failure was documented in NASA’s Mars Climate Orbiter Mishap Investigation Board Phase I Report. The report traced the loss to the failed translation of English units into metric units in a ground software file.
Every calculation downstream could look orderly while carrying the wrong meaning.
Volume measures traffic, not understanding
Now put a bank operations lead in front of a polished automation dashboard.
The chatbot handled 100,000 queries this month. The chart rises. The total looks substantial. Absa has reported that its chatbot fields 100,000 queries monthly, alongside the use of AI coding tools by 1,400 developers.
That figure tells us the system received substantial traffic. It does not reveal what happened inside each exchange.
Did the customer get the right answer? Did the system understand why an account was frozen? Did it recognize when a caller moved from English into Twi because the issue became personal, urgent, or difficult to explain? Did the conversation end with a clear next step, or did the customer repeat the problem and then give up?
A query count cannot answer those questions because it measures an event, not its meaning. The same problem appears in voice automation when teams treat connected calls, completed calls, and minutes consumed as evidence of successful service.
Those are useful operational measures. They still need outcome measures beside them.
A frozen account exposes the limits quickly
A customer calling about a frozen account may begin in English because that is how the bank describes the issue. Then the customer may switch into Twi to explain who sent the money, why the transfer matters, or what happened before access was restricted.
The language switch carries information. If the voice agent misses it, translates it poorly, or forces the caller back into English, the call can remain technically active while the actual problem goes unresolved.
This is why a customer’s switch between Twi and English during a dispute deserves evaluation as part of the task itself. Recognition accuracy matters, but the test must go further. The agent needs to preserve intent across both languages, respond in a form the caller understands, and avoid inventing policy or account details.
A useful evaluation set would include real task shapes such as:
- A caller explains the same frozen-account problem first in English, then in Twi.
- A caller uses English banking terms inside a Twi sentence.
- A caller repeats the issue after receiving an incomplete answer.
- A caller asks for a human because the explanation remains unclear.
- A caller withdraws consent or asks not to receive another automated call.
The review should score whether the agent understood the reason for the call, gave an allowed response, recorded the outcome correctly, and respected the caller’s choice. Call duration alone cannot provide that evidence.
Call truth needs more than a completion status
A “completed” status usually describes the telephony lifecycle. It may mean the connection opened and later closed without a transport failure. It does not prove that the customer understood the response, consented to the call, or achieved the intended outcome.
That distinction matters for banking, campaigns, collections, and support desks. An operations team needs a record connecting the phone event to the conversational event: what the caller asked, which language was used, whether the agent resolved or escalated the issue, and whether an opt-out occurred.
Asenda Talk is being built around that distinction. Its current early-access platform supports configurable voice agents, native Twi speech recognition and synthesis fine-tuned in-house, and a telephony webhook pipeline designed for call-truth tracking. Consent, opt-out, and audit records are part of the call model rather than notes added after a campaign.
The platform also uses metered per-minute billing with an operator-controlled real-money gate. Outbound calling remains gated while the live telephony-provider decision is unresolved. That limitation should stay visible in any evaluation. A credible pilot tests what exists today and labels what remains on the roadmap.
For compliance work, the same discipline applies: a completed call does not prove consent.
Build the dashboard around decisions
Keep the query total. Operations teams still need capacity, cost, and adoption measures. Then add the measures that determine whether the system is safe and useful.
Track task resolution by language. Review Twi-to-English switches separately. Measure escalation after misunderstanding, repeated caller turns, unsupported answers, opt-outs, and cases where the recorded outcome conflicts with the transcript. Sample calls involving account restrictions and have bilingual reviewers judge whether the explanation remained accurate and natural.
NASA’s Mars Climate Orbiter had numbers, software, and reporting. The missing protection was a reliable check that both sides meant the same thing. A bank evaluating voice AI faces a smaller-stakes version of that measurement problem every time a successful call count hides a failed conversation.
The next dashboard review should begin with one recorded exchange: a caller explaining a frozen account in Twi. Ask what the agent understood, what it answered, and what the audit trail can prove.
Comments
No comments yet.