Asenda Talk
Professional man multitasking, talking on the phone while using a laptop in an office.

Photo by Vitaly Gariev on Pexels

When an applicant answers partly in Twi and partly in English, the workflow should retain the original audio, transcript each language as spoken, and mark uncertain meaning for human review. It should never turn an ambiguous phrase into a confirmed administrative fact.

Consider an illustrative composite. At 4:40 p.m. in a small Accra office, Efua is reviewing an application call with a paper cup of tea cooling beside her keyboard. The applicant answers a question about household income in English, pauses, then qualifies the answer in Twi.

The system records a clean English value in the application. The Twi qualification is missing. If Efua approves the record, the applicant could be assessed using a statement they never made.

Preserve the answer before interpreting it

The first responsibility of a voice workflow is preservation. Keep the call audio, speaker turns, timestamps, and language changes together. If the applicant says a number in English and adds a condition in Twi, both parts belong to one response.

Translation can help a reviewer understand the answer, but it should remain a derived layer. The source recording and original-language transcript need to stay available. A reviewer should be able to distinguish what the applicant said from what the system inferred.

This matters because code-switching often carries meaning. A speaker may use English for a formal term, then switch to Twi to explain an exception, correct the question, or express uncertainty. Removing that second part can reverse the practical meaning of the response.

The same problem appears when a system marks a bilingual call complete while unresolved language remains inside it. What Really Happened When the Bilingual Call Was Marked Complete? examines why a completion state needs evidence behind it.

Keep uncertainty out of the facts table

A speech model can produce several kinds of uncertainty. It may be unsure which words were spoken, which language a phrase belongs to, or how the phrase relates to the question. Those are separate problems, and the record should expose them separately.

Suppose the applicant appears to say, “Yes, but only when…” before continuing in Twi. A workflow should not save “yes” as the final answer while treating everything after it as incidental. The safe record might contain the original response, a proposed interpretation, an uncertainty marker, and a review state. Only a verified interpretation should populate a field used for eligibility or another consequential decision.

This separation protects both the applicant and the reviewer. It also creates a usable audit trail. Later, someone can see the audio segment, the language detected, the proposed transcript, any correction, and who approved the final field.

A green checkmark without that history proves little.

Asenda Talk is being built around native Twi speech recognition and synthesis fine-tuned in-house, with Vapi orchestrating the assistant runtime. Its telephony lifecycle pipeline tracks what happened during a call, while consent, opt-out, and audit records provide evidence for review. The platform is in active early access. Native language work and call-truth tracking are built today, while feature parity with established voice-agent platforms remains in progress.

Put consequential decisions behind review

Governments should not rely on AI to decide who receives a social grant. The same principle applies to any workflow where a transcript may affect access to money, services, employment, credit, or support.

The voice agent can collect an answer. Speech recognition can propose words. A translation layer can offer an interpretation. None of those steps should silently become the decision.

Design the boundary explicitly. If a response switches language near a number, date, denial, correction, consent statement, or eligibility condition, route it for review. Show the reviewer the relevant audio rather than forcing them to search through the full call. Require a reason when they replace the proposed interpretation. Preserve both versions.

Testing should also include deliberate language switches at the hardest point in the call. A polished greeting proves little if the workflow loses meaning when the applicant becomes uncertain. Twi Voice Agent Testing: What a Caller’s Switch to English Reveals offers a related way to test those transitions.

Make the unresolved state visible

Efua returns to the original audio. The Twi phrase changes the answer from a fixed amount to an amount that depends on irregular work. Because the workflow retained the source response and flagged the mismatch, she leaves the administrative field unconfirmed and sends it for review.

The application is no longer resting on a false certainty. On Efua’s screen, the transcript still shows both languages, the proposed interpretation sits beside the recording, and the consequential field remains visibly unresolved.

That is the safer standard for a bilingual voice workflow: preserve first, interpret second, verify before a machine-generated answer becomes a fact.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.