Asenda Talk

The missed appointment wasn’t a scheduling failure. It was a language failure. The caller spoke Twi. The voice agent only understood English.

Here is how it happened. A clinic in Kumasi ran an automated reminder system to confirm antenatal appointments. The system was built on an English-centric voice agent, the kind that works well in call centres in London or Texas. The patient, a pregnant woman who had traveled in from a village outside the city, answered the call. She needed to confirm she was coming in for her scan. She spoke naturally, in the language she thinks in. The agent couldn’t parse it. It asked her to repeat herself. It asked her to press a number. It asked her to speak in English. She didn’t have that option, because her English, practical but hesitant, failed her under pressure. The call ended with no confirmation.

The cost of a missed confirmation

She arrived at the clinic the next day, but the slot had been reallocated because the system marked her as a no-show. She waited hours for a rescheduling that pushed her scan back two weeks. For a routine check, that is an inconvenience. For a high-risk pregnancy, those two weeks are the difference between catching a problem early and catching it too late.

That is the real cost of an English-only voice agent in a Twi-speaking country. The technology didn’t fail loudly. It failed quietly, politely, with a pleasant automated voice asking the caller to try again. The system’s dashboard showed no error. The call was logged as “completed.” The patient was simply, invisibly excluded.

Why the agent couldn’t understand

This isn’t about the agent having a small vocabulary. It’s about the difference between hearing words and understanding intent. An English-centric agent, or a wrapper that bolts Twi onto a third-party speech engine, recognizes a narrow band of expected phrases. It is built for the call it predicts: a confident speaker who says “yes” or “no” cleanly into the phone.

Real callers don’t do that. They say, “Me prɛ sɛ mekɔ hɔ ɛkyena” (I want to go there tomorrow), or they explain that their transport failed, or they ask a question the script never anticipated. The agent has no model for that. It hears phonemes that don’t match its training data and defaults to the only thing it has: the English fallback. The patient isn’t being stubborn. She’s being abandoned by a system that wasn’t built for her.

There’s a documented precedent for what happens when a system’s assumptions don’t match reality. In 1999, NASA lost the Mars Climate Orbiter because one team used metric units and another used imperial. The spacecraft flew too close to Mars and burned up in the atmosphere. The engineers didn’t notice the error because each half of the system was internally consistent. The mismatch only appeared at the boundary, where the two systems had to talk to each other. The orbiter cost $125 million. It was destroyed by a unit conversion, a detail everyone assumed someone else had handled.

This is the same failure mode. The English agent is internally consistent. The Twi speaker is internally consistent. The boundary between them is where the signal dies. As the NASA postmortem showed, the failure isn’t in either system. It’s in the assumption that they can communicate.

What a Twi-native agent does differently

The fix isn’t a bigger English model. It’s a model that actually processes Twi as Twi, not as broken English. At Asenda Talk, we built our speech recognition and synthesis in-house, fine-tuned on Twi specifically, rather than wrapping a third-party API designed for English. This means the agent can understand the patient’s intent, not just her isolated words. When she says she’s coming tomorrow, the agent confirms the appointment. When she says her transport failed, the agent reschedules instead of marking her as a no-show.

This is early-access software. We’re honest about that. The platform is reaching feature parity with established players, and outbound calling is gated behind a telephony-provider decision that hasn’t gone live yet. But the core technology, native Twi recognition and synthesis, is built and working today. The infrastructure that matters, the call-truth tracking, the consent and audit trail for every call, the metered billing you control, is in place.

The patient the dashboard didn’t see

The clinic’s dashboard showed a completed call and an unconfirmed patient. The two facts didn’t connect in the system. But they were the same event. The patient wasn’t unreachable. She was reachable, in the only language she could genuinely respond in. The system just wasn’t listening.

The Mars Climate Orbiter didn’t fail because of a dramatic malfunction. It failed because two perfectly reasonable systems couldn’t translate between each other. The missed appointment failed the same way. The solution is the same in both cases: build the boundary so the translation actually works.

Asenda Talk

A self-serve platform for building and running voice AI agents, built on native African-language speech (Twi, with more languages in progress) instead of a wrapper around a third-party voice API.

Try Asenda Talk

Comments

No comments yet.