A reliable handoff test checks whether the agent carries the caller’s meaning, constraints, consent and unresolved request from Twi into English. Supervisors should score the full conversation state before and after the switch, rather than judging transcription or pronunciation alone.
Consider an illustrative test call with Ama, a composite caller in Accra. At 4:40 p.m., she is holding a handwritten account reference and explains in Twi that a payment appears twice, but she does not want either transaction cancelled until someone confirms which entry is valid.
The agent replies, and Ama switches to English: “So what will you do now?” If the agent treats that sentence as a new request, it may cancel the wrong transaction or promise an action Ama never approved. Her account record could be changed before a supervisor catches the error. For one tense beat, the whole outcome depends on whether the agent preserved the instruction given in Twi.
Define the meaning that must survive the switch
Before running the call, write a compact intent record that describes what a correct agent must retain. For Ama’s scenario, that record might say:
- The caller reports a possible duplicate payment.
- The caller wants an investigation.
- The caller has withheld permission to cancel either entry.
- The next response must explain the review step or offer human escalation.
- The agent must not interpret the English question as consent.
This record becomes the answer key. It also keeps reviewers from grading by instinct. A fluent English response can still be wrong if it drops the condition expressed in Twi.
Include at least one constraint, one unresolved detail and one safety-sensitive instruction in every test. A simple request such as asking for opening hours will rarely expose a broken handoff. A request with a conditional instruction will.
The key distinction is covered more broadly in Twi and English Code-Switching: Why Voice Agents Must Preserve Context: recognizing each language separately does not prove that meaning survives between them.
Run the same handoff in controlled variations
Use a fixed script, then vary only the point of the language switch. Start one call in Twi and move to English immediately after the caller states the main request. In the next call, switch after a number, name or account reference. In another, switch after the caller refuses an action or asks to stop future calls.
Keep the expected meaning constant. This lets the supervisor identify whether failure follows the language change, a particular phrase or a wider problem in the agent’s reasoning.
For each run, capture five checkpoints:
- What did the caller ask for?
- What action did the caller permit?
- What action did the caller refuse?
- What remains unresolved?
- What should the agent do next?
Then compare the answers at three points: after the Twi segment, immediately after the first English sentence and at the end of the call. A passing result requires consistency across all three.
Do not repair a failed run by mentally filling in what the agent probably meant. Score the response it produced and the action it attempted. If the spoken reply sounds correct but the recorded call state loses the refusal, the test fails.
Separate speech quality from intent preservation
A supervisor needs distinct scores for recognition, response quality and action state. Combining them into one “good call” rating hides the failure you are trying to find.
First, check whether the Twi speech was captured accurately enough to recover the request. Next, check whether the English response reflects the complete request. Finally, inspect the resulting action and call record. The agent should carry forward consent, opt-out status and pending escalation without silently rewriting them during the switch.
Asenda Talk’s Twi speech recognition and synthesis are fine-tuned in-house, while Vapi orchestrates the assistant runtime. The platform also records telephony lifecycle events, consent, opt-out status and an audit trail. Those components give reviewers several places to compare what was heard, what the agent said and what the system recorded.
The product remains in active early access. Outbound calling is still held behind an operator-controlled real-money gate while the live telephony-provider decision remains unresolved. Supervisors can use this test structure in controlled evaluations now; they should not treat a successful staged call as proof that live outbound operations are ready.
Make one failure block the next stage
Ama’s test turns at the moment the agent answers: “I’ll review both entries and won’t cancel either one without your confirmation.” The English sentence preserves the request, the constraint and the next step. Her call record still shows that no cancellation was authorized.
Now change one detail. If the agent says, “I’ll cancel the duplicate,” the handoff fails, even if every word was transcribed correctly. If the spoken response respects Ama’s instruction but the stored record marks cancellation as approved, it also fails. The source conversation and the action record must agree.
This is especially important for opt-outs. A Twi refusal cannot disappear because the caller later asks an unrelated question in English. When records disagree, the safe operational response is to pause the affected calling flow and review the source evidence, as described in Twi Opt-Out Compliance: Why Supervisors Must Pause Calls When Records Disagree.
End each test with a written pass or fail decision tied to the answer key. For Ama, the final line should be plain: investigation requested, cancellation withheld, no consent inferred, next step preserved. That is the handoff working where it matters.
Comments
No comments yet.