How to test a call translation tool before you buy it
Vendor demos are run on clean audio, simple sentences and the language pairs the tool handles best. Here's a test protocol that finds the failure modes in an hour, using your own calls.
Prakash Vakhesa · October 3, 2026 · 5 min read
Disclosure: we make TellAcross, one of the tools you might be evaluating. This protocol is designed to find weaknesses — including ours. If a vendor objects to you running it, that tells you something.
Every call translation demo works. They're run on good microphones, in quiet rooms, with prepared sentences, in the two or three language pairs the product handles best.
Your actual calls are none of those things. Here's how to find out what happens then, before you've signed anything.
Test 1: your real language pair, in both directions
The single most common disappointment. Vendors advertise language counts, and a count tells you nothing about your specific pair — or about direction.
Do this: run the same conversation twice, once in each direction. English → Portuguese and Portuguese → English are different code paths with different quality, and many tools are noticeably weaker inbound than outbound.
Watch for: a tool that's excellent for your language into English and mediocre coming back. That's a common asymmetry and it's the direction your customer experiences.
Test 2: someone with an accent the tool didn't expect
Speech recognition is trained on distributions, and regional accents sit in the tail.
Do this: test with a speaker whose accent actually reflects your customers — a Scottish supplier, a Quebecois client, a speaker of Indian English, someone from the specific region you sell into. Not a neutral broadcast accent.
Watch for: whether recognition degrades gracefully or catastrophically. A tool that mishears one word in twenty is usable. One that loses the sentence is not.
Test 3: your own vocabulary
Every industry has terms that general models get wrong: part numbers, drug names, legal terms, product SKUs, the name of your company.
Do this: write down the twenty terms that recur in your calls and work them into a test conversation naturally.
Watch for: whether proper nouns survive at all, and whether numbers stay exact. Numbers are the ones that cost money.
Test 4: bad audio, on purpose
This is the test vendors least want you to run, and it's the most predictive.
Do this: call from a phone on mobile data, not office wifi. Have someone speak from across the room. Add background noise — a café, a factory floor, a car. Have one participant on a laptop's built-in microphone.
Watch for: how the tool behaves when it can't hear well. Does it say so, or does it confidently produce a fluent, wrong sentence? Confident errors are far more dangerous than visible failures, because nobody catches them.
Test 5: interruption and overlap
Demos use polite turn-taking. Real conversations don't.
Do this: talk over each other. Interrupt mid-sentence. Change the subject abruptly. Have three people on the call.
Watch for: whether the tool recovers or falls behind and stays behind. Latency that accumulates across a call is a different problem from latency that's constant.
Test 6: what the other person has to do
Your prospect, patient or supplier will not install an app, create an account, or troubleshoot a permissions dialog to take your call.
Do this: send the join link to someone outside your organisation and watch them join without helping. Time it. Use a phone, not a laptop.
Watch for: anything that requires a download, a login, or a settings change. Every step here is people who don't join.
Test 7: a genuinely long call
Short demos hide resource problems.
Do this: run one call for forty-five minutes.
Watch for: quality drifting, latency creeping, battery drain on mobile, and whether anything reconnects cleanly if the network hiccups.
Test 8: the money questions
Ask directly: How is a minute counted — connected time, or speech time? Is there a per-call minimum? What happens when we exceed the plan? Is it per-seat or per-use? Can we see usage per user before the invoice?
Watch for: per-seat pricing on a tool you'll use in bursts. That's how translation budgets get quietly wasted — you pay for twelve seats to serve four calls a week.
Test 9: the data questions
These matter more than most buyers realise, and the answers should be immediate.
Is call audio stored, and for how long? Are transcripts stored? Which sub-processors see the audio — the speech recognition, translation and voice vendors are usually separate companies. Where is data processed geographically? Can we delete a call and everything derived from it? Is a DPA available?
Watch for: vagueness. A vendor who can't name their sub-processors hasn't thought about your compliance position, and in a translated call there's more data than people expect — original audio, source transcript, translated text and often synthesised speech, each a separate record.
Test 10: failure
Ask: what happens when translation fails mid-call? Does the call continue untranslated, or does it end?
For anything operational, graceful degradation matters more than peak quality. A tool that drops to a default voice or shows an error while the call continues is far better than one that stops.
How to score it
Don't score on average quality — score on worst case. You'll remember the call where it failed, not the twenty where it worked.
Weight it by what you actually do. A hospital front desk should weight Test 2 and Test 6 heavily. An importer running defect calls should weight Test 3 and Test 4. A sales team should weight Test 1 and Test 5.
And run the same protocol on every tool you're comparing, with the same script, on the same day. Vendor demos aren't comparable to each other; your test is.
One honest note
If you run this on TellAcross you'll find things it does well and things it doesn't. We're browser-based with nothing for the other side to install, priced per minute, and we publish what we store. We're not the best choice for scheduled, high-stakes interpretation — that's still a human interpreter's job, and any vendor telling you otherwise is overselling.
Ten free minutes a month is enough to run most of this protocol without talking to anyone.