In-browser machine translation: I measured 30 phrases, found 30% wrong in meaning, and changed the engine
Our free translator ran a small machine-translation model directly in the browser (a Bergamot / Opus-MT style model compiled to WebAssembly). The appeal is privacy: the text never leaves the device. The downside is quali
Our free translator ran a small machine-translation model directly in the browser (a Bergamot / Opus-MT style model compiled to WebAssembly). The appeal is privacy: the text never leaves the device. The downside is quality. I finally stopped guessing and measured it on the live site with 30 phrases.
The test
30 phrases in both directions between English and Russian, grouped by type. I marked each result as correct, correct with small flaws, or wrong in meaning.
Result: 43% correct, 26% with small flaws, 30% with a changed meaning. Speed was fine: the first translation takes about 11 s (model download and load), after that about 1.7 s per phrase.
What worked and what did not:
- Connected prose and business text: good, in both directions.
- Technical text: it invented words ("driver" became "Π²ΠΎΠ΄ΠΈΡΠ΅Π»Ρ" (a vehicle driver) instead of a software driver).
- Idioms and casual speech: 5 out of 5 failed. "I'll hit you up" came out as "I will hit you", which reads as a threat.
- Ambiguity: "the bank refused the loan" was inverted.
- A legal sentence, "shall not disclose ... to any third party", came out with the direction of the prohibition reversed. For a translator that advertises use on contracts, that was the one that worried me most.
The pattern is simple: it fails exactly on what a random visitor tries first, an idiom or a casual phrase.
Trying LLMs as a replacement
I ran the same 30 phrases through 8 models over an API (240 requests). The whole benchmark cost about $0.013. The hard set was 17 phrases where the in-browser model had failed.
| Model | Hard phrases right (of 17) | Median latency |
|---|---|---|
| Gemma 26B (instruction-tuned) | 17 | 1.46 s |
| GPT-5.4 mini | 15 | 1.06 s |
| Llama 3.3 70B | 13 | 0.92 s |
| DeepSeek V4 Flash | 13 (+2 empty) | 2.85 s |
| Qwen3 30B A3B | 11 | 1.96 s |
Two things surprised me:
- The cheapest model was the worst. Qwen3 30B had the lowest price and the lowest score. It translated "KlΓΆΓe" as cutlets and "Vorspeisen" as first courses. "Newer and bigger" is not the same as "better for this task"; I had bet on Qwen before measuring and was wrong.
- Reasoning models are a trap for translation. One of them returned empty text for all 30 phrases, because it spent the whole token budget on thinking and never reached the answer. Raising the limit just means paying for the most expensive tokens.
I shipped the winner. The cost came out to roughly $0.16 per million source characters, and it fixed the threatening "hit you" case and the inverted legal sentence.
The trade-off I had to accept
Moving translation to a server model means the text now leaves the device, which was the original selling point. The client falls back to the local model if the server call fails. If privacy is your main constraint, the small local model is still the right tool. If quality on everyday text matters more, it is not good enough.
You can compare for yourself: the browser translator and the English to Russian page. What phrases break your favourite translator?
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.