Translation Mixer.

September 26, 2026 · By Jeremy Lemley / Lemley Tech · 10 min read

The Same Sentence Costs Five Times More in Russian Than in Chinese

behind the scenes
technology

The Same Sentence Costs Five Times More in Russian Than in Chinese

Here is one sentence, sent through four languages and back, twice. Same input, same number of hops, same afternoon, same engine.

Chain A: Russian, German, Spanish, French.

Please put the milk back in the icebox before it spoils.
→ Please return the milk to the refrigerator before it goes bad.

Chain B: Chinese, Japanese, Korean, Thai.

Please put the milk back in the icebox before it spoils.
→ To prevent spoilage, put the milk back in the refrigerator.

Both are good round trips. Chain A billed 348 characters. Chain B billed 164.

Nothing about the request explains that. The input was identical and the chain length was identical. The difference is entirely in which languages were in the middle, because translation APIs bill by the character and languages do not agree about how many characters a thought takes.

I have been running Translation Mixer since the backend's first commit on February 28, 2026. Most of what I have learned in those seven months is some version of that sentence: the thing you are billed for is not the thing you thought you were buying.

All numbers below are real, collected September 26, 2026, with Azure serving. Language limits and quota values are the ones in the running code.

What a chain actually bills

When you send a sentence through six languages, that is not one API call. It is seven: one for each hop, and one to bring it home.

It used to be eight. The first version asked the engine what language you had typed, in its own separate call, which bills the whole input an additional time before any translating happens. Neither engine charges for detection when it rides along on a translation, so now the first hop just asks for both at once and the meter starts on real work.

The part that people don't realize is that the cost of a translation is determined by the length of each API call, and they are not always the same length as what you originally typed. The text changes length at every stop. By hop four you are paying for a sentence you never wrote.

So the bill for a chain is the input, plus the running length of the text at every stop along the way. If your chain runs through languages that use more words or complex characters, you pay for that at every subsequent hop.

Paste HTML, pay twice

Both engines count every character you send, including markup and whitespace. Microsoft's documentation says so outright: HTML and XML tags are counted.

There is no setting that gets you out of this. Both engines have an HTML mode that leaves tags alone instead of translating them, and in that mode the tags are still billed. Google's own answer is that tags are not translated but are still counted as charged characters. You are paying for the markup either way. The only question is whether it survives.

The mixer sends text in plain text mode, which is deliberate, because HTML mode treats line breaks as collapsible whitespace and throws them away. A poem or a shopping list would come back as one run-on paragraph. Plain text keeps your line breaks, and the price is that the engine treats a tag as something to translate.

Here is the same sentence twice, once clean and once wrapped in a single paragraph tag, through the same four languages.

PlainWrapped in <p class="intro">
Characters in2647
Billed across four hops110233

Twenty one characters of markup on a twenty six character sentence, and by the end of the chain it had more than doubled the bill.

You do not get working HTML back either. Watch the tag decay:

Russian: <p class="intro">
German: < Unterricht="Einführung">
Spanish: < Lecciones="Introducción">
French: < Leçons="Introduction">
English: < Lessons="Introduction">

Russian kept it intact. German translated the attribute name, turned class into Unterricht, and dropped the p entirely. Every language after that translated the translation. The tag went in as markup and came home as a remark about lessons.

So if you are copying out of a web page, paste the text and not the source. Stripping the markup before you send it is the only thing that actually saves you money. Switching modes just decides whether you get your tags back for the same price.

Thirteen languages, one sentence

Same 56-character English sentence, one hop out, measured on arrival.

LanguageCharactersVersus English
Russian75134%
German74132%
Spanish73130%
French66118%
Finnish57102%
Hindi56100%
Vietnamese5191%
Arabic4275%
Thai4275%
Hebrew4071%
Korean2545%
Japanese2036%
Chinese1425%

Chinese rendered the whole thing as 请把牛奶放回冰箱,免得变质。 Russian used 75 characters to say the same thing. That is a 5.4x spread on one sentence, and it is a real 5.4x on the invoice.

The pattern is mostly writing system rather than language. Scripts where one character carries a whole word or syllable are cheap. Alphabets are expensive, and alphabets with long compounds and grammatical endings are the most expensive of all. Hebrew and Arabic land in the middle because they mostly skip the vowels.

None of this correlates with translation quality, which is the genuinely annoying part. Chinese is the cheapest language on that list and one of the most destructive in our damage ranking. You do not get to pay less for worse results. You pay less for shorter ones.

The estimate is wrong in both directions

The app shows you a cost estimate before it runs a big chain. The formula is the obvious one: input length times the number of hops, plus one for the trip home.

Here is what the two chains actually cost, hop by hop. Every number is the length of the text going into that call.

Chain AChain B
Hop 1 (from English)5656
Hop 275 (ru)14 (zh)
Hop 379 (de)24 (ja)
Hop 473 (es)27 (ko)
Hop 5 (back to English)65 (fr)43 (th)
Billed348164
Estimated280280

One estimate, the same for both chains, and it is 20% too low on chain A and 71% too high on chain B.

The reason is the one assumption the formula has to make: that every hop is the same length as your input. It never is. On chain A the languages write long, so every hop after the first costs more than the formula budgeted. On chain B they write short, so the formula charges you for roughly two sentences you never sent.

There is no version of this that is accurate in advance. The only honest estimate would need to know how long your sentence becomes in Russian before it has asked Russian.

Worth noticing in chain B: Chinese collapsed the sentence to 14 characters, and then it grew back. Japanese 24, Korean 27, Thai 43, English 59. The saving is front-loaded and it leaks away. Put the cheap language last and you save almost nothing, because you already paid for the long versions on the way there.

Two vendors and a calendar

Google's Cloud Translation free tier is 500,000 characters a month. That sounds enormous until you multiply it by a telephone game. At six hops, a 500-character paragraph costs about 3,500 characters, and more than that if the languages write long. The free tier is not 500,000 characters of user input. It is closer to 70,000.

So the app runs two vendors. Google serves everything until month-to-date usage crosses 450,000 characters, at which point Azure becomes the preferred engine and Google is kept only for language pairs Azure cannot serve. At 490,000 Google is dropped from routing entirely.

The 10,000 characters between those two numbers are not a safety margin for the bill. They are a safety margin for my own cache. Routing decisions are cached for up to five minutes, so the hard cap has to sit far enough below the real ceiling that a stale cache cannot spend past it.

The counter is keyed by year and month, so the switch back to Google resets on the first automatically. There is no cron job and nothing to forget. That is the single piece of this system I am happiest with, and it is also the least clever.

The vendor that fails by succeeding

On August 1, with Azure serving, every chain with Swahili in the middle started coming back as nonsense. Not an error. Nonsense that looked like a translation.

HopWhat came back
Japanese牛乳は冷蔵庫に戻してください。
SwahiliRudisha maziwa kwenye jokofu.
FinnishRudisha maziwa näkee jokofu.
Arabicروديشا مازيوا ناكي جوكوفو.
EnglishRudisha Maziwa Naki Jokofu.

Swahili was fine. The hop out of Swahili was not. Finnish passed the text through nearly untouched, Arabic spelled it out phonetically in Arabic script, and English spelled that back into Latin letters. The user gets a plausible-looking result that is actually their sentence transliterated twice.

No exception, no error code, no fallback, HTTP 200. Swahili is documented as supported. The response was structurally perfect and completely wrong.

I now treat "the vendor returned something" and "the vendor worked" as different questions. Auto-detection was what was going wrong, so every hop after the first one passes its source language explicitly. The first hop has to guess because the language you typed in genuinely is unknown, and that is the only place in the chain where guessing is allowed.

Your own numbers move when the vendor does

The worst consequence of running two engines is not cost. It is that the experiments on this blog are data, and the data changes depending on who is serving.

When we ranked 20 languages by how badly they damage a paragraph, Mongolian came first under Azure and Japanese came first under Google. So every post here names its engine and its date, including this one. It is a small discipline that I resented adding and have never once regretted.

What I would tell someone building this

Count the unit your vendor bills, not the unit that is convenient to count. Assume a successful response can still be a failed one. Write down which vendor produced any number you plan to publish. And when a limit exists to protect you from strangers, check whether it is currently protecting you from your own users, because mine was.

Try it

Here is the expensive chain, loaded and ready:

→ Run the 348-character version

Then swap the languages for zh-CN,ja,ko,th and watch the same sentence cost roughly half as much. The output stays good either way, which is the part I still find slightly unfair.


Related: Google Translate API Rate Limits Explained is the practical version of this post, including what to do when you hit a limit. We Round-Tripped One Paragraph Through 20 Languages has the damage scores that stubbornly refuse to correlate with any of the costs above. How Translation Mixer Was Born is why any of this exists.

Try it yourself →

Send your own sentence through the translation telephone game and see what comes back.