Translation Mixer.

August 7, 2026 · By Jeremy Lemley / Lemley Tech · 6 min read

Finnish and Hungarian Don't Break Translators. Everyone Assumes They Do.

linguistics
language

Finnish and Hungarian Don't Break Translators. Everyone Assumes They Do.

Measured on Google Translate, August 2026. Scores from this kind of test shift with the engine, as we found the hard way.

Finnish has fifteen grammatical cases. Hungarian has eighteen, depending on who's counting. Both languages build enormous single words by stacking suffixes onto a root, and both show up constantly on lists of the world's most complex languages. If you wanted to guess which languages would destroy a sentence in machine translation, these are the obvious picks.

We tested that assumption and it turned out to be wrong. In our 20-language damage ranking, Finnish placed 19th out of 20, meaning almost nothing happened to the text. Hungarian placed 15th, below German and Vietnamese, neither of which has any of this machinery.

The languages that actually wrecked the paragraph were Japanese, Korean, and Chinese. None of them are agglutinative in the Finnish sense. What they have in common is the opposite habit: they leave information out.

What Agglutination Actually Does

Agglutinative languages express with suffixes what English expresses with separate words. The standard demonstration is a single word that requires a whole phrase in English.

We ran a few through the mixer:

WordLanguageMachine translation
taloissanikinFinnish"even in my houses"
házaimbanHungarian"in my houses"

Thirteen letters become four English words. The Finnish word contains the root talo (house), a plural, a case meaning "in," a possessive meaning "my," and a clitic meaning "even," all stacked in order. Hungarian does the same thing with a different set of suffixes.

The machine handled both correctly. It also handled the long showpiece words that Finns and Hungarians use to intimidate foreigners:

Epäjärjestelmällistyttämättömyydellänsäkään (Finnish, 43 letters)
Machine translation: "Not even with its lack of systematization."

Çekoslovakyalılaştıramadıklarımızdanmışsınızcasına (Turkish, 49 letters)
Machine translation: "As if you were one of those we couldn't make Czechoslovakian."

Both are accurate. The Turkish one is genuinely impressive, since it unpacks a causative, a negation, a plural, a possessive, a past-tense evidential, and a comparative suffix into a coherent English clause.

Turkish deserves a caveat, because it is a partial exception to this post's argument. It handles long words well, but it placed 7th of 20 on the paragraph test, well above Finnish and Hungarian. Turkish also drops subjects, which is the trait that dominates the top of that table, so it gets pulled in both directions at once. Its agglutination is not what costs it.

Hungarian's famous entry is the one case where the machine gave up on the detail:

megszentségteleníthetetlenségeskedéseitekért (Hungarian, 44 letters)
Machine translation: "for your unholy acts"

That is the gist and very little else. The word is usually glossed as something like "for your repeated pretending to be undesecratable," and the suffixes carrying "repeatedly," "pretending to," and "cannot be" all collapsed into "unholy." In fairness, this word is a party trick rather than something Hungarians write, so there is no natural usage to learn from. It fits the rule below: the longer the suffix stack, the likelier one of them is doing work English has no slot for.

Why Long Words Are Easy

The reason agglutination doesn't break translation is that it makes information more explicit, not less. Every suffix in taloissanikin is a fact stated out loud: plural, location, possession, emphasis. English states the same facts using prepositions and pronouns. The two languages package identical information differently, and converting between packaging formats is exactly what these systems are good at.

Compare that to the languages at the top of the damage ranking. When Japanese omits the subject of a sentence, the information isn't packaged differently, it's absent. When Mongolian uses a pronoun that carries no gender, there is nothing for the machine to convert. It has to guess, and a guess is where meaning goes to die.

Complexity for a human learner and difficulty for a machine translator are two different measurements. Finnish is brutal to learn because you have to memorize fifteen case endings and their interactions. That same explicitness is what makes it safe to translate. The machine doesn't have to memorize anything under pressure.

Where They Do Fail

Agglutination isn't the weak point, but these languages have one, and it showed up clearly.

We sent the same sentence through Finnish and Hungarian:

In: I wonder whether I should have told her about the letter before she left.
Through Finnish: I wondered if I should have told him about the letter before he left.
Through Hungarian: I wonder if I should have told him about the letter before he left.

Both changed her to him, twice. Finnish uses hän for any third person and Hungarian uses ő, neither of which specifies gender. The translation into those languages is correct and the information simply stops existing at that point. Coming back to English, which requires a gendered pronoun, the machine guessed and both times it guessed male.

This is the same failure that hit the grandmother in our ranking experiment, where four of the nine genderless-pronoun languages we tested lost her gender on the way back. Hungarian was among them, demoting her to "it." Finnish recovered the right pronoun in that longer paragraph, most likely because the word "grandmother" appeared earlier and gave the model something to anchor on. Remove that context, as the single sentence above does, and Finnish misses too.

The other thing that gets lost is smaller and harder to notice. Finnish has clitic particles that add shades of meaning to a word, and they tend to evaporate:

In: Ehkä en kirjoittaisikaan sinulle, jos tietäisin osoitteesi.
Out: Ehkä en kirjoittaisi sinulle, jos tietäisin osoitteesi.

The English in between was "Maybe I wouldn't write to you if I knew your address," which is close enough. But the original kirjoittaisikaan carries -kaan, roughly "after all" or "not even then." The returning Finnish dropped it. The sentence still works and a native speaker would notice the difference immediately.

So the loss profile for agglutinative languages is narrow and specific. Gender disappears if the sentence involves a person. Fine-grained particles get sanded off. Everything structural comes back intact.

The Practical Version

If you're assembling a language chain and want maximum chaos, skip Finnish and Hungarian. They will hand your sentence back looking much the way it left. Reach for the subject-droppers instead, or better, put two of them next to each other so the second one can misread the first one's guesses.

If you're translating something you actually care about, the reverse advice applies, with one exception. If your text is about a specific person, check the pronouns on the way back. That's the one place these languages reliably lose something, and it's easy to miss because the rest of the sentence looks perfect.

More broadly, the lists that rank languages by case count are answering a question about students, not about machines. The hard part of translation isn't handling elaborate grammar. It's reconstructing what a language chose not to say.

Test It Yourself

The pronoun sentence is loaded and runs through both languages back to back:

→ Watch her become him, twice

Try it with a sentence about someone you know. The grammar will survive. The person might not.


Related: We Round-Tripped One Paragraph Through 20 Languages is the full experiment this post pulls from. Why Machine Translation Still Struggles With Japanese covers the language that finished at the other end of the table. And Lost in Translation: The Loanwords That Weren't tests what happens to borrowed vocabulary in the same setup.

Try it yourself →

Send your own sentence through the translation telephone game and see what comes back.