
The hardest machine translation error to catch is often a sentence that reads well. Nothing looks broken. The grammar is clean. The meaning has simply moved.
Imagine a support instruction that says a user can recover a file unless the retention period has expired. A translation drops the exception. The sentence is now shorter and perfectly fluent, but it promises something the product cannot do.
That is why a quality process needs more than a fluency check. It needs different checks for different failure modes, with a clear route from a suspected issue to a confirmed correction.
Terminology: plausible is not always approved
A model can choose a reasonable everyday word where the product requires a specific term. It can also translate the same feature name differently across pages. Neither problem necessarily looks like bad writing.
Use a maintained glossary with definitions, approved equivalents, forbidden variants and examples of where a term applies. An automated check can flag a missing approved term or a prohibited one. It should account for the target language’s inflection and allowed variants where the tooling supports them.
A reviewer still needs to judge whether the rule applies in that sentence. A glossary entry for “workspace” as a product feature should not force the same wording into a paragraph about physical office space.
Meaning: watch conditions, quantities and scope
Review small words with large consequences: not, only, until, unless, may and must. Check who performs an action, what the action affects, and which conditions limit it. An omission can change an instruction without damaging its grammar.
Automate comparisons of numbers, units, placeholders and tags. Then handle legitimate differences explicitly. A date may change format; a measurement may have an approved conversion. A mismatch is evidence to inspect, not automatic proof of an error.
For higher-risk content, compare source and target at the level of the claim. Does each permission, restriction and exception survive? A second model can suggest places to look, but agreement between two outputs does not establish correctness.
Tone: the sentence can be accurate and still wrong for the situation
A payment failure message should help someone recover. A playful translation may preserve the literal information while sounding dismissive. A formal pronoun can also be inconsistent with the rest of a product.
Give the translation process concrete examples of acceptable tone. “Friendly” is too broad. A short collection of approved error messages tells a translator or model much more.
Automated checks can spot banned phrases and some inconsistent forms of address. A fluent reviewer should judge whether the message suits the audience, the product and the moment. Keep personal preference separate from a breach of the agreed style.
Context: isolated strings invite confident guesses
“Archive” could name a location or describe an action. “Free” could refer to price or availability. A short string may not contain enough information to choose.
Include stable string identifiers, nearby text, screenshots or screen references, and descriptions of variables. When an LLM translates a batch, verify that it has not merged segments, added explanatory prose or removed content while making the result read smoothly. Validate the returned structure before accepting the text.
If the source itself is ambiguous, route a question to its owner. Asking a model to be more confident cannot supply missing product knowledge.
Give each check a defined job
A practical sequence starts with structural validation: can the file be parsed, are required keys present, and have placeholders survived? Follow that with terminology and consistency checks. Then route meaning, tone and unresolved context to a qualified reviewer according to the content’s risk.
Record findings against the exact source and target revision. Let reviewers accept, reject or amend a flag, and save their reason. A checker that repeatedly raises irrelevant warnings needs tuning; otherwise reviewers learn to dismiss the queue.
Measure both sides of the system. Review a sample of flagged content to see how many findings are useful. Also review a sample of unflagged content to look for missed errors. Passing automated checks should never be presented as proof that every sentence is correct.
The aim is to spend human attention where judgment changes the outcome. Let software find broken placeholders and suspicious patterns quickly. Give people enough context to decide whether the translation says what the reader needs it to say.


