
Two vendors receive the same feedback: “The translation scored 92.” One treats that as a good result. The other asks which errors were counted, how much text was reviewed and why a stylistic preference cost five points.
The number looks precise. The agreement behind it is missing.
A useful quality score begins with shared definitions and traceable evidence. The calculation comes later. You will still have disagreements, but you can make them specific enough to resolve.
Decide what the evaluation is for
An evaluation used to improve a draft can tolerate open questions. A release decision needs an agreed acceptance rule. A vendor comparison needs comparable content and review conditions. Trying to use one score for all three purposes creates avoidable confusion.
Write down the content type, audience, language pair, review scope and intended decision before reviewing. Also state which source revision and reference materials apply. A vendor should not be scored against a glossary updated after delivery without that change being acknowledged.
Use a small, shared set of error categories
Start with categories that reviewers can reliably distinguish. Accuracy covers changes to meaning, including omissions and additions. Terminology covers the use of required domain or product terms. Other useful categories include linguistic conventions, style and locale conventions.
The MQM error typology provides a shared vocabulary and more detailed subtypes. Choose the subset that fits the work, and document it. More categories do not automatically produce better decisions.
Give each category an example and a boundary. A preferred synonym is not necessarily a terminology error. A wording change is not proof that the original was wrong. Record optional improvements separately from confirmed defects.
Define severity by effect on the reader
Severity describes the consequence, not how strongly the reviewer dislikes the sentence. The same surface error can have different effects in different contexts.
For an illustrative project rubric, a minor issue might distract without changing the intended meaning. A major issue might materially mislead the reader or obstruct a task. A critical issue might create a consequence the project has explicitly defined as unacceptable for release. These definitions need examples from the actual content.
Agree how to handle repeated errors, overlapping findings and issues caused by the source. If one mistranslation also sounds awkward, counting it twice can exaggerate the problem. Keep a clear policy for choosing the primary finding.
Attach every finding to a segment
Each scored issue should identify the source and target version, the affected span, the category, the severity and a short explanation of the impact. Include the relevant reference or rule where one exists.
“Major accuracy error” is a label. “The target removes the condition that limits when a refund is available” gives the vendor something to verify. A proposed correction can help, but the explanation should stand without it.
Allow the vendor to respond to that finding. Keep the original annotation, the response and the final adjudicated result. Rejected findings must stop contributing to the score.
Make the arithmetic reproducible
Here is a deliberately simple example, not a universal MQM formula. Suppose a project assigns one penalty point to a minor error and five to a major error, and reports penalty points per 1,000 reviewed source words.
A 2,000-word sample contains four confirmed minor errors and two confirmed major errors. That produces 14 penalty points: four plus ten. Divide by 2,000 and multiply by 1,000. The result is seven penalty points per 1,000 words. Lower is better under this example rule.
Publish the weights, denominator, counting rules and acceptance threshold alongside the score. Specify any critical-error release rule separately so an average cannot hide a serious issue. MQM supports defined scoring approaches; consult the MQM scoring models when designing the project’s model.
Calibrate before comparing vendors
Have reviewers independently assess a shared sample, then discuss the disagreements. Are they detecting different issues, choosing different categories, or assigning different severities to the same issue? Each problem needs a different adjustment.
Repeat calibration when the content, guidance or reviewer group changes. Keep a small set of settled examples as a reference.
Report the sample size and selection method. A targeted sample of difficult strings should not be compared directly with a random sample of routine text. A score from a small sample is a finding about that sample, with limited grounds for judging the whole delivery.
The most useful outcome is not a number nobody questions. It is a number every participant can reconstruct, challenge with evidence and use to decide what needs to improve.


