Process & operationsこの記事を日本語で読む

How to evaluate a translation vendor with a scorecard you can defend

You have used the same translation vendor across three titles. The most recent delivery felt worse than the first one — more terminology drift, more questions arriving the night before the deadline, corrections that came back applied to exactly the lines you listed and nowhere else. But when the renewal conversation starts, all you can bring is a feeling, and the vendor arrives with a different feeling that is just as sincere as yours.

A vendor scorecard exists to end that stalemate. It is not a ranking system, and it is not a file you build to justify firing someone. It is a small set of things you record every time a delivery lands, so that six months later you can look at a record instead of arguing about who remembers what.

The hard part is not choosing what to measure. It is choosing measures cheap enough that they actually get recorded. Anything that takes an hour of analysis per delivery will quietly stop being filled in by the second month, and a half-populated scorecard is worse than none because it invites conclusions drawn from the deliveries someone happened to have time for.

What a scorecard is for, and what it is not for

A scorecard does three useful things. It makes feedback concrete, so that a review meeting is about specific counts rather than adjectives. It gives renewal and scope decisions a basis that survives a change of staff on either side. And it lets you compare two vendors working on the same title without the comparison collapsing into whoever the producer likes better.

There is one thing it must not do: produce a single composite quality number. The moment you average error counts, delivery slips, and turnaround into one figure, you have hidden the only information you had. Every weighting you could choose is arbitrary, and once a vendor knows the formula, the number becomes the thing being optimized rather than the work. Keep three separate axes — quality, delivery, responsiveness — and read them side by side, permanently.

It is also worth being honest about what the scorecard is measuring. It measures the collaboration, not the vendor in isolation. A large share of what looks like vendor weakness on paper turns out to be a glossary that was never supplied, a source file with no context columns, or an answer to a translator's question that took your team eleven days. Design the categories so that this can show up, because a scorecard that can only ever blame the other side will eventually be ignored by both.

Quality: count error types, not just error counts

A single error count is nearly useless. It moves with the size of the delivery, with which reviewer happened to be assigned, and with how strict that reviewer was feeling. Two numbers that were produced by different reviewers on different volumes cannot be compared, and putting them next to each other on a slide implies they can.

What survives comparison is a breakdown by category and severity, loosely normalized against how much new or changed content the delivery contained. The categories matter because they point at different causes and different owners.

  • Meaning — the translation says something the source does not say, or misses a condition, negation, or number
  • Terminology — a term deviates from the glossary you supplied, or the same term is rendered two ways within one delivery
  • Tone and register — technically accurate but wrong for the character, the platform, or the age rating
  • Format — placeholders, markup tags, escape sequences, or line-break markers damaged or dropped
  • Length and fit — strings that will not fit the space they were budgeted for
  • Missing — untranslated, empty, or partially delivered entries

Severity is a second, independent axis: would this block release, should it be fixed before ship, or is it a preference. Recording severity separately from category stops a pile of cosmetic notes from looking like a crisis and stops one shipped mistranslation from being buried in an average.

The largest source of noise in the quality axis is the reviewer's own taste. Separate wrong from I would have phrased it differently, and only count the first. Preference edits are still worth logging, but as a signal about your style guide rather than about the vendor: if a reviewer rewrites a third of the lines to match a voice that was never written down anywhere, the gap is in your kit, and next quarter's terminology count will tell you whether writing it down helped.

One more distinction is worth the effort of making. Terminology that deviates from a glossary you supplied is the vendor's to fix. Terminology that is inconsistent where no glossary ever existed is yours. Splitting those two apart, even roughly, is what turns the quality axis from an accusation into a work item.

Delivery: dates, and the quality of the questions you receive

The delivery axis is the easiest to record and the most commonly recorded badly. Logging only the delivered date tells you less than logging three things: the promised date, the actual date, and whether a slip was announced in advance or discovered by you when the files were not there. A vendor who tells you on Tuesday that Friday will slip to Monday is a manageable partner. A vendor who lets Friday pass in silence is a scheduling risk regardless of how small the slip was, because you cannot plan around information you do not get.

The more informative half of this axis is the query log, and almost nobody keeps it. Record how many questions the vendor asked, when they arrived, and roughly what kind they were. Questions are the visible evidence that someone is actually reading your content rather than processing it.

Good questions have a recognizable shape: who is speaking to whom and in what relationship, what values a variable can take, whether a string has a length limit, whether two source lines that look identical are meant to be identical. A vendor who spots an inconsistency in your own source text has read more carefully than most of your team. Warning signs have a shape too. Zero questions on a text-heavy narrative title means someone is guessing, not that the brief was perfect. A cluster of questions in the final forty-eight hours means the work started late, whatever the delivery date says.

Finally, track how many of your answers actually made it into the delivered files. An answered question that the delivered line ignores is a defect in the vendor's internal handoff, and it is the kind of thing that never appears in a quality review because the line reads fine in isolation.

Responsiveness: what happens after you send corrections

The third axis is where long-term partners separate themselves, and it only becomes visible after you have sent a correction batch. Record the turnaround time, but treat it as the least interesting number of the three that follow.

The first is whether corrections stick. If an issue you reported and had confirmed as fixed reappears in a later delivery, that is the single most informative figure on the whole sheet. It says the fix never reached the file the vendor works from, or that a later step overwrote it — a broken loop in their process, not a translation problem. A vendor with a moderate error rate and zero recurrence is more valuable than one with a low error rate and a habit of losing fixes.

The second is scope of fix. When you report that a term was translated wrongly in four places, does the next delivery fix those four, or every instance of the same term in the project? A vendor who searches is maintaining your product with you. A vendor who patches exactly the lines listed is complying with a ticket, and you will be finding the rest of that term for the next year.

The third is whether the correction changed anything upstream. Ask whether the glossary and the translation memory were updated as a result, and ask for evidence rather than assurance. Corrections that do not flow back into those assets are guaranteed to be re-made on the next title.

It is also worth logging disagreement, positively. A vendor who pushes back on a correction with a reason — this reads as a formal register in this locale, this term is a platform requirement, this length limit cannot be met without dropping information — is doing the job you are paying for. A vendor who silently accepts every change is not protecting you from your own reviewers.

Putting it on one sheet

The whole scorecard should fit in one row per delivery per language, filled in during the review you were already doing. The per-language part is not optional: vendors subcontract, and quality varies by language pair far more than it varies by company. A single blended figure for a vendor across nine languages can hide one language pair that has been quietly failing since the first milestone.

# one record per delivery, per language
delivery:   2026-04-12   language: fr   new/changed words: 8,400

schedule    due 04-10, delivered 04-12
            slip announced in advance: no

quality     meaning 3 | terminology 9 | tone 2 | format 0 | missing 1
            blocking 1 | should-fix 12 | preference 5
            (preference logged, never scored)
            of the 9 terminology issues: 7 against supplied glossary,
                                         2 in terms we never defined

queries     14 total, 9 arrived in the final 48h
            1 answered question not applied in the delivery

follow-up   correction batch returned in 6 days
            fix scope: reported lines only
            2 issues recurring from the previous delivery

Review the sheet on a cadence, not per delivery. A single delivery is noisy — one difficult chapter, one reviewer having a bad week, one milestone that arrived with half the context missing. Patterns become readable across a quarter or a milestone group, and reading it more often than that tends to produce whiplash decisions.

Using the results without turning them into a weapon

Share the scorecard with the vendor from the beginning. A scorecard kept secret is a grievance file; a scorecard agreed in advance is a shared instrument. Agreeing on the categories before the first delivery also removes the argument you would otherwise have later, when a count you consider a mistranslation is described by them as a stylistic choice.

Then read the patterns for what they point at. A high terminology count against a glossary you never supplied is a work item for your team, and the fix is a kit, not a discount. Recurrence points at the vendor's internal review loop and should be raised as a process question with a named owner on their side. Late-clustering queries point at their scheduling. Undisclosed slips are the one pattern that is a genuine red flag on its own, because it is about communication rather than capacity, and capacity problems are fixable while a habit of silence usually is not.

When the question really is whether to switch, weigh the sheet against what switching costs. A new vendor starts without the translation memory, the glossary, and — most expensively — without translators who have read your game. Check what your contract says about handing back the memory and term base before you assume that history transfers, because a vendor change that leaves those assets behind resets years of accumulated consistency and will show up as a terminology spike in your first delivery from the new partner.

The most common outcome of running a scorecard honestly, though, is not a switch. It is that the vendor you already have improves, because for the first time both sides are looking at the same page instead of trading impressions across a table.

Related articles