What is a CAT tool? What it does to your game strings and your quote
A CAT tool, short for computer-assisted translation tool, is the editor a professional translator works in. It is not machine translation. It cuts your text into segments, shows each one beside an empty target field, and fills in what it can from two databases the translator maintains: a translation memory of sentences translated before, and a termbase of agreed terms. The translator still writes or approves every line. The software supplies reuse, consistency checks, and a count of how much of your file is genuinely new.
That count is the part that reaches you as a developer. Before quoting, the tool analyses your file and sorts every segment into categories: repeated inside the file, identical to an earlier translation, similar to one, or new. Quotes are commonly built on that breakdown, and the way you write and format your source strings moves lines between those categories. This article walks through what happens to a string file inside the tool, computes match percentages for game strings with two different algorithms, counts repetitions in a sample string table, and ends with the file changes and questions worth settling before your next handoff.
What happens to your string file inside a CAT tool
Screens and feature names differ between products, but the sequence is broadly the same. Two neighbouring terms get mixed up with it. Machine translation produces a draft on its own; a CAT tool can be set up to show a machine suggestion next to the memory match, but a person decides. A translation management system organises projects, files and people; some products bundle both, so the labels overlap in everyday use.
- Import through a file filter. The tool reads your CSV, JSON, XLIFF or resource file and decides what is translatable text and what is structure, such as keys, IDs and markup. Misclassification is where damage starts: a key column read as text gets translated, and a placeholder not recognised as code becomes ordinary characters the translator can edit or delete.
- Segmentation. Each string is cut into units, usually sentences, by rules that belong to the tool. As an illustration only, the sentence segmenter built into JavaScript (Intl.Segmenter, which follows ICU rules rather than any CAT tool's rules) splits “Talk to Dr. Mira in the lab. She has your map.” into three pieces, cutting after Dr., while it keeps “Wait... is that a ghost?” and “Level 2.5 unlocked.” intact. Each tool has its own rules and abbreviation lists, so ask the translator or read the analysis to learn how your text was cut.
- Lookup. Each segment is compared with the translation memory. An identical earlier source returns its stored translation as an exact match, a similar one returns a fuzzy match with a percentage, and anything else is new. Termbase entries are highlighted where they appear.
- Translation and checks. The translator edits in a two-column grid. Built-in checks can flag a missing placeholder or tag, a number that differs from the source, or a term that ignores the termbase.
- Export and memory update. The tool writes the target file back in the original format and adds each confirmed segment to the memory. That growing memory is why the second job on the same game tends to cost less than the first. What the memory file itself contains is covered in TMX and TBX, the translation memory and terminology formats.
Intl.Segmenter("en", { granularity: "sentence" }) -- ICU rules, not a CAT tool's
"Talk to Dr. Mira in the lab. She has your map."
-> ["Talk to Dr. ", "Mira in the lab. ", "She has your map."]
"Wait... is that a ghost? Run!"
-> ["Wait... is that a ghost? ", "Run!"]
"Level 2.5 unlocked. Well done."
-> ["Level 2.5 unlocked. ", "Well done."]How a fuzzy match percentage can be computed: ten strings, two algorithms
A fuzzy match percentage is a similarity score between your new source and a source already in the memory. The formula is each tool's own, and many do not publish it. To show which kinds of edits move the number, the table uses one textbook method, character-level edit distance (the fewest single-character insertions, deletions and substitutions that turn one string into the other) turned into a score as 1 minus the distance divided by the longer length. Beside it is the ratio from Python's difflib SequenceMatcher, a different matching algorithm from the standard library. Neither is claimed to be what any particular CAT tool does.
- Changing only a number barely moves the score: 3 to 5 gold coins is 95.7 percent under both methods. The stored translation still needs its number corrected, which is the edit a fuzzy match exists for, and why a tool's number check matters.
- Renaming a placeholder from {0} to {count} cost 17 points by edit distance and 11 by difflib, although the translation would be identical. Mixing %s and {0} for the same value cost about the same. Once both sides were masked so any placeholder counts as one symbol, both pairs scored 100 percent. Whether a given tool treats your placeholders as protected tags like that depends on its file filter settings, so ask before the job starts.
- Short strings are fragile. Swapping door for gate drops Open the door. to 71.4 percent, while the same swap inside a 59-character sentence stays at 93.2 percent. In a UI label, one word is most of the string.
- Reordering is where the algorithms disagree most. Save your progress before quitting. against Before quitting, save your progress. scored 19.4 percent by edit distance and 50.7 percent by difflib: the same pair with the same meaning, landing in very different places depending on the math.
- Invisible differences count. A dropped full stop, or a line break stored as LF in one file and CR LF in another, is one edit each, so two strings that look identical on screen may not be exact matches.
- The lesson is not the specific numbers. Match categories in a quote describe how one tool scored your text, and edits that change nothing for the player, such as renamed placeholders, mixed line endings or reworded duplicates, can move a line out of a discounted category. Where each band starts and what it costs is set by the translator's quote, not by a standard, so read those figures there and check the tool's own documentation if you need the scoring rules.
Source already in memory New source Edit dist. difflib
You found 3 gold coins. You found 5 gold coins. 95.7% 95.7%
You found 3 gold coins. You found 12 gold coins. 91.7% 93.6%
You found {0} gold coins. You found {count} gold coins. 82.8% 88.9%
You found %s gold coins. You found {0} gold coins. 88.0% 89.8%
Press A to continue. Press A to continue 95.0% 97.4%
Press A to continue. Press B to continue. 95.0% 95.0%
Open the door. Open the gate. 71.4% 71.4%
The ancient door is sealed. ... The ancient gate is sealed. ... 93.2% 93.2%
(59 characters; the rest reads: Find the three keys to open it.)
Save your progress before Before quitting, save your 19.4% 50.7%
quitting. progress.
Line one\nLine two Line one\r\nLine two 94.4% 97.1%
With every placeholder replaced by one symbol before comparing:
You found {0} / {count} gold coins. -> identical, 100%
You found %s / {0} gold coins. -> identical, 100%Repetitions: the lines you would otherwise pay for twice
A repetition is a segment that occurs more than once in the same job. The translator handles the first occurrence and the tool offers the same translation for the others, so analysis reports usually list repetitions apart from new words. To see how much of a game file that can be, here is an 18-row string table of the sort a small game has in its menus, shop and dialogs, counted with the script at the end of this section.
- As written, 5 of the 18 rows repeat an earlier row, carrying 12 of the 76 words, or 15.8 percent. Continue appears under three keys; Options, Quit to Title and You do not have enough gold. under two each.
- The two loot lines say the same thing but name their placeholder differently, so they count as different text. Rename {count} to {0} and the table has 6 repeated rows holding 17 of 76 words, 22.4 percent, with nothing changed for the player.
- A repetition is not always the same translation. Continue on the title menu and Continue on the game-over screen may deserve different wording in some languages, and a tool that copies the first translation to every occurrence cannot know that. The keys menu.continue and gameover.continue are the context the translator needs, so make sure they see the key column. String key naming covers keys that carry that context.
- Count your own file before a quote arrives. The script works on any list of key and text pairs; a large share of repeats is a reason to ask how repetitions are priced and whether they are reviewed.
strings.json (the sample table)
menu.continue Continue
menu.new_game New Game
menu.options Options
pause.continue Continue
pause.options Options
pause.quit Quit to Title
gameover.continue Continue
gameover.quit Quit to Title
confirm.quit Are you sure you want to quit? Unsaved progress will be lost.
confirm.overwrite Are you sure you want to overwrite this save? This cannot be undone.
shop.buy Buy
shop.sell Sell
shop.not_enough You do not have enough gold.
craft.not_enough You do not have enough gold.
loot.gold You found {0} gold coins.
loot.gold_chest You found {count} gold coins.
quest.accept Are you sure you want to accept this quest?
npc.blacksmith.greet Welcome, traveler. Need something forged?
node -e '
const rows = JSON.parse(require("fs").readFileSync("strings.json", "utf8"));
const words = (s) => s.trim().split(/\s+/).length;
const seen = new Set(); let total = 0, repeated = 0;
for (const [key, text] of rows) {
total += words(text);
if (seen.has(text)) repeated += words(text); else seen.add(text);
}
console.log(repeated + " of " + total + " words are repetitions");'
# 12 of 76 words are repetitionsFive changes to your source file before the next handoff
None of these needs new tooling. Each one removes a difference the matcher sees and the player does not.
- Settle each recurring phrase on one wording. Quit to Title, Return to Title and Back to title menu are three new segments; one wording used three times is one new segment and two repetitions. Sorting the source column and reading it top to bottom finds most near-duplicates.
- Use one placeholder syntax with a consistent naming scheme, and tell the translator which syntax it is so their file filter can protect it. A placeholder the tool recognises cannot be mistyped and stops dragging down match scores; one it reads as plain text can do both. The placeholder section of preparing a localization kit lists what to write down.
- Normalise line endings and trailing whitespace before export. They are invisible to you and visible to the matcher.
- For an update, send the new and changed rows under stable keys, together with the previous translations. Unchanged text should match exactly either way, but a smaller file is quicker to check and simpler to price. Do not rename keys between versions to tidy up: the memory matches on source text, while your import matches on keys.
- Agree the exchange format with the translator. A plain CSV works when both sides agree on columns and escaping; XLIFF carries placeholders as tags and notes as metadata at the cost of an export step, and XLIFF files explained shows what that file looks like.
Questions to ask a translator who quotes from a CAT tool
These turn the analysis report from a figure you accept into one you understand. The answers vary between translators and tools, which is exactly why they are worth asking in writing.
- Which match categories and bands do you use, and what rate applies to each? Get the band edges and the rates in the quote itself.
- Are exact matches and repetitions reviewed in context, or only confirmed? For short interface strings, the cheaper option can cost you in the build.
- Will the translation memory built on my project be delivered to me, and in which format? If you hold it, a later translator can reuse it.
- Is machine translation used inside the tool, and does my text leave your machine for it? Your confidentiality terms may have an opinion.
- How is my source language counted? Word counts that split on spaces do not work for Japanese: 金貨を3枚手に入れた。 comes out as one word, although it has 11 characters. Ask whether you are charged per word or per character and how the tool counts.
- What else do you need besides the file: screenshots, character limits, a glossary? Then plan a check of the translated build itself, because the CAT tool never sees how the text looks in the game; localization QA covers that step.