What is LQA? Localization QA, linguistic QA, and how to run one
LQA stands for localization quality assurance: a review pass that checks translated text inside the running game, not only in the file the translator delivered. Some teams call it linguistic QA. It exists because a translation can be correct as a sentence and broken as a piece of software, and because the reverse is also true.
The term is used loosely, and that causes real problems in contracts. Some vendors mean only the language review. Some studios mean only the on-device hunt for clipped and garbled text. A quote that says LQA may or may not include a build, a native reviewer, or a retest after fixes. Settle which of those you are buying before you compare prices.
This article is written for the person who has to organize the pass rather than perform it: how to scope it, where it goes in the schedule, how to write tickets a programmer and a translator can both act on, how to decide what blocks a release, and which checks are cheap enough to run on every build.
Bilingual review, in-build linguistic QA, and functional QA
Three activities get filed under the same acronym. They catch different bugs and need different people, so naming which one you mean is the first useful thing you can do.
Bilingual review and in-build linguistic QA both need a reviewer at native level in the target language who also plays the game. Functional and format QA needs someone who can drive a build and recognize missing, clipped, or garbled text, which is visible without reading the language at all. That is why a team can usefully run its own functional pass in languages nobody on it speaks.
The split also gives you a test for a proposal. If a quote says LQA and the deliverable is a comment column in a spreadsheet, you are buying bilingual review. If nobody has asked you for a build, a save file near the content, or a way to file bugs, it is not an in-build pass. Neither is wrong, but a release plan that assumed the other one will be short. What each of the three actually catches:
- Bilingual review, also called translation review: a second linguist compares source and target in the delivered file. It catches mistranslation, omission, invented content, terminology drift between files, and the wrong register. It needs no build, so it can start the day the translation lands, and it cannot see anything that only exists on screen.
- Linguistic QA in the build: the same judgment applied where the player meets the text. A term that is accurate but wrong for this character. A line that reads well alone and repeats the line before it. A joke that works in the script and not in the two seconds the game gives it.
- Functional and format QA: whether the string behaves as software, regardless of whether the wording is good. The functional categories listed further down are exactly its scope.
Scoping a pass: a fixed route, risk weighting, one stress language
Full coverage, meaning every screen in every language, is affordable for a small game and for almost nothing else. Choose a scoping rule deliberately instead of running out of budget halfway through the second language.
- A fixed route. One scripted playthrough that reaches every UI surface, every system message you can trigger, and the opening hours of story, run identically in each language. Identical routes make languages comparable: if one language shows six overflow bugs and another shows none, that is a fact about the text rather than about who tested.
- Risk weighting. Full coverage for anything the player cannot avoid or cannot undo: the first hour, tutorials and control prompts, store and purchase text, error messages, save and delete confirmations, age and legal text. Sampled coverage for late-game flavour text and collectibles.
- One stress language in full. Layout bugs are properties of the layout, not of the language, so the target that is longest and least like the source surfaces most of them. German and Russian stress width; Japanese and Chinese stress wrapping and font coverage.
Test cases, the device matrix, and when to run each half
Write test cases as states with a path, not as an instruction to look at the menus. Each case names the screen or event, how to reach it, and what to check: reach the first shop, try to buy something you cannot afford, confirm the refusal message fits one line and names the currency correctly. Cases written that way survive into the retest, and a second tester reproduces the first tester's coverage.
Language is one axis of the matrix. The other is everything that changes how text renders. Platforms ship different font fallbacks, so a string that renders on one shows empty boxes on another. Include UI scale and aspect ratio, the operating system's language and region settings, since number and date formats usually follow the OS rather than the game, and controller versus keyboard prompts, which substitute different glyph labels into the same string. You need every axis exercised somewhere and the riskiest combinations exercised directly, not every cell.
Before any of it, run pseudo-localization: replace source text with accented, lengthened stand-ins and play. Anything still readable in the original is hardcoded, and anything clipped will clip worse in a real language. One build removes a whole class of bug from a pass you are paying for, and getting hardcoded strings out of the code covers the extraction side.
Scheduling is where a single LQA week before submission usually fails, because the two halves have opposite needs. Linguistic review produces rewrites, so it has to happen while text can still change, before text freeze. The functional pass produces layout and rendering bugs, and those are only stable after freeze. Run it earlier and you rediscover the same overflow every time a line is edited.
pseudo-localization pass early build, before real text kit and glossary to translators source frozen enough to quote bilingual review in the file, before import import + automated checks every build from here on in-build linguistic + functional release candidate, after text freeze fix round text changes reopen layout questions retest the tickets, plus screens fixes touched platform submission one full fix cycle of slack before this
A bug report template for localization issues
A localization ticket that omits the source string or the exact screen gets closed as cannot reproduce or working as intended. The audience is unusual: a programmer who does not read the target language, and a translator who cannot run the build. One ticket has to serve both.
- String key. This is what turns a fix into a one-line edit. If testers cannot see keys, fix that before the pass: a debug overlay that prints the key of the string under the cursor removes a round trip from every ticket. Naming that holds up is covered in string key naming that survives growth.
- Source text and current text, always both. The programmer needs the current string to find it; the reviewer needs the source to judge the fix.
- Expected and suggested fix, kept apart. For a functional bug the expectation is behaviour, such as fitting one line at this size. For a linguistic bug it is wording. Merging the fields makes a programmer arbitrate a translation argument.
- Category and severity, both filled in by the reporter. Category routes the ticket to a translator or a programmer; severity decides whether it blocks the build. A ticket with neither gets triaged by whoever has the least context.
- Screenshot for anything visual, and a build identifier on every ticket. Without the build you cannot tell a live bug from one already fixed, and the retest becomes a second full pass.
Title: [ja] Settings/Audio - volume label clipped Language / locale: ja-JP Build: 1.4.2-rc3, Steam, Windows 11, UI scale 100% Where: Settings > Audio, first row label Steps: 1. Title > Settings 2. Audio tab 3. read row 1 String key: ui.settings.audio.master Source text: Master volume Current text: マスターボリューム調整 Expected: label fits the column at this size without clipping Suggested fix: 音量 Category: functional / truncation Severity: major Screenshot: attached Also occurs in: pause menu > Audio (same key)
Categories and severity: what blocks a release
Classify every ticket on two axes, because they answer different questions. Category answers who fixes it. Severity answers whether you ship without the fix. Agree both lists before the pass, not during it: two testers with the same list produce results you can compare, and two testers improvising produce two opinions.
If you want an established framework instead of a house list, MQM (Multidimensional Quality Metrics) is the common reference: an error typology organized by dimension, such as accuracy, fluency, terminology, style, and locale conventions, crossed with severity levels and combined into a score per unit of text. Read its current published specification before quoting a threshold or a weight, because the typology has been revised and a score compares two deliveries only if both used the same version.
Boundary cases are where a written list earns its keep: is a clipped word in an optional codex entry major or minor, and does your answer change when the clipped word is a character's name? The list below gives the linguistic categories first, then the functional ones.
- Mistranslation: the target says something the source does not.
- Omission, or content the source never contained.
- Terminology inconsistency: one concept rendered several ways.
- Wrong register: formality, honorifics, or a character speaking out of voice.
- Grammar and spelling.
- Unnatural phrasing that parses but does not read as written by a person.
- A sentence made nonsensical by the order its variables are substituted in.
- Truncation, overflow, and overlap, including text clipped by a platform's safe area or colliding with art.
- A variable token exposed to the player, or broken so it no longer substitutes.
- Garbled characters, where the bytes are fine but the decoder was wrong: reading garbled text back to the original.
- Empty boxes where characters should be, which is a font coverage problem rather than an encoding one: why characters show as empty boxes.
- Lines breaking where the language forbids it, such as a Japanese line starting with a closing bracket or comma.
- Locale formats: decimal separator, date order, currency position, sorting.
- Strings left untranslated or copied from the source, and damaged inline markup.
- Subtitles out of step with the audio they belong to, or over the line limit for their language: subtitle timing and line limits across languages.
critical the player is blocked, misled, or cannot read the text
- a variable token shown raw in a tutorial instruction
- a negation dropped from a warning or a purchase screen
- garbled or boxed-out text in a language you ship
- untranslated UI in a language the store page advertises
- wrong price, wrong age rating wording, wrong legal text
major visibly wrong, the player still gets through
- a clipped word whose meaning survives
- one game term translated three different ways
- wrong formality level for a character
- a date shown in the source locale's order
minor a competent reviewer would have written it differently
- punctuation width, spacing, a smoother phrasingWhat a person must judge, what to outsource, what to automate
The line between what a script can check and what needs a person is not difficulty. It is whether the question has one right answer independent of who asks it.
What is left for a person is everything that depends on knowing what the text is for: whether an accurate translation is the right one in this place, tone and character voice, cultural fit, humour, wordplay, names, whether a line reads as something a person wrote, and anything about what surrounds the text on screen.
Outsource the judgment you cannot perform. If nobody on the team reads the language well enough to tell native from merely correct, no process fixes that, and what actually changes going from English to Japanese shows how far past vocabulary the gap runs. Keep in house the automated checks, the functional build pass, triage and severity decisions, and the retest. What you buy outside is judgment in a language, not testing labour, so send everything that judgment needs: the build, the route, the glossary, the severity list, the ticket template, and the localization kit you prepared before sending text.
These have one right answer, so they belong on every build across every string rather than in a pass you pay a person for:
- Variable token parity between source and target: same tokens, same count, and for positional formats the order the code expects. Plural and gender forms have their own rules, covered in the ICU MessageFormat guide.
- Glossary compliance per language: the approved term present, the banned term absent.
- Length against a per-string limit, in the unit the layout uses. Master volume is 13 characters and 13 UTF-8 bytes; the Japanese マスターボリューム調整 is 11 characters but 33 bytes, and it takes roughly 22 half-width columns on screen. A byte limit rejects text that fits, a character limit passes text that overflows, so hold limits in rendered width for the font you ship.
- Untranslated detection: target identical to source, empty, or still holding a fallback value.
- Consistency both ways: one source string with two translations, and two source strings with one translation.
- Encoding validity, characters outside the shipped font's coverage, damaged tags, stray control characters, doubled and trailing spaces.
the smallest version still worth calling LQA one fixed route, run identically in every shipped language the checks above on every build, across every string one in-build pass per shipped language, on a release candidate tickets in the format above, with a build id and a string key severity criteria agreed before anyone starts looking a fix round and a retest inside the schedule, before submission