Unity localization: which approach to pick, and what breaks
Most Unity projects should use the official Localization package, and the exceptions are narrow: a game with a couple of hundred strings and no plans to swap non-text assets, or a project that cannot take on the Addressables dependency the package brings. If you are deciding right now, that is the answer.
Everything after this is the part a manual does not cover, because a manual documents the package rather than the choice of approach or the places the work goes wrong. Editor menu paths, component names and inspector layouts change between package versions, so this article names the manual chapter to read rather than guessing at UI details. The parts you can verify without opening Unity at all, the file format and the locale identifiers, were run and checked before being written down here.
Three ways to localize Unity text, and how to choose
The official Localization package gives you a list of Locale assets, String Table collections that map a key to one string per locale, Asset Table collections that do the same for sprites, audio and video, Smart Strings for plural and conditional formatting, and CSV and XLIFF import and export. It loads tables through Addressables, which is both its strength and its main new failure mode. Read the manual's installation and quick start chapters for the current setup steps.
Writing your own loader means a CSV or JSON file parsed at startup into a dictionary, plus one lookup function. It is genuinely small, and for a game whose text fits on one screen it is the right call. You are then responsible for everything the package would have given you: plural rules, the editor tooling to find a missing key, and the import step that catches a broken file before it ships.
Asset store packages sit between the two, a genre framework like Naninovel included. Some are excellent, and the cost is maintenance risk: when a new Unity version lands you wait for the author, and a package that stops being updated becomes your code to maintain. Judge one on how recently it shipped and on whether you could replace it in a week.
Pick by what has to change per locale, not by string count alone. Answer these four questions, then read them against the table underneath:
- Does anything besides text change per locale? If yes, you want Asset Tables, which in practice means the package.
- Does the player switch language without restarting? That needs a change-notification mechanism, not just a dictionary swap.
- Can your build pipeline run an extra content build step? If not, Addressables is a problem rather than a feature.
- Will someone outside the team edit the text? Then the exchange format matters more than the runtime, and CSV or XLIFF export becomes the deciding feature.
Situation Approach
--------------------------------------------------------------------------
Under ~200 strings, solo, text only Own CSV or JSON loader
UI + dialogue + item names, still growing Official Localization package
Localized sprites, audio or video Official package (Asset Tables)
Plural or gender agreement needed Official package (Smart Strings)
or an ICU library of your own
Cannot add the Addressables dependency Own loader
Console certification, several SKUs Official package (explicit,
auditable locale list)
Switching language mid-session required Official package, or your own
change-notification layerThe exchange table, and why a column-count check is not enough
Whatever you choose, a translator receives a table. The package's CSV export produces a key column, a numeric identity column and one column per locale, and it can be configured to carry comment columns for translator notes as well; confirm the exact header names and column options in the manual's CSV chapter for the version you have installed. The shape is what matters, and an example is at the end of this section.
The obvious checks on that file are structural: does every row have as many fields as the header, and is every quote closed. Both are necessary and neither is sufficient, which a deliberately damaged row demonstrates. Take a cell where someone escaped an inner quotation mark the way a programmer would rather than the way CSV requires, writing one quote instead of two, so the field reads as a quoted string containing quoted text.
A standard parser accepts it silently. The row still yields the right number of fields, and the file has no unbalanced quote at the end, so both structural checks pass. What comes out is the sentence with the quotation marks deleted. A player sees an item name with no quotes around it, in one language, on one line, and nothing anywhere reported a problem.
A placeholder comparison, by contrast, catches real damage immediately. Removing one placeholder from a Japanese cell while leaving it in the English source produced a mismatch on exactly that row, because the sorted placeholder sets no longer agree. That check costs a few lines and finds the class of bug that becomes a runtime format exception.
So compare content, not only structure. The column design and the send-a-diff workflow are covered in the Unity CSV import and export workflow; key naming that survives an English copy edit is in designing string keys; and getting strings out of scenes in the first place is in externalizing hardcoded text. The three checks worth automating are these:
- Field count per row equals the header. A raw comma typed into an unquoted cell shows up here and nowhere else.
- The set of placeholders is identical across every locale column in a row. Order may differ, and should: the Simplified Chinese line below puts the price before the item on purpose.
- The count of characters that carry meaning, quotation marks and brackets among them, did not change between export and re-import.
Key,Id,en,ja-JP,zh-Hans,pt-BR
menu.settings.title,10001,Settings,設定,设置,Configurações
shop.buy_confirm,10002,"Buy {0} for {1} gold?","{0}を{1}ゴールドで購入しますか?","用 {1} 金币购买 {0}?","Comprar {0} por {1} de ouro?"Name locales in BCP 47, and do not trust SystemLanguage alone
Locale identifiers are the one part of this you can verify exactly, so verify them. Two separate operations in any modern JavaScript runtime answer two separate questions. Canonicalization shows which forms are real, which get rewritten, and which are rejected outright. Maximization, which applies the standard likely-subtags data, shows what a tag means once its blanks are filled in. The table runs both over the tags a game normally needs, which is why one tag appears in each half with a different result.
Underscores are the most common mistake, because file naming habits and .NET culture code habits both encourage them. They are not a stylistic variant; a tag with an underscore is rejected. Legacy codes are the second trap: the .NET-era abbreviations for simplified and traditional Chinese are not valid tags at all, while the old two-letter codes for Hebrew and Indonesian are silently rewritten to their modern forms. Case, on the other hand, is not a difference. A lowercase script subtag is normalized, so compare tags case-insensitively rather than treating two spellings as two locales.
The subtler result is the second half of the table, which is maximization rather than canonicalization. A language-only tag is never neutral: the likely-subtags data resolves Portuguese to Brazil and Traditional Chinese to Taiwan, which is why Traditional Chinese is unchanged in the first half and gains a region in the second. If you ship a locale named just Portuguese, something downstream will decide which Portuguese you meant, and you will not be the one deciding. Name the region yourself whenever two variants of a language are meaningfully different to players.
Canonicalization (Intl.getCanonicalLocales) ja-JP -> ja-JP zh-Hans -> zh-Hans zh-Hant -> zh-Hant left alone by this operation pt-BR -> pt-BR es-419 -> es-419 Latin American Spanish, region code 419 zh-hans-cn -> zh-Hans-CN case is normalized, not a real difference pt_BR -> rejected underscore is not a valid separator zh-CHS -> rejected legacy .NET code, never valid BCP 47 iw -> he legacy code silently rewritten in -> id same Filling in the blanks (likely subtags, Intl.Locale.maximize) pt -> pt-Latn-BR Portuguese resolves to Brazil zh-Hant -> zh-Hant-TW and Traditional Chinese to Taiwan es -> es-Latn-ES ja -> ja-Jpan-JP
Five things that break after the package is installed
First, strings that work in the editor and come out blank in a player build. The package loads tables as Addressable assets, so a player build needs an Addressables content build, and in the editor the default play mode script reads straight from the asset database, which is exactly why the editor hides the problem. Confirm the play mode script names and the build menu path in the Addressables documentation. The symptom is empty labels rather than an exception, so a build can pass a smoke test on any screen nobody has translated yet.
Second, tofu. A font asset renders only the glyphs in its character table, so adding Japanese, Korean or Chinese to a project whose font was built from a Latin character set produces blank boxes, not an error. Decide deliberately between a dynamic font asset that adds glyphs at runtime and a static atlas built from the text you actually ship, and configure a fallback chain for the characters that slip through. The mechanism is explained in why missing glyphs render as boxes, and the harder case of a stylized face covering two scripts is in pixel fonts for Japanese and English.
Third, translators breaking plural and conditional syntax. It matters because plural rules are not a singular-plural pair: English has two plural categories, Japanese and Simplified Chinese have one, Russian and Polish have four, Arabic has six. In Russian, 1 and 21 select the same category while 11 selects a different one, so a rule written as one-is-singular-everything-else-is-plural is wrong at 21, 31 and 101. Give translators the categories their language actually has, and validate the syntax on import rather than at runtime. The message format itself is covered in the ICU MessageFormat guide.
Fourth, concatenation freezing word order. A sentence assembled from parts in code cannot be reordered by a translator, and in Japanese the order almost always has to change. Put the whole sentence in one entry with numbered placeholders.
Fifth, runtime locale selection. Setting the selected locale updates components that subscribe to the change event; text your own script assigned once does not update, and neither does a string you cached. Test switching mid-session on every reachable screen, not only from the title menu, and persist the choice so it survives a restart. Use the system language as a first guess and never as the stored answer: Unity reports it as a language enum with no region, so Portuguese arrives with no way to tell Brazil from Portugal, and Chinese is distinguished only as simplified or traditional, which is a script distinction rather than a regional one. Check the SystemLanguage page in the scripting reference for the full list, and read how BCP 47 language tags are structured before mapping it onto your locale names.
Run a pseudo-locale pass before any translation exists
A pseudo-locale is a fake language generated from your source text: every letter is replaced with an accented lookalike, the string is padded to simulate a longer translation, and the whole thing is wrapped in brackets. Placeholders are skipped so they survive intact. Generating one takes a few lines in any language, and the official package has its own pseudo-locale feature described in the manual.
Three defects become visible in one playthrough. Any text still in plain unaccented letters never went through the lookup, which is a far more reliable audit than searching your scenes for string literals. Padding that overflows its container marks the layouts that cannot take a longer translation. And a missing closing bracket marks a field that clips, which is otherwise invisible whenever the last word happens to fit.
Two rules keep the pass trustworthy. Pick a padding ratio and keep it, because a fixed ratio makes results comparable between builds while a per-string ratio makes an overflow look like a regression. And never pad inside a placeholder: split the string on the placeholder pattern, transform only the literal parts, and rejoin. Padded at 40 percent of the visible characters, real UI strings come out like this:
Settings -> 【Šéttîñgš····】 8 -> 14 chars
Continue -> 【Çøñtîñüé····】 8 -> 14
Buy {0} for {1} gold? -> 【Büý {0} før {1} gøld?······】 21 -> 29
Are you sure you want to abandon this quest? -> 【Åré ýøü šüré ýøü wåñt tø åbåñdøñ thîš qüéšt?··················】 44 -> 64The minimum setup, and what to add later
A first pass does not need the whole system. It needs enough that adding the second language is an import rather than a rewrite.
- Minimum: one source locale and one target, every player-facing string in one table with a stable key, numbered placeholders for anything assembled at runtime, a font verified inside the game for each script, a language setting the player can change, and one pseudo-locale pass.
- Add when the text grows past a few hundred entries: an export and reimport round trip on a diff rather than the whole table, and a placeholder check that runs on import.
- Add when non-text assets differ per locale: Asset Tables, and a written list of which assets are localized.
- Add when a language needs plural or gender agreement: Smart Strings or an ICU library, plus the plural categories for each target language written down where translators can see them.
- Add before a console submission: an explicit locale list you can hand to certification, and a build that fails rather than ships when a table is missing an entry.