How to externalize hardcoded strings for game localization
Hardcoded text cannot be exported, cannot be counted for a quote, and cannot be handed to a translator without handing over the codebase. Getting it out works best in one direction: find the literals mechanically, decide where they have to land in your engine, replace them, then prove nothing was missed by running a pseudo-localized build before a translator ever opens the file. This article gives an extraction pattern and an honest account of what it misses, the target shape in four engines, the six defects worth fixing while the text is already in your hands, and the three checks that catch what the move broke.
Find the literals with a regex, then filter by call site
Extraction is two stages, and the second one matters more. Stage one is a regex over your source files that captures every double-quoted literal. Stage two decides which of those a player can actually see. Nothing about a string's text tells you that reliably, so the filter has to work on two different signals: the shape of the literal, and the function it is passed to.
Shape rules remove the obvious non-text. A literal that is all lowercase with underscores is an identifier. One that is dotted with no spaces is a key. One containing a slash or a scheme prefix is an asset path. One that is all uppercase with underscores is a constant. Run that over the three snippets below and thirteen literals drop to seven.
The seven survivors still include three log messages, and they look exactly like UI text: ordinary sentences in ordinary English. Shape cannot separate them, because there is no shape difference to find. What does separate them is the call. A literal passed to a logging function is not player-facing whatever it reads like, so a second filter listing the logging calls your codebase actually uses takes seven down to four. Those four are the real work.
Two limits are worth knowing before you trust the output. Short player-facing labels get dropped: lowercase 'ok' and 'next' look like identifiers, uppercase 'OK' looks like a constant, and a progress label like '3/5' looks like a path. Meanwhile an internal name like 'ItemName' survives every rule and lands in your candidate list. A one-character literal such as the exclamation mark glued onto the end of a concatenated sentence slips past any minimum-length guard you set, so the fragment gets reported without its punctuation. The extraction gives you a worklist to read, not an answer to act on.
One case is much easier than the rest. If your source language is Japanese, a single test for kana and kanji separates player-facing text from keys, paths and identifiers with no shape rules at all, because none of those ever contain CJK characters.
// Stage 1 — every double-quoted literal, two characters or more
/"([^"\\\n]{2,})"/g
// Input: three snippets, one per language
// C# ShowDialog("Not enough gold to buy this item.");
// Debug.Log("purchase rejected: insufficient funds");
// Analytics.Send("shop_purchase");
// ShowToast("You bought " + item.name + "!");
// icon = Resources.Load<Sprite>("UI/Icons/gold.png");
// state = "IDLE";
// C++ UE_LOG(LogQuest, Warning, TEXT("no quest available for this actor"));
// ShowSubtitle(TEXT("Take this sword. You will need it."));
// const FName Tag = TEXT("quest.intro.greeting");
// GDScript push_warning("save slot is full")
// label.text = "Game saved."
// var path = "res://saves/slot1.tres"
// emit_signal("save_completed")
// Stage 1 result: 13 literals.
// Stage 2a — shape filters drop 6:
// "shop_purchase" "save_completed" lowercase identifier
// "quest.intro.greeting" dotted key, no spaces
// "UI/Icons/gold.png" "res://..." asset path
// "IDLE" uppercase constant
// Stage 2b — literals passed to a logging call drop 3 more:
// "purchase rejected: insufficient funds"
// "no quest available for this actor"
// "save slot is full"
// Final candidates (4):
// "Not enough gold to buy this item."
// "You bought " <- a fragment: needs a code change, not a move
// "Take this sword. You will need it."
// "Game saved."Where the text has to land in Unity, Unreal, Godot and RPG Maker
In Unity, the Localization package's unit is a String Table Collection: one entry key per string, one column per locale. Code can hold a LocalizedString field, or a LocalizeStringEvent component on the object can drive the text property of a Text or TextMeshPro component so no script touches the string at all. Values with inserted data go through Smart Strings rather than manual concatenation. Confirm the current component names in the String Tables and Localized Strings chapters of the Localization package documentation, and see how Unity's localization system fits together for the surrounding setup.
Unreal's split is at the type level, and this is the most useful thing to know about it. FText is the localizable text type; FString is a mutable string the gather step does not collect. Wrapping a literal in LOCTEXT or NSLOCTEXT registers it with a namespace and key so the gather commandlet can find it in source. A string built at runtime with FText::FromString is an FText but never existed as a literal in source, so it is not gathered and will never appear in your translation targets. That is how text goes missing in a project whose types all look correct. The Text Localization chapter of the engine documentation lists which sources the gather step reads; Unreal's localization pipeline covers the dashboard side.
Godot keeps the call at the object level: tr() with a key, resolved against translation resources registered under Project Settings. The source can be a CSV imported as translations or gettext PO files, described in the Internationalizing Games and Importing Translations chapters of the Godot documentation. Its untranslated behaviour is convenient for testing, because tr() with an unknown key returns the key itself. A missing entry shows on screen as the literal text ui.continue rather than as a blank or as silently correct English.
RPG Maker MV and MZ are a different case worth naming, because the conclusion people draw from it is usually wrong. The text is already outside the code, in the JSON files under the data folder. That is not the same as being ready to translate. Those files interleave text with numeric database fields, address records by internal id rather than by a key you chose, and carry control codes inside the message text itself. External storage is a precondition for translation, not a substitute for a keyed table. Localizing an RPG Maker project goes through what that means in practice, and what has to react when the player changes language is a separate list.
// Unity — key into a String Table Collection
var text = LocalizationSettings.StringDatabase
.GetLocalizedString("UI", "shop.error.insufficient_gold");
// Unreal — gathered because the literal is inside the macro
FText Msg = NSLOCTEXT("Shop", "InsufficientGold",
"Not enough gold to buy this item.");
// NOT gathered: no literal exists in source for the gather step to find
FText Bad = FText::FromString(BuildMessageAtRuntime());
// Godot — returns the key itself when the entry is missing
label.text = tr("shop.error.insufficient_gold")
// RPG Maker MV/MZ — already in data/*.json, but addressed by record id
// and mixed with non-text fields, not by a key you chose:
{ "id": 4, "name": "Potion", "price": 50, "description": "Restores 50 HP." }Fix these six defects while the text is in your hands
The move is the cheapest moment to fix everything else that will block a translator, because you are editing every call site anyway. Six things are worth catching on the way past.
- Concatenation. A sentence assembled from fragments at runtime cannot be reordered, because the translator only ever sees the pieces. Translate the whole sentence with placeholders inside it.
- Plurals. Appending an s, or choosing between two strings on whether the count equals one, encodes English grammar into the code. Several languages need more than two forms.
- Case conversion in code. Uppercasing or lowercasing a translated string is locale-sensitive, and in some languages it is either wrong or a no-op.
- Text baked into images. Signage, a logo with a word in it, button art with the label drawn on. No string table reaches those; each needs a per-locale asset.
- Enum names shown raw. Calling ToString on an enum and printing the result puts an identifier on screen, and an identifier has no translation to look up.
- Debug and shipping strings sharing a call path. If one function shows both, your call-site filter cannot tell them apart and neither can the translator reading the export.
Case conversion and concatenation deserve their own warning
Case conversion fails quietly, which is why it survives review. Uppercasing the letter i gives I in English but a dotted capital in Turkish, and lowercasing I gives i in English but a dotless one in Turkish, so a locale-blind uppercase on a Turkish menu produces words Turkish readers see as misspelled. The German word for street uppercases from six characters to seven, because the sharp s becomes two letters, which breaks any layout that reserved room by measuring the original. And in Japanese the operation does nothing at all: uppercasing the four-character word for continue returns the same four characters. A design that leans on all-caps for emphasis silently loses that emphasis in a large share of your locales. If a string should read as uppercase, store it uppercase in the table for the locales where that is correct.
Concatenation is the defect that looks most like a shortcut. Reusing one word for gold and one for item to assemble a sentence saves a couple of table rows and costs you word order permanently, in every language whose grammar puts those pieces in a different sequence. Write the sentence as one entry with named placeholders and let the translator move them. When the sentence also varies by count or gender, that is what a message format is for, and ICU MessageFormat expresses both inside a single entry rather than forcing a branch in code.
// Fragile — word order is baked into the join, and the translator
// only ever receives "You received " and " x" as separate rows
"You received " + itemName + " x" + count
// One entry, placeholders named, order owned by the translator
en: "You received {itemName} x{count}"
ja: "{itemName} を {count} 個手に入れた"
// Enum printed raw: the player sees an identifier, not a word
label.text = damageType.ToString(); // -> "Fire"
label.text = t("damage.type." + damageType); // -> looked up per locale
// Locale-sensitive case, measured:
// "i" uppercased en -> "I" tr -> dotted capital I
// "I" lowercased en -> "i" tr -> dotless i
// "straße" uppercased 6 chars -> 7 chars
// "つづける" uppercased unchanged, 4 charsThe migration order, and why pseudo-localization comes first
Five steps, in this order. Inventory: run the extraction, read the candidate list, and mark each entry as player-facing, internal, or a fragment of a sentence. The fragments are the ones that need a code change rather than a move, so counting them early tells you the real size of the job. Key design: settle the naming scheme once, before you have two thousand keys, because renaming keys later throws away every translation-memory match that was stored against them. Naming rules for string keys covers the trade-offs. Replace: change the call sites one screen or one system at a time, keeping the game runnable after each batch, so a regression is attributable to a small diff.
Pseudo-localize: generate a fake locale from your source table and play the game in it. This is the step teams skip and the one that earns back its time, because it answers three questions at once that no code review answers. Text that appears on screen without the markers is still hardcoded. Text that overflows its box under padding will overflow in German or Russian too. A line showing two bracket pairs is a sentence still being assembled from fragments at runtime.
A useful pseudo-locale does three things. It substitutes accented lookalikes for ASCII letters, so the text stays readable while being unmistakably not English. It pads each string to roughly 1.4 times its original length. And it wraps the result in brackets, so truncation is visible at both ends rather than only at the right. Placeholders must pass through untouched: if the braces in an entry come out transformed, your own generator has corrupted the string and the run tells you nothing about the layout. In the output below, the eight characters of Continue become fourteen, which is where a fixed-width button first breaks.
Then check, and hand off. A translator receives the source table plus context, never the code; what goes into a localization kit lists the rest of that package.
// Transform: accent map + pad to 1.4x + wrap in brackets.
// Placeholders are split out first and passed through unchanged.
// chars -> chars source pseudo-locale
8 -> 14 "Continue" [Ĉöñţíñûé!!!!]
11 -> 18 "Game saved." [Ğáɱé šáṽéđ.!!!!!]
33 -> 49 "Not enough gold to buy this item." [Ñöţ éñöûğĥ ğöĺđ ţö ƀûý ţĥíš íţéɱ.!!!!!!!!!!!!!!]
32 -> 40 "You received {itemName} x{count}" [Ýöû řéĉéíṽéđ {itemName} ẋ{count}!!!!!!]
// Japanese source: accents do not apply to kana or kanji, so mark the
// string and pad toward the length the translation will need (~2x here)
4 -> 10 "つづける" 【つづけるーーーー】
2 -> 6 "設定" 【設定ーー】
19 -> 40 "ゴールドが足りないため購入できません。" 【ゴールドが足りないため購入できません。ーーーーーーーーーーーーーーーーーーー】Three checks that catch what the move broke
Run these against the tables rather than the game, on every export, because all three are mechanical and none needs a translator's judgment.
Target identical to source. The most common defect after a move is an entry copied into the target locale and never translated, which renders as source-language text inside a build that otherwise looks finished. Comparing the two tables key by key finds every one in a single pass. Expect legitimate matches: OK is OK in many locales, and proper nouns stay put. Treat the output as a list to confirm, and record the confirmed entries so they stop reappearing in the next report.
Placeholder sets match. Extract every placeholder token from source and target and compare them as sets. Three failures show up. A token dropped in translation means the value never renders, so the player sees a sentence missing its number. A token invented in translation asks the formatter for something the code does not pass. And a token whose name was case-changed is the sneakiest of the three, because a lowercased itemname looks correct to a human proofreader and matches nothing at all to the formatter.
Keys are unique. Check this mechanically, because nothing else will tell you. A JSON string table containing the same key twice does not fail to parse: the parser keeps the last occurrence and discards the first without a warning. A three-entry file with one key repeated loads as two entries, and the translation you can see when you open the file is not the one the game uses. Most CSV importers behave the same way. Compare the key count in the file against the entry count after load, and investigate any difference.
What still needs a person is whether each translated line fits its box, reads naturally, and matches the situation it appears in, which is the pass localization QA exists for. The three checks above only guarantee that the table you hand over is structurally the table the game will read.
// 1. Duplicate key — JSON.parse throws nothing and keeps the last value
{"shop.buy":"Buy","shop.sell":"Sell","shop.buy":"Purchase"}
-> 2 entries loaded; shop.buy = "Purchase" (the "Buy" row is gone)
// 2. Placeholder sets, source vs target
ok {count} {itemName} vs {count} {itemName}
FAIL {count} {itemName} vs {count} {itemname} case-changed name
FAIL {count} {itemName} vs {itemName} {count} dropped
// 3. Target identical to source
ui.continue "Continue" / "つづける" ok
ui.quit "Quit" / "Quit" untranslated
ui.ok "OK" / "OK" confirm, then record as intended