Getting startedこの記事を日本語で読む

What video game localization covers, measured with real commands

Localization is the work of making a game correct for a place. Translation is its most visible part and not the whole of it: the same screen has to hold text of a different length, print numbers and dates the way that market writes them, obey that language's rules for counting and ordering, and declare which locale it is running in at all.

This article draws the boundary of that work in numbers rather than adjectives, because the useful question before committing to a language is not whether localization is hard but which parts of your project it reaches. Every value below came out of a command run on one machine — Node v22.13.0 with ICU 76.1, CLDR 46.0 and Unicode 16.0 — and the commands are printed so you can rerun them on your own strings. Where the answer depends on your engine or your store rather than on a published standard, this says so instead of guessing.

The words change length, and length is three different numbers

Start with the part everyone expects. Forty articles here exist as matched English and Japanese pairs. Taking each article's median of Japanese characters per English word and lining the forty up gives 2.19 in the middle, 1.66 at the low end and 2.58 at the high end. Read the units before the number: characters and words are different units, so this is not a claim that text grows 2.19 times. What it does establish is that the relationship is neither fixed nor mild. The loosest article sits more than fifty percent above the tightest, on prose written by the same hands to the same brief, which means a layout that assumes one constant is assuming something this corpus does not show.

Then there is a subtler problem: length itself has no single definition. The same five-character Japanese label counts as 5, 5 or 15 depending on which count you ask for, and the byte count changes again with the encoding you write it in.

  • The Japanese label is 5 UTF-16 code units, 5 code points and 15 UTF-8 bytes. A rule of at most 12 characters accepts it when JavaScript measures it with String.length, and rejects it when a 12-byte column or a fixed-size buffer measures the same string. Neither measurement is wrong; they answer different questions.
  • The emoji inverts the trap. It is one code point and two UTF-16 code units, so JavaScript reports a length of 2 for a single glyph. Any loop that walks a string one index at a time can cut it into two halves that are not characters at all.
  • Bytes depend on the encoding as much as on the text. Those five Japanese characters are 15 bytes in UTF-8, 10 in CP932 and 10 in UTF-16, while the English word is 5, 5 and 10. A byte budget that does not name its encoding does not mean anything, and a CJK language is not uniformly more expensive than Latin — in UTF-16 the Japanese label and the English one cost exactly the same.
  • Japanese writes without spaces between words, so any wrapping, truncation or ellipsis that splits on spaces has nothing to split on. The boundaries exist as a language rule rather than as characters in the text: ICU's segmenter finds seven word-like pieces in a line containing no space. Getting breaks to land in the right places is its own problem, covered in line breaking and wrapping in CJK.
node -e 'for (const s of ["Start","ゲーム開始","👾"]) console.log(JSON.stringify(s), s.length, [...s].length, Buffer.byteLength(s,"utf8"));'
# string          UTF-16 units   code points   UTF-8 bytes
  "Start"              5              5             5
  "ゲーム開始"           5              5            15
  "👾"                  2              1             4

/usr/bin/python3 -c "
for s in ['Start','ゲーム開始']:
    print(repr(s), len(s.encode('utf-8')), len(s.encode('cp932')), len(s.encode('utf-16-le')))"
# string        UTF-8   CP932   UTF-16
  'Start'          5       5      10
  'ゲーム開始'      15      10      10

node -e 'const seg = new Intl.Segmenter("ja", { granularity: "word" });
console.log([...seg.segment("セーブデータを削除しますか")].filter(x => x.isWordLike).map(x => x.segment).join(" | "));'
# セーブ | データ | を | 削除 | し | ます | か      (no space anywhere in the input)

Numbers, dates and money reformat with nothing translated

A large share of the work contains no translation at all. Hand one number, one date and one price to four markets' formatting rules and four different strings come back with the digits unchanged. A translation file cannot fix this category, because there is no source text to replace — the strings are produced by your code at runtime.

  • The decimal point and the grouping separator trade places. The English 1,234,567.89 is 1.234.567,89 in German and in Brazilian Portuguese. Feeding that back through a naive parseFloat returns 1.234 — no error, no warning, a value wrong by six orders of magnitude. Any place where a player types a number, or where a number survives a round trip through text, has to know which locale wrote it.
  • French groups with U+202F, a narrow no-break space, rather than the ordinary space you would type. A fixture or an equality check written with a plain U+0020 will not match the output even though the two are visually identical in most fonts.
  • A short numeric date changes meaning, not only appearance. The same day prints as 3/4/26 for en-US and 04/03/2026 for en-GB, so a numeric short form shown to a mixed audience is ambiguous by construction rather than by accident.
  • Currency carries its own rounding rule. The Japanese yen has no minor unit, so 1234.5 renders as 1,235 while the identical amount is 1,234.50 in US dollars. The yen symbol that comes back is U+FFE5, the fullwidth form, not the U+00A5 you may have typed into a test. The symbol also sits after the number in German and French rather than before it.
  • What a store actually charges in each currency is a separate matter from formatting — regional prices are set in the store's own pricing tools and are not derived from a conversion in your code. Check that store's official documentation for how its regional pricing works before you display any price yourself.
node -e 'const n = 1234567.89;
for (const l of ["en-US","de-DE","fr-FR","ja-JP","pt-BR"]) console.log(l, "|", new Intl.NumberFormat(l).format(n));'
# en-US | 1,234,567.89     de-DE | 1.234.567,89     fr-FR | 1 234 567,89
# ja-JP | 1,234,567.89     pt-BR | 1.234.567,89

node -e 'const d = new Date(Date.UTC(2026,2,4));
for (const l of ["en-US","en-GB","de-DE","ja-JP"]) console.log(l, "|", new Intl.DateTimeFormat(l,{dateStyle:"short",timeZone:"UTC"}).format(d));'
# en-US | 3/4/26     en-GB | 04/03/2026     de-DE | 04.03.26     ja-JP | 2026/03/04

node -e 'for (const [l,c] of [["en-US","USD"],["de-DE","EUR"],["ja-JP","JPY"],["fr-FR","EUR"]])
console.log(l, c, "|", new Intl.NumberFormat(l,{style:"currency",currency:c}).format(1234.5));'
# en-US USD | $1,234.50     de-DE EUR | 1.234,50 €
# ja-JP JPY | ¥1,235        fr-FR EUR | 1 234,50 €   (yen sign is U+FFE5, not U+00A5)

node -e 'console.log(parseFloat("1.234.567,89"), Number("1.234.567,89"));'
# 1.234 NaN

Counting, sorting and capitalising live in your code, not in the copy

Some of a language's rules never appear in a translated file, because they are decisions your code is making on the translator's behalf. Three of them break quietly, and all three are visible in one command each.

  • English distinguishes two plural categories and Japanese one, while Polish and Russian have four and Arabic six. A label assembled as a count followed by the word items cannot be translated into those languages at all — it has to become a message with one case per category, which is the problem ICU MessageFormat was designed for.
  • The categories are not a matter of the number being large or small. In Polish and Russian 22 selects few while 5 and 0 select many. Special-casing 1 and treating everything else as the plural form is correct in English and wrong in both of them, and it is wrong in a way that only a speaker will notice.
  • Uppercasing the letter i gives I in English and a dotted capital in Turkish; lowercasing I gives a dotless one. A button style that uppercases its label, or a case-insensitive comparison written without a locale, silently turns one Turkish letter into another. If you uppercase anything for display, pass the locale.
  • Sorting is a language rule as well. The same five words come out in a different order under the German and Swedish collators, because German files the umlauted letters with their base letters and Swedish files them at the end of the alphabet. The plain code-point sort happens to match Swedish for this list and is still wrong for German — matching one locale by accident is not the same as being sorted.
node -e 'for (const l of ["en","ja","pl","ru","ar","fr"])
console.log(l, "|", new Intl.PluralRules(l).resolvedOptions().pluralCategories.join(" "));'
# en | one other            ja | other
# pl | few many one other   ru | few many one other
# ar | few many one two zero other                fr | many one other

node -e 'for (const n of [0,1,2,5,22]) console.log("pl", n, new Intl.PluralRules("pl").select(n),
"| ru", new Intl.PluralRules("ru").select(n), "| en", new Intl.PluralRules("en").select(n));'
#  0 -> pl many | ru many | en other
#  1 -> pl one  | ru one  | en one
#  2 -> pl few  | ru few  | en other
#  5 -> pl many | ru many | en other
# 22 -> pl few  | ru few  | en other

node -e 'console.log(JSON.stringify("i".toLocaleUpperCase("en")), JSON.stringify("i".toLocaleUpperCase("tr")), JSON.stringify("I".toLocaleLowerCase("tr")));'
# "I" "İ" "ı"    -- two inputs, three different letters

node -e 'const names = ["Zoe","Ärger","Apfel","Ost","Öl"];
console.log("codepoint |", [...names].sort().join(" "));
for (const l of ["de","sv"]) console.log(l, "        |", [...names].sort(new Intl.Collator(l).compare).join(" "));'
# codepoint | Apfel Ost Zoe Ärger Öl
# de        | Apfel Ärger Öl Ost Zoe
# sv        | Apfel Ost Zoe Ärger Öl

The locale tag you ship decides things you never chose

On a store page a language is one word. Inside the system it becomes a tag, and an incomplete tag gets completed for you. Asking ICU to expand four bare language subtags to their most likely full form shows exactly what a platform assumes when you do not say.

  • Ship the bare tag pt and the assumed audience is Brazil; ship zh and it is Simplified script in mainland China; ship es and it is Spain rather than any Latin American market. These completions come from CLDR's likely-subtags data — a published table of what a bare tag most probably means — and not from anything about your project. If the assumption does not match what you translated, nothing in the pipeline will tell you, because the tag is valid either way.
  • The region is not decoration. Brazilian and European Portuguese are separate locales, and the difference shows up before any word is translated: pt-BR groups digits with dots and pt-PT with a space. Picking one tag to serve both is a decision with visible consequences, not a formality.
  • The script subtag is a third axis. Simplified and Traditional Chinese are different writing systems rather than dialect labels, and zh-Hant maximizes to Taiwan while zh-Hans maximizes to mainland China. Which subtags exist and in what order they go is the grammar described in BCP 47 language tags; the data that fills in the blanks, including these defaults, is Unicode CLDR.
  • The identifiers your game engine and your store accept are lists of their own, separate from both standards and from each other. A tag can be perfectly valid BCP 47 and still not be an option in a store's language list. Check the engine's official localization documentation and the store's own supported-languages help page rather than assuming the standard is the interface.
node -e 'for (const t of ["pt","zh","es","en"]) console.log(t, "->", new Intl.Locale(t).maximize().toString());'
# pt -> pt-Latn-BR      zh -> zh-Hans-CN      es -> es-Latn-ES      en -> en-Latn-US

node -e 'for (const l of ["pt-BR","pt-PT","zh-Hans","zh-Hant"])
console.log(l, "|", new Intl.NumberFormat(l).format(1234567.89), "|", new Intl.Locale(l).maximize().toString());'
# pt-BR   | 1.234.567,89 | pt-Latn-BR
# pt-PT   | 1 234 567,89 | pt-Latn-PT
# zh-Hans | 1,234,567.89 | zh-Hans-CN
# zh-Hant | 1,234,567.89 | zh-Hant-TW

What is left is not strings at all

Length, formatting, grammar and locale tags cover the parts a program can measure. The remainder of the scope is the material a translation file cannot reach, and it is worth taking inventory of early, because its cost is counted in art and audio hours rather than in words.

  • Text drawn into images: signs inside a texture, a title rendered into a sprite sheet, a label baked into an animation or a cutscene. No string table touches any of it, and each one becomes a new asset per language.
  • Fonts. A typeface chosen for Latin text does not necessarily contain Japanese, Korean or Cyrillic glyphs. Browsers and native UI toolkits hide that behind a per-script fallback chain; a game engine draws text through its own system and typically starts with one font and an empty fallback list, so until you fill that list in the gap shows as a visible box. Pixel and hand-drawn typefaces are the hardest case, since a matching family may not exist at the size the UI was designed around.
  • Strings the code does not expose. Text written directly into a script or a scene is invisible to a translator, so finding it is a prerequisite for the whole job rather than a later cleanup — the mechanics are in externalizing hardcoded strings.
  • Voice lines, subtitle timing, age ratings and store page metadata each run on their own process, and platform requirements differ per store and per territory. Read the official submission documentation for the platforms you are actually shipping on; summaries age badly and requirements are not uniform.
  • Some of this a scan finds before anything runs: hardcoded strings turn up in a source search, missing per-language assets in a file listing. Clipped text, boxes instead of glyphs and a date in the wrong order appear only once a build runs under that locale, which is why checking a localized build is a distinct activity from reviewing the translation — see what localization QA (LQA) is.

The scope of localization, then, is not a list of languages. It is the answer to one question about your own project: how much of what this game communicates lives somewhere a translator cannot edit? Text length, number and date formatting, plural and sorting rules, and the locale tag itself are all measurable today, on the strings you already have, with the commands above. The parts that are not measurable — art, audio, store requirements — are the ones worth counting by hand before you promise a date.

Related articles