English to Japanese: what actually changes in your game text
Japanese is usually the first language that breaks an assumption your English build never had to state out loud. Two measurements set the scene. Across eight common UI strings, the Japanese wording came to 50 graphemes against English's 108, or 46 percent of the character count. Measuring rendered width in monospace columns, with each full-width character counted as two cells, the same strings came to 100 against 108, or 93 percent. So Japanese text is much shorter to store and about as wide to draw, and the shortest button labels get wider rather than narrower. The rest of this guide covers that measurement, the engine work it implies, the decisions Japanese forces you to make explicitly, and the handoff that keeps you from doing the translation twice.
Character count halves, rendered width does not
The pattern inside those totals matters more than the totals. Long sentences do shrink: the quit prompt drops from 30 columns to 22, and the key prompt from 25 to 24. Short labels move the other way. Save goes from 4 columns to 6, and both Continue and New Game go from 8 to 10. A title menu made of four-to-eight character English labels is the most likely place in your game for Japanese to overflow, which is the opposite of what the character count predicts.
Two caveats keep this honest. First, the column count assumes every Latin character occupies half of a full-width cell, and a proportional font does not behave that way. Measuring advance widths in Arial, the mean lowercase letter is 0.49 em, close to the assumption, but capitals and spaces push real labels past it: Save measures 4.6 half-cells rather than 4, and New Game measures 10.0 rather than 8, which is exactly the width of はじめから. In Verdana, whose mean lowercase advance is 0.56 em, New Game comes out wider than the Japanese. The direction of the finding holds for the shortest labels, since セーブ is 32 percent wider than Save even in Arial, but the size of the gap is font-dependent. Treat the table as a ratio to reason with rather than a substitute for measuring in your own font. Second, the wording is a choice, not a constant. Settings can be 設定 at four columns or 環境設定 at eight, and a translator will pick the one that fits your box only if you tell them what the box is.
The practical form of that is a width budget carried in the string table, expressed in whatever unit your renderer measures in, plus a review pass over the built game rather than the spreadsheet. Sorting your strings by English length and checking the shortest ones first finds more defects per minute than reading from the top.
The figures below are measured, not estimated. Grapheme counts come from a Unicode grapheme segmenter; the column count treats each full-width character as two cells and each Latin character as one, which is what a monospace terminal or a fixed-cell bitmap font does. The Japanese side uses wordings a Japanese game would plausibly ship, not literal glosses.
string graphemes columns japanese graphemes columns
Save 4 4 セーブ 3 6
Continue 8 8 つづきから 5 10
New Game 8 8 はじめから 5 10
Settings 8 8 設定 2 4
Inventory 9 9 持ち物 3 6
Not enough gold. 16 16 お金が足りません。 9 18
Press any key to continue 25 25 何かキーを押してください 12 24
Are you sure you want to quit? 30 30 ゲームを終了しますか? 11 22
----------------------------------------------------------------------------------------------------
total 108 108 50 100
of English 46% 93%Line breaking, fonts, and the encoding you inherit
Japanese does not separate words with spaces, so a renderer that only breaks on whitespace has no break opportunity at all inside a Japanese paragraph. The line runs past the edge of the box, or gets truncated at a fixed character count. Engines that instead break between any two characters then need the next rule, because unrestricted breaking is also wrong: Japanese typesetting forbids certain characters at the start of a line, including closing brackets, 、 and 。 and the small kana in ゃゅょっ, and forbids opening brackets at the end of one. Those character classes are tabulated in JIS X 4051, the Japanese standard for text composition, and an engine that ignores them will eventually show a line beginning with a lone 。. The symptom and its fix are covered in Japanese text wrapping in the wrong place, and the underlying algorithm in line breaking for CJK text.
Fonts are the other prerequisite, and the coverage question has concrete numbers behind it. The JIS X 0208 code table holds 2,965 kanji in level 1 and 3,390 in level 2, 6,355 in total, while the modern subset most fonts aim at is the 2,136 characters of the 常用漢字表 published as a cabinet notification. Those are very different targets. A font limited to the common-use set will render your menus perfectly and then fail on a character name, because personal-name kanji largely live outside it. A missing glyph shows as a fallback box, or silently swaps in a system font with different metrics, which is why empty boxes are a font problem rather than an encoding problem. For low-resolution art, holding kanji and English in one pixel font adds a legibility floor of its own, since kanji need more vertical pixels than Latin letters to stay readable at all.
If any part of your pipeline predates Unicode you also inherit an encoding limit. CP932, the Microsoft extension of Shift_JIS, encodes 6,716 CJK ideographs. That is more than JIS X 0208 but still a closed set: characters outside it, 𠮟 and 𩸽 among them, cannot be represented at all, and an export step that quietly substitutes them corrupts names instead of reporting an error. Keep the whole text path in UTF-8 and convert only at the boundary, with the difference between Shift_JIS and CP932 in mind, because the two names get used interchangeably and do not describe the same mapping.
Placeholders, word order, and how Japanese counts things
Japanese puts the verb at the end of the clause, so a message assembled as a fixed prefix plus a value plus a fixed suffix will not survive translation with its parts in that order. Named placeholders the translator can move are the requirement; string concatenation in code is the thing that prevents it.
Japanese also has no grammatical plural that agrees with a number. Asking a runtime for the plural categories of Japanese returns exactly one, other, where English returns one and other. A Japanese message therefore needs a single form, and a translation format that demands separate singular and plural entries will simply receive the same sentence twice. That is worth knowing before you design the message format, and it is one of the cases ICU MessageFormat exists to express without forking the string per language.
What Japanese has instead is counters. A number is followed by a word chosen for the kind or shape of the thing counted: 本 for long thin objects and bottles, 枚 for flat things, 人 for people, 匹 for small animals, and 個 across a wide range of generic objects. A single template on the English pattern of {count} items cannot be reused across item types, so either carry the counter alongside the item definition or write the message per type. Guessing produces text that is grammatical and still reads as foreign.
Digit grouping matches English, but abbreviated magnitudes do not. Japanese groups large numbers in units of ten thousand, so compact formatting of 12,345,678 yields 1235万 in Japanese against 12M in English. A hand-rolled helper that divides by a thousand and appends K hands a Japanese player a number they have to convert in their head. Keep numerals half-width inside Japanese sentences, and leave the grouping to the platform's locale-aware number formatter.
en "You found {item}!" ja "{item}を手に入れた!"
en "{count} potions left" ja "ポーション残り{count}本"
en "{name} joined the party." ja "{name}が仲間になった。"The decisions Japanese forces and English let you skip
Japanese picks a register per sentence and holds it. The polite です・ます style and the plain だ・である style are both correct, but mixing them inside one screen reads as carelessness. Interface and system text is conventionally polite while narration often runs plain, and whichever you choose belongs in a style guide written before translation starts, because changing it afterwards is a rewrite rather than an edit.
Characters need more than that. English I is one word; Japanese makes the speaker choose among several first-person pronouns, and the choice carries age, gender, formality and self-image, as do sentence-final particles and verb endings. No translator can infer this from a single line of English dialogue. A character sheet listing each speaker's pronoun, politeness level, and one sample line in their voice is the highest-value document you can supply, and it is far cheaper to write than to reverse out of a finished translation.
Loanwords are a project decision rather than a lookup. セーブ or 保存, アイテム or 道具, ダメージ or 損傷 are all defensible; the only wrong answer is using two of them for one concept on different screens. Fix each in a glossary. Punctuation follows the same logic with firmer conventions: 、 and 。 replace the comma and period, 「」 marks speech and 『』 a title or a quote inside a quote, and an ellipsis uses the three-dot leader …, conventionally doubled. ASCII commas and periods inside Japanese prose read to a Japanese player the way a missing space after a period reads in English. Numerals and Latin abbreviations, by contrast, normally stay half-width in a Japanese sentence, so the rule is per character class and not global.
Two more decisions are layout rather than text. If your audience includes children, decide where furigana, the small kana printed above a kanji to give its reading, are needed, and budget line height for them. And vertical setting, with columns running top to bottom and right to left, is common in print but unusual in game interfaces. Decide early if you want it, because it changes the text engine rather than the strings: punctuation, brackets and small kana all have to rotate or reposition.
What to hand over, and what only the build will tell you
The glossary, the character sheet and the width budget go out before the first line is translated, not after the first review. So does the context attached to each string, which is where most avoidable rework originates. Per string, supply:
- The speaker, keyed to their entry in the character sheet
- Register: polite or plain, and whether the string is interface text or spoken dialogue
- The screen or scene, so the translator can picture where it appears
- Maximum width in the unit your renderer measures, plus how many lines the box allows
- Every placeholder, with what it holds and a realistic sample value
- Whether the string may wrap at all, and whether it is reused anywhere else
Failures that reach players
Assembling that kit before you send the text is what turns a translation into a localization. Then play the build in Japanese. The spreadsheet catches wording; only the running game catches a clipped label, a line that breaks before 。, a name rendered in a fallback font, or a register that shifts between two adjacent menus. That pass is what localization QA means, and it needs a native reader with the build in their hands rather than a proofread of the source file. Keep the roles separate: one person translates, a second reviews the Japanese, and the developer checks integration. A developer who cannot read Japanese can still catch every failure below, which is why the integration check deserves its own slot in the schedule.
- Machine-translated strings pasted in one at a time, so sentence endings alternate between です・ます and だ inside a single menu
- Clipped or overflowing labels in exactly the places the source string looked shortest
- A fallback box in a character's name, because the font covers the common-use kanji set and not name kanji
- ASCII commas, periods and exclamation marks inside otherwise full-width prose
- Half-width katakana used to save space, which reads as a legacy system rather than a design choice
- A fully translated game behind a store page, screenshots and trailer captions still in English