Japanese text wraps in the wrong place — a line starts with 、or 」
A Japanese text box wraps and something reads as wrong even to someone who cannot say why: a line starts with a lone 、 or 。, a closing bracket 」 sits by itself at the start of a line, or a small kana like っ or ょ opens a line on its own. Nothing is mistranslated and nothing is misspelled. The line break itself landed in a position Japanese typography forbids.
The rule being broken has a name — kinsoku shori (禁則処理), the line-breaking prohibitions — and it is specific enough to check mechanically. Three different faults produce nearly identical screenshots: a renderer with no prohibition table at all, a renderer whose table is switched to its permissive setting, and a newline character hard-coded into the source string and carried through translation. Each has a different fix, so the first job is telling them apart.
One property of this bug matters more than any of the three causes, and it is why these reports get closed as unreproducible. Whether a string violates the rules depends on the exact number of characters that fit on a line. The same sentence can be clean at ten characters per line, broken at twelve, and clean again at thirteen. A wide debug view proves nothing about the narrow box that ships.
Which characters may not start a line, and which may not end one
Kinsoku shori is usually described as two rules. It is really three. Line-start prohibition (行頭禁則) forbids a set of characters from beginning a line; the break has to move earlier so they stay attached to the previous line. Line-end prohibition (行末禁則) forbids opening brackets from ending a line, because an opening bracket separated from what it opens reads as an error. No-separation (分離禁止) covers character pairs that must not be split from each other at all, such as a two-character ellipsis or a two-character dash.
The authoritative enumeration is the Japanese Industrial Standard JIS X 4051, the standard for line composition of Japanese documents, in its clause on line-breaking rules. The sets tabulated at the end of this section are representative, not exhaustive, and implementations differ at the edges — particularly over whether a full-width space and the middle dot belong in the line-start set.
Within that line-start set, the brackets can be computed and the small kana cannot. Closing brackets carry the Unicode general category Close_Punctuation and opening brackets Open_Punctuation, so a renderer identifies both without a table. But the small tsu っ is U+3063 and the ordinary tsu つ is U+3064: both have the category Other_Letter, both have the Alphabetic property, and no Unicode property exposed to a regular expression in a browser separates them. Nor does any other property help: Modifier_Letter and Extender select ー and the iteration mark 々 but never match a small kana, which is Other_Letter without Extender exactly like ordinary kana.
So the small-kana prohibition can only come from an explicit list of code points that somebody typed in. That is why a renderer handles 、 and 」 perfectly and still lets っ open a line. When you file this bug, naming which group is affected tells the engineer whether the fix is a setting or a missing table entry. For why word-based wrapping cannot be adapted to Japanese in the first place, see line breaking and wrapping for CJK text.
Line-start prohibition (must not begin a line) — representative set
punctuation 、 U+3001 。 U+3002 , U+FF0C . U+FF0E
closing brackets 」 U+300D 』 U+300F ) U+FF09 〕 U+3015
】 U+3011 } U+FF5D 〉 U+3009 》 U+300B
separators ・ U+30FB : U+FF1A ; U+FF1B
sentence enders ! U+FF01 ? U+FF1F
small kana っ U+3063 ゃ U+3083 ゅ U+3085 ょ U+3087
ぁ U+3041 ぃ U+3043 ぅ U+3045 ぇ U+3047
ぉ U+3049 ゎ U+308E
marks ー U+30FC 々 U+3005 ゛ U+309B ゜ U+309C
Line-end prohibition (must not end a line) — representative set
opening brackets 「 U+300C 『 U+300E ( U+FF08 〔 U+3014
【 U+3010 { U+FF5B 〈 U+3008 《 U+300A
No-separation (the pair stays together)
…… (U+2026 twice) ‥‥ (U+2025 twice) ―― (U+2015 twice)Why the same sentence is correct at one width and wrong at another
The experiment that makes this bug tractable: take a Japanese string, wrap it at a fixed number of characters per line, flag any line whose first character is in the line-start set, and repeat for every width from 8 to 24. Three sample strings wrapped naively, with no prohibition handling at all, produce the table at the end of this section.
The result worth sitting with is sample C. It is an ordinary system message, and with no prohibition handling whatsoever it renders correctly at fifteen of the seventeen widths tested. Two widths break it. If your text box happens to fit nine or nineteen characters you ship a visible typographic error; at every other width in that range you see nothing wrong and conclude the renderer is fine. Sample B is clean at thirteen of seventeen and fails at twelve. So a clean visual pass is weak evidence, and a violation you cannot reproduce is usually a width difference rather than a flake.
The effective width is also not a number you set anywhere. It is the box width divided by the advance width of the glyphs, so a font substitution or a point-size change on one platform silently shifts every break position in the game. The arithmetic of how many characters actually fit is worked through in pixel fonts for Japanese and English, and a font that silently substitutes also gives you missing-glyph tofu boxes.
Naive fixed-width wrap, no prohibition handling, widths 8-24
A 43 chars 「セーブデータを上書きしますか?」とっさに判断できないので、もう一度確認してください。
violates at widths 8, 9, 14, 15, 16, 18, 21 clean at 10 of 17 widths
w=14 -> line 4 starts with 。
w=15 -> line 2 starts with ?
w=16 -> line 2 starts with 」
B 32 chars ここは危険だ、先に進むなら回復アイテムを持っていったほうがいい。
violates at widths 8, 12, 21, 24 clean at 13 of 17 widths
w=12 -> ここは危険だ、先に進むな / ら回復アイテムを持ってい / ったほうがいい。
line 3 starts with っ
w=13 -> clean
C 32 chars 所持金が足りません。装備を売却してから、もう一度お試しください。
violates at widths 9 and 19 only clean at 15 of 17 widthsWhere a renderer is allowed to break Japanese at all
It is tempting to conclude that Japanese, having no spaces, gives a wrapping algorithm nothing to work with. The opposite is true, and the distinction changes how you read the bug. Under the Unicode line-breaking model a break is permitted between most ideographs and kana, so a Japanese string offers a break opportunity at very nearly every character boundary. Kinsoku shori is the subtraction step that removes the forbidden ones. A renderer that breaks Japanese at a random-looking position has not failed to find a break point; it has failed to exclude one.
Word boundaries are a separate question from break opportunities. Word segmentation for Japanese is dictionary-based, and browsers expose it directly. Running the standard segmentation interface over a save-overwrite prompt at word granularity returns the eight segments shown at the end of this section, next to the English equivalent split on spaces.
In that output the dictionary keeps とっさ intact, so its internal っ never lands at a boundary. Breaking only at these positions would give far fewer and far more natural line breaks than breaking between arbitrary characters, which is roughly what phrase-aware wrapping does. But this is word segmentation, not line breaking: they are separate algorithms in the same internationalisation libraries, and segmenting words does not by itself prohibit anything. Correctness comes from the prohibition rules; phrase-aware wrapping on top is polish, worth requesting for headings and short labels.
Intl.Segmenter("ja", { granularity: "word" })
input セーブデータを上書きしますか? (15 characters)
output セーブ / データ / を / 上書き / し / ます / か / ?
input 「セーブデータを上書きしますか?」とっさに判断できない。
output 「 / セーブ / データ / を / 上書き / し / ます / か / ? / 」
/ とっさ / に / 判断 / でき / ない / 。
English, "Do you want to overwrite the save data?"
split on spaces Do / you / want / to / overwrite / the / save / data?The setting to change, by implementation
In Unity's TextMesh Pro the prohibition sets are project assets, not hard-coded behaviour. Its settings expose line-breaking character files: one listing characters that may not be left at the end of a line, one listing characters that may not begin a line. An incomplete file is a plausible cause of a partial failure. Confirm the current field names and file locations in the official TextMesh Pro documentation's settings section before editing, because these have moved between versions.
Unreal Engine bundles the International Components for Unicode library for internationalisation, and its text layout uses ICU break iteration to decide line breaks, so Japanese prohibition handling is largely inherited rather than configured per project. If breaks are wrong there, suspect the text's culture or locale settings and any custom text layout before suspecting the break algorithm; the official Unreal localization documentation is where to confirm what applies to your version.
RPG Maker MV and MZ are a different situation entirely: the default message window does not wrap automatically, so breaks are authored by hand and there is no algorithm to configure. Kinsoku shori becomes an authoring responsibility, and an import responsibility too — a script that replaces Japanese text with a translation without re-authoring the break positions produces violations mechanically. If automatic wrapping came from a plugin, the prohibition behaviour is that plugin's, so verify it. The surrounding constraints are in the RPG Maker localization guide.
In CSS the property is line-break, and its initial value is auto: the specification lets the user agent pick the restrictions and vary them by line length, explicitly permitting a looser set for short lines. Of the named values, only loose allows a break before small kana and the prolonged sound mark; normal and strict both forbid it, along with breaks before iteration marks and inside inseparable pairs like the two-character ellipsis. Closing punctuation and closing brackets never begin a line under these; only the separate anywhere value allows breaks regardless. So a line starting with 」 is a rendering or markup fault, while a line starting with っ means either loose is set or auto has chosen a looser ruleset — most likely in the narrow boxes a game UI uses. Pin an explicit value instead of leaving it at auto; strict is the most stringent, and normal is equivalent for these groups.
Two neighbouring properties get reached for and do not help: overflow-wrap acts only when an unbreakable run would overflow, and word-break with keep-all suppresses breaks between letters, so only a Japanese run containing no punctuation overflows rather than wrapping. Check the current definitions in the CSS Text specification's line-breaking section.
.jp-text {
/* pin these; line-break's initial auto lets the UA relax on short lines */
line-break: strict; /* forbids a break before small kana and ー, as normal does */
word-break: normal; /* keep-all suppresses breaks between letters */
overflow-wrap: normal; /* break-word does nothing here */
white-space: normal; /* pre / pre-wrap would preserve source newlines */
}Hard newlines in translated strings, measured
The third cause is a newline character typed into the source string to make the layout look right in the original language. It survives translation unchanged, and because a newline is an absolute position while wrapping is relative to the box, the two go out of sync the moment anything changes. The measurement at the end of this section takes a 21-character Japanese line with a newline placed by hand for a ten-character box, and lays the same text out without it.
At the width the break was authored for, both versions are identical and the newline looks harmless. One character narrower, the authored break stops being the widest fit and the remainder spills onto a line of its own: a one-character orphan at width 9, two characters at width 8. The line count goes from three to four at every narrower width, which is how an authored break turns into text overflowing a fixed-height window.
So keep newline characters out of strings that get translated, and let the layout wrap them with prohibition handling enabled. If a break is genuinely required — a two-line title, a deliberately staged line of dialogue — store it as per-locale data so each language gets its own position. If the real constraint is that the text must fit, express it as a character budget for the translator and leave the breaking to the layout, because Japanese and English differ in length in both directions depending on the sentence, as English to Japanese game localization goes into.
text 危険だ、ここから先は一人で行くしかないぞ。 (21 chars)
manual break placed after ここから先は (character 10)
box width 10 no manual break: 3 lines with manual break: 3 lines
box width 9 no manual break: 3 lines with manual break: 4 lines
危険だ、ここから先 / は / 一人で行くしかない / ぞ。
box width 8 no manual break: 3 lines with manual break: 4 lines
危険だ、ここから / 先は / 一人で行くしかな / いぞ。
box width 7 no manual break: 3 lines with manual break: 4 lines
危険だ、ここか / ら先は / 一人で行くしか / ないぞ。A check you can run, and what only the device shows
The detection itself is one regular expression applied to already-wrapped lines: build a character class from the line-start set and test each rendered line against it. The important constraint is that this needs rendered lines, not source strings. Line breaks exist only after layout, so the check belongs in a debug overlay that dumps wrapped lines, or in a screenshot comparison, not in a pass over your translation spreadsheet.
What you can check statically: translated strings containing newline characters, whether the prohibition setting is enabled at all on each rendering surface, and whether the character files contain the full small-kana set.
What no static check reaches is runtime substitution. A player name, an item name, or a number inserted into a template changes the string's length, and therefore every break position after it, so a message that is clean with a three-character name can violate the rules with a six-character one. Test templates with the longest plausible substitution rather than the placeholder. Where this sits in a wider review process is described in what localization QA is.
- Run the check over wrapped lines at the narrowest box in the game, then at each width the shipping UI actually uses
- Check the three groups separately: closing punctuation, closing brackets, and small kana plus ー — a pass on the first two says nothing about the third
- Confirm opening brackets are not left at line ends, which is the same bug seen from the other side
- Substitute the longest plausible player name, item name and number into every template before judging it clean
- Re-check after any font, point size or box size change, and test with the shipping font on each target platform — all of these move the effective column count