CJK line breaking: how Japanese, Chinese and Korean differ

A layout engine does not break lines at words. It breaks them at break opportunities, and every writing system supplies those differently. Japanese and Chinese offer one between almost every pair of characters. Korean is written with spaces between word units, so the spaces are the opportunities. Thai has no interword spaces at all, so a renderer without a Thai dictionary cannot wrap Thai text. The setting that looks correct in Japanese breaks Korean mid-word and pushes Thai out of its box.

What decides all of this is the Unicode line breaking algorithm, UAX #14. Every character carries a line break class, and each boundary is decided from the pair of classes around it. Japanese prohibition rules, the integrity of Korean syllables and Thai dictionary lookup are all expressed inside that one mechanism. This article compares what the four languages actually need, shows which parts of the mechanism are tailorable, and gives you checks you can run on strings before anything renders. For the Japanese prohibition rules themselves, see why Japanese line breaks come out wrong.

Break opportunities are not word boundaries

Start by separating the two ideas. Where a word ends and where a line may break are different questions. Running the same short instruction in four languages through the word segmenter that browsers and Node.js expose shows each language's assumption directly.

  • Japanese and Chinese are segmented by dictionary, with no spaces to fall back on. The segmentation is also not clean: the run above splits ください into くだ and さい, and 開いて into 開, い and て. Never use a word segmenter's output as your list of break positions. Deciding where a line may break is a different algorithm's job.
  • Korean is spaced. Four word units come back, and the three spaces between them are returned as their own segments. Those spaces are where the language wants the line to break.
  • Thai has zero spaces, and the 27 characters resolve into 7 units only because a dictionary said so. Without one there is not a single break position to be found.
Intl.Segmenter(locale, { granularity: "word" })
run on Node.js 24.18.1 / ICU 78.3

ja  設定を開いて言語を選んでください         16 chars / 11 word units
    設定 / を / 開 / い / て / 言語 / を / 選 / んで / くだ / さい
zh  请打开设置并选择语言                     10 chars /  6 word units
    请 / 打开 / 设置 / 并 / 选择 / 语言
ko  설정을 열고 언어를 선택하세요             16 chars /  4 word units
    설정을 / 열고 / 언어를 / 선택하세요   (3 spaces, returned as separators)
th  เปิดการตั้งค่าแล้วเลือกภาษา              27 chars /  7 word units
    เปิด / การ / ตั้ง / ค่า / แล้ว / เลือก / ภาษา   (no spaces at all)

How implementations decide: the UAX #14 pair table

In UAX #14 every character carries a Line_Break property. Ideographs are ID, closing punctuation is CL or CP, opening punctuation is OP, characters that may not start a line are NS, exclamation and question marks are EX, space is SP, non-breaking glue is GL, Hangul jamo are JL, JV and JT, and composed Hangul syllables are H2 and H3. The algorithm itself is defined as numbered rules applied in order, with a pair table summarising them. Its cells distinguish a break that is allowed, one allowed only where a space intervenes, and one that is prohibited.

Japanese prohibition rules fall straight out of that table. ID followed by CL, CP, EX or NS is a prohibited pair, so punctuation and closing brackets cannot land at the start of a line. Nothing may break after OP, so an opening bracket cannot be left stranded at the end of one. There is no separate prohibition engine running afterwards; the table simply says so.

The tailorable part is explicit too. Small kana and the prolonged sound mark carry the class CJ, Conditional Japanese Starter. It resolves to NS by default, keeping those characters off the start of a line, and an implementation may instead resolve it to ID, which allows them there. In CSS that choice is the line-break property, and only loose takes the permissive side. Both normal and strict prohibit a break before small kana and the prolonged sound mark; what separates those two is whether a break may fall before CJK hyphen-like characters such as U+301C and U+30A0, which strict prohibits. The initial value is auto, which leaves the restriction set to the user agent and lets it vary with line length. Confirm the classes in the line break class and pair table sections of UAX #14, and the property values in the line-break section of CSS Text Module Level 3.

Width properties are worth reading at the same time. The values below were confirmed with Python's unicodedata module (UCD 13.0.0). W and F both occupy two columns, so they cost the same when you work out how much fits on a line. A is the one that hurts.

  • An ambiguous-width character is one column in some environments and two in others. The horizontal ellipsis is one, and so are the quotation marks used in simplified Chinese. Any UI that computes line width on a fixed column assumption is off by a column wherever they appear, and the tail of the line overflows. Character width in bitmap faces is covered in pixel fonts for Japanese and English.
  • Line break class and width are separate properties. Being two columns wide is not what makes a character breakable on either side. Read Line_Break to know where a break may fall, East_Asian_Width to know how much fits.
U+3001 IDEOGRAPHIC COMMA                      Po  East_Asian_Width=W
U+3002 IDEOGRAPHIC FULL STOP                  Po  W
U+FF0C FULLWIDTH COMMA                        Po  F
U+300C LEFT CORNER BRACKET                    Ps  W
U+300D RIGHT CORNER BRACKET                   Pe  W
U+30FC KATAKANA-HIRAGANA PROLONGED SOUND MARK Lm  W
U+3063 HIRAGANA LETTER SMALL TU               Lo  W
U+301C WAVE DASH                              Pd  W    <- where normal and strict differ
U+30A0 KATAKANA-HIRAGANA DOUBLE HYPHEN        Pd  W    <- where normal and strict differ
U+002E FULL STOP (ASCII)                      Po  Na
U+2026 HORIZONTAL ELLIPSIS                    Po  A    <- ambiguous width
U+201C LEFT DOUBLE QUOTATION MARK             Pi  A    <- same

Chinese wraps like Japanese; the punctuation is what differs

Han characters are class ID in both simplified and traditional Chinese, so wrapping behaves essentially as it does in Japanese. If per-character wrapping is correct for Japanese in your renderer, it is correct for Chinese. What differs is not the breaking but the characters you are given and how they are drawn.

  • The comma is a different character. Japanese uses U+3001, while both simplified and traditional Chinese use U+FF0C. The full stop U+3002 is shared. A Chinese file that arrives carrying Japanese commas wraps identically, so nothing looks wrong until someone reads the screen. The only defence is deciding which punctuation each language may use and checking for the rest.
  • Taiwanese typography centres the full stop and comma in the em box, where Japanese and mainland Chinese place them at the bottom left. The code point is the same; which glyph you get comes from the locale forms in the font and from the language you tag the text with. Check your font's own documentation for the locale forms it carries, and see how BCP 47 language tags are built for the tagging side.
  • Runs of Latin letters and digits inside Chinese text (a product name, a version number, a figure) have to stay whole. A blanket setting that allows a break between any two characters splits those runs too.
Comma, full stop and quotation marks (code points confirmed)

ja        、U+3001   。U+3002   「」U+300C U+300D
zh-Hans   ,U+FF0C   。U+3002   “”U+201C U+201D
zh-Hant   ,U+FF0C   。U+3002   「」U+300C U+300D
ko         ,U+002C    .U+002E   (Latin punctuation, words spaced)

Korean breaks at spaces, yet the defaults break it mid-word

Korean separates word units with spaces, and as the segmentation above shows, those units come out cleanly. Under the CSS default of word-break: normal, however, Hangul is treated as CJK and breaks are allowed between syllables as well. So the language offers a correct break position, the space, and the layout breaks in the middle of a word anyway. Korean readers see it immediately.

The property that stops it is word-break: keep-all, which restricts breaking to spaces and explicit break positions. It has a side effect: a long unspaced run, a URL-like string or a long proper noun, then cannot break at all and overflows its container. Pair keep-all with overflow-wrap, set to break-word or anywhere, as the emergency valve. Confirm the exact definitions in the word-break and overflow-wrap sections of CSS Text Module Level 3.

There is a second trap that hand-rolled wrapping code walks into. Hangul syllables are normally single composed characters, but NFD normalisation decomposes them into jamo.

  • Wrapping code that slices by code point splits a syllable of NFD Korean into its jamo. What renders is not 한 but its decomposed parts, with a bare consonant left at the start of a line.
  • Counting grapheme clusters keeps the syllable intact in both forms, as the run above shows. UAX #14 also prohibits breaking inside a jamo sequence, so a conforming implementation is safe. The risk sits in hand-written code that cuts strings by UTF-16 unit or code point.
  • NFD Korean does turn up in practice, for instance in files that came through macOS. Normalising to NFC on import settles it. The wider problem of code points not matching what a reader sees appears again in emoji and surrogate pairs in game text.
"한글"
  NFC: 2 code points  U+D55C U+AE00
  NFD: 6 code points  U+1112 U+1161 U+11AB U+1100 U+1173 U+11AF
                      (CHOSEONG HIEUH / JUNGSEONG A / JONGSEONG NIEUN ...)
  graphemes (Intl.Segmenter granularity "grapheme", ko): 2 in both forms

Thai and other spaceless scripts need a dictionary

Thai has no interword spaces; a space marks a phrase or sentence boundary instead. The example above is 27 code points with zero spaces, six of them nonspacing marks stacked above or below, which counts as 21 grapheme clusters. UAX #14 assigns Thai, Lao, Khmer and Myanmar characters the class SA, complex context, and states that resolving them needs a dictionary or morphological analysis. The algorithm on its own cannot wrap Thai.

  • A renderer with no Thai dictionary either treats the text as one long unbreakable run and never wraps, or breaks it at arbitrary positions. Both are obvious on screen. Check whether your text system does Thai line breaking in the line breaking or text shaping section of its official documentation; systems built on ICU generally do.
  • Do not substitute break-anywhere behaviour for the dictionary. It wraps, but it wraps mid-word, and the result is not readable Thai.
  • Watch line height as well as width. Marks stack above and below the base characters, so a line box sized for Latin text clips the upper tone marks. If characters render as empty rectangles instead, that is font coverage rather than breaking: see tofu boxes and missing glyphs.
เปิดการตั้งค่าแล้วเลือกภาษา

  code points 27 / spaces 0 / nonspacing marks (category Mn) 6 / graphemes 21
  the marks: U+0E34 U+0E31 U+0E49 U+0E48 U+0E49 U+0E37

What to set per language, and what to check before release

On the settings side, three properties divide the work. word-break decides whether a break may fall inside a word (normal, break-all, keep-all). line-break decides whether a break may fall before small kana and similar characters (auto, loose, normal, strict). overflow-wrap (break-word, anywhere) is the last resort when nothing fits. Their jobs are separate, so none of them substitutes for the others. Game engines have their own settings rather than CSS: Unity's text rendering (TextMesh Pro) and Unreal Engine's text display each expose line breaking and CJK handling options, so confirm them in the relevant section of the official documentation rather than in a forum answer.

Strings that still carry a hand-placed newline from the source language can be checked before anything renders. The expressions below were run, and the results shown are the ones they produced.

  • Put halfwidth brackets and punctuation in the character class as well as fullwidth ones. Halfwidth parentheses inside Japanese text are common, and a class listing only fullwidth forms lets them through. The first version of the expression above did exactly that and missed the halfwidth example.
  • The Korean check can only tell you that a newline sits between two Hangul syllables with no space. It cannot distinguish a break placed mid-word from a break that replaced the space. Both come from a newline typed into the source string, so flagging both for review is right, but neither is proof of a bug on its own.
  • The stronger fix is upstream: do not embed newlines in translatable strings at all. A break position chosen for the source word order and sentence length means nothing after translation. What to agree on when handing strings over is covered in putting a localization kit together.
  • On the device, look at the longest translation in each language in the narrowest layout you ship: a handheld, a phone held upright, a small dialogue box. Wrapping failures depend on width and do not appear on a wide screen.
  • Keep one Korean string containing a long unspaced run, such as a long proper noun or a code, for the overflow that keep-all introduces. Show Thai in a fixed-height box and check that the marks above and below the base characters are not clipped.
  • Confirm that the language tag on the text reaches the renderer. Traditional Chinese punctuation position and the CJ resolution both depend on it, so a dropped tag silently disables the settings you chose. How scripts and regions are identified leads into what Unicode and CLDR each cover.
Per language
  ja / zh  leave word-break at the default; line-break: strict for the
           conservative prohibition set
  ko       word-break: keep-all plus overflow-wrap for unspaced runs
  th       a text system with a dictionary; no property setting substitutes

Static check on source strings (run, with results)
  const startForbidden = /\n[、。,.,.))」』】〕}]!!??:;:;ー・…っゃゅょ]/u;
  const endForbidden   = /[((「『【〔{[]\n/u;
  const korNoSpace     = /[가-힣]\n[가-힣]/u;

  "セーブデータを\n、上書きします。"      startForbidden: match
  "ロードしています\nっと…"              startForbidden: match
  "確認してください(\nはい/いいえ)"      endForbidden:   match (halfwidth)
  "確認してください(\nはい/いいえ)"    endForbidden:   match (fullwidth)
  "セーブデータを\n上書きします。"        no match
  "언어를\n선택하세요"                    korNoSpace:     match
  "설정을 열고 \n언어를 선택하세요"        korNoSpace:     no match

Related articles