QA & troubleshootingこの記事を日本語で読む

Why emoji break string length: surrogate pairs in game text

A player types one emoji into a name field and something downstream goes wrong. The name comes back with a replacement glyph where the emoji was, or the emoji is simply gone. A limit of sixteen characters rejects a name that shows eight. A family emoji sent in chat arrives as three separate people standing in a row. None of this is a font problem and none of it is a database encoding problem. Somewhere in the path, code is counting or cutting the string in a unit that is not the unit the player sees.

Four units are in play, and for emoji they all disagree: the UTF-16 code unit, the Unicode code point, the grapheme cluster, and the UTF-8 byte. Every bug in this family is one function assuming that two of those four are the same number. What follows measures all four for eight real strings, shows seven specific ways a careless cut breaks them, gives the per-language table of what length actually returns, and ends with the decision you have to make about how much emoji to accept.

One emoji, measured in four disagreeing units

Before anything else the four units have to be separated, because the bug always lives in the gap between two of them.

  • Code point: the number Unicode assigns to a character, written U+ plus hexadecimal digits. Everything above U+FFFF is the supplementary planes, which is where most emoji live.
  • UTF-16 code unit: the 16-bit slot that JavaScript, Java and C# use internally. A code point above U+FFFF takes two of them, and that two-unit form is a surrogate pair. Indexing and length in those languages count these slots.
  • Grapheme cluster: what a reader calls one character. It can be several code points, and Unicode's text segmentation rules define where the boundaries fall.
  • UTF-8 byte: what goes on the wire and into most files, at one to four bytes per code point.
string                            .length  code points  graphemes  UTF-8 bytes
"A"                                     1            1          1            1
"あ"  U+3042                             1            1          1            3
"😀"  U+1F600                            2            1          1            4
"👍🏽"  U+1F44D U+1F3FD                    4            2          1            8
"👨‍👩‍👧"  man ZWJ woman ZWJ girl             8            5          1           18
"🇯🇵"  U+1F1EF U+1F1F5                    4            2          1            8
"é"  U+00E9 (precomposed)               1            1          1            2
"é"  U+0065 U+0301 (e + accent)         2            2          1            3

// Node 24, in column order: s.length, [...s].length,
// [...new Intl.Segmenter(undefined, { granularity: "grapheme" }).segment(s)].length,
// Buffer.byteLength(s, "utf8")

What the numbers mean, and seven ways a cut breaks them

The grapheme column is the only one that matches what a player would count, and the only one that is 1 for every row. Everything else swings by a factor of eighteen. Two rows deserve attention even if your game never allows emoji. The hiragana row is one code unit but three UTF-8 bytes, so a byte limit and a character limit already disagree with no emoji anywhere. The two accented rows are the same visible word spelled two ways, as one precomposed code point or as a plain e followed by a combining acute accent. They are different strings with different lengths, and a name typed on one input method will not match the same name typed on another unless you normalise first.

The family row is the extreme case. Eight UTF-16 code units for one visible character means a field limited to sixteen code units holds exactly two family emoji, where the same limit holds sixteen letters. If your character limit exists for layout reasons, it has stopped describing the layout. The boundary rules that produce the grapheme column, and the data that lists which emoji sequences exist, are published by the Unicode Consortium; what Unicode and CLDR actually provide covers where those tables come from and why two library versions can segment the same string differently.

The block below is real Node 24 output rather than illustration. Seven distinct failures, each one shipped by somebody.

// 1. Cutting inside a surrogate pair
"😀".slice(0, 1)         // "\ud83d" — one half, length 1
"😀".substring(0, 1)     // the same half
"😀".charAt(0)           // "\ud83d", charCodeAt(0) -> 55357
"😀".codePointAt(0)      // 128512 = U+1F600 — reads the whole pair

// 2. The half does not survive a trip through UTF-8
Buffer.from("😀".slice(0, 1), "utf8")  // ef bf bd -> U+FFFD, the emoji is gone
"😀".slice(0, 1).isWellFormed()        // false

// 3. Cutting between two flags produces a third country
"🇬🇧🇷🇺".length           // 8 — two flags, four regional indicators
"🇬🇧🇷🇺".slice(2, 6)      // "🇧🇷" — B + R, which is the flag of Brazil
"🇯🇵".slice(0, 2)         // "🇯" — one regional indicator, a boxed letter in most fonts

// 4. Cutting a modifier loose
"👍🏽".slice(2, 4)         // "🏽" — the modifier alone, a colour swatch in most fonts
[..."👍🏽"].reverse().join("")   // "🏽👍" — 2 grapheme clusters now, not 1

// 5. Reversing
"a😀b".split("").reverse().join("")  // "b\ude00\ud83da" — both halves broken
[..."a😀b"].reverse().join("")       // "b😀a" — safe for a plain pair
[..."👨‍👩‍👧"].reverse().join("")        // "👧‍👩‍👨" — one cluster still, different family

// 6. An orphaned combining mark attaches to whatever precedes it
"café".normalize("NFD").slice(4)  // "\u0301" — the accent, alone
"Name" + "\u0301"                 // renders as "Namé" — it lands on the e of Name

// 7. The regex dot counts code units unless you set the u flag
"😀".match(/./g).length   // 2
"😀".match(/./gu).length  // 1
"😀".replace(/./g, "*")   // "**"
"😀".replace(/./gu, "*")  // "*"

Three of those failures are worse than a broken glyph

The flag case is the one to remember, because it is silent and it is wrong rather than broken. A flag emoji is two regional indicator symbols, one per letter of a region code. Cut a string between two flags and the surviving indicators pair up with each other: the United Kingdom followed by Russia, cut at code unit 2, becomes the flag of Brazil. There is no error, no replacement glyph and no broken shape. A truncated chat line or a trimmed player name simply shows a country nobody typed.

The disappearing emoji is the second. A lone surrogate is not a valid character, so encoding it to UTF-8 substitutes something else, and what you get depends on the runtime. Node and .NET produce the three bytes of U+FFFD, the replacement character that renders as a dark diamond with a question mark inside. Java produces a single ASCII question mark. Either way the original emoji is unrecoverable the moment the string is written to storage or sent over the network, even if the cut that caused it was a display-only truncation upstream.

The third is the combining mark, and it is the failure that never looks like a Unicode problem. Split a decomposed accented character and the accent has no base left, so it attaches to whatever character happens to precede it in its new context. Join a truncated chunk onto a label and the accent lands on the last letter of the label. The symptom is a stray accent on an unrelated word, several fields away from the code that cut the string.

What length returns in each language, and what your column stores

The unit your code counts is decided by the language, not by you, and moving a limit from client to server, or from gameplay code to a web dashboard, usually changes the unit silently. The table below measures the same eight strings in six languages.

Swift is the outlier worth knowing about: String.count is grapheme clusters, so it reports 1 for every row, and Swift compares the two spellings of the accented word as equal without being asked. Bridge the same string to NSString and length is 8 for a family emoji, because that API is UTF-16 again. Python 3 counts code points, so a slice can never split a surrogate pair, but a skin-tone modifier or a ZWJ sequence still splits. C++ std::string counts bytes, which makes a byte-oriented substring the most dangerous of all: it can land inside a multi-byte sequence for ordinary Japanese text, not only for emoji.

The database is the other place the unit changes. In MySQL, utf8mb3 stores at most three bytes per character, so a four-byte emoji cannot go into such a column at all, while utf8mb4 allows four. Which name is an alias for which, and what the server default is, varies between versions, so confirm the current behaviour in the character set chapter of the MySQL reference manual rather than trusting a value inherited from an older schema. Separately, check whether your column limit is declared in characters or in bytes, and whether an index on it has a byte-length cap, because a limit that was comfortable for ASCII has a quarter of the headroom once four-byte characters are allowed.

string           JS .length  Java length()  C# .Length  Python len()  Swift .count  C++ .size()
"A"                       1              1           1             1             1            1
"あ"                       1              1           1             1             1            3
"😀"                       2              2           2             1             1            4
"👍🏽"                       4              4           4             2             1            8
"👨‍👩‍👧"                       8              8           8             5             1           18
"🇯🇵"                       4              4           4             2             1            8
"é" U+00E9                1              1           1             1             1            2
"é" e + U+0301            2              2           2             2             1            3

// UTF-16 code units: JavaScript, Java, C#      code points: Python 3
// grapheme clusters: Swift                     UTF-8 bytes: C++ std::string
// Measured on Node 24, OpenJDK 17, .NET 9, CPython 3.9, Swift 6.3, clang C++17.

// The same broken half, encoded to UTF-8:
//   Node 24 and .NET 9  ->  ef bf bd   (U+FFFD)
//   OpenJDK 17          ->  3f         (ASCII question mark)

Counting, cutting and validating safely

The fix is one rule: never index a user-facing string by a number you did not get from a segmenter. In JavaScript that is Intl.Segmenter with grapheme granularity. C# has StringInfo text elements, which report 1 for every row of the table above. Swift needs nothing extra. Java is the exception: on OpenJDK 17, BreakIterator with a character instance reported 5 for the family emoji and 2 for a flag, so it does not group emoji sequences. Verify the behaviour on your JDK, or use an ICU-based segmenter.

Server-side validation needs a second look, because well-formedness is necessary but not sufficient. A split surrogate pair is detectable. A dangling joiner, a loose skin-tone modifier and a single regional indicator are all well-formed UTF-16 and pass that check, so a hostile client can post them and a naive length check will accept them. Check the grapheme count as well, and reject a string whose last cluster ends in a joiner or consists only of a modifier.

A filter written as a script allowlist has its own trap. A pattern permitting Latin, kana and Han rejects a decomposed accented name, because the combining mark is in the Mark category and belongs to none of those scripts. Normalise to NFC before the filter runs, or permit Mark explicitly. Note also that NFC leaves emoji ZWJ sequences untouched, so normalising never shortens or simplifies them.

const seg = new Intl.Segmenter(undefined, { granularity: "grapheme" });
const count = (s) => [...seg.segment(s)].length;

// Cut on a cluster boundary, never inside one
function cut(s, max) {
  const out = [];
  for (const { segment } of seg.segment(s)) {
    if (out.length >= max) break;
    out.push(segment);
  }
  return out.join("");
}
cut("🇬🇧🇷🇺Ku👨‍👩‍👧", 3)      // "🇬🇧🇷🇺K" — every prefix is well-formed
"🇬🇧🇷🇺Ku👨‍👩‍👧".slice(0, 5)  // "🇬🇧\ud83c" — not

// Well-formedness catches only the split pair
("Yuki" + "😀".slice(0, 1)).isWellFormed()  // false — reject
"Yuki👨\u200d".isWellFormed()                // true — dangling joiner passes
"Yuki🏽".isWellFormed()                      // true — loose modifier passes
"Yuki🇯".isWellFormed()                       // true — lone flag letter passes

// "Does it contain emoji" is easy to get wrong
/\p{Emoji}/u.test("1")                   // true — digits carry the Emoji property
/\p{Extended_Pictographic}/u.test("🇯🇵")  // false — flags are not pictographic
/^\p{RGI_Emoji}$/v.test("👨‍👩‍👧")           // true — a recommended sequence
/^\p{RGI_Emoji}$/v.test("👨\u200d👩")      // false — a cluster, but not recommended

// Normalise before you store or compare
"Nie\u0301" === "Nié"                                     // false
"Nie\u0301".normalize("NFC") === "Nié".normalize("NFC")    // true

// Intl.Segmenter, isWellFormed and the v flag are recent additions to the
// language — confirm your runtime supports them before relying on them.

Rendering, storage routes, and how much emoji to accept

Rendering is the last mile. A colour emoji font has to be present and reachable through the font fallback chain, or the cluster your code preserved perfectly arrives as an empty box; why characters render as empty boxes covers diagnosing that. In Unity, inline emoji in TextMeshPro come from a sprite asset rather than from the text font, and the mapping from characters to sprites is worth confirming in the sprite chapter of the TextMeshPro documentation for your version. For Unreal, check the text and fonts section of the official documentation rather than assuming a system emoji font exists in a packaged build. A grapheme cluster must also never be split across a line break, which is the same boundary problem one layer up; line breaking and wrapping for CJK text covers those rules.

Then trace the route the string takes. Player names pass through save files, network messages, server logs, analytics events, moderation queues and CSV exports, and a single hop that is not UTF-8 end to end loses the emoji permanently; the common causes of garbled text and how to fix them covers what a legacy code page in the middle of that chain does. Test the whole round trip, not the field in isolation.

How much emoji to accept is a product decision, not a technical one, and it is cheaper made before the field ships. The workable positions are to accept any well-formed cluster, to accept a restricted subset, or to reject emoji in identity fields while allowing them in chat. Narrowing the set narrows the surprises: a single code point emoji costs two code units and needs no sequence support, a ZWJ sequence costs eight code units for a family of three and eleven for a family of four and renders as one glyph only if the renderer recognises that specific sequence, and a cluster that is valid but not a recommended sequence is drawn as its separate components.

Flags need their own decision. A region flag in a player name is read by other players as a statement, and some titles decline regional indicators for that reason alone rather than for any technical one. Skin-tone modifiers are usually accepted wherever the base emoji is, because rejecting them selectively is hard to justify to players. Whatever you choose, apply the rule once on the server and let the client mirror it, so a client that has not updated cannot create data the server considers invalid.

Finally decide what happens when a cluster cannot be rendered. Substituting a placeholder at display time keeps the original in storage, so the name comes back intact after a font or engine update; stripping at storage time destroys it forever. Then put the awkward inputs in your test data permanently, including the hostile ones a real client can send.

// Test data for any field with a length limit
A                  plain ASCII baseline
あ                  1 code unit, 3 UTF-8 bytes
😀                  surrogate pair
👍🏽                  base emoji + skin-tone modifier
👨‍👩‍👧                  ZWJ sequence, 8 code units, 18 bytes
🇯🇵                  regional indicator pair
🇬🇧🇷🇺                  two flags — a cut between them yields 🇧🇷
é (e + U+0301)     combining mark, 2 code points
\ud83d              a lone high surrogate, as a hostile client would send it
👨\u200d             a dangling zero-width joiner

// Pass condition: the value survives truncation, storage, retrieval and
// re-render unchanged, or is rejected with a message — never silently altered.

Related articles