Tofu boxes in game text: missing glyphs, not broken encoding
An empty rectangle where a character should be means one thing: the font has no drawing for that code point. The string itself is almost certainly correct. So before touching an encoding setting, inspect the value of the string instead of the screen. If the code points are the ones you expect, the data is fine and the renderer simply had nothing to draw, which makes this a font problem with a font fix.
That triage takes a minute. The slow part is everything after it: which font to bundle, how large it is allowed to be, and which characters to pre-generate into the atlas. To put real numbers on those decisions I wrote a small cmap reader and scanned ten fonts installed on one macOS machine, measuring what each one actually claims to cover. Every figure below comes from that scan or from arithmetic on it.
Tofu, mojibake, and U+FFFD are three different bugs
All three arrive as damaged-looking text and only the first is about fonts. What separates them is the string value, which is why inspecting it is the entire diagnosis.
Take the Japanese word for settings, whose code points are U+8A2D and U+5B9A. As UTF-8 the bytes are e8 a8 ad e5 ae 9a; as CP932 they are 90 dd 92 e8. Hand the UTF-8 bytes to a CP932 decoder and out comes the sequence 險ュ螳 followed by a replacement character: plausible-looking characters that are the wrong ones. Hand the CP932 bytes to a UTF-8 decoder and out comes U+FFFD, U+0752, U+FFFD, because most of those bytes are not valid UTF-8 sequences. Both of those changed the data. When a font is at fault, the code points after the failure are still U+8A2D and U+5B9A.
The box you see is itself a glyph from some font. On Apple platforms it usually comes from a small system font called LastResort, which measures 2.5 KB, contains seven glyphs, and claims the whole Unicode code space of 1,114,112 code points through a cmap format 13 subtable whose four ranges resolve to just three of those glyphs. An engine with no equivalent draws the primary font's own .notdef box instead, or nothing at all.
- Copy the broken text out of the running build, or log its code points. Correct code points plus a broken display means a font problem, full stop.
- Wrong-but-readable characters mean the bytes were decoded with the wrong codepage. U+FFFD means the bytes were not valid in the encoding that read them. Both happened before rendering.
- Uniform boxes across a whole line usually mean the font has no coverage for that script at all. Boxes on scattered characters mean a partial coverage gap, which is the more common case.
symptom string value cause ------------------------------ --------------------- --------------------------------- empty or outlined boxes correct code points font has no glyph for them readable but wrong characters different code points bytes read as the wrong codepage U+FFFD diamond or box U+FFFD in the data bytes invalid in that encoding
What glyph coverage actually measures
Every font carries a cmap table: a lookup from code point to glyph index. If a code point is absent from cmap, or maps to glyph 0, the font is telling the renderer it cannot draw that character. That claim is exactly what system fallback logic and engine queries read, so cmap coverage is the number worth measuring. It is a claim rather than a guarantee, since a font can map a code point to a blank glyph, but a gap in cmap is always a real gap.
Here is what those ten fonts cover. The percentages are against repertoires generated from codec tables rather than typed by hand: 6,355 JIS X 0208 kanji across levels 1 and 2, 11,172 Hangul syllables from U+AC00 to U+D7A3, and 6,763 GB2312 hanzi. Sizes are for the whole file, the code-point column counts the Unicode subtables only, and where a file is a collection of several faces the glyph count is the first face only.
Four results in that table settle most font arguments. A standard Latin UI font covers nothing CJK at all, so Japanese in it is a solid wall of boxes rather than a scattered gap. Every CJK font measured covers zero of the 11,172 Hangul syllables, the two Chinese fonts included, so Korean never arrives as a side effect of adding Japanese or Chinese. The Korean font fails the same way in reverse: all 11,172 syllables, but only 64.3 percent of JIS X 0208, missing 522 of the 2,965 level-1 kanji and 1,749 of the 3,390 level-2 kanji. And the broadest single font on the machine, covering 38,917 code points, still has no glyph for U+1F3AE VIDEO GAME or U+26A0 WARNING SIGN, which are exactly the characters a modern UI drops into a tooltip.
One more trap hides inside a 100 percent score. The Simplified Chinese font in that table covers every JIS X 0208 kanji, yet it draws many of them with mainland shapes, because one Han code point has region-specific forms. That passes a coverage check and fails a Japanese review. Treat coverage as the floor, not the goal; for keeping shapes right in a low-resolution UI, choosing pixel fonts for Japanese and English goes further.
font faces size glyphs cmap cps kana JIS Hangul GB2312 --------------------- ----- --------- -------- --------- ------- ------- ------- ------- Arial 1 755.1 KB 3381 2792 0.0% 0.0% 0.0% 0.0% Menlo 4 2.1 MB 3157 2727 0.0% 0.0% 0.0% 0.0% Osaka 1 3.5 MB 8121 7318 98.9% 100.0% 0.0% 49.3% YuGothic Medium 1 10.5 MB 23060 15914 100.0% 100.0% 0.0% 66.9% Hiragino Sans GB W3 4 22.4 MB 29352 29318 100.0% 100.0% 0.0% 100.0% PingFang 24 74.6 MB 49533 33258 96.0% 99.6% 0.0% 99.7% Apple SD Gothic Neo 18 52.8 MB 18662 18067 96.0% 64.3% 100.0% 39.7% Apple Color Emoji 2 183.2 MB 3844 1469 0.0% 0.0% 0.0% 0.0% Arial Unicode MS 1 22.2 MB 50377 38917 98.9% 100.0% 100.0% 100.0% LastResort 1 2.5 KB 7 1114112 100.0% 100.0% 100.0% 100.0%
Swapping in a CJK font breaks other languages
The usual reaction to a wall of boxes is to replace the UI font with one big CJK font and move on. The scan says that trades one tofu for another. Osaka, the smaller of the two Japanese fonts above, has no glyph for the dotless i or the s with cedilla that Turkish needs, and none for the d with stroke, u with horn, a with breve and dot below, e with circumflex and acute, or u with dot below that Vietnamese needs. The larger Japanese font handles Turkish but still misses four of those Vietnamese letters. The Simplified Chinese font additionally lacks the capital C with cedilla. The Latin-only font covers all of them.
The size trade-off is real but smaller than the folklore, as long as you compare single faces. The Latin font is 755 KB for 3,381 glyphs. The two Japanese faces are 3.5 MB for 8,121 glyphs and 10.5 MB for 23,060 glyphs, so five to fourteen times the Latin font rather than a hundred times. The outlier is the color emoji font: 183 MB for 1,469 code points, because it stores bitmap images at several sizes. Shipping emoji as a font is a decision with an install-size number attached, and how emoji and surrogate pairs behave in game text covers the string-handling half.
So the shape that works is not one font. It is a primary font per script plus an explicit ordered fallback list, which is what the operating system does for you and what your engine does not.
The OS has a fallback chain; your engine does not
Browsers and native UI toolkits hold a per-script fallback chain and walk it whenever the primary font misses a glyph. That is why tofu is rare outside games. Game engines draw text through their own systems, and those start with exactly one font and an empty fallback list until you fill it in.
In Unity, a TextMesh Pro font asset holds an atlas of rendered glyphs, generated either statically at import or dynamically at runtime. Fallback is two-tier: a per-asset Fallback Font Assets list, then a project-wide list in the TextMesh Pro settings, consulted in that order. Both are assets, so a fallback that works in the editor works in the build. Check the Font Asset Creator and fallback sections of the TextMesh Pro documentation for the generation modes and the lookup order.
In Unreal, the mechanism is a composite font: a default typeface plus sub-typefaces bound to character ranges, so Latin, Japanese and Korean can each point at a different file inside one font asset. The composite font section of the Unreal Engine documentation covers how ranges are declared and which caching mode applies.
In Godot, font resources carry a Fallbacks array that the engine walks in order. The fonts chapter of the Godot documentation describes the array and the dynamic caching behavior.
In all three, a chain only reaches what you shipped. A fallback that resolves to a font installed on the authoring machine but absent from the build produces boxes on a clean test device and nowhere else, which is the single most common route this bug takes to players. For the surrounding project setup, see the Unity localization overview and the Unreal Engine localization overview.
Sizing the atlas, and the subsetting trap
A static atlas is generated from a character set you choose, and the tempting choice is the set of characters that appear in today's translation files. It holds until a player types a name, opens chat, or a patch adds a string with a character nobody pre-generated. Each of those renders as a box in a build that passed QA.
Size that set with real counts. Japanese needs the 6,355 JIS X 0208 kanji plus 176 kana; Simplified Chinese 6,763 GB2312 hanzi; Traditional Chinese 13,061 Big5 hanzi as accepted by the codec I scanned; Korean 11,172 Hangul syllables. The overlap is smaller than people assume: JIS X 0208 and GB2312 share only 3,330 characters, which leaves 3,433 hanzi that adding Chinese to a Japanese build genuinely adds. The union of all three Han repertoires is 16,289 characters, and ASCII plus kana plus those three plus Hangul comes to 27,732 distinct code points.
Turn that into atlas pages. At a 32-pixel cell with 2 pixels of padding, a 4096 by 4096 sheet holds 120 by 120, or 14,400 glyphs, so a 27,732-character set needs two sheets. At a 48-pixel cell the same sheet holds 6,561 glyphs and the set needs five. That arithmetic, not intuition, decides whether the atlas fits your texture budget.
The policy that survives launch: pre-generate the fixed UI set statically, and put anything a player can type or receive on a dynamic atlas with a fallback font behind it. Where dynamic generation is not available on a target platform, pre-generate the full repertoire for every shipped language rather than the subset today's strings happen to use.
Check it before you ship
A missing glyph is detectable in code, which means it belongs in a build check rather than in playtest notes.
If the string value turns out to be wrong rather than merely undrawable, you are in the other failure mode, and the causes and fixes for mojibake is where to continue. For getting the Japanese text right in the first place, see localizing an English game into Japanese.
- Scan the cmap of every bundled font against the character set of every string in every language, and fail the build on a gap. It needs no rendering and takes a few dozen lines, which is how the table above was produced.
- Use the engine's own query where one exists, since it reflects the atlas and the fallback list rather than just the file. TextMesh Pro exposes a call that returns the characters an asset cannot draw, and Godot fonts expose a per-character query. In Unreal, verify the ranges declared in the composite font instead.
- Cover the characters players generate, not only the ones you translate: display names, chat, and any text pulled from a platform account, which can contain any script at all.
- Test on a clean device with no development fonts installed, in every shipped language, and look hardest at screens that mix scripts, such as a leaderboard of names from several regions.
- When a language is added, run the scan before the translation lands. A gap found then is a font decision; the same gap found in certification is a patch.