Mojibake cheat sheet: what é, ’ and 縺 mean, and how to undo them
You are looking at text that should say something and instead says café, or ’, or 縺ゅ↑縺�, or a run of half-width katakana. This page is a lookup sheet for that moment, organized by what is on your screen rather than by encoding theory: find the row that looks like your text, and it tells you what the original was, what encoding the bytes are really in, and what encoding they were mistakenly read as.
That pair is all you need, because every mojibake pattern is the same accident: bytes written in one encoding were decoded with another. Once the pair is known the fix is mechanical. Take the garbled text, turn it back into bytes using the wrong encoding it was read as, then decode those bytes with the encoding the file is actually in. Three steps: match the pattern, read the pair, re-decode. The reversal is spelled out below.
One rule before you start. If what you see is � (the Unicode replacement character, U+FFFD) or plain question marks in place of characters, the bytes have already been thrown away; nothing on this sheet can bring them back, and the only recovery is an older copy of the file. A single � at the end of an otherwise recognizable pattern is different: the display could not show the last byte, and the file itself is intact. Always reverse from the file, never from text copied off the screen.
Latin-script patterns: é, ’, © and 
Almost every garbled accented letter in a European language is UTF-8 read as Windows-1252, or as ISO-8859-1, which browsers treat as the same thing. UTF-8 stores é, ü, ñ and their relatives as two bytes, the first of which is 0xC3. Windows-1252 is a single-byte encoding, so it shows that first byte as à and the second byte as whatever letter or symbol sits at that position. The result is the Ã-plus-one-character pair people paste into search engines: é is é, ü is ü, ñ is ñ.
- Ã plus one character: an accented Latin letter. The text was typed correctly and is stored as UTF-8; whatever displayed it assumed Windows-1252. Fully recoverable.
- †plus one character: typographic punctuation. Curly quotes, the ellipsis, the en dash and the euro sign are three bytes in UTF-8 beginning with 0xE2 0x80, which Windows-1252 shows as â€; the third byte becomes ™ for the apostrophe, œ for the opening quote, ¦ for the ellipsis. The closing double quote ends in a byte Windows-1252 leaves undefined, so it shows as � or as nothing. ’, the most searched fragment in this family, is simply a curly apostrophe.
-  plus a symbol: characters such as © ° « » and the non-breaking space sit in the range where the first UTF-8 byte is 0xC2, and Windows-1252 shows 0xC2 as Â. A stray  before a symbol, or an  where a space should be, is this pattern.
-  at the very start of a file and nowhere else: the UTF-8 byte order mark, three bytes some editors prepend, displayed as Windows-1252. The rest of the file is usually fine. Strip the mark rather than re-decoding; the article 'UTF-8 BOM problems in game files: only the first line breaks' covers what it damages.
what you see → original | actual encoding → read as café → café | UTF-8 → Windows-1252 über → über | UTF-8 → Windows-1252 niño → niño | UTF-8 → Windows-1252 ’ → ’ | UTF-8 → Windows-1252 “ â€� → “ ” | UTF-8 → Windows-1252 … → … | UTF-8 → Windows-1252 – → – | UTF-8 → Windows-1252 € → € | UTF-8 → Windows-1252 © → © | UTF-8 → Windows-1252 ° → ° | UTF-8 → Windows-1252  (file start) → (nothing) | UTF-8 byte order mark → Windows-1252
Japanese patterns: 縺 繧 繝 runs, ã�‚ pairs, half-width kana and 銇傘仾
Japanese has more distinct signatures than any other script because three legacy encodings are still in circulation alongside UTF-8: Shift_JIS, its Windows variant CP932, and EUC-JP. Every pairing of them looks different, and the common ones are below.
- Runs of 縺, 繧 and 繝 with half-width katakana sprinkled in (縺ォ, 繝シ): UTF-8 kana read as Shift_JIS. Every hiragana in UTF-8 begins with 0xE3 0x81 or 0xE3 0x82, and in Shift_JIS those two bytes are the kanji 縺 or 繧; katakana gives 繧 or 繝, and kanji become dense unrelated kanji such as 譌・譛ャ. By far the most common Japanese pattern, and fully recoverable. A trailing � is the display giving up on the last byte, not damage.
- ã followed by � or a symbol, three characters per kana (ã�‚, ã‚“, æ—¥): the same UTF-8 text read as Windows-1252. Kana produce ã groups; kanji produce æ, ç, è or é groups. The � in the middle is a byte Windows-1252 does not define, and browsers usually show nothing there. This is what a Western-locale machine shows for Japanese UTF-8.
- Curly quotes, ƒ, Œ and plain ASCII brackets mixed together (“ú–{Œê, ƒQ�[ƒ€): a Shift_JIS file read as Windows-1252. Shift_JIS lead bytes land on the Windows-1252 punctuation row (‚ ƒ „ … † ‡ ˆ ‰ Š ‹ Œ ‘ ’ “ ” • – —), and its second bytes include plain ASCII, so { [ Q and lowercase letters appear inside the noise. This is the spreadsheet-export case: a CSV saved on Japanese Windows, opened on a Western-locale machine.
- Half-width katakana in pairs (、「、ハ、ソ, ・ォ・ソ・ォ・ハ): EUC-JP read as Shift_JIS. EUC-JP uses bytes 0xA1 to 0xFE for every Japanese character, and in Shift_JIS most of that range is single-byte half-width katakana, so each original character becomes two half-width kana, with an occasional � or kanji where the second byte is higher.
- ¤ or ¥ before every other character (¤¢¤Ê¤¿, ¥²¡¼¥à): EUC-JP read as Windows-1252. Hiragana all share the lead byte 0xA4 (¤) and katakana 0xA5 (¥), which makes this the easiest to spot; kanji give pairs like ÆüËܸì without the repeating symbol.
- Plausible-looking Chinese characters where Japanese should be (銇傘仾, 鏃ユ湰, 偁側偨, 擔杮岅): the file was opened by a tool defaulting to GBK, the Chinese Windows code page. UTF-8 read as GBK turns three bytes into one and a half characters, so the count no longer matches and a � often lands at the end; Shift_JIS read as GBK turns each character into exactly one hanzi with no � at all, which looks deceptively clean. Both are recoverable.
- Mostly � with an occasional { or ASCII letter (���Ȃ�, ���{��): Shift_JIS read as UTF-8. Shift_JIS byte pairs are rarely valid UTF-8, so the decoder discards them and throws most of them away. What survives is unpredictable: a second byte that happens to be ASCII shows as itself, which is where the { and Q come from, while two bytes belonging to neighbouring characters can accidentally form a valid sequence and appear as an unrelated accented letter such as the Ȃ above. If the file was merely displayed this way, reopen it as Shift_JIS. If it was saved in this state, the discarded bytes are gone.
what you see → original | actual encoding → read as
縺ゅ↑縺� → あなた | UTF-8 → Shift_JIS (CP932)
縺薙s縺ォ縺。縺ッ → こんにちは | UTF-8 → Shift_JIS (CP932)
譌・譛ャ隱� → 日本語 | UTF-8 → Shift_JIS (CP932)
譚ア莠ャ → 東京 | UTF-8 → Shift_JIS (CP932)
ã�‚ã�ªã�Ÿ → あなた | UTF-8 → Windows-1252
ã�“ã‚“ã�«ã�¡ã�¯ → こんにちは | UTF-8 → Windows-1252
日本語 → 日本語 | UTF-8 → Windows-1252
“ú–{Œê → 日本語 | Shift_JIS → Windows-1252
“Œ‹ž → 東京 | Shift_JIS → Windows-1252
ƒQ�[ƒ€ → ゲーム | Shift_JIS → Windows-1252
、「、ハ、ソ → あなた | EUC-JP → Shift_JIS
・ォ・ソ・ォ・ハ → カタカナ | EUC-JP → Shift_JIS
¤¢¤Ê¤¿ → あなた | EUC-JP → Windows-1252
¥²¡¼¥à → ゲーム | EUC-JP → Windows-1252
ÆüËܸì → 日本語 | EUC-JP → Windows-1252
銇傘仾銇� → あなた | UTF-8 → GBK
鏃ユ湰瑾� → 日本語 | UTF-8 → GBK
偁側偨 → あなた | Shift_JIS → GBK
擔杮岅 → 日本語 | Shift_JIS → GBK
���Ȃ� → あなた | Shift_JIS → UTF-8 (lossy)
���{�� → 日本語 | Shift_JIS → UTF-8 (lossy)
�サソ (file start) → (nothing) | UTF-8 byte order mark → Shift_JISChinese, Korean and Cyrillic patterns
The same logic covers every other script, and the tell is the first byte of each character. UTF-8 puts Chinese characters (and Japanese kanji) under 0xE4 to 0xE9, Korean hangul under 0xEA to 0xED, and Cyrillic under 0xD0 to 0xD1. Windows-1252 renders those as ä å æ ç è é, as ê ë ì í, and as Ð Ñ respectively, so the lead letter of each group tells you which script was garbled.
- Chinese: groups led by ä å æ ç è é (䏿–‡) are UTF-8 read as Windows-1252, three characters per hanzi. Accented uppercase letters with no à (ÖÐÎÄ) are GBK or GB2312 read as Windows-1252, two letters per hanzi, because GBK uses two bytes above 0x80 for each character. Wrong hanzi with a � every other position (涓�鏂�) are UTF-8 read as GBK.
- Korean: groups led by í, ê or ì (한êµì–´) are UTF-8 hangul read as Windows-1252. Hangul begins with 0xEA to 0xED, so the groups start with ê ë ì í rather than ã. A legacy EUC-KR file read as Windows-1252 looks like the GBK row above rather than the UTF-8 row: accented letters and symbols in pairs.
- Cyrillic: alternating Ð and Ñ (Привет) is UTF-8 read as Windows-1252, because UTF-8 Cyrillic letters start with 0xD0 or 0xD1. Accented lowercase Latin of exactly the same length as the original word (Ïðèâåò) is Windows-1251 read as Windows-1252, one letter for one letter. Alternating Р and С (Привет) is the mirror image: UTF-8 read as Windows-1251, where 0xD0 and 0xD1 are Р and С.
what you see → original | actual encoding → read as 䏿–‡ → 中文 | UTF-8 → Windows-1252 ÖÐÎÄ → 中文 | GBK → Windows-1252 涓�鏂� → 中文 | UTF-8 → GBK 한êµì–´ → 한국어 | UTF-8 → Windows-1252 Привет → Привет | UTF-8 → Windows-1252 Ïðèâåò → Привет | Windows-1251 → Windows-1252 Привет → Привет | UTF-8 → Windows-1251
Double-encoded patterns: é, ’ and 日
Sometimes the pattern is longer than anything above: à where a single à should be, or †at the start of every piece of punctuation. That is the Latin-script pattern applied twice. The text was read as Windows-1252, saved as UTF-8 in that garbled state, and then read as Windows-1252 again. The mechanism, and why it tends to happen in pipelines with a database in the middle, is covered in the article 'ã� and ’: fixing double-encoded UTF-8 text'. Here is only how to recognize it.
- Recognize it by the prefixes. à (à then ƒ) is a doubled Ã;  is a doubled Â; †is a doubled â€. The string is roughly twice as long as a once-garbled version, and every character in it is a valid Windows-1252 character, so nothing shows as �.
- Reverse it by applying the same reversal twice: encode as Windows-1252, decode as UTF-8, and repeat. Doing it once leaves you with the once-garbled text, which can look like progress. Check that the result is real words before you stop.
- Three layers exist but are rare. If the second pass still shows à pairs, run it a third time.
original → garbled once → garbled twice café → café → café ’ → ’ → ’ … → … → … 日本 → 日本 → 日本
How to reverse it once you know the pair
The fix is the same for every row on this sheet. Take what you see, turn it back into bytes using the encoding it was wrongly read as, and decode those bytes with the encoding the file is actually in. In pseudo-code, with each pair written as actual → read-as:
- In practice you rarely need code. Most editors can reopen a file with a chosen encoding: pick the actual encoding from the pair and the text is correct. If a tool misread the file and exported a new one, set the actual encoding at import and export again. The two-step re-encode is only for when the original is gone and all you have is the garbled string.
- Never fix mojibake with find-and-replace. Each garbled fragment is one byte shared by many different characters; replacing é with é fixes one letter, leaves the others, and can corrupt text that legitimately contains that fragment. Re-decoding fixes every character at once because it addresses the cause.
- Reverse the file, not the screen. Text copied from a browser or terminal may already have lost bytes to � or to invisible control characters. Go back to the file or the database field.
- Confirm on a known word. After reversing, look at a string you know should read 東京 or café. If it now shows a different garbled pattern from this sheet, you had the pair backwards, or there is a second layer.
- Give up when the garbled state was saved with loss: � anywhere in the file, question marks where characters should be, or a reversal that produces a fresh � partway through. Find an earlier copy instead: version control, the original delivery, or the spreadsheet the export came from. The article 'Fix mojibake: read the garbled text, find the encoding, decode it' walks through that recovery.
// Pattern 縺ゅ↑縺� pair: UTF-8 → Shift_JIS
const bytes1 = encodeWith(garbled, "Shift_JIS"); // back to the original bytes
const fixed1 = decodeWith(bytes1, "UTF-8"); // "あなた"
// Pattern café pair: UTF-8 → Windows-1252
const bytes2 = encodeWith("café", "Windows-1252");
const fixed2 = decodeWith(bytes2, "UTF-8"); // "café"
// Pattern Ïðèâåò pair: Windows-1251 → Windows-1252
const bytes3 = encodeWith("Ïðèâåò", "Windows-1252");
const fixed3 = decodeWith(bytes3, "Windows-1251"); // "Привет"
// Double-encoded café: run the Windows-1252 block twice.
// First pass gives "café", second pass gives "café".Why game and localization pipelines produce these
A CSV exported from a spreadsheet on Japanese Windows is often written in Shift_JIS (CP932), because that is the system's legacy code page, while everything downstream expects UTF-8. That one default explains most 縺 runs and most “ú–{Œê strings that come back from translators. The article 'Shift_JIS vs CP932: the differences that break Japanese text' covers which encoding name means which table.
Tools that never state an encoding fall back to the operating system's code page: Windows-1252 on Western systems, CP932 on Japanese systems, GBK on Chinese systems. The same UTF-8 file therefore garbles three different ways depending on whose machine opened it, which is why a file can be fine for the developer and broken for the translator.
Databases and older string APIs with a single-byte default are where double encoding is born: text goes in through one wrong assumption and comes out through a correct one, producing valid UTF-8 of the wrong characters. If your pattern is in the doubled section, look at the middle of the pipeline, not at its two ends.