Fix mojibake: read the garbled text, find the encoding, decode it
Mojibake is a decoding mismatch, not damage: the bytes on disk are almost always the exact bytes that were written, and some program read them back under the wrong encoding label. Nothing was lost, so the fix is usually a single command — the whole job is working out which two encodings were involved.
This article hands you that identification step. Below is a lookup table built by encoding three sample strings and decoding those same bytes under eight wrong labels, the visual signatures that let you name a mismatch from the garbled text alone, the commands that reverse each case, and the two situations where the original really is gone. Every string and byte value here came out of a command you can rerun on your own machine.
Eight misread pairs, generated so you can match yours against them
Each row was produced the same way: encode a sample with one codec, then hand those bytes to a different codec. The encode side used the Python standard library codecs. Six of the eight pairs printed identically from Python and from a TextDecoder in Node, on both Node versions installed here. Two did not, and they are the reason not to trust one tool. The Windows-1252 rows follow the WHATWG windows-1252 index: Node 24.18.1 reproduces both, Python's cp1252 reproduces 日本語 but marks the undefined byte 81 with a replacement character where a WHATWG decoder leaves the invisible control character noted under こんにちは, and Node 22.13.0 renders the whole 80 to 9F range as bare C1 control characters so prints neither row. The EUC-JP row is the second disagreement, and the last bullet below describes it. Find your garbled text in the right-hand column and you have named both encodings at once.
- The replacement character at the end of 譌・譛ャ隱 is not noise. The nine UTF-8 bytes of 日本語 end in 9E, which CP932 treats as a lead byte with nothing following it, so the decoder gives up on the last byte only. An odd trailing replacement character usually means the run ended mid-sequence, not that the file is broken.
- café cannot be encoded in CP932 at all — Python raises UnicodeEncodeError on the é. A pipeline that writes Shift_JIS therefore has no way to carry accented Latin text, which is worth knowing before you route a European language through a Japanese-locale export.
- Python and Node disagree most sharply on the EUC-JP row. From the same nine bytes of 日本語, Python recovered a second CJK character that Node gave up on, and the two printed a different number of replacement characters. Treat the count of replacement characters as a hint about where a sequence broke, never as evidence of how much broke.
Three samples, encoded once and decoded under eight wrong labels.
UTF-8 bytes read as CP932 / Shift_JIS
こんにちは -> 縺薙s縺ォ縺。縺ッ
日本語 -> 譌・譛ャ隱�
café -> cafテゥ
UTF-8 bytes read as Windows-1252
日本語 -> 日本語
café -> café
こんにちは -> ã“ã‚“ã«ã¡ã¯ (four invisible U+0081 control characters dropped here)
UTF-8 with a byte-order mark, read as Windows-1252
café -> café
UTF-8 bytes read as GBK
日本語 -> 鏃ユ湰瑾�
UTF-8 bytes read as EUC-JP
日本語 -> �ユ�� (three invisible control characters dropped; Python differs)
CP932 bytes read as UTF-8
こんにちは -> ����ɂ���
日本語 -> ���{��
EUC-JP bytes read as CP932
こんにちは -> 、ウ、ヒ、チ、マ (a private-use character, dropped here, sits after the second 、)
UTF-16LE with a byte-order mark, read as UTF-8
こんにちは -> ��S0�0k0a0o0
café -> ��c a f � (each gap is a NUL byte)
Regenerate any row yourself:
python3 -c "print('こんにちは'.encode('utf-8').decode('cp932', errors='replace'))"
python3 -c "print('日本語'.encode('utf-8').decode('cp1252'))"
node -e "process.stdout.write(new TextDecoder('euc-jp').decode(new TextEncoder().encode('日本語')))"Read the signature: what the shape of the garbling tells you
The table generalises into a handful of rules. You can usually name the mismatch from one line of garbled text, before you touch the file.
- A repeating 縺, 繧 or 繝 every second or third character means UTF-8 read as Shift_JIS or CP932. Hiragana all start with the UTF-8 pair E3 81 or E3 82, and CP932 maps E3 81 to 縺, so the same kanji reappears wherever the byte alignment lands on that pair — four times in 縺薙s縺ォ縺。縺ッ.
- Runs of æ, è and ã mixed with curly quotes, dashes and  mean UTF-8 read as a single-byte Western encoding. Japanese produces æ and è because its lead bytes are E6 and E8; Latin text with accents produces the familiar à pairs such as café and äö.
-  at the very start of the very first line and nowhere else is a UTF-8 byte-order mark (bytes EF BB BF) read as a single-byte encoding. The rest of the file can be perfectly fine, which is what makes this one easy to misdiagnose as a stray header.
- Nothing but replacement characters, with the odd ASCII punctuation mark surviving, means you are decoding a legacy double-byte file as UTF-8 — the opposite direction from the cases above. The survivors are trail bytes that happen to land in ASCII: 日本語 in CP932 is 93 FA 96 7B 8C EA, and that 7B is why a lone { appears in ���{��.
- Almost nothing but half-width katakana such as 、 ウ ヒ チ マ means an EUC-JP file read as Shift_JIS. EUC-JP puts both bytes of every kanji and kana in the A1 to FE range, while ASCII stays single-byte and half-width katakana carries an 8E prefix. Shift_JIS assigns A1 to DF to single-byte half-width katakana, so most of those bytes surface one at a time as katakana. Bytes above DF are the exception: Shift_JIS reads them as lead bytes and swallows the byte after them, which is where the private-use character hidden in the table came from.
- Characters separated by gaps you cannot select means UTF-16 read as UTF-8. For ASCII the second byte of every unit is a NUL, so the gaps look blank. Japanese instead leaves a visible 0 after each garbled character, because the high byte of U+30xx is 0x30 — S0, k0 and a0 in the table are those zeroes, and they are real characters, not artefacts.
- Dense unfamiliar Chinese characters mean UTF-8 read as GBK: 鏃 and 瑾 turn up where the original was kanji, and 銇 repeats through hiragana. Same cause as the CP932 case, different legacy table.
- æ—¥ and similar three-character clusters mean the damage happened twice: text was garbled, saved as UTF-8, then read under the wrong label again. See double-encoded UTF-8 and how to detect it — it reverses in two passes rather than one.
- Empty squares or boxes where every other character renders correctly is not mojibake at all. The decoding worked and the font has no glyph, which is a different problem with a different fix: tofu boxes and missing glyphs.
Reverse it with iconv, and mind the encoding names
Once you know the pair, recovery is mechanical. If you still have the original file, decode it with the correct label and you are done. If all you have is the garbled text, you re-encode it with the wrong codec that produced it, which regenerates the original bytes, then read those bytes correctly. Both directions are one iconv invocation.
The trap is that encoding names are not interchangeable. CP932 and SHIFT_JIS are separate codecs in iconv, not aliases, and a recovery that works under one aborts under the other. The same applies to CP1252 and ISO-8859-1: the characters that mojibake puts in the 80 to 9F range exist in Windows-1252 but not in ISO-8859-1, and iconv will quietly substitute rather than fail. Both traps are shown below with the output they actually produce.
- Anything that came off a Windows machine or a Japanese-locale spreadsheet wants CP932, which is Microsoft's superset of Shift_JIS. It carries the circled numbers, Roman numerals and full-width tilde that strict Shift_JIS rejects — the divergences are catalogued in Shift_JIS and CP932 pitfalls.
- In the ISO-8859-1 run above, iconv exited 0 and still produced the wrong bytes: it turned the em dash into a hyphen (2D), the œ into oe (6F 65) and the ž into z (7A). A successful exit status is not proof that a recovery worked. Always eyeball the recovered text.
- In stock Node, decoding is easy and re-encoding is not. TextDecoder handles shift_jis, euc-jp, windows-1252, gbk, big5, euc-kr, windows-1251 and koi8-r, and its shift_jis label follows the WHATWG Encoding Standard, whose index covers the CP932 additions rather than strict JIS X 0208 — the byte pair 87 40 decodes to ① instead of failing. TextEncoder, however, ignores its argument and always emits UTF-8 — constructing one with shift_jis still reports an encoding of utf-8.
- Buffer.from(text, 'latin1') is the usual Node one-liner for reversing Western mojibake, and it works only inside the Latin-1 range. It recovers café to café, but on 日本語 it truncates the em dash at U+2014 down to byte 14 and produces garbage. Use iconv, or a real encoder library, for the Windows-1252 case.
# You still have the original file: decode it with the right label. iconv -f CP932 -t UTF-8 original.csv > fixed.csv # You only have the garbled text: re-encode it with the codec that broke it. # UTF-8 that was read as CP932 -- fixed.csv comes out as valid UTF-8. iconv -f UTF-8 -t CP932 garbled.csv > fixed.csv # UTF-8 that was read as Windows-1252. iconv -f UTF-8 -t CP1252 garbled.csv > fixed.csv # Trap 1: CP932 and SHIFT_JIS are different codecs. printf '~' | iconv -f UTF-8 -t CP932 | xxd -p # 8160 printf '~' | iconv -f UTF-8 -t SHIFT_JIS # iconv: iconv(): Illegal byte sequence printf '〜' | iconv -f UTF-8 -t CP932 # iconv: iconv(): Illegal byte sequence printf '〜' | iconv -f UTF-8 -t SHIFT_JIS | xxd -p # 8160 printf '①' | iconv -f UTF-8 -t CP932 | xxd -p # 8740 printf '①' | iconv -f UTF-8 -t SHIFT_JIS # iconv: iconv(): Illegal byte sequence # Trap 2: ISO-8859-1 silently mangles what CP1252 recovers. printf '日本語' | iconv -f UTF-8 -t CP1252 | xxd -p # E697A5E69CACE8AA9E -> 日本語 printf '日本語' | iconv -f UTF-8 -t ISO-8859-1 | xxd -p # E62DA5E66F65ACE8AA7A, exit status 0
The two cases where the original text is really gone
Reversal works because the original bytes still exist somewhere. Two things destroy them, and both are worth checking for before you spend an afternoon on a file that cannot be saved.
The first is the replacement character. When a decoder hits bytes it cannot interpret, it emits U+FFFD, which carries no record of what it replaced. Re-encoding that text substitutes a question mark or fails outright, and the information is gone. The demonstration below loses exactly one character of three: 日本語 becomes 日本 plus rubble, because a single trailing byte had already been discarded. The second is a save. If a garbled file is written out again as UTF-8, what lands on disk is a new and entirely valid encoding of the wrong characters. Nothing about those bytes is invalid any more, so no tool can tell they were ever wrong.
The practical consequence: when you find mojibake, do not open the file in an editor and save it, and do not run an import that rewrites it. Take a copy of the raw bytes first. If the pipeline has already re-saved, your recovery path is the upstream source — the translator's delivery, the previous commit, the database dump — not the file in front of you.
# One byte has already become U+FFFD. Re-encoding cannot put it back.
python3 -c "
u = '日本語'.encode('utf-8').decode('shift_jis', errors='replace')
print(u) # 譌・譛ャ隱 plus a replacement character
b = u.encode('cp932', errors='replace')
print(' '.join('%02X' % x for x in b)) # E6 97 A5 E6 9C AC E8 AA 3F
print(b.decode('utf-8', errors='replace')) # 日本 then U+FFFD and ? -- the 語 is gone
"
# Lossy in one step, with no error raised at all:
python3 -c "print('日本語'.encode('cp1252', errors='replace'))" # b'???'Pin the encoding at every boundary so it stops recurring
Mojibake happens at handoffs — wherever a file moves between a tool, a system or a person. Auto-detection is good but not perfect, and one wrong guess that gets saved is how a recoverable problem becomes a permanent one. The fix that holds is to state the encoding explicitly at every boundary instead of letting each tool decide.
- Files: settle on UTF-8 and write down whether you use a byte-order mark. A mark makes some spreadsheets open a CSV correctly on a double-click, and makes some parsers read a stray  into the first column name. Pick one convention per pipeline and record it in the localization kit you hand to translators.
- Databases: the connection charset matters as much as the column charset, and a mismatch between them is the classic source of double-encoded text. MySQL's character-set documentation describes utf8mb3 as a three-byte subset that cannot store four-byte characters, and the utf8 alias has pointed at different things across versions, so name utf8mb4 explicitly wherever emoji or rarer kanji can appear — see emoji and surrogate pairs in game text.
- Editors: turn off encoding auto-detection where you can, and set it per project. The moment an editor guesses wrong and you save, the original bytes are gone.
- Game engines: check what the importer expects rather than assuming. If the garbling appears only in the running game and not in the source file, the encoding is fine and the problem is downstream — garbled Japanese text and the Windows system locale covers that case.
- Regression test: put a row containing こんにちは, café and an emoji into your own fixture data and assert on it after every import and export. A boundary that breaks encoding will break that row visibly, on the day it breaks, instead of in a shipped build. For faster eyeball checks, keep a cheat sheet of mojibake patterns next to the test.