Mojibake decoder

日本語版はこちら

Paste text that came out as é, ’, 縺ゅ↑, ‚ ‚È‚½ or Ïðèâåò and this page works out which encoding it was read with, reverses it, and shows the original — in your browser, with nothing sent anywhere.

Try an example:

Decoded in your browser. Nothing is sent anywhere.

Results update as you type. Every common encoding pair is tried and the most text-like result is listed first.

How the decoder works

Mojibake is not corruption. The bytes are almost always intact; they were decoded with the wrong encoding. So the fix is mechanical: encode the garbled text again with the encoding it was wrongly read as, then decode those bytes with the encoding they really were. This page does not know either side, so it tries every plausible pair and ranks the results by how much they look like real text — everyday kanji and hangul rather than rare ones, no leftover control characters or replacement marks, no single foreign character glued into a Latin word. The top result is usually right; when two look plausible, the label under each one tells you which story fits where the text came from.

Everything runs in your browser. The text you paste is never sent to a server, so support tickets, unreleased dialogue and customer data are safe to paste.

When text cannot be recovered

If the garbled text contains � (the Unicode replacement character) or question marks where letters should be, the program that produced it had already thrown those bytes away, and no decoder can bring them back. The rest of the string can still be recovered, which is often enough to tell what the original said and where the pipeline went wrong.

Double-encoded text — garbled text that was saved and then garbled again — shows up as longer runs such as é or ’. The decoder applies the reversal twice for the top results, and you can feed any result back in with "Decode this again".

What the result tells you about your pipeline

The encoding pair under each result is the actual diagnosis. "UTF-8 read as Shift_JIS" means a Japanese file was opened by a tool that assumed the Windows Japanese code page — what spreadsheet software on Japanese Windows does with a UTF-8 CSV. "UTF-8 read as Windows-1252" is the classic web and database mistake: a connection or column that still defaults to Latin-1. "Shift_JIS read as Windows-1252" means a file exported from a Japanese tool was opened on a non-Japanese system. Fix the boundary that made the wrong assumption, re-export, and the string-level repair becomes unnecessary.

Fixing the file, not just the string

Use this page to identify the pair, then re-decode the whole file with the right encoding instead of pasting strings one by one. Find-and-replace on garbled text misses combinations nobody listed and can corrupt text that merely looked similar. For localization CSVs, the encoding checker on this site reports which encoding a file really is before it reaches your build.

Related reading

Other free tools

All free tools