ã and ’: fixing double-encoded UTF-8 text
If your text shows é where an accented letter belongs, ’ where a curly apostrophe belongs, or a run of ã where Japanese belongs, the bytes were read once as Windows-1252 or Latin-1 and then saved as UTF-8 again. That is double-encoded UTF-8: two separate mistakes, each of which reported success, leaving a file that passes every UTF-8 validity check and is still wrong.
The rest of this article is the part that is hard to find: a lookup table built by running the corruption rather than describing it, a count of how many layers you are looking at, one command that repairs a file, and the exact character classes where that command truncates your data and returns a non-zero exit status you have to check for.
The two steps, byte by byte
A UTF-8 character outside ASCII is a run of two to four bytes. A single-byte encoding such as Windows-1252 assigns a character to almost every byte value on its own and has no concept of a multi-byte run, so it turns each byte of one UTF-8 character into a separate letter or symbol. That first mistake is ordinary mojibake, covered in the causes and fixes for garbled text.
Double encoding adds a second step: that already-wrong string of separate characters is written out as UTF-8. Every wrong character is a legitimate Unicode character, so this step cannot fail. What lands on disk is a longer, perfectly well-formed UTF-8 file whose contents are the wrong characters.
The part almost no explanation makes explicit is that what you see next depends on the lens you look through, and that is what tells you how deep the damage goes. Read the double-encoded bytes correctly as UTF-8 and you see the single-step form, so the file looks like an ordinary one-mistake problem. Read the same bytes as Windows-1252 and you see the doubled form. In other words, é and ’ and ã mean two things are wrong at the same time: the stored bytes are double-encoded, and whatever is displaying them is also guessing the encoding wrong.
original café
bytes: 63 61 66 C3 A9
step 1 those bytes read as Windows-1252
C3 -> Ã A9 -> ©
string is now café (5 characters, no error raised)
step 2 that 5-character string written out as UTF-8
bytes: 63 61 66 C3 83 C2 A9 <- this is the file you have
reading it back, two ways:
as UTF-8 -> café looks like one mistake
as Windows-1252 -> café the doubled form you searched forA lookup table from the garbage back to the character
Every row below was produced by encoding the character as UTF-8, decoding those bytes as Windows-1252, encoding the result as UTF-8, and decoding that as Windows-1252 again. Nothing in it is transcribed from another page. Braces mark a character that renders as nothing, a box, or a stray space in most viewers, which matters later.
Use the table as a depth gauge. Find your garbage in the middle column and the data is one layer deep. Find it in the right-hand column and it is two, and the repair has to be applied twice.
Length is the other tell, and it is sharper than it looks. One accented Latin letter is two bytes in UTF-8, so it becomes two characters at one layer and four at two, or five for 17 of the 64 code points in U+00C0 to U+00FF, 16 of them uppercase letters: É doubles to five characters. A curly apostrophe, an en dash, an ellipsis or a euro sign is three bytes, and so is one Japanese character: three characters at one layer, then six to eight at two. Across the 20,992 ideographs in U+4E00 to U+9FFF the two-layer count is six for 11,421 of them, seven for 8,126 and eight for 1,445. That is why a doubled Japanese string looks catastrophic while a doubled French string merely looks odd: the same rule applied to text where nearly every character is affected.
character one layer two layers
-------------------------------------------
é é é
ü ü ü
ñ ñ ñ
’ (apostrophe) ’ ’
– (en dash) – –
“ (open quote) “ “
… (ellipsis) … …
€ (euro) € €
あ ã{81}‚ ãÂ{81}‚
日本 日本 日本
ゲーム ゲーãƒ{A0} ゲーãƒÂ{A0}
{81} = U+0081 {A0} = U+00A0 no-break spaceWhy Japanese breaks the repair tools and French does not
Windows-1252 leaves five byte values unassigned: 0x81, 0x8D, 0x8F, 0x90 and 0x9D. Implementations disagree about what to do with them, and that disagreement is the single biggest practical difference between repairing European and Japanese text.
Decoders that follow the WHATWG Encoding Standard, which is what browsers and the TextDecoder API in Node use, map those five bytes to the control characters U+0081, U+008D, U+008F, U+0090 and U+009D. The round trip is therefore lossless. The iconv command shipped with macOS rejects them as an illegal byte sequence instead. Other iconv builds may differ, so test yours before trusting it with a repair.
How often does that matter? Counting characters whose UTF-8 encoding contains one of the five bytes: 67 of the 86 hiragana in U+3041 to U+3096, 3,115 of the 20,992 ideographs in U+4E00 to U+9FFF, but only 5 of the 90 katakana in U+30A1 to U+30FA, and only 5 of the 64 code points in U+00C0 to U+00FF. The five Latin ones are Á, Í, Ï, Ð and Ý, all uppercase. So almost any sentence containing hiragana is affected, while lowercase French, Spanish or German text usually is not.
Two consequences follow, and both cost people hours. First, single-layer Japanese mojibake contains invisible characters. Breaking ゲームを保存 one layer produces an 18-character string holding U+009D, U+00A0 and U+00AD, which a terminal or a browser renders as nothing or as a plain space. Copy that out of a log and paste it into a fixer and you have already destroyed the bytes the repair needs, so the repair reports failure on text that was recoverable. Work on the file, never on a copy-paste of what the screen showed you.
Second, the Encoding Standard's label table treats iso-8859-1 and latin1 as aliases of windows-1252, so a page or an XML declaration announcing ISO-8859-1 is actually decoded as Windows-1252 by every conforming browser. Do not assume the label in the header tells you which of the two did the damage. The cheat sheet of mojibake patterns is the faster way to identify the pair from the output.
Repairing a file, and where the obvious command truncates
The repair mirrors the damage in reverse. Take the wrong characters, encode them back into bytes with the single-byte encoding that produced them, and the bytes that come out are the original UTF-8. As a whole-file operation that is one iconv invocation, converting from UTF-8 to Windows-1252.
On a one-layer English file that works exactly as advertised. Running it on a file holding the one-layer forms of an apostrophe, an em dash, an e-acute, an i-diaeresis and a euro sign returned the original bytes byte for byte, with exit status 0.
On Japanese it does not. The same command on a one-layer file holding ゲームを保存, the six Japanese characters for save the game, printed iconv: iconv(): Illegal byte sequence, exited 1, and still wrote an output file: 14 bytes of the 19 expected, which are 18 for the text plus a newline. The cut falls inside a character, so the partial output is not even valid UTF-8. It stopped at the third byte of 保, which is 0x9D. A two-layer English file fails the same way, because one layer of repair leaves a closing curly double quote in the text and that character's third UTF-8 byte is also 0x9D.
Two rules follow. Check the exit status and the output size on every run. And never substitute ISO-8859-1 as the target when the damage was Windows-1252. On macOS that conversion silently transliterates: U+201A became a backtick, U+0192 became the letter f, U+02DC became a tilde and a curly apostrophe became 0xB4, all with exit status 0, producing a file that looks repaired and is not.
When iconv refuses, use a decoder that follows the Encoding Standard. The snippet below builds the reverse mapping from the decoder itself, so it never guesses, and it uses a strict UTF-8 decode as its own guard: a row that was never double-encoded fails to decode and is returned untouched instead of being damaged. Run it in a loop and it recovered one-layer and two-layer English and Japanese exactly, and left clean ASCII, clean Japanese and a correct café alone.
Apply it per cell rather than per file when a CSV is in a mixed state, which is the usual case. Looping it while it keeps returning a different string peels one layer per pass, so you do not have to decide in advance whether you are looking at one layer or two.
One caution about looping: stop as soon as a pass returns null, and keep the original value for any cell where the first pass already returns null. Repairing a cell that was never broken is the one way this process loses data that no later pass can recover.
# whole file, one layer, European text: works
iconv -f UTF-8 -t WINDOWS-1252 broken.csv > fixed.csv
echo $? # must be 0 -- 1 means truncated output
# same command, Japanese: exit 1, 14 of 19 bytes written, cut mid-character
# NEVER do this instead -- silently transliterates, exit 0:
# iconv -f UTF-8 -t ISO-8859-1 broken.csv > fixed.csv
# ---- per-cell version for Node or the browser ----
const WIN1252 = new TextDecoder("windows-1252");
const toByte = new Map();
for (let b = 0; b < 256; b++) toByte.set(WIN1252.decode(Uint8Array.of(b)), b);
// Sanity check: a conforming decoder maps byte 0x83 to U+0192, not U+0083.
// If this is false, the runtime's windows-1252 is not the standard one.
WIN1252.decode(Uint8Array.of(0x83)).codePointAt(0) === 0x192;
// Returns the text one layer less broken, or null if it was not
// double-encoded through Windows-1252 at all.
function undoOneLayer(text) {
const bytes = [];
for (const ch of text) {
const b = toByte.get(ch);
if (b === undefined) return null; // never came from Windows-1252
bytes.push(b);
}
try {
return new TextDecoder("utf-8", { fatal: true }).decode(Uint8Array.from(bytes));
} catch {
return null; // not double-encoded -- leave it alone
}
}
undoOneLayer("Don’t stop") // "Don’t stop"
undoOneLayer("Don’t stop") // null (already correct)Finding the boundary that did the second encode
Repairing the data without finding the boundary means repairing it again next week. Four places account for most cases, and each leaves a different signature.
A database connection character set that disagrees with the column is the classic one, and it comes in two shapes that need opposite responses. In the first, the server treats incoming UTF-8 bytes as single-byte characters and re-encodes them on the way in, which is exactly step two, so the column really does hold double-encoded bytes. Reading that row back over a single-byte connection looks correct only because the read undoes the write, and the stored data still has to be repaired. In the second, the column holds the original UTF-8 and only the declaration on the read path is wrong, so changing the setting is the whole fix and there is nothing to repair. Dump the column's raw bytes as hex and compare them with the UTF-8 of the text you expect: that, not a round trip through a connection, is what tells the two apart. Setting names differ by engine and version, so check your engine's own manual chapter on character sets and collations.
A spreadsheet round trip is the second. Opening a UTF-8 CSV in a tool that guesses a single-byte encoding, then saving it back as UTF-8, performs both steps in one sitting.
An HTTP response or a form submission with no character set in the Content-Type header is the third. The client falls back to a locale-dependent default, which is Windows-1252 across most Western locales, and a header that does declare ISO-8859-1 resolves to Windows-1252 as well. The signature here is that the stored data is clean and only the display is wrong, which means there is nothing to repair once the header is fixed.
An editor's save-with-another-encoding command is the fourth, and the signature is unmistakable: only the files one person touched are affected, and only from one commit onward. The most useful signal across all four is uniformity. A file where some rows are doubled and others are clean points at a per-row or per-cell path, so look at the code that writes individual values, not at a whole-file conversion. When the surviving characters in a Japanese file are the katakana rather than the hiragana, suspect a code page conversion rather than double encoding, and check the Shift_JIS and CP932 pitfalls and garbled Japanese from the Windows locale instead.
Making the second encode impossible
Declaring UTF-8 explicitly at every boundary is the real fix, and the boundary that bites is never one of the two visible ends. Audit the middle: the connection string, the response header, the file open call, the export routine. An end-to-end test that round-trips one string containing an accented Latin letter, the hiragana あ and one emoji covers every case above: those are two, three and four UTF-8 bytes, and あ contains the unassigned byte 0x81 while 237 of the 768 characters from U+1F300 to U+1F5FF contain one too.
Add a content check that does more than confirm the bytes are valid UTF-8, because double-encoded data always is. Scan for the marker sequences instead: the two-character run starting with U+00C3 followed by U+0192, and the three-character run U+00E2 U+20AC, are reliable enough to fail a build on, and they cost nothing to test. Repair once, store the repaired values, and never repair on read.
Finally, keep the pre-repair file until someone has read the repaired text in every affected language. The strict decode above protects individual cells, but only a human reading the output can confirm that the recovered text is the text that was supposed to be there.