Shift_JIS vs CP932: the differences that break Japanese text
Shift_JIS and CP932 are two different mapping tables that share almost every byte. A file a Japanese Windows tool labels Shift_JIS is almost always CP932, also called Windows-31J, and the disagreement between them is small enough to pass casual testing and specific enough to corrupt production data. Running both tables against every possible two-byte sequence gives the exact size of the gap: exactly six byte pairs decode to different Unicode characters, 457 assigned characters exist only in CP932, and one byte value, 0x5C, is shared between the backslash and the second half of 42 assigned double-byte characters.
The short answer for a pipeline: decode with a converter you have named CP932 explicitly, store UTF-8, and verify by round-tripping the specific characters below rather than a sentence of ordinary kanji, which survives either table. Every number here came from running the conversion in Python 3.9, Node 24 and the iconv shipped with macOS, so each one is reproducible on your own machine.
Exactly six byte pairs decode to two different characters
Decoding every two-byte sequence with Python's strict shift_jis codec, which implements the JIS X 0208 table, and again with its cp932 codec, which implements Microsoft's, leaves six sequences that are valid under both tables and mean something different in each. Six, not a vague handful — the full set is in the table below.
The wave dash at 0x8160 is the pair that reaches real game text, because Japanese uses it for ranges and approximation: chapter numbering, level bands, price ranges. The minus sign and the fullwidth hyphen-minus at 0x817C are the sneakier pair, nearly identical to the eye and never equal to a string comparison.
The two tables also fail in opposite directions on the way in, and that asymmetry matters more than the table itself. CP932 substitutes silently: encoding U+301C produces 0x8160, and decoding 0x8160 produces U+FF5E, so the character changes identity during a round trip with no error raised anywhere. Strict Shift_JIS refuses instead: encoding U+FF5E raises an error outright, so a build step that writes strict Shift_JIS crashes on text that came from a Windows tool. The crash is the better outcome, because the silent substitution is the one that reaches players.
The downstream symptom is that string equality stops working across a conversion boundary. Two strings that both display as a range separator compare unequal, which makes a translation memory miss, a duplicate-detection pass treat one string as two, and a diff report a change nobody made. When text looks right but behaves wrong, the neighbouring failure modes are collected in the causes of mojibake and how to fix them.
# decode: same two bytes, two different characters bytes strict shift_jis cp932 (Microsoft) 8160 U+301C 〜 WAVE DASH U+FF5E ~ FULLWIDTH TILDE 8161 U+2016 ‖ DOUBLE VERTICAL LINE U+2225 ∥ PARALLEL TO 817C U+2212 − MINUS SIGN U+FF0D - FULLWIDTH HYPHEN-MINUS 8191 U+00A2 ¢ CENT SIGN U+FFE0 ¢ FULLWIDTH CENT SIGN 8192 U+00A3 £ POUND SIGN U+FFE1 £ FULLWIDTH POUND SIGN 81CA U+00AC ¬ NOT SIGN U+FFE2 ¬ FULLWIDTH NOT SIGN # encode: what each codec will accept U+301C shift_jis -> 8160 cp932 -> 8160 (decodes back as U+FF5E) U+FF5E shift_jis -> UnicodeEncodeError cp932 -> 8160 U+00A5 shift_jis -> 5C cp932 -> UnicodeEncodeError
457 characters exist only in CP932, and 396 sit in two places
CP932 fills code points JIS X 0208 left undefined with two blocks: the NEC special characters in the 0x87 lead-byte row, and the IBM extended characters at lead bytes 0xED and 0xEE. Leaving out the user-defined area described below, 845 two-byte sequences decode under cp932 and raise an error under strict shift_jis, covering 457 distinct characters. A further 1,880 sequences at lead bytes 0xF0 through 0xF9 are the user-defined area, where a studio or a font vendor could install its own glyphs; those map to Unicode private-use code points and mean nothing outside the system that defined them.
These are the characters usually called 機種依存文字, platform-dependent characters, and none of them are exotic. Circled numbers appear in ranking screens and numbered dialogue. 髙 and 﨑 are ordinary surname kanji that show up in a credits roll. A strict Shift_JIS encoder cannot write a single one of them, which is the usual reason a credits export fails while the rest of the script converts cleanly.
The harder problem in the same region is duplication. 396 characters are reachable from two or more different byte sequences — 394 from two, and ¬ and ∵ from three — because the NEC-selected IBM extension block at lead bytes 0xFA to 0xFC re-encodes characters that already exist elsewhere. 髙 is both 0xEEE0 and 0xFBFC. The Roman numerals occupy 0x8754 to 0x875D and again 0xFA4A to 0xFA53.
Converters disagree about which form to write, and the disagreement is easy to reproduce. Python's cp932 encoder chose 0xEEE0 for 髙; the iconv shipped with macOS, asked for CP932, chose 0xFBFC. Both outputs are valid CP932 and display identically. They are not byte-identical, so a checksum, a byte-level diff, or a content hash used for change detection will report a difference that does not exist. Compare CP932 files after decoding them to Unicode, never as bytes.
character shift_jis cp932 shift_jisx0213 ① U+2460 circled digit one error 8740 8740 ㈱ U+3231 partnership sign error 878A 878A № U+2116 numero sign error 8782 8782 ℡ U+2121 telephone sign error 8784 8784 Ⅰ U+2160 roman numeral one error 8754 8754 髙 U+9AD9 surname kanji "taka" error EEE0 error 﨑 U+FA11 surname kanji "saki" error ED95 9892 纊 U+7E8A first IBM ext kanji error ED40 EDB5
The 0x5C problem: why 表示 can turn into 侮ヲ, and only sometimes
CP932 mixes single and double-byte characters, so a byte's meaning depends on what preceded it. Deriving the ranges by brute force rather than trusting a summary gives lead bytes 0x81 to 0x84, 0x87 to 0x9F, 0xE0 to 0xEA, 0xED to 0xEE and 0xF0 to 0xFC; trail bytes 0x40 to 0x7E and 0x80 to 0xFC; and single-byte half-width katakana at 0xA1 to 0xDF.
Because the trail range starts at 0x40, which ASCII characters can hide inside a double-byte character is fully decidable, and the answer is a short list worth memorising. It settles most tooling arguments on the spot. Splitting raw CP932 bytes on a comma, a tab or a newline is safe, which is why naive CSV readers mostly survive Japanese files. Splitting on a pipe, treating 0x5C as an escape introducer, or scanning for brace-delimited placeholders at byte level is not safe. Printf-style placeholders are safe because the percent sign sits below 0x40; brace-style placeholders are not, because 0x7B and 0x7D each appear as the second byte of 40 different characters. Identifier-shaped text hides this best: every ASCII letter and the underscore are valid trail bytes while digits are not, so scanning snake_case keys byte by byte through CP932 data is unsafe even where scanning for a comma is fine.
42 assigned characters carry 0x5C as their second byte, and the common ones are everyday vocabulary: ソ 0x835C, 十 0x8F5C, 申 0x905C, 能 0x945C, 表 0x955C, 予 0x975C, 構 0x8D5C, 圭 0x8C5C, 貼 0x935C, and the horizontal bar 0x815C. Feeding those bytes through a C-style unescaper, the kind that turns a backslash plus one character into that character, produces the results at the bottom of this section — run for real rather than reasoned about.
One of those rows comes back unchanged, and that is why the bug survives testing. 表 at the end of a string leaves its 0x5C with no following byte to swallow, so the same character corrupts or does not depending on where it sits — provided the unescaper keeps a trailing lone backslash rather than dropping it. Any test fixture that happens to end in one of those 42 characters passes.
The yen sign is the same byte seen from the other side. Python's strict shift_jis encodes U+00A5 to the single byte 0x5C, while 0x5C decodes back to U+005C, so that mapping is deliberately one-way; cp932 refuses U+00A5 outright. One byte, two glyphs: drawn with a Japanese-locale font it is a yen sign, elsewhere a backslash, which is also why Windows paths in Japanese screenshots look like they are separated by yen signs. The rendering side of this is covered in garbled Japanese text and the Windows system locale.
the second byte is always 0x40-0x7E or 0x80-0xFC, so:
CAN be the second byte — byte-level scanning is unsafe:
@ A-Z [ \ ] ^ _ ` a-z { | } ~
CANNOT be the second byte — byte-level scanning is safe:
TAB CR LF space ! " # $ % & ' ( ) * + , - . /
0-9 : ; < = > ?
a C-style unescaper (backslash + one char becomes that char) over real CP932 bytes:
表示 955C8EA6 -> 侮ヲ
ソート 835C815B8367 -> メ[ト
十分 8F5C95AA -> 助ェ
能力 945C97CD -> 迫ヘ
構造 8D5C91A2 -> 国「
データ表 8366815B835E955C -> データ表 (unchanged)Which encoding name means which table in your toolchain
Encoding names are not consistent between tools, and the trap is that a tool can accept the string Shift_JIS and hand you Microsoft's table anyway. Decoding 0x8160 is a one-command test for any converter: U+301C means the JIS table, U+FF5E means Microsoft's.
- iconv on macOS: SHIFT_JIS, SJIS, MS_KANJI and SHIFT_JISX0213 all return U+301C, while CP932 and WINDOWS-31J return U+FF5E. Despite the name, MS_KANJI is the JIS table here — do not read it as an alias for CP932.
- SHIFT_JISX0213 is not a superset of CP932 and is not a substitute for it. It refuses 髙 outright, and where it does cover a character it may place it elsewhere: it encodes 纊 as 0xEDB5 and 﨑 as 0x9892, against CP932's 0xED40 and 0xED95.
- Node: TextDecoder with the label shift_jis returns U+FF5E and decodes ①, ㈱, 髙 and 﨑, because the WHATWG Encoding Standard defines that label as Microsoft's table. The labels shift-jis, sjis, windows-31j, x-sjis, ms_kanji and csshiftjis all resolve to the same decoder, and the label cp932 throws a RangeError. TextEncoder only produces UTF-8, so the platform reads CP932 and cannot write it.
- Python: shift_jis and cp932 are separate codecs behaving as shown above, so the codec name you type decides the table. There is no ambiguity, only a choice that is easy to make by accident.
- Java: Shift_JIS and windows-31j are distinct charsets. Check the supported-encodings page in the official documentation for your JDK rather than assuming, because which charsets are required and which are optional has changed between versions.
- Excel on Japanese Windows: the plain comma-delimited CSV option writes the system ANSI code page, which is CP932 on a Japanese system, and recent versions add a separate UTF-8 CSV option. Confirm which save-as entries your version offers in the official documentation, because the menu wording differs by release.
Converting safely: one direction, and a round trip you run
None of these tools is wrong. They document different tables under overlapping names, which is why a name in a spreadsheet dialog or a filename is a hypothesis and the decode of 0x8160 is the evidence. Convert at the boundary, in one direction, and verify with characters that can fail rather than characters that cannot.
- Decode incoming files with a converter named CP932 or Windows-31J explicitly. It reads everything strict Shift_JIS reads plus the 457 extension characters, so the permissive choice is the correct one for real-world files.
- Normalize to UTF-8 immediately and keep CP932 bytes out of every later stage. Add the byte-order mark EF BB BF only when a spreadsheet on Japanese Windows must open the file by double-click; omit it when a parser reads the file, or the mark arrives as U+FEFF glued to your first column name.
- When a target genuinely requires CP932, encode and immediately decode the result, then compare character by character against the original and report the differing indices. Running that on 表示ソート十分〜①㈱髙 100% gives exactly one difference, at index 7, where U+301C became U+FF5E.
- Decide once what happens to U+301C and write the decision down. Either normalize it to U+FF5E before encoding so the change is deliberate, or reject it at import. Finding it during a round trip and shrugging is how both forms end up mixed in one database.
- Detect rather than trust: try UTF-8 first, because Japanese CP932 text is almost never valid UTF-8. Treat that as strong evidence and not proof — 990 two-byte sequences decode as both CP932 and UTF-8 — and keep the declared encoding as a tiebreak for short fields.
- Never accept the absence of an exception, or an unchanged byte count, as verification. CP932's silent substitutions raise nothing and change no lengths.
The pattern behind all of it
The lesson outruns Japanese text. A declared encoding name states an intent, a superset is exactly where low-frequency corruption hides, and the only trustworthy check is a round trip you actually ran on the characters capable of failing. The same discipline applies to double-encoded UTF-8 and how to recover it, and a wider symptom index lives in the mojibake patterns cheat sheet. If you are adding Japanese to a project now, starting in UTF-8 removes this whole class of problem, and the rest of that work is described in what changes when you localize a game into Japanese.