CSV encoding converter (UTF-8 BOM ⇄ Shift_JIS)
日本語版はこちらDrop a CSV or any text file. The page works out whether it is UTF-8, Shift_JIS, EUC-JP or UTF-16, shows the first lines decoded, and saves a copy in the encoding you pick — all in your browser, with nothing sent anywhere.
- Detects UTF-8 (with or without BOM), UTF-16LE/BE, Shift_JIS (CP932) and EUC-JP
- UTF-8 with BOM — the one Excel on Japanese Windows opens correctly by double-click
- UTF-8 without BOM — what Unity, Unreal, Godot, Git and most build tools expect
- Shift_JIS (CP932) and EUC-JP for legacy Japanese tools, with a list of every character they cannot hold
- Shows the line endings (CRLF / LF) so a round trip through Windows does not surprise your diff
- Only the encoding changes: columns, quotes and newlines are written back exactly as read
Why Excel garbles a UTF-8 CSV
Excel on Japanese Windows assumes a CSV is in the system code page — Shift_JIS (CP932) — unless the file starts with a byte order mark. A UTF-8 file without a BOM is therefore decoded as Shift_JIS, and every Japanese character becomes two or three wrong ones: 縺薙s縺ォ縺。縺ッ instead of こんにちは. The bytes are fine; only the first assumption was wrong.
Adding the three BOM bytes (EF BB BF) tells Excel the file is UTF-8, and nothing else about the file changes. That is the whole fix for the most common complaint, and it is the default this page suggests for a UTF-8 file.
Which encoding to pick
UTF-8 with BOM is for a file a person will open in Excel. UTF-8 without BOM is for a file a program will read: Unity's localization importer, Unreal's string tables, Godot, gettext tools and most CSV libraries treat a BOM as part of the first column name, which is how a key called id is born. Shift_JIS is only for tools that cannot read anything else — it has no room for emoji, most of the Unicode symbols, or the characters a game adds beyond JIS X 0208, and this page lists every character the conversion would lose before you download.
How detection works
A BOM decides UTF-8, UTF-16LE and UTF-16BE outright. Without one, the bytes are decoded as strict UTF-8 first; a valid UTF-8 file almost never happens by accident, so that is certain. Only when the bytes are not UTF-8 does it come down to Shift_JIS versus EUC-JP, and both decode most byte pairs into something, so the two results are compared on how much real Japanese they contain. That case is marked as a guess — check the preview before converting.
Checking a whole translation file, every time it changes?
Loclint checks your CSV or JSON files per project — placeholders, tags, missing or untouched strings, length limits, duplicate keys — plus an AI review pass, and tracks which issues are new, fixed, or back since the last run. The free plan needs no card.