CSV encoding converter (UTF-8 BOM ⇄ Shift_JIS)

日本語版はこちら

Drop a CSV or any text file. The page works out whether it is UTF-8, Shift_JIS, EUC-JP or UTF-16, shows the first lines decoded, and saves a copy in the encoding you pick — all in your browser, with nothing sent anywhere.

  • Detects UTF-8 (with or without BOM), UTF-16LE/BE, Shift_JIS (CP932) and EUC-JP
  • UTF-8 with BOM — the one Excel on Japanese Windows opens correctly by double-click
  • UTF-8 without BOM — what Unity, Unreal, Godot, Git and most build tools expect
  • Shift_JIS (CP932) and EUC-JP for legacy Japanese tools, with a list of every character they cannot hold
  • Shows the line endings (CRLF / LF) so a round trip through Windows does not surprise your diff
  • Only the encoding changes: columns, quotes and newlines are written back exactly as read

Why Excel garbles a UTF-8 CSV

Excel on Japanese Windows assumes a CSV is in the system code page — Shift_JIS (CP932) — unless the file starts with a byte order mark. A UTF-8 file without a BOM is therefore decoded as Shift_JIS, and every Japanese character becomes two or three wrong ones: 縺薙s縺ォ縺。縺ッ instead of こんにちは. The bytes are fine; only the first assumption was wrong.

Adding the three BOM bytes (EF BB BF) tells Excel the file is UTF-8, and nothing else about the file changes. That is the whole fix for the most common complaint, and it is the default this page suggests for a UTF-8 file.

Which encoding to pick

UTF-8 with BOM is for a file a person will open in Excel. UTF-8 without BOM is for a file a program will read: Unity's localization importer, Unreal's string tables, Godot, gettext tools and most CSV libraries treat a BOM as part of the first column name, which is how a key called id is born. Shift_JIS is only for tools that cannot read anything else — it has no room for emoji, most of the Unicode symbols, or the characters a game adds beyond JIS X 0208, and this page lists every character the conversion would lose before you download.

How detection works

A BOM decides UTF-8, UTF-16LE and UTF-16BE outright. Without one, the bytes are decoded as strict UTF-8 first; a valid UTF-8 file almost never happens by accident, so that is certain. Only when the bytes are not UTF-8 does it come down to Shift_JIS versus EUC-JP, and both decode most byte pairs into something, so the two results are compared on how much real Japanese they contain. That case is marked as a guess — check the preview before converting.

Checking a whole translation file, every time it changes?

Loclint checks your CSV or JSON files per project — placeholders, tags, missing or untouched strings, length limits, duplicate keys — plus an AI review pass, and tracks which issues are new, fixed, or back since the last run. The free plan needs no card.

Related reading

Other free tools

All free tools