Character code inspector (code points and bytes)

日本語版はこちら

Paste any string — a key that will not match, a line that is one character "too long", a value that looks identical to another but is not — and this page lists each character with its code point, its bytes in four encodings, and a note when it is something invisible or easily mistaken.

Try an example:

Inspected in your browser. Nothing is sent anywhere.

  • Unicode code point (U+XXXX) and the bytes in UTF-8, UTF-16LE, Shift_JIS (CP932) and EUC-JP for every character
  • Four lengths side by side: visible characters, code points, UTF-16 units (JavaScript / C# length) and bytes
  • Fullwidth brackets, commas, percent signs and digits that look like ASCII but match nothing
  • Ideographic spaces, no-break spaces, zero-width spaces and joiners, byte order marks, soft hyphens
  • Emoji and other characters that take two UTF-16 units — the ones that count as 2 in JavaScript and C#
  • Characters Shift_JIS or EUC-JP cannot represent, marked before they turn into ?

Why the same string has four different lengths

A string has no single length. JavaScript, C# and Java report UTF-16 units, so an emoji counts as 2 and a family emoji with joiners as 11. Python 3, Rust and Swift report code points, where the emoji is 1 and the family is 7. A database column or a packet limit counts bytes — 4 for the emoji in UTF-8, none in Shift_JIS, which has no emoji at all. And what a player sees is graphemes: one family. A "20 characters max" rule means something different in each layer, and this page shows all four so the rule can say which it means.

The characters that look identical and are not

Japanese input produces fullwidth punctuation by default: ( instead of (, % instead of %, 012 instead of 012, and a ideographic space (U+3000) where a plain space was wanted. All of these render like their ASCII twins and match none of them, which is how a placeholder written as {count} reaches players as literal text and why a key with a trailing ideographic space never resolves. Copy-paste from documents adds the invisible ones: no-break spaces from web pages, zero-width spaces from rich-text editors, a byte order mark at the start of a value that came from a file. The table highlights every one of these so the character that is actually different is the one you are looking at.

Reading the byte columns

The hex bytes are what a file in that encoding contains for the character. A dash means the encoding has no code for it, so saving the string in that encoding will replace the character with ? — which is how a Shift_JIS export silently drops the ① and ㈱ characters of the Windows extensions, or any emoji. Compare the UTF-8 column with what a hex dump of your file shows at the same position to confirm what the file actually is, rather than what its name claims.

Checking a whole translation file, every time it changes?

Loclint checks your CSV or JSON files per project — placeholders, tags, missing or untouched strings, length limits, duplicate keys — plus an AI review pass, and tracks which issues are new, fixed, or back since the last run. The free plan needs no card.

Related reading

Other free tools

All free tools