File formats & standardsこの記事を日本語で読む

BCP 47 language tags: what ja-JP, zh-Hans, and pt-BR actually mean

A BCP 47 language tag is a language subtag, optionally followed by a script, a region and a few rarer parts, joined with hyphens: ja, ja-JP, zh-Hans, zh-Hant-HK, pt-BR. The short answer to the question that brings most people here is that Chinese is split by script and Portuguese and Spanish are split by region, because script decides which characters appear on screen while region decides which words are used.

The rest of the confusion is mechanical, and you can settle it in a terminal instead of in a specification. Any runtime with a complete Intl implementation will canonicalize a tag and reject one that is not a tag at all. Every tag output shown below came from running Node 22 rather than from memory.

How a tag is built, and which registry each subtag comes from

BCP 47 is not a single document. It is a Best Current Practice label that currently points at two RFCs: RFC 5646 defines the syntax of tags, and RFC 4647 defines how tags are matched against each other. Neither one contains the list of valid codes. That list is the IANA Language Subtag Registry, and it is the only place worth checking when you are unsure whether a code exists.

Subtags appear in a fixed order, and anything out of order is not a tag no matter how sensible it looks.

  • Language, first and required. Two letters when ISO 639-1 has a code (ja, en, zh, pt), three letters from ISO 639-2 or ISO 639-3 otherwise (fil for Filipino, ceb for Cebuano).
  • Script, four letters, from ISO 15924. Hans is Han Simplified and Hant is Han Traditional. Include a script subtag only when the language is genuinely written in more than one script and you need to tell them apart.
  • Region, two letters from ISO 3166-1 alpha-2 (JP, US, BR), or three digits from the UN M.49 list for areas larger than a country. 419 is Latin America and the Caribbean, which is why es-419 exists.
  • Variant, individually registered, five to eight alphanumerics or a digit followed by three alphanumerics. The 1996 in de-DE-1996 marks the German orthography reform. You will almost never need one.
  • Extension, a single-letter singleton followed by its own subtags. The u- extension carries Unicode locale settings such as calendar and numbering system, so th-TH-u-nu-thai asks for Thai digits. The t- extension carries transform and source-language information.
  • Private use, x- followed by whatever you like. Valid syntax, and meaningless outside your own system.
language[-script][-region][-variant...][-extension...][-x-private]

zh-Hant-HK
 |   |   |
 |   |   +-- region, ISO 3166-1 alpha-2: Hong Kong
 |   +------ script, ISO 15924: Han Traditional
 +---------- language, ISO 639-1: Chinese

es-419             region from UN M.49, not a country
de-DE-1996         registered variant: 1996 orthography
th-TH-u-nu-thai    u- extension: Thai numbering system
en-US-x-private    x- private use: meaningful only to you

Canonical form and deprecated codes, straight from the runtime

Do not hand-normalize tags. Intl.getCanonicalLocales applies the case conventions and the alias table in one call and throws on input that is not a tag. RFC 5646 recommends lowercase for the language, Titlecase for the script and uppercase for the region, which is why zh-Hant-HK is spelled that way. That recommendation is a readability convention rather than a rule of the grammar, because tags are compared case-insensitively: getCanonicalLocales collapses ja-jp, JA-JP and ja-JP into the single tag ja-JP.

The four short replacements in that output are deprecations you will still meet in old data and in spreadsheets from clients. iw was Hebrew before he, in was Indonesian before id, ji was Yiddish before yi, and mo was Moldavian before ro. If a column header says iw, that column is Hebrew.

Two lines deserve care. tl becoming fil comes from CLDR's alias data rather than from a Preferred-Value in the IANA registry, so confirm the registry entry before telling anyone that a tag is formally deprecated. And no stays no: it is the Norwegian macrolanguage, while nb for Bokmal and nn for Nynorsk are separate live tags rather than replacements for it. If your translator delivered Bokmal, no is not the right name for that file.

What the function does not do is check that a code exists. It validates shape only, so jp and english both pass. Pairing it with maximize and with supportedLocalesOf turns it into a real validity check: CLDR has no likely subtags for jp, so maximize leaves it untouched, and the runtime has no data filed under it, so a lookup match drops it. Treat the last of those as a statement about locale data coverage rather than about tag validity.

$ node -e 'for (const t of ["ja-jp","JA-JP","zh-hans-cn","sr-latn","en-us-x-private",
    "iw","in","ji","mo","tl","no"]) console.log(t, "->", Intl.getCanonicalLocales(t)[0])'
ja-jp            ->  ja-JP
JA-JP            ->  ja-JP
zh-hans-cn       ->  zh-Hans-CN
sr-latn          ->  sr-Latn
en-us-x-private  ->  en-US-x-private
iw               ->  he
in               ->  id
ji               ->  yi
mo               ->  ro
tl               ->  fil
no               ->  no

Intl.getCanonicalLocales(["ja-jp","JA-JP","ja-JP"])  ->  ["ja-JP"]

Intl.getCanonicalLocales("jp")                       ->  ["jp"]
Intl.getCanonicalLocales("english")                  ->  ["english"]
new Intl.Locale("jp").maximize()                     ->  jp
Intl.NumberFormat.supportedLocalesOf(["ja-JP","zh-Hant-TW","jp","xx-YY"],
  { localeMatcher: "lookup" })                       ->  ["ja-JP","zh-Hant-TW"]

// all of these throw RangeError: Incorrect locale information provided
ja_JP   en-USA   ja-JPN   zh-CHS   zh-Hans-Hant   x-private   i-klingon

What a bare tag means, and why zh-Hans is not zh-CN

A one-subtag tag like zh or pt is not ambiguous to a localization library, because it has a documented default. Intl.Locale.maximize fills in the script and region that CLDR considers most likely.

Read that first block as a list of defaults you are accepting silently. Shipping zh means Simplified Chinese for mainland China. Shipping pt means Brazilian Portuguese rather than European. Shipping sr means Cyrillic, so a Latin-script Serbian delivery has to say sr-Latn. Shipping ar resolves to Egypt, which matters for date and number formatting more than for the strings themselves.

minimize is the mirror image and is mostly a trap. zh-Hans-CN minimizes all the way down to zh, which is exactly the ambiguous name you were trying to avoid, while zh-Hant-TW minimizes to zh-TW rather than zh-Hant because the algorithm drops the script before it drops the region. That is a large part of why zh-CN and zh-TW are everywhere: they are the shortest forms that still round-trip, so tools that minimize keep producing them. Use minimize to test whether two tags mean the same thing, never to name your resource files.

This gives a precise answer to zh-Hans versus zh-CN. They are not synonyms, and they answer different questions. zh-Hans says the text is written in Simplified Han characters and says nothing about where the reader lives. zh-CN says the reader is in mainland China and leaves the script to be inferred. Tag by script, because the script is a property your text genuinely has: a Simplified build read in Singapore or Malaysia is still zh-Hans, and under region tagging you would either mismatch those players or invent a region tag for every place Simplified is read. For how these defaults get decided in the first place, see what CLDR is and how locale data is decided.

new Intl.Locale(x).maximize()
  zh  ->  zh-Hans-CN      pt  ->  pt-Latn-BR      sr  ->  sr-Cyrl-RS
  ja  ->  ja-Jpan-JP      es  ->  es-Latn-ES      ar  ->  ar-Arab-EG
  en  ->  en-Latn-US      ko  ->  ko-Kore-KR      az  ->  az-Latn-AZ

new Intl.Locale(x).minimize()
  zh-Hans-CN  ->  zh        zh-Hant-TW  ->  zh-TW      zh-Hant-HK  ->  zh-HK
  ja-JP       ->  ja        pt-BR       ->  pt         en-Latn-US  ->  en

The locale list most games actually ship

This is a working starting list rather than a standard. Every tag in it is already canonical, checked by running all 25 through getCanonicalLocales and confirming that none of them changed, so it can be pasted into a config file without reformatting. The names beside them are the English names CLDR uses, taking CLDR's short region form where it has one, so Hong Kong rather than Hong Kong SAR China.

The Chinese rows are the decision that matters, and zh-Hans plus zh-Hant covers most releases. Add zh-Hant-HK only when you have Hong Kong vocabulary actually reviewed, not as a copy of the Taiwan file, because it differs from zh-Hant-TW in everyday words. es-419 is the pragmatic single Latin American Spanish when you cannot fund es-MX and es-AR and es-CO separately, and the differences between Latin American and Spain Spanish covers what changes in the text. pt-BR comes before pt-PT for almost every game, and what to watch for in Brazilian Portuguese covers the specifics.

The region subtag on en-US, ja-JP, de-DE and the rest carries almost nothing for string selection. Drop it if you are starting fresh, because ja and ja-JP select the same strings. The genuinely split pairs are the exception: en-US against en-GB, pt-BR against pt-PT, es-ES against es-419, fr-FR against fr-CA. There the region subtag is doing real work.

en-US       American English                  pl-PL   Polish (Poland)
en-GB       British English                   tr-TR   Turkish (Türkiye)
ja-JP       Japanese (Japan)                  ru-RU   Russian (Russia)
ko-KR       Korean (South Korea)              it-IT   Italian (Italy)
zh-Hans     Simplified Chinese                de-DE   German (Germany)
zh-Hant     Traditional Chinese               fr-FR   French (France)
zh-Hans-CN  Chinese (Simplified, China)       fr-CA   Canadian French
zh-Hant-TW  Chinese (Traditional, Taiwan)     ar      Arabic
zh-Hant-HK  Chinese (Traditional, Hong Kong)  th-TH   Thai (Thailand)
pt-BR       Brazilian Portuguese              vi-VN   Vietnamese (Vietnam)
pt-PT       European Portuguese               id-ID   Indonesian (Indonesia)
es-ES       European Spanish
es-419      Latin American Spanish
es-MX       Mexican Spanish

Platform names that look like language tags but are not

The most common cause of broken language switching is assuming that a platform speaks BCP 47. Several important ones do not, and they disagree with each other. Two rules keep this survivable. Hold one canonical BCP 47 tag as the identity of each language inside your own pipeline, and convert to each platform's spelling at the edge through a single table you can read at a glance. Then copy every row of that table out of the platform's own current documentation rather than from memory or from any article, including this one. These lists change when platforms add languages, and one wrong string means one language silently never loads.

  • Steam's API language codes are English words rather than tags, and no mechanical transformation gets you from a BCP 47 tag to them, so the lookup table is mandatory rather than a convenience. For the store-facing side, see what counts as a supported language on Steam.
  • Unity has two systems that do not match. The SystemLanguage enumeration is English words with no region axis at all, and its older Chinese member predates the Simplified and Traditional pair. The Localization package instead uses LocaleIdentifier, built on the .NET CultureInfo names, which are BCP 47-shaped such as ja-JP and zh-Hans. Mapping between the two is your code's job, and the exact enumeration members are worth reading in the SystemLanguage page of the Unity scripting reference.
  • Unreal uses culture names that are BCP 47-shaped (ja, zh-Hans, pt-BR, es-419) and names the folders under a localization target after them, so the tag and the directory name are the same string. Confirm the culture list for your engine version in the localization section of the Unreal documentation.
  • Android resource qualifiers are their own notation, and the two halves of it are easy to mix up. A region alone takes an r prefix, while anything carrying a script needs the BCP 47 form that puts b+ in front and replaces the hyphens with plus signs. The b+ form also needs a recent enough minimum API level, so check the qualifier table in the Android app resources documentation. See how strings.xml and resource qualifiers work.
  • Apple platforms name localization directories after the language ID with an lproj suffix. These are close to BCP 47, though very old projects can still carry legacy English names such as Japanese.lproj.
BCP 47      Steam        Unity SystemLanguage   Android           Apple
zh-Hans     schinese     ChineseSimplified      values-b+zh+Hans  zh-Hans.lproj
zh-Hant     tchinese     ChineseTraditional     values-b+zh+Hant  zh-Hant.lproj
ja-JP       japanese     Japanese               values-ja         ja.lproj
ko-KR       koreana      Korean                 values-ko         ko.lproj
pt-BR       brazilian    Portuguese             values-pt-rBR     pt-BR.lproj
es-419      latam        Spanish                values-b+es+419   es-419.lproj
es-ES       spanish      Spanish                values-es         es.lproj

Copy every row from the platform's own current documentation before shipping it.

Naming files and keys so that fallback actually fires

Matching itself is defined in RFC 4647 and has two modes. Filtering returns every tag that starts with the requested prefix. Lookup returns a single best match by truncating subtags from the right until something is found, skipping single-character subtags as it goes. Lookup is what a resource loader wants: a request for zh-Hant-HK tries zh-Hant-HK, then zh-Hant, then zh, then your hardcoded default.

The second result below is the whole problem. zh-Hans is not a prefix of zh-Hant-HK, so truncation walks straight past it to zh and then to nothing. Fallback only works down a chain of prefixes, which means your file names have to be those prefixes. These are the mistakes that produce that situation:

  • Splitting Chinese at zh only. A single zh file forces Simplified and Traditional readers onto the same strings, and the symptom is wrong characters rather than a missing file, so nobody files a bug until review. Ship zh-Hans and zh-Hant, and add a bare zh only as a deliberate default.
  • Splitting Portuguese or Spanish at pt or es only. The strings load and nothing looks broken, while Portuguese players read Brazilian spellings or the reverse. Ship pt-BR with pt-PT, and es-ES with es-419.
  • Case mismatch between the tag and the file system. Tag comparison ignores case, but your loader's string comparison does not, and neither does a case-sensitive file system. Directory names, CSV column headers and dictionary keys are compared byte for byte, so normalize once at the boundary with getCanonicalLocales, store that form, and name the files with it.
  • Underscores. ja_JP is a POSIX-style locale identifier, not a BCP 47 tag: strict parsers reject it and getCanonicalLocales throws on it. Convert underscores to hyphens at the boundary instead of letting both spellings circulate.
  • Inventing codes. jp is not Japanese. It passes the syntax check, maximizes to nothing, and resolves to nothing, which is the worst combination because no error appears. The language is ja and JP is the region.
  • Over-specifying with no broader file behind it. If your only Traditional Chinese file is named zh-Hant-TW, a zh-Hant-HK request finds nothing at all. Name files at the level you actually distinguish.
  • Fixing the file names and letting the keys drift. The two are one discipline, and designing string keys that survive localization covers the key side.
function lookup(requested, available) {
  const parts = requested.toLowerCase().split("-");
  const have = new Map(available.map((t) => [t.toLowerCase(), t]));
  while (parts.length > 0) {
    const candidate = parts.join("-");
    if (have.has(candidate)) return have.get(candidate);
    parts.pop();
    if (parts.at(-1)?.length === 1) parts.pop(); // skip a singleton
  }
  return null;
}

lookup("zh-Hant-HK", ["en", "zh-Hant", "ja"])  ->  "zh-Hant"
lookup("zh-Hant-HK", ["en", "zh-Hans", "ja"])  ->  null
lookup("ja-JP", ["en", "ja"])                  ->  "ja"
lookup("th-TH-u-nu-thai", ["en", "th"])        ->  "th"

Related articles