File formats & standardsこの記事を日本語で読む

What is Unicode CLDR? The locale data behind Intl and ICU

CLDR, the Unicode Common Locale Data Repository, is the versioned dataset that records how each locale writes dates, times, numbers, currency, plurals, lists, relative times, names and sort order. It is data, not code. ICU, the International Components for Unicode, is the library that reads that data, and the formatting APIs you already call are front ends over it: JavaScript's Intl, Android's formatters, Apple's Foundation locale APIs, .NET's globalization stack. When you format a date, you are querying CLDR.

Every output in this article was printed on one machine, so the version is part of the evidence. Node reports what it bundles, and these outputs come from Node 24.18.1 with ICU 78.3, CLDR 48.0, Unicode 17.0 and time-zone database 2026b. A different build carries different numbers and can print a different string for the same call. Read the outputs as samples of what CLDR decides, not as constants to assert in a test.

node -e 'console.log(process.versions.icu, process.versions.cldr, process.versions.unicode)'
// 78.3 48.0 17.0

Dates and calendars: field order, separators and invisible marks

Rendering a date requires answers that have nothing to do with translation: field order, separators, whether the month is a number or a name, whether the hour runs to 12 or 24. CLDR holds a pattern per locale and per style, so one call is correct for locales nobody wrote code for.

Three details in that output are worth naming. The Arabic short date is eleven UTF-16 code units for nine visible characters, because the pattern inserts U+200F RIGHT-TO-LEFT MARK immediately before each slash so the date reads correctly inside left-to-right text. Code that measures the string, truncates it to fit a label, or compares it against a literal will be wrong by those marks. The numbering system is locale data too: ar-EG resolves to the arab system, so the digits are U+0660 to U+0669 rather than ASCII. And the calendar is an axis separate from the language. Ask for ja-JP-u-ca-japanese and the same instant renders in the Japanese imperial calendar.

Regional variants are not cosmetic. de-AT prints Jänner where de-DE prints Januar, a spelling a hand-written German month list will not contain. Clock and week conventions travel with the locale as well: en-US resolves to the h12 hour cycle and reports Sunday as the first day of the week, while en-GB resolves to h23 and reports Monday. The accessor differs too: Node 22.13.0 exposes only a weekInfo getter, and 24.18.1 adds a getWeekInfo method alongside it. A calendar widget with a hardcoded Monday-first grid is wrong in the United States, and a Sunday-first one is wrong in Britain.

const d = new Date(Date.UTC(2026, 7, 15, 13, 5));
const opt = { timeZone: "UTC" };

new Intl.DateTimeFormat("ja-JP", opt).format(d); // 2026/8/15
new Intl.DateTimeFormat("en-US", opt).format(d); // 8/15/2026
new Intl.DateTimeFormat("en-GB", opt).format(d); // 15/08/2026
new Intl.DateTimeFormat("de-DE", opt).format(d); // 15.8.2026

// dateStyle: "long"
// ja-JP  2026年8月15日      en-US  August 15, 2026
// en-GB  15 August 2026     de-DE  15. August 2026
// ar-EG  ١٥ أغسطس ٢٠٢٦

// timeStyle: "short"
// ja-JP  13:05   en-GB  13:05   en-US  1:05 PM

// the Arabic short date carries two invisible bidi marks
const ar = new Intl.DateTimeFormat("ar-EG", opt).format(d);
ar.length; // 11 code units for nine visible characters
[...ar].map((c) => c.codePointAt(0).toString(16));
// 661 665 200f 2f 668 200f 2f 662 660 662 666
// one U+200F sits immediately before each slash, at index 2 and 5

new Intl.DateTimeFormat("ar-EG").resolvedOptions().numberingSystem; // arab
new Intl.DateTimeFormat("ja-JP-u-ca-japanese", { dateStyle: "long", ...opt }).format(d);
// 令和8年8月15日

new Intl.Locale("en-US").getWeekInfo(); // { firstDay: 7, weekend: [6, 7] }
new Intl.Locale("en-GB").getWeekInfo(); // { firstDay: 1, weekend: [6, 7] }

Numbers and currency: three kinds of space and a rounding rule

Grouping and decimal separators are the most familiar piece of CLDR, and they still hold more variation than most codebases assume. de-CH groups with an apostrophe and uses a period for the decimal, so it disagrees with de-DE on both. It is also the sharpest proof that the version is part of the answer: on Node 22.13.0, with ICU 76.1 and CLDR 46.0, that separator is U+2019 RIGHT SINGLE QUOTATION MARK, and on ICU 78.3 with CLDR 48.0 it is U+0027 APOSTROPHE. hi-IN groups the top of the number in twos rather than threes, matching the lakh and crore system. ar-EG uses neither comma nor period: it groups with U+066C ARABIC THOUSANDS SEPARATOR and separates the decimal with U+066B.

Then there are the spaces you cannot see. In this single build, fr-FR groups with U+202F NARROW NO-BREAK SPACE and de-AT groups with U+00A0 NO-BREAK SPACE. Both render as a gap; neither is U+0020. The euro symbol in de-DE is preceded by U+00A0 as well. Any code that compares a formatted number against a string it built with an ASCII space, or splits on whitespace to parse one back, fails on locales it never tested and looks correct in a diff.

Currency adds its own data on top. JPY has zero fraction digits, so 1234.5 formats as a rounded 1,235, and the symbol CLDR gives for ja-JP is U+FFE5 FULLWIDTH YEN SIGN rather than U+00A5. KWD has three fraction digits. Symbol placement is data, not a rule you can generalise: the euro follows the number in de-DE and fr-FR, the dollar precedes it in en-US. Percent behaves the same way, prefixing in tr-TR and suffixing in en-US. Ask formatToParts when you need to style the pieces yourself.

const n = 1234567.89;
new Intl.NumberFormat("en-US").format(n); // 1,234,567.89
new Intl.NumberFormat("de-DE").format(n); // 1.234.567,89
new Intl.NumberFormat("de-CH").format(n); // 1'234'567.89
new Intl.NumberFormat("hi-IN").format(n); // 12,34,567.89
new Intl.NumberFormat("fr-FR").format(n); // looks like 1 234 567,89
new Intl.NumberFormat("de-AT").format(n); // looks like 1 234 567,89

// same build, two different invisible spaces
[...new Intl.NumberFormat("fr-FR").format(n)].map((c) => c.codePointAt(0).toString(16));
// 31 202f 32 33 34 202f 35 36 37 2c 38 39
[...new Intl.NumberFormat("de-AT").format(n)].map((c) => c.codePointAt(0).toString(16));
// 31 a0 32 33 34 a0 35 36 37 2c 38 39

const cur = (loc, currency) =>
  new Intl.NumberFormat(loc, { style: "currency", currency }).format(1234.5);
cur("en-US", "USD"); // $1,234.50
cur("ja-JP", "JPY"); // ¥1,235      U+FFE5, and JPY rounds to 0 digits
cur("de-DE", "EUR"); // 1.234,50 €  with U+00A0 before the symbol
cur("en-US", "KWD"); // KWD 1,234.500

new Intl.NumberFormat("tr-TR", { style: "percent" }).format(0.42); // %42
new Intl.NumberFormat("en-US", { style: "percent" }).format(0.42); // 42%

Plural categories: a singular and a plural string is not enough

CLDR assigns every language a set of plural categories, and the set is not two. Japanese has exactly one, other, which is why a Japanese string never needs a count-dependent variant. English has one and other. Russian has one, few, many and other. Arabic has six: zero, one, two, few, many and other. Treat that as a set, not a sequence, because the array order also changes between versions.

The selection is a rule over the number, not a count of it. In Russian, 21 and 101 select one, while 0, 5, 11 and 100 select many. Polish has the same four categories and selects many for 21. So two languages with identical category lists still need different strings for the same number, and a UI that stores one singular and one plural string cannot be filled in for either, no matter how the translator tries.

French is the case that catches English-speaking teams, because select(0) returns one. Zero is singular in French. So is 1.5. French also has a many category that fires only for round large numbers: 1000000 selects many and 1100000 selects other. Ordinals are a separate rule set again, with their own categories: English ordinal selects one for 21 and other for 11, which is exactly the 21st and 11th distinction.

The practical consequence is that the count-dependent choice has to live inside the translated string, where the translator can add or remove branches per language, rather than in code that picks between two strings. That is what the plural argument in ICU MessageFormat exists for, and CLDR's plural rules are the data it selects against.

const cats = (loc) => new Intl.PluralRules(loc).resolvedOptions().pluralCategories;
cats("ja"); // [ "other" ]
cats("en"); // [ "one", "other" ]
cats("ru"); // [ "one", "few", "many", "other" ]
cats("ar"); // [ "zero", "one", "two", "few", "many", "other" ]

// Intl.PluralRules(loc).select(n)
// n:      0      1     2      3      11     21     101    1.5
// ja      other  other other  other  other  other  other  other
// en      other  one   other  other  other  other  other  other
// ru      many   one   few    few    many   one    one    other
// pl      many   one   few    few    many   many   many   other
// ar      zero   one   two    few    many   many   other  other
// fr      one    one   other  other  other  other  other  one

new Intl.PluralRules("fr").select(1000000); // many
new Intl.PluralRules("fr").select(1100000); // other

const ord = new Intl.PluralRules("en", { type: "ordinal" });
[1, 2, 3, 4, 11, 21].map((n) => ord.select(n));
// [ "one", "two", "few", "other", "other", "one" ]

Names, list joining, relative time and sort order

CLDR also holds the vocabulary you would otherwise type into a constants file. The English name of region TR in this build is Türkiye; CLDR's release notes record that rename in CLDR 42. The same data answers CZ with Czechia and SZ with Eswatini, both names that replaced older ones. A hand-maintained country list goes stale silently, and it is only correct in one display language: the same region asked for in Japanese is トルコ. Language names carry region qualifiers too, which is where the distinctions in Brazilian Portuguese and Latin American versus European Spanish show up as data rather than opinion.

List joining is a pattern, not a comma. en-US inserts the serial comma before the conjunction and en-GB does not. German joins with und, Spanish with y, French with et. Japanese uses the ideographic comma and no conjunction word at all. So a loot message assembled as items.join(", ") plus " and " plus the last item is wrong in most locales, and subtly wrong in British English. Relative time is patterned the same way, irregular words included.

Sort order is the one people discover last. A default code-unit sort puts accented words after z, because it compares code points. German collation sorts ä with a; Swedish and Finnish sort ä and ö after z as separate letters. Spanish sorts ñ between n and o. Passing a collator's compare to sort fixes all three, and the numeric option fixes the Item1, Item10, Item2 ordering that shows up in every asset browser.

const reg = (loc) => new Intl.DisplayNames([loc], { type: "region" }).of("TR");
reg("en"); // Türkiye
reg("ja"); // トルコ
new Intl.DisplayNames(["en"], { type: "language" }).of("pt-BR"); // Brazilian Portuguese
new Intl.DisplayNames(["ja"], { type: "language" }).of("pt-BR"); // ポルトガル語 (ブラジル)
new Intl.DisplayNames(["ja"], { type: "currency" }).of("JPY");   // 日本円

const items = ["Sword", "Shield", "Potion"];
const and = (loc, list) => new Intl.ListFormat(loc, { type: "conjunction" }).format(list);
and("en-US", items); // Sword, Shield, and Potion
and("en-GB", items); // Sword, Shield and Potion
and("de", items);    // Sword, Shield und Potion
and("ja", ["剣", "盾", "薬草"]); // 剣、盾、薬草
new Intl.ListFormat("ja", { type: "disjunction" }).format(["剣", "盾", "薬草"]);
// 剣、盾、または薬草

const rtf = (loc) => new Intl.RelativeTimeFormat(loc, { numeric: "auto" });
rtf("en").format(-1, "day"); // yesterday
rtf("ja").format(-1, "day"); // 昨日
rtf("de").format(-3, "day"); // vor 3 Tagen
rtf("ja").format(2, "week"); // 2 週間後

const words = ["zebra", "ärger", "apfel", "Öl", "osten"];
[...words].sort();                                     // apfel osten zebra Öl ärger
[...words].sort(new Intl.Collator("de").compare);      // apfel ärger Öl osten zebra
[...words].sort(new Intl.Collator("sv").compare);      // apfel osten zebra ärger Öl
["nube", "ñu", "nuez"].sort(new Intl.Collator("es").compare); // nube nuez ñu
["Item10", "Item2", "Item1"].sort(new Intl.Collator("en", { numeric: true }).compare);
// Item1 Item2 Item10

CLDR is the data, ICU is the library, every platform pins a version

CLDR is published as XML in the Locale Data Markup Language, with a JSON distribution alongside it, on a release train of roughly two versions a year. ICU compiles that into a binary data bundle and exposes the APIs. Everything downstream picks an ICU build and freezes it until it upgrades.

Browsers each bundle their own ICU, so the same JavaScript can format differently in two browsers and in two versions of one browser. Node exposes what it carries through process.versions, and a small-icu build has English data only. Android has exposed ICU4J directly as android.icu since API level 24, and the platform formatters behind a resource file such as the one described in Android strings.xml sit on the same data. Apple platforms format through Foundation's locale APIs over ICU. .NET has used ICU by default since .NET 5 on every platform, and a Windows application can be switched back to the older NLS APIs through a runtime configuration switch, as the .NET globalization and ICU documentation describes.

Game projects inherit this twice. The C# side of a Unity or Godot project gets whatever ICU the .NET runtime on that platform carries, while Unreal ships an ICU build with the engine, documented in its internationalization pages. So one game can format the same number two ways on two platforms, not because of a bug, but because the CLDR versions differ. If two clients or a client and a server must agree byte for byte, format once and send the resulting string.

Two habits follow. Never assert an exact formatted string in a test: assert the parts from formatToParts, or assert the behaviour you care about, because the next ICU bump is allowed to change the string. And canonicalise locale tags instead of comparing them as text, since Intl.getCanonicalLocales turns JA-jp into ja-JP and zh-hans-cn into zh-Hans-CN. Unknown tags fall back silently to the runtime's default locale, not to English: zz-ZZ resolves to en-US on this machine and to ja-JP when LANG says so, which makes a typo in your locale list look like a working locale rather than an error. Checking resolvedOptions().locale is how you find out what you actually got, and the shape of the tags themselves is covered in BCP 47 language tags.

new Intl.DateTimeFormat("de-AT").resolvedOptions().locale; // de-AT, real data
new Intl.DateTimeFormat("zz-ZZ").resolvedOptions().locale;
// en-US here, but it follows the runtime default: LANG=ja-JP gives ja-JP

Intl.getCanonicalLocales(["JA-jp", "zh-hans-cn", "PT-br"]);
// [ "ja-JP", "zh-Hans-CN", "pt-BR" ]

// how much data one build carries
Intl.supportedValuesOf("timeZone").length;        // 418
Intl.supportedValuesOf("currency").length;        // 162
Intl.supportedValuesOf("numberingSystem").length; // 78
Intl.supportedValuesOf("calendar").length;        // 18

Where to stop delegating, and the failures that reach players

Delegate everything that is a convention of the reader's locale, which is more than most projects hand over: the clock, the first day of the week, numbers, currency, percent, units, plural and ordinal selection, list joining, relative time, the sort order of any list a player reads, and the names of languages, regions and currencies. There is no version of your own judgement that beats the data here.

Write your own formatting when the string belongs to the game's voice rather than the reader's locale. Stat numbers in a fantasy interface where you deliberately want no grouping separator, an in-world calendar with invented month names, damage numbers you want monospaced, a currency that does not exist. Even then, a NumberFormat with useGrouping set to false beats concatenation, because it keeps one code path instead of two.

The dividing line is a question you can actually answer: would a native reader of that locale call the format wrong? If yes, it is the locale's decision and CLDR has the answer. If the only authority is your art direction, it is yours. In both cases keep the format out of the translated string, so the translator receives a pattern and the runtime fills in the number.

When the line is drawn in the wrong place, these are the failures that reach players, in the order they are usually found:

  • A number or date assembled by concatenation. It looks right in the developer's locale and stays wrong everywhere else, because every separator in it was a guess.
  • A plural system designed around a two-category language. The code picks between a singular and a plural string, and Russian, Polish and Arabic have no way to be correct inside that shape.
  • A comparison or parse against a formatted string. The no-break space, the narrow no-break space, the bidi marks and the non-ASCII digits all defeat an equality check that passed in testing.
  • A hand-maintained list of language, region or currency names. It is stale the moment a rename lands, and it exists in only one display language.
  • An assumption that two platforms agree. The same locale on two ICU versions can print two strings, so a cross-platform screenshot comparison fails on something that is not a bug.

Related articles