File formats & standardsこの記事を日本語で読む

ICU MessageFormat syntax: plurals, select, and CLDR categories

ICU MessageFormat is a template syntax for translatable messages, defined by the ICU project (International Components for Unicode). It puts the branching a sentence needs — plural forms, grammatical gender, ordinals, number and date formatting — inside the translatable string itself, so the code around it never has to know any language's grammar.

One fact explains most of the syntax: a plural branch is not selected by the number. It is selected by the CLDR plural category that the number maps to in the target language. English uses two categories, Japanese one, Russian four, Arabic six. That is why "You have 1 items" is a grammar bug rather than a typo, and why the number of branches in the Russian version of a message is decided by Russian, not by whoever wrote the English.

What follows is the whole syntax with examples, a table of the categories each language actually uses that was generated from the CLDR data shipped with a JavaScript runtime, the quoting rule that trips people up, what a translator must be told so they do not break the braces, and the cases where a plain placeholder is the better answer.

Arguments, types and styles: the syntax underneath everything

Every argument is a name in braces. Add a second field and it gains a type; add a third and the type gains a style.

The number, date and time types carry no formatting rules of their own. They hand the value to the target locale's CLDR formatting data, which is the entire point: one template, and the separators, digit shapes and field order all change per locale without the template changing. The examples below were produced by asking a JavaScript runtime for the same values in several locales. One detail defeats naive comparisons: the French group separator there is a narrow no-break space, U+202F, not an ordinary space, so a check that expects a plain space fails against correct output.

Newer ICU versions also accept a skeleton after a double colon, which describes the output you want instead of a fixed pattern. Whether you can use one depends on the ICU version your library links against, so confirm it in that library's documentation first.

Because braces are the syntax, a message that has to display a literal brace needs quoting. In ICU's default mode an ASCII apostrophe starts quoted text only when the character after it is one that needs quoting, a doubled apostrophe is always one literal apostrophe, and an apostrophe in ordinary prose is left alone. Some implementations offer a stricter mode in which every apostrophe quotes. The Quoting and Escaping section of the ICU documentation is the reference for the rule, and your library's own documentation says which mode it runs.

{playerName}                      argument, substituted as-is
{score, number}                   type
{ratio, number, percent}          type plus style
{releasedOn, date, long}          date/time styles: short medium long full
{startsAt, time, short}
{price, number, ::currency/JPY}   number skeleton (newer ICU only)
{count, plural, ...}              the branching types
{gender, select, ...}
{rank, selectordinal, ...}

One template, different locales (CLDR output via Intl):
  {n, number}       1234567.5   en-US  1,234,567.5
                                de-DE  1.234.567,5
                                fr-FR  1 234 567,5  (separator: U+202F)
                                ar-EG  ١٬٢٣٤٬٥٦٧٫٥
  {d, date, short}              en-US  9/12/26
                                ja-JP  2026/09/12
                                de-DE  12.09.26
  {d, date, long}               en-US  September 12, 2026
                                ja-JP  2026年9月12日
  {t, time, short}              en-US  3:04 PM
                                ja-JP  15:04

Literal braces and apostrophes:
  '{'          renders as   {
  '{count}'                 {count}   (displayed, not substituted)
  ''                        '
  It's fine                 It's fine  (apostrophe left alone)

plural: exact matches, categories, the # token and offset

A plural argument lists branches, and two kinds of selector can appear in it. An exact selector, an equals sign followed by a number, matches that value literally. A keyword selector matches a CLDR category. Exact selectors are tried first, which is how you get a special empty-state sentence without pretending that zero is a grammatical category in English. Every plural argument needs an other branch: it is the fallback and ICU treats it as required.

Inside a branch, the # token is replaced by the number, formatted for the locale exactly as the number type would format it. Typing the digits by hand instead of using # is a frequent mistake. It produces an unseparated number in every locale, and it removes the one token a translator is allowed to move.

The offset feature subtracts a constant before the category is chosen. It exists for phrasings where the displayed count and the count that drives the grammar differ, such as a sentence that names one person and counts the rest. Two details are easy to get backwards. Keyword selection and the # token both use the value after the offset is subtracted. Exact selectors are matched against the input value before subtraction, which is what makes the equals-one and equals-two branches below readable. The Plural Formatting section of the ICU documentation carries the canonical example of this behaviour.

{count, plural,
  =0    {Your inventory is empty}
  one   {# item}
  other {# items}
}

offset: the sentence names one guest, the grammar counts the rest
{guests, plural, offset:1
  =0    {Nobody is coming}
  =1    {{host} is coming}
  =2    {{host} and one other person are coming}
  other {{host} and # other people are coming}
}
  guests = 1  ->  =1 matches the input value; the offset is not applied to it
  guests = 4  ->  other branch, and # renders 3 (4 minus the offset)
  no one branch is needed here: in English the only input whose
  offset-adjusted value lands in the one category is 2, and =2
  already matches it exactly. Russian would need more branches.

The plural categories each language actually uses

Which categories a language uses, and which numbers land in them, is CLDR data. You can read it yourself instead of trusting an article, including this one: JavaScript runtimes implement Intl.PluralRules on top of ICU, so it answers with the same rules. The table below came from asking it for each locale's category list and then for the category of specific numbers. The details of that data set are covered in Unicode CLDR explained.

Five conclusions from it change how you write messages.

Japanese has a single category, other. A Japanese translation with a one branch is not wrong so much as dead code: nothing ever selects it. Chinese and Korean behave the same way. If your source language is Japanese, the plural problem is invisible until the English build.

Russian needs four branches, and one does not mean one. Twenty-one is in the one category, twenty-two is in few, and five, eleven and one hundred are in many. A Russian string that only has a one branch and an other branch reads as broken for most numbers.

French puts zero in the one category, so the French singular covers both zero and one. Hardcoding the English two-form rule gives French the wrong form for zero. Polish splits similar-looking numbers differently again: twenty-two is few while twenty-five is many.

Arabic uses all six categories, including zero and two. It is the language that makes the case for the format on its own, because no arrangement of conditionals over an English source produces it.

Ordinals are a separate rule set from cardinals. English needs two cardinal categories but four ordinal ones, and the exceptions are not where a naive rule would put them: eleventh, twelfth and thirteenth fall in other, while twenty-first is back in one and one hundred and eleventh is in other.

Cardinal category of n, by locale (Intl.PluralRules, ICU/CLDR data):

  n          en      ja      ru      ar      pl      fr
  0          other   other   many    zero    many    one
  1          one     other   one     one     one     one
  2          other   other   few     two     few     other
  5          other   other   many    few     many    other
  11         other   other   many    many    many    other
  21         other   other   one     many    many    other
  22         other   other   few     many    few     other
  25         other   other   many    many    many    other
  100        other   other   many    other   many    other
  1000000    other   other   many    other   many    many

Categories a locale actually uses (cardinal):

  en   one other
  ja   other
  ru   one few many other
  ar   zero one two few many other
  pl   one few many other
  fr   one many other

Ordinal categories:

  en   one two few other     1st 2nd 3rd 4th ... 11th 12th 13th ... 21st
  ja   other
  ru   other

select and selectordinal, and the cost of nesting them

select is the general branching form. It matches an argument's value against names you choose, with other required as the fallback, and those names come from your code rather than from CLDR. Grammatical gender is the usual reason to reach for it: in many languages a verb form, article or adjective agrees with the subject, and the English source gives the translator nothing to agree with.

selectordinal looks like plural and behaves like it, except that it selects from the ordinal rule set rather than the cardinal one. Placements, rankings and day-of-month strings need it.

Arguments nest. A select branch can contain a plural, and that is how a single message handles gender and count at once. It is also where MessageFormat begins costing more than it saves. A translator three levels in has to reconstruct which combination of conditions produced the fragment in front of them, a reviewer cannot see every output without exercising every combination, and one missing closing brace breaks the message at runtime rather than at build time unless something validates it first.

The cheaper shape is usually to split: one message per gender, chosen by a key in code, each holding a flat plural. It duplicates a little text, and in exchange every branch is independently readable, translatable and testable. Keep nesting for messages whose branches genuinely interact. That decision is much easier to make when keys already carry the distinction, which is one more argument for the kind of scheme described in string key naming design.

{gender, select,
  female {She picked up the sword}
  male   {He picked up the sword}
  other  {They picked up the sword}
}

{rank, selectordinal,
  one   {#st place}
  two   {#nd place}
  few   {#rd place}
  other {#th place}
}

Nested — valid, but hard to translate and to review:
{gender, select,
  female {{count, plural, one {She found # key}  other {She found # keys}}}
  male   {{count, plural, one {He found # key}   other {He found # keys}}}
  other  {{count, plural, one {They found # key} other {They found # keys}}}
}

Split — same output, three flat messages:
found.female = {count, plural, one {She found # key}  other {She found # keys}}
found.male   = {count, plural, one {He found # key}   other {He found # keys}}
found.other  = {count, plural, one {They found # key} other {They found # keys}}

What a translator needs to be told, and what review can script

A MessageFormat string is code and prose in the same field, and a translator who was never told which half is which will reasonably edit both. Three instructions prevent almost all of the damage.

The structure is fixed. Argument names, type keywords, braces and the # token have to survive unchanged. Translating an argument name, or replacing # with a digit, breaks the message.

The set of categories is not fixed, and changing it is the translator's job rather than a defect. Russian needs branches English does not have, Japanese needs fewer, and a translator who can only fill in the branches the English happened to contain cannot produce a correct Russian string. The file format and the workflow both have to allow adding and removing branches.

Word order moves the # token. The number does not have to stay where English put it, and in many languages it cannot. The same freedom applies to the position of every other argument inside a branch.

Most mechanical failures are detectable by script before a human opens the file, which is the cheapest review you will ever run:

  • Braces balance, and every plural and select argument has an other branch.
  • Every argument name in the source appears in the translation, spelled the same way.
  • No branch has a literal digit where the source had a # token.
  • Every keyword selector is a category the target language actually uses, so a stray one branch in Japanese or a missing few branch in Russian is caught.
  • Nesting depth has not increased relative to the source.
Instruction sheet to ship with the strings

  Do not change    {count, plural, ...}   {playerName}   #   type keywords
  Do change        the text inside each pair of braces
  Must change      the set of selectors, to match your language
                     ja  other
                     ru  one few many other
                     ar  zero one two few many other
  May move         # and {arguments} to wherever your word order needs them

When to reach for MessageFormat, and what supports it

MessageFormat earns its complexity when the grammar of the target language depends on a runtime value: counts inside a sentence, gender agreement, ordinals. It earns nothing when the value is merely substituted. A message that is only a player name followed by a fixed verb gains a parser and no correctness.

Two alternatives often beat a plural argument outright. Avoiding the sentence is the first: a label and a number in separate cells, or a count in parentheses after a noun, has no grammar to get wrong in any language, and cramped UI frequently prefers it anyway. Separate keys per case is the second, as above. Both are worth considering before the strings even leave the code, which is the stage covered in externalizing hardcoded strings.

Support is uneven enough to check before you commit a project's strings to the syntax. ICU itself provides MessageFormat in its Java and C libraries, which is the reference implementation. Several JavaScript internationalization libraries implement the same syntax, sometimes a subset of it, and the honest answer for any particular one is the version and feature table in its own documentation rather than a general claim here.

Neighbouring formats solve the same problem differently. Android string resources use a plurals element whose quantity keywords come from the same CLDR set, but the file syntax is not MessageFormat, as the guide to Android strings.xml sets out. The gettext PO format selects a numbered form through an expression in the file header instead of a keyword in the string, described in the gettext PO file format. Game engine localization systems generally do not accept MessageFormat strings as they stand: some ship their own plural syntax over the same CLDR categories, some ship none and leave the branching to your code or a library. Check the engine's localization documentation rather than assuming either way.

Related articles