File formats & standardsこの記事を日本語で読む

Android strings.xml localization: escaping, format args, plurals

A translated strings.xml can be perfectly valid XML and still fail the Android build. It can also compile cleanly and then crash the app the first time the string is drawn. That gap is where almost every Android localization bug lives, and you cannot see it by looking at the file.

Android keeps interface text out of code: each user-visible string is declared once in res/values/strings.xml and referenced by name from layouts and Kotlin or Java. Send that file for translation, get it back, and the damage falls into two groups with very different consequences. The lists and tables below come from running the Android resource compiler (aapt2 2.19, build-tools 34.0.0) over deliberately broken files, and from running the same format strings through Java's String.format.

What the Android build catches, and what it lets through

aapt2 is the tool Gradle runs over your resources, so its verdict is the build's verdict. Each damaged string below was compiled on its own. The split is sharp, and it is the single most useful thing to know before you brief a translator.

  • Build fails: an unescaped apostrophe, reported as unescaped apostrophe in string, followed by not a valid string
  • Build fails: a mangled tag such as a space inside an opening bracket, reported as xml parser error: not well-formed
  • Build fails: an xliff:g element whose namespace declaration was stripped, reported as unbound prefix
  • Build fails: two substitutions with no positional index, reported as multiple substitutions specified in non-positional format
  • Build fails: a quantity value that is not a recognised keyword, reported as item in plural has invalid value 'singular' for attribute 'quantity'
  • Ships: a format argument dropped from the translation, so the number silently vanishes from the sentence
  • Ships: two arguments whose types no longer match the code, which throws when the string is formatted
  • Ships: a space inserted inside a string specifier, which throws at format time, while the same space on a number is a legal flag that silently pads the value
  • Ships: a lost backslash in a line break, so two lines render as one run of text
  • Ships: a full-width percent sign, which prints as itself instead of introducing an argument

The pattern is that structural mistakes stop the build and semantic ones do not. Everything in the second group reaches players. So the review that matters is not whether the file parses, but whether every placeholder in every language file matches the default file one for one, in count and in type.

The failures that ship are ordinary Java exceptions, which is useful because they are easy to recognise in a crash report. Swapping an integer and a string specifier raises IllegalFormatConversionException. Adding a third argument to a string the code only supplies two for raises MissingFormatArgumentException. Breaking the syntax of a specifier raises UnknownFormatConversionException, and the message names the stray character it tried to read as a conversion. A space directly before a string conversion raises FormatFlagsConversionMismatchException instead, because a space is a flag that numbers accept and strings do not.

Escaping rules the XML validator does not enforce

Ampersand and angle brackets need XML entities, and any XML validator will tell you so. Android then adds a second layer on top of XML, and that layer is what translators break, because a generic validator says the file is fine. Running xmllint over a file whose apostrophes are all bare reports no error at all, while aapt2 refuses to compile it.

Here is what the compiler actually stores, dumped from a linked package. The left column is the source text inside the element, the right column is the string the app receives.

source in strings.xml          compiled value
  Loading...                   "Loading..."
"  Loading...  "               "  Loading...  "
Now     loading                "Now loading"
Line one\nLine two             two real lines
Press <b>Start</b> to continue "Press Start to continue" + bold over chars 6-10
You\'ll lose progress          "You'll lose progress"
"You'll lose progress"         "You'll lose progress"

The whitespace rows surprise people. Leading and trailing whitespace is trimmed, and a run of internal spaces collapses to one. A translator who pads a label to line it up with the label above it loses the padding, and nothing warns anyone. Wrapping the whole value in straight double quotes preserves the spacing exactly; those quotes are markup, not text, and never appear on screen.

The apostrophe has two fixes and they are equivalent: a backslash before it, or double quotes around the whole value. Compiled, both produce byte-identical text. For French, Italian, Portuguese and Catalan this is not an edge case but a property of the language, so pick one convention, write it in the translator brief, and expect to fix a first delivery that ignores it.

The last row matters for a different reason. Tags like b, i and u are not stripped and are not literal text either: the compiler records them as style spans over character offsets in the plain string. Tell translators the tags must survive, must stay balanced, and may move to wrap the words that carry the emphasis in their language, which is often not the same position as in English.

Format arguments: why a positional index is not optional

Android formats these strings through Java's String.format, so the specification to read is String.format's, not an Android-specific one. Two consequences follow.

First, a literal percent sign must be doubled. aapt2 catches most bare ones, because the percent and the character after it count as a second, non-positional substitution and the file is refused. The case it misses is a percent at the very end of the string: HP %1$d% compiles, and then throws UnknownFormatConversionException at format time, naming the percent itself as the conversion it could not read.

Second, positional indices. aapt2 refuses a string that holds two substitutions with no index, which means indices are mandatory rather than a matter of taste as soon as a sentence has two arguments.

<string name="pickup">You picked up %1$s (x%2$d).</string>
<string name="pickup_ja">%1$sを%2$d個 手に入れた。</string>
<string name="hp">HP %1$d%%</string>

String.format("%1$s picked up %2$d", "Ether", 3)  ->  Ether picked up 3
String.format("%2$d x %1$s", "Ether", 3)          ->  3 x Ether
String.format("%1$s %s %s", "a", "b")             ->  a a b

That third line is the trap. A specifier without an index keeps its own counter, and positional specifiers do not advance it, so the first bare one goes back to argument one. Mixing the two forms in a single string produces a duplicated argument if the types happen to agree and an exception if they do not. Never leave one bare specifier in a string that otherwise uses indices.

If a string genuinely is not a format string and needs a literal percent pair, the formatted attribute set to false on the element turns the check off and the value stays literal. Use it for text you pass to something other than String.format, not to silence a warning on a real format string.

Plurals: which quantity keywords each language actually uses

A plurals element supplies one string per grammatical category instead of an if-else around a count. The quantity attribute is not a free label: it must be one of zero, one, two, few, many or other, and which of those a given language uses is decided by the CLDR plural rules. aapt2 rejects an invented keyword, but it does not check that the keywords present are the ones the language needs, and it accepted a plurals block with no other item at all. The compiler is not the safety net here.

The categories below were read out of Node's Intl.PluralRules, which resolves against the same CLDR data.

categories in use          languages
other                      Japanese, Korean, Chinese, Thai, Vietnamese, Indonesian
one other                  English, German, Dutch, Turkish
one many other             French, Spanish, Portuguese, Italian
one few many other         Russian, Ukrainian, Polish, Czech
zero one two few many other Arabic

Russian:  1=one  2=few  5=many  21=one  100=many  101=one  102=few
French:   0=one  1=one  2=other  1000000=many

So a Japanese translator who deletes the item marked one is correct to do so. Japanese has no separate singular form and every count selects other, which is also why a count in Japanese usually needs a counter word rather than a plural ending. Leaving an English-shaped pair in the Japanese file is harmless dead weight. The reverse is the real bug: shipping only one and other to Russian leaves counts of 2, 5 and 102 selecting categories nobody wrote, and French treats zero as one, so a French string for an empty list should read in the singular.

Two limits are worth stating. Plurals handle number agreement only, not gender or grammatical case, and a sentence that needs those wants a richer message syntax such as ICU MessageFormat and its plural and select arguments. And the category a number falls into is data, not intuition, so when you need to know which numbers map where, read it from the rules rather than guessing: what CLDR decides for each locale covers where that data comes from.

Qualifier directories and the default values you cannot skip

Translations are parallel files at the same relative path, each in a res directory whose name carries a qualifier. The region part takes an r prefix, and the spelling is not negotiable. Compiled side by side, the plain tag form fails with invalid configuration.

res/values/strings.xml             default - required
res/values-ja/strings.xml          Japanese
res/values-pt-rBR/strings.xml      Portuguese (Brazil) - compiles
res/values-pt-BR/strings.xml       error: invalid configuration 'pt-BR'
res/values-b+sr+Latn/strings.xml   Serbian in Latin script - compiles

For anything beyond language plus region, a script subtag for instance, the qualifier switches to the b plus form with subtags joined by plus signs. Confirm the accepted shapes in the official App resources documentation, in the chapter on supported configuration qualifiers, before you invent a folder name; the tag itself follows the standard described in BCP 47 language tags and how they are built.

The default directory with no qualifier is mandatory, and its failure mode deserves a look. Link a package that has a Japanese file but no default and aapt2 does not error. It warns that it is removing the resource without a required default value, and the string is then gone from the built package entirely, so the name no longer exists to be referenced. A warning scrolls past in build output; a missing resource is a crash on every device in every language.

The opposite gap is quiet by design. A key present in the default file and missing from a language file falls back to the default text, so an untranslated string shows English instead of failing. The lint check named MissingTranslation is what reports that, and it is the only thing that will, so keep it enabled and read it. Suppress it per string when the omission is deliberate, and mark strings that must never be translated as not translatable.

<resources xmlns:tools="http://schemas.android.com/tools"
           xmlns:xliff="urn:oasis:names:tc:xliff:document:1.2">
    <string name="event_id" translatable="false">boss_defeated_01</string>
    <string name="debug_note" tools:ignore="MissingTranslation">Debug build</string>
    <string name="save_slot">Save slot <xliff:g id="slot" example="3">%1$d</xliff:g></string>
</resources>

That last line is the one piece of protective markup worth adopting as a habit. An xliff:g element wraps a placeholder and tells a translation tool to treat its contents as untouchable, and the example attribute shows the translator what a real value looks like so they can judge word order. It needs the namespace declaration on the resources element; without it the build fails with an unbound prefix error, which is how a stripped declaration announces itself.

One more shape to watch. Items inside a string-array have no name attribute, so order is their only identity. A translator who reorders or drops an item shifts every index after it, and if that array feeds a picker whose selection you store by position, the stored value now means something else. Either state the order constraint in the brief or build the array from individually named strings.

How much of this matters if you ship a Unity, Unreal or Godot game

For a game built in an engine, almost none of your in-game text goes through strings.xml. Dialogue, item names, skill descriptions and menu labels live in the engine's own localization tables and are read by the engine at runtime, not by the Android resource framework. Nothing in this article applies to them.

What does live in strings.xml on an Android build is the OS-facing set: the app label under the icon, permission rationale text, notification channel names, shortcut labels, and any string in an Android plugin or activity you wrote yourself. It is usually a short list, often under twenty entries, and it is disproportionately visible, because it appears on the home screen and inside system dialogs rather than three menus deep.

The practical split is three buckets with three owners. The engine's tables are the bulk of the work, and how Unity organizes locales, tables and string references describes that side. strings.xml is a small separate deliverable with the checklist above. Store listing text is a third place again, entered in the console and never in either file. If your text is still sitting in scenes and scripts, none of the three is your first job: getting hardcoded text out of code is.

Related articles