File formats & standardsこの記事を日本語で読む

XLIFF files explained: 1.2 vs 2.0, and what breaks on the round trip

XLIFF is an XML format for moving translatable text between systems. Your pipeline extracts strings into it, a translator or a translation tool fills in the target text, and your pipeline merges the result back into your own files. It is an OASIS standard, and the first practical fact about it is that the two versions you are likely to be handed are 1.2, standardized in 2008, and 2.0, standardized in 2014. They use different element names, different workflow states and different inline markup, and a tool that reads one will not necessarily read the other. Later revisions in the 2.x line, 2.1 and 2.2, have been published as OASIS standards as well, but tool support in practice still concentrates on 1.2 and 2.0.

If a translation vendor asked you for XLIFF, or a tool returned a file your importer rejected, the parts you need are below: one complete minimal document in each version holding the same two strings, the version differences that cause real failures, the specific ways a round trip loses data, and what to put in writing when you hand the file over. Every XML example here was checked for well-formedness with a command-line XML parser, and the damaged examples were checked to confirm they fail.

A complete minimal XLIFF, in 1.2 and in 2.0

Most introductions show a fragment. A fragment will not tell you whether your importer works, because the parts that break first are the root element, the namespace and the language attributes, and none of those appear in a fragment. Below are two complete documents carrying the same two strings: a tutorial line with a placeholder, and a combat line with a bold span.

Side by side, the mapping is short.

  • The root element is named xliff in both versions, but the namespace differs — it ends in document:1.2 or document:2.0 — and the version attribute has to agree with the namespace. Importers that trust one and ignore the other are a common source of an unhelpful unsupported-file error.
  • The language pair is declared in different places. 1.2 puts source-language and target-language on the file element. 2.0 puts srcLang and trgLang on the root element, where srcLang is required and trgLang is needed as soon as any target exists.
  • The translatable container is trans-unit in 1.2, wrapped in a mandatory body element inside file. In 2.0 it is unit, sitting directly inside file with no body wrapper, and the text lives one level deeper, inside segment.
  • Both versions pair a source element with a target element, and by convention tools leave source untouched rather than overwriting it. That pair is what a merge step compares, and the reason an XLIFF round trip can be checked at all.
  • The id attribute on trans-unit or unit is the join key between the file and your own data. Make it your string key rather than a row number, so that inserting a line in your source file does not renumber everything after it.
dialogue.1.2.xlf
<?xml version="1.0" encoding="UTF-8"?>
<xliff xmlns="urn:oasis:names:tc:xliff:document:1.2" version="1.2">
  <file original="dialogue.json" source-language="en" target-language="ja" datatype="plaintext">
    <body>
      <trans-unit id="tutorial.open_inventory">
        <source>Press <ph id="1">%s</ph> to open your inventory.</source>
        <target state="translated"><ph id="1">%s</ph> を押してインベントリを開きます。</target>
        <note from="developer">%s is a key name. Do not translate or reorder.</note>
      </trans-unit>
      <trans-unit id="combat.weapon_broke">
        <source>Your <g id="2" ctype="bold">Iron Sword</g> broke.</source>
        <target state="needs-review-translation"><g id="2" ctype="bold">鉄の剣</g>が壊れました。</target>
      </trans-unit>
    </body>
  </file>
</xliff>

dialogue.2.0.xlf
<?xml version="1.0" encoding="UTF-8"?>
<xliff xmlns="urn:oasis:names:tc:xliff:document:2.0" version="2.0" srcLang="en" trgLang="ja">
  <file id="dialogue" original="dialogue.json">
    <unit id="tutorial.open_inventory">
      <notes>
        <note category="developer">The placeholder is a key name. Do not translate or reorder.</note>
      </notes>
      <originalData>
        <data id="d1">%s</data>
      </originalData>
      <segment state="translated">
        <source>Press <ph id="1" dataRef="d1"/> to open your inventory.</source>
        <target><ph id="1" dataRef="d1"/> を押してインベントリを開きます。</target>
      </segment>
    </unit>
    <unit id="combat.weapon_broke">
      <segment state="initial">
        <source>Your <pc id="2">Iron Sword</pc> broke.</source>
        <target><pc id="2">鉄の剣</pc>が壊れました。</target>
      </segment>
    </unit>
  </file>
</xliff>

What changed between 1.2 and 2.0, and why your importer cares

The jump from 1.2 to 2.0 is not a syntax refresh. The two versions model a segment differently, and the differences land exactly where a careless converter drops information.

That fifth difference does more work than it looks like it does. Asking an XML tool for the text content of the 2.0 source element above returns the string with the placeholder absent, because the placeholder characters live in originalData and not in the text. The same query against the 1.2 document returns the placeholder inline, as part of the string. So anything that treats a segment as plain text — a spell checker, a machine translation pass, a find and replace — can damage the 1.2 placeholder and cannot reach the 2.0 one. If you can emit either 1.2 or 2.0, that is the strongest single argument for 2.0.

  • Structure. 1.2 nests file, body, trans-unit, then source and target. 2.0 nests file, unit, segment, then source and target, and one unit may hold several segments plus ignorable elements for the untranslatable text between them.
  • Where state lives. In 1.2 the state attribute sits on the target element. In 2.0 it sits on segment. An importer that looks for state on target reads nothing at all out of a 2.0 file. Confirm the placement in the target element section of the 1.2 specification and the segment section of the 2.0 specification before you write the reader.
  • State vocabularies. 2.0 defines four values — initial, translated, reviewed and final — plus a subState attribute for tool-specific refinement. 1.2 defines a longer list that mixes progress with the reason work is needed, including new, needs-translation, needs-review-translation, translated, signed-off and final, and it carries a separate approved attribute on trans-unit. Read the state attribute definition in the 1.2 specification before mapping values: the conversion is lossy, because several 1.2 values have no distinct 2.0 equivalent and collapse onto initial.
  • Inline markup. 1.2 offers g for a span, x for a standalone placeholder, ph for inline code, and the bpt and ept pair for a code whose two halves are separate. 2.0 replaces that set with pc for a span, the sc and ec pair for a span whose ends are separated or overlap another span, and ph for a standalone placeholder.
  • Where the original code text lives. In 1.2 the original markup can sit as the content of ph, inside the translatable text. In 2.0 it moves out into an originalData element on the unit, and the inline element points at it with dataRef.
  • Notes. 1.2 attaches note elements directly and labels their author with a from attribute. 2.0 wraps them in a notes container and classifies them with category instead.

What breaks on the round trip, and what a parser catches for free

Well-formedness is a floor, not a goal, but it is a cheap gate: running an XML parser over every returned file costs nothing and catches a specific family of damage before it reaches your importer. Three deliberately broken files produce three distinct errors.

Those three are, in order: a game string containing an unescaped ampersand, written straight into XLIFF by an exporter that forgot to escape it; an edit that deleted the closing half of an inline pair, which cascades into mismatch errors all the way up to the root element; and a file saved as Shift_JIS while its declaration still claims UTF-8. Note what the third case is not. Declaring Shift_JIS and saving as Shift_JIS parses cleanly and reads back correctly. The failure is the disagreement between the declaration and the bytes, which is what happens when an editor configured for a local codepage re-saves a file it was never meant to touch.

The failures that matter more are the ones a parser waves through.

  • CDATA around the source text. It is well-formed, and it silently defeats the point of inline markup: wrap a string in CDATA and its markup becomes ordinary characters, so the translator sees raw tags and any of them can be mistyped. Checked against a parser, a CDATA-wrapped source returns its markup as part of the text content, where a real inline code would not.
  • Unknown vendor namespaces. A tool may attach its own namespaced attributes and elements, such as an internal segment key. These are well-formed, and a conforming reader is free to ignore them — which means your own round trip can quietly drop the data the other tool expects to see on the next pass. Do not build anything on top of an extension you did not put there yourself.
  • Unstable ids. Nothing in the format stops an exporter from renumbering units between runs. If ids change, every match against the previous version fails, and a returned file either merges into the wrong rows or is counted as entirely new work. Pin ids to your own string keys, and add a test that fails when an id changes while the source text does not.
  • Whitespace. XML character data is preserved as written, so a leading space or a trailing newline in source is part of the string. Some tools normalize it and some do not. If whitespace is meaningful in your strings, set xml:space to preserve and then verify that the value survived, rather than trusting that it will.
  • Inline order and count. Well-formedness does not care that a translator moved a span across a placeholder or dropped one id of a pair, as long as the tags still nest. Compare the set of inline ids in target against source, per segment, in your own importer — that check is yours to write, not the format's to enforce.
$ xmllint --noout broken-amp.xlf
broken-amp.xlf:6: parser error : xmlParseEntityRef: no name
        <source>Save & Quit</source>
                      ^

$ xmllint --noout broken-halfpair.xlf
broken-halfpair.xlf:7: parser error : Opening and ending tag mismatch: pc line 7 and target
        <target><pc id="2">鉄の剣が壊れました。</target>
  ... four further mismatch errors, ending in "Premature end of data in tag xliff line 2"

$ xmllint --noout broken-cp932.xlf
broken-cp932.xlf:13: parser error : Input is not proper UTF-8, indicate encoding !
Bytes: 0x82 0xF0 0x89 0x9F
  ... followed by the offending line and a caret marker

Where XLIFF sits next to Unity, Unreal and a spreadsheet

Engines do not agree on this. Unity's Localization package can import and export XLIFF for its String Tables; the import and export part of the package manual, under Localization Tables, documents XLIFF alongside CSV, and it states which XLIFF versions a given package release handles — check it for the version you actually have installed rather than assuming. Unreal Engine does not use XLIFF as its interchange format at all: its localization pipeline exports and imports Portable Object files through the Localization Dashboard, so a vendor who only sends XLIFF means a conversion step on your side. The PO file format covers what survives that conversion and what does not.

A spreadsheet is not the wrong answer for every project. It wins when you own both ends, the string count is small, and the people editing are comfortable in it. It loses as soon as inline markup, per-segment workflow state or per-segment context has to travel, because a cell has no defined place for any of that without a convention only your team knows. If a spreadsheet is where you are today, a CSV workflow that survives updates is the cheaper next step; XLIFF starts to earn its plumbing once an outside party is doing the translating.

XLIFF also gets confused with the two formats next to it. XLIFF carries the job: this source, this target, this state, for this release. Translation memory and terminology are separate, reusable across jobs, and have their own formats — TMX for translation memory and TBX for terminology.

What to state when you hand the file over

Most XLIFF problems are agreement problems rather than format problems. Send the answers to these with the file, in writing rather than in a chat message.

The rest of the handoff — builds, screenshots, glossary, the context a line needs to be read correctly — is the same package whatever format you choose, and what to send before you send the text lists it. XLIFF carries the strings. It cannot carry the reason a line reads the way it does.

  • The version and namespace you produce, and the version you will accept back. Say 2.0 or 1.2 explicitly. A vendor's tool may convert between them silently, and you want to discover that before the deadline rather than during the merge.
  • The language tags, in valid form. Both versions expect standard language tags, which use hyphens and not underscores, and case is normalized: ja-JP is a tag and ja_JP is not, pt-br means pt-BR, zh-hans means zh-Hans. What ja-JP, zh-Hans and pt-BR actually mean covers how far to go with region and script subtags.
  • Which placeholders exist and what they expand to. List every placeholder form in the file, say what each one becomes at runtime, and say whether reordering is allowed. Put it in a note on the unit as well, because that is the only place the information travels attached to the string.
  • Length limits per string, on the unit that has them.
  • Where each string appears. A key name is not context. One sentence about the screen, the speaker, and whether the text is a button label or narration prevents most of the ambiguity that otherwise comes back as a question, or worse, as a confident wrong choice.
  • The encoding, and that you will reject a file whose declaration disagrees with its bytes.
  • Your id contract: ids are stable, they are the join key, and a returned file with changed ids gets rejected rather than merged.

Related articles