File formats & standardsこの記事を日本語で読む

TMX vs TBX: translation memory and terminology file formats explained

A TMX file is an export of a translation memory: pairs of a source sentence and its translation, as they were confirmed during real translation work. A TBX file is an export of a termbase: one entry per concept, holding the approved term for that concept in each language. Both are open XML interchange formats. Neither is a string file you ship — you cannot load a TMX into your game, and nothing in it tells your engine what to display. You load it into a translation tool.

That is why a TMX handed over at the end of a project is so easy to file away and forget. It is also the part of the delivery with the longest life, because it is the only part that makes the next project cheaper. This article shows a minimal complete example of each format, explains why the same file produces different match percentages in different tools, says where a memory helps a game project, where it does not, and how these files break.

What a TMX file contains: a minimal complete example

A TMX document has one root element, one header, and one body of translation units. The header is not decoration. The TMX 1.4 specification lists a set of attributes it must carry, and two of them, srclang and segtype, decide whether your matches will land at all.

  • srclang — the language the units are considered to be translated from. The spec also defines a wildcard value for memories with no single direction; its header section gives the exact literal.
  • segtype — the granularity the exporting tool used: block, paragraph, sentence or phrase. This is the attribute everyone skips and the one that quietly sets your match rate.
  • adminlang — the language of the administrative notes in the file, which need not be either translation language.
  • creationtool and creationtoolversion — which tool wrote the file. Worth reading first when an import goes wrong, because it tells you whose dialect you are holding.
  • datatype and o-tmf — what the original content was, and which internal memory format it came from.
<?xml version="1.0" encoding="UTF-8"?>
<tmx version="1.4">
  <header
    creationtool="example-exporter"
    creationtoolversion="1.0"
    segtype="sentence"
    o-tmf="plain"
    adminlang="en"
    srclang="en-US"
    datatype="plaintext"/>
  <body>
    <tu tuid="1">
      <tuv xml:lang="en-US">
        <seg>Are you sure you want to leave the dungeon?</seg>
      </tuv>
      <tuv xml:lang="ja-JP">
        <seg>ダンジョンから出てもよろしいですか?</seg>
      </tuv>
    </tu>
    <tu tuid="2">
      <tuv xml:lang="en-US">
        <seg>You found <ph type="item">{0}</ph>.</seg>
      </tuv>
      <tuv xml:lang="ja-JP">
        <seg><ph type="item">{0}</ph> を見つけた。</seg>
      </tuv>
    </tu>
  </body>
</tmx>

Inside the body, a tu is one translation unit. It holds one tuv per language, and each tuv holds a seg with the actual text. The language sits on the tuv as xml:lang, which XML defines as an IETF language tag — the same tags you use everywhere else, so how BCP 47 language tags are built applies here unchanged. The second unit above shows a placeholder wrapped in an inline element, one of the small family TMX has for formatting and variables inside a segment.

What a TBX file contains, and why a two-column sheet cannot replace it

TBX is organized around concepts, not sentences. One term entry is one thing in the world that needs a name. Inside it there is a language section per language, and inside that, one term section per term. A single language can hold more than one term for the same concept, each with its own status — which is how you record that one rendering is approved and another must not be used any more, something a memory of sentences that happened cannot express at all.

Before writing anything that reads a TBX, check what the file itself declares. TBX exists in more than one version and in several dialects, TBX-Basic and TBX-Core among them, and they differ in element names and in which notes are allowed. The example below is the TBX-Basic-shaped nesting from the 2008-era standard. The 2019 revision keeps the same nesting but renames the root element and the concept, language and term elements. Confirm the version and dialect in the root element, and the permitted values for a status note in that dialect's picklist section of the specification.

<?xml version="1.0" encoding="UTF-8"?>
<martif type="TBX-Basic" xml:lang="en">
  <martifHeader>
    <fileDesc>
      <sourceDesc><p>Game UI and economy glossary</p></sourceDesc>
    </fileDesc>
  </martifHeader>
  <text>
    <body>
      <termEntry id="c001">
        <langSet xml:lang="en">
          <descripGrp>
            <descrip type="definition">The premium currency spent on cosmetic items.</descrip>
          </descripGrp>
          <tig>
            <term>Gem</term>
            <termNote type="partOfSpeech">noun</termNote>
            <termNote type="administrativeStatus">preferredTerm-admn-sts</termNote>
          </tig>
        </langSet>
        <langSet xml:lang="ja">
          <tig>
            <term>ジェム</term>
            <termNote type="administrativeStatus">preferredTerm-admn-sts</termNote>
          </tig>
          <tig>
            <term>宝石</term>
            <termNote type="administrativeStatus">deprecatedTerm-admn-sts</termNote>
          </tig>
        </langSet>
      </termEntry>
    </body>
  </text>
</martif>

A glossary sheet with one row per term is enough for one language pair and forty terms. It stops being enough in four specific ways:

  • There is no standard place for a definition, a usage note or an approval status, so those become free-text comments no tool can act on
  • One column per language stops scaling after a handful of languages, and adding one means restructuring every row
  • Nothing enforces one entry per concept, so an English word used for two different things becomes two rows nobody can tell apart
  • No tool can run a terminology check against it automatically, so the check stays a manual read by a person who has to remember every decision

How matching really works: exact, fuzzy, and the segmentation trap

An exact match means the tool normalized both strings and found them identical. What normalizing covers is the tool's decision, not something TMX specifies, and that is where sentences that look the same quietly fail to match. A full-width question mark (U+FF1F) and an ASCII question mark are different code points, so two Japanese sentences that are indistinguishable on screen are not equal as strings; NFKC normalization folds them together and NFC does not. The character が written as one precomposed code point and as か followed by a combining voiced sound mark is the same story, except that here NFC alone is enough to make the two match.

Fuzzy matching is where the misunderstanding is most expensive, because the percentage lives in the tool and not in the file. TMX stores sentence pairs; it stores no match rates. Give the identical TMX to two tools and they will report different percentages for the same lookup, because each has its own edit-distance definition and its own weighting of punctuation, numbers and inline tags. A leverage figure in a quote therefore describes a tool, not your asset, so ask which tool produced each analysis before comparing two quotes.

Segmentation is the failure that makes a healthy memory look empty. The segtype attribute in the header says what one unit is. A memory exported at paragraph granularity, looked up sentence by sentence, yields almost no exact matches even though every one of your sentences sits somewhere inside it. Segmentation rules have their own exchange format, SRX, for the case where two tools must agree on where a sentence ends — worth asking about if matches drop right after a tool change.

Inline tags are the last piece. The elements that mark formatting and variables inside a segment are what tell the next tool that a value belongs in a given position. A tool that flattens them to plain text on import raises the apparent match rate and destroys that information in the same step, which is one of the ordinary ways a translated line ships without its placeholder. If your strings carry variables, keep the tags: ICU MessageFormat placeholders and plurals covers what is inside them, and the XLIFF format for handing over a translation job is what you send out for the work itself, as opposed to the memory you keep.

Where a memory pays off in a game project, and where it does not

A translation memory earns its keep when the same sentences come back. In game work that is a specific and recognisable set of situations:

  • A sequel or spin-off reusing the systems, menus and item descriptions of the previous title
  • Patch and update cycles, where most of a file is unchanged and only the new lines need a translator
  • Live-service content, where event copy follows the same frame every time and only names and numbers change
  • A studio with several titles sharing platform and interface boilerplate: settings screens, controller prompts, store and legal wording

It does almost nothing in two other situations, and pretending otherwise is how a bad translation gets approved. The first is a catalogue of very short interface fragments. On a two-word string a 100 percent match is a hazard rather than a saving, because the same source word is a verb on one screen and a label on another, and the memory has no way to tell you which one it recorded. The second is anything strongly context-dependent: a line whose translation depends on who is speaking, a reply whose register is set by the previous line, a character whose voice is deliberately unlike everyone else's. A perfect match taken from another speaker is a regression with a green checkmark next to it.

That is also the clean division of labour between the two formats. The memory answers whether a sentence has been translated before. The termbase answers what a thing is called — and for short fragments and named entities, the termbase plus a usage note in the source file is what actually helps, while the memory is close to useless. Context notes therefore belong in the handoff, not a follow-up email; see what goes into a localization kit.

Owning the asset: what to ask for at delivery, and what to check

Start with the agreement: it has to say that the memory and the termbase are yours, and that wording should be checked by someone qualified in your jurisdiction. Once that is settled, four things are worth doing.

  • Put the exports in the deliverables in writing, before work starts, and ask for them at every milestone instead of only at the end. A memory received once, on the last day, is one nobody notices is missing units.
  • Store the raw export next to your source strings, in the same repository or at least the same backup. A memory that exists only inside a tool account is one cancelled subscription away from gone.
  • When you change tools, export, import, then compare unit counts before you cancel the old one. A silent drop on import is ordinary, and it is only cheap to notice on the day it happens.
  • Decide what travels. A memory export is a record of everything ever translated, including what has not shipped: unannounced titles, character names, plot beats, store copy for an unannounced region. TMX also carries per-unit attributes for who created and who last changed a unit, and in practice those often hold a translator's account name. Treat the file like a database dump before sending it to a new vendor, and strip what should not leave.

Five ways these files break, and what one command can and cannot catch

Two of the five are ordinary XML failures that a well-formedness check finds in a second. The command xmllint --noout memory.tmx ships with macOS and comes with the libxml2 tools on Linux; exit status zero means well-formed, and a failure names the line and column.

  • Declared encoding against actual bytes. A file declaring UTF-8 whose Japanese segment is really CP932 fails the check, and the parser prints the offending bytes so you can identify the encoding. This is the same family of problem as mojibake in game text and Shift_JIS and CP932 handling.
  • An unescaped ampersand or angle bracket in an interface string. A segment reading Save & Exit is a parser error, and it is what you get when XML was assembled by string concatenation instead of a serializer.
  • Language tag variation, which the check does not catch. A case difference is harmless, because JA-jp and ja-JP canonicalize to the same tag. But ja and ja-JP canonicalize to two different strings, so a tool comparing them literally sees two languages and files half your units under a language you did not ask for. An underscore form such as ja_JP is not a language tag at all, and a file containing one still passes a well-formedness check, because the content of xml:lang is not validated.
  • Inline tags dropped on a round-trip. Nothing in a well-formedness check notices, because a file that lost its placeholders is still perfectly well-formed.
  • Tool-specific extensions. Both formats let a tool attach its own properties, and anything the next tool does not recognise is dropped without comment. Round-trip a copy and compare counts before you trust an export.
$ xmllint --noout memory.tmx
memory.tmx:7: parser error : Input is not proper UTF-8, indicate encoding !
Bytes: 0x8C 0xAE 0x82 0xF0
      <tuv xml:lang="ja"><seg>������������B</seg></tuv>
                              ^

$ xmllint --noout glossary.tmx
glossary.tmx:6: parser error : xmlParseEntityRef: no name
      <tuv xml:lang="en"><seg>Save & Exit</seg></tuv>
                                    ^

Well-formed is not the same as importable. The check tells you the XML parses, and nothing about whether the header carries the attributes your tool requires, whether the language tags are the ones it expects, or whether the dialect is one it reads at all. The import log is the second thing to read, and the unit count the tool reports afterwards is the number to compare against the count in the file. The engineering half of the same pipeline is covered by the PO format and what localization QA actually checks.

Related articles