PO file format explained: msgid, msgstr, plurals, POT and MO
A PO file is a plain-text catalog. It is a list of entries, each pairing one source string (msgid) with one translation (msgstr). POT is the same file with every msgstr left empty, the template a project extracts from its own source code. MO is the compiled binary that the running program actually opens. You edit the PO, you compile it to MO, and the program never reads your PO file at all.
That last sentence is where most PO trouble lives. A translation can sit in the file, look finished to everyone who opens it, and still never reach the screen. This article walks one complete catalog field by field, shows how many plural forms each language really needs, and lists the failure modes. Every claim about tool behaviour below was reproduced against GNU gettext's own utilities and read back through a gettext runtime, and the observed result is stated rather than summarised.
Every part of a PO file, in one validated example
Below is a complete Japanese catalog, not a fragment. It carries the header entry, all six comment markers, a plural entry, two entries disambiguated by msgctxt, a multi-line string with escapes, a fuzzy entry and an obsolete entry. Checking it with msgfmt reported no errors, only warnings that two optional header fields were absent, and compiling it produced a catalog a runtime could read.
# Japanese translations for the demo package. # Copyright (C) 2026 Example Studio # msgid "" msgstr "" "Project-Id-Version: demo 1.0\n" "POT-Creation-Date: 2026-09-12 10:00+0900\n" "PO-Revision-Date: 2026-09-12 11:30+0900\n" "Language: ja\n" "MIME-Version: 1.0\n" "Content-Type: text/plain; charset=UTF-8\n" "Content-Transfer-Encoding: 8bit\n" "Plural-Forms: nplurals=1; plural=0;\n" #. Tutorial overlay. %s is a key name such as I. #: src/ui/tutorial.c:42 #, c-format msgid "Press %s to open your inventory." msgstr "%s キーで持ち物を開きます。" #: src/ui/inventory.c:118 #, c-format msgid "%d item found" msgid_plural "%d items found" msgstr[0] "%d 個見つかりました" #: src/ui/dialog.c:7 msgctxt "button" msgid "Close" msgstr "閉じる" #: src/ui/map.c:210 msgctxt "distance" msgid "Close" msgstr "近い" #: src/ui/quest.c:64 msgid "" "The bridge is out.\n" "Ask the ferryman about the \"north road\"." msgstr "" "橋が落ちています。\n" "渡し守に「北の道」について聞いてください。" #: src/ui/menu.c:31 #, fuzzy #| msgid "Save game" msgid "Save progress" msgstr "ゲームを保存" #~ msgid "Old tutorial text" #~ msgstr "古いチュートリアル文"
The header, the comment markers, and how strings are quoted
The header is a real entry whose msgid is the empty string; its msgstr holds the metadata, one line per field. Two of those fields are load-bearing. Content-Type declares the character encoding the rest of the file is written in, and Plural-Forms declares how the runtime chooses between plural slots. The rest is bookkeeping that the tools warn about but do not require. The exhaustive field-by-field definition is the PO Files chapter of the GNU gettext manual.
Strings are C-style double-quoted literals. A backslash before n is a newline, before a double quote a literal quote, before a backslash a literal backslash. A string too long for one line is written as an empty first string followed by continuation lines, which are joined with nothing between them, so the line break in the file is not part of the value. Whitespace inside the quotes is literal, which means a trailing space left in the source code becomes part of the key.
The comment markers above a msgid are not interchangeable. Each means something specific and tools route them differently, so a translator note written with the wrong marker can be discarded on the next extraction.
- # followed by a space is a translator comment, written by hand and carried across re-extraction
- #. is an extracted comment, written by the developer in the source code for the translator to read
- #: is a reference, naming the file and line the string was extracted from
- #, holds flags. c-format marks a printf-style string so tools validate its placeholders; fuzzy marks the entry as needing human review
- #| records the previous msgid, written when a source string changes so the translator can see what it used to say
- #~ is an obsolete entry, kept commented out after its source string disappeared
How many plural forms your language needs, and two traps in the table
Plural-Forms declares two things: nplurals, the number of msgstr slots an entry must have, and plural, a C-like integer expression that maps a count to a slot index. The translator fills slots and the expression picks one at runtime. Getting the number wrong is caught only if you ask for it. Declaring nplurals=3 while an entry carried only msgstr[0] made msgfmt in checking mode report a fatal error and exit 1, yet it still wrote the MO file, which decompiled straight back to the broken entry. Without a checking flag the same file compiled silently and exited 0.
You do not have to write these expressions yourself. gettext ships a table of them and writes the header for you when you initialise a language from a template. Two things about that table matter more than the expressions do.
First, gettext's counts are not the plural category counts in the Unicode locale data. Asking a JavaScript runtime for its CLDR-backed categories returns four for Russian (one, few, many, other) and three for French (one, many, other), while gettext writes three for Russian and two for French. The categories gettext leaves out are the ones its integer-only expressions cannot reach, such as fractional counts and compact large numbers. So read nplurals as gettext's own count and never assume it equals a CLDR category count, because what CLDR standardises is a separate model. If you need the full category set, or gender and case selection inside one message, that is what ICU MessageFormat is for.
Second, the initialiser did not know Arabic or Simplified Chinese. It wrote a file with no Plural-Forms header at all and left the two plural slots it had copied from the template. Checking that file passed, because the slots were still empty. Fill them and the same missing header turns into a fatal error naming both the absent nplurals and the absent plural attribute. The defect is invisible when you add the language and surfaces at the translator's desk instead, so add the header by hand for any language the table does not cover, and confirm the coverage in your own gettext version, because the table changes between releases. Here is what it produced, verbatim:
en nplurals=2; plural=(n != 1); de, es nplurals=2; plural=(n != 1); fr, pt_BR nplurals=2; plural=(n > 1); ja, ko nplurals=1; plural=0; cs nplurals=3; plural=(n==1) ? 0 : (n>=2 && n<=4) ? 1 : 2; pl nplurals=3; plural=(n==1 ? 0 : ...); ru nplurals=3; plural=(n%10==1 && n%100!=11 ? 0 : ...); ar, zh_CN no Plural-Forms header written at all
Four ways a PO file silently loses a translation
Fuzzy entries do not reach the program. The catalog above has one fuzzy entry with a complete Japanese msgstr. Compiling it and reading it back through a gettext runtime returned the English source string, not the Japanese. msgfmt drops fuzzy entries from the MO, and its own statistics counted that file as five translated messages and one fuzzy translation.
Editing source text throws the translation away. I changed one character in a source string, a period to an exclamation mark, regenerated the template and merged. With the merge tool's default fuzzy matching, the old translation was carried forward and marked fuzzy, which by the rule above means the screen shows English. With fuzzy matching turned off, the entry came back empty and the old translation was pushed into an obsolete block. This is the real price of keying translations by source text: msgctxt splits collisions apart, but the msgid is still part of the key, so if your English gets edited often, weigh a stable identifier instead, as in designing string keys. Merging with the previous-msgid option helps in both cases, because it writes the old source text above the changed entry as a #| comment.
A charset mismatch produces mojibake with no error at all. I changed the header charset to ISO-8859-1 and left the file's bytes as UTF-8. Checking the file exited clean and said nothing about it. The compiled catalog then returned the Japanese for Close as nine Latin-1 characters where three Japanese ones belonged, one character per byte of the original, and re-saving that string as UTF-8 grew it from nine bytes to eighteen. This is ordinary double encoding, which is worth learning to recognise on sight: why UTF-8 read twice produces those characters and how to identify and undo mojibake.
Re-wrapping makes diffs that mean nothing. I wrote one long entry on a single line and ran it through the catalog tool: the file went from seven lines to eleven with identical content, and compiling both versions produced byte-identical MO files. So when a tool hands a file back with every long entry re-flowed, the whole diff is noise. The no-wrap option lives on the commands that write PO files, not on the compiler: msgcat, msgmerge, msgattrib, xgettext and msginit all take it, while msgfmt rejects it outright because it writes a binary catalog and has no wrapping to control. Pick one convention and apply it on both sides of the handoff.
Where you meet PO files, and what to check before handing one over
gettext grew up in the C world and travelled with it, so PO files turn up in two kinds of place. Some ecosystems use gettext natively: the C and C++ projects it was written for, desktop Linux stacks, PHP-based publishing platforms, and Python, whose standard library ships a module that reads MO files directly. That module is what I used to check the runtime behaviour described above. Elsewhere PO is an interchange format a tool chose. Unreal Engine's localization dashboard exports and imports PO while its packaged build uses its own compiled form, a split covered in the Unreal localization pipeline. Godot accepts PO as one of its two translation formats, alongside CSV. Support beyond those varies by engine and by version, and several visual-novel engines use their own script-based translation files instead, so read the engine's own localization documentation rather than assuming PO will work.
Before a catalog leaves your hands, six checks catch almost everything above, and four of them are things no tool will volunteer. The rest of the material that travels with the files is in what belongs in a localization kit.
- Compile every PO file with msgfmt in checking mode and fail the build on a non-zero exit code. Plural-slot and placeholder mismatches both come back as fatal errors, which I confirmed on purpose-broken files, but msgfmt writes the MO anyway, so a step that ignores the exit code ships the broken catalog.
- Read the two headers that matter. The Content-Type charset must match the file's actual bytes, and Plural-Forms must be present and correct for the language. Neither is checked for you in the case that hurts.
- Count the fuzzy entries and decide what happens to them. The statistics option prints the number, and the attribute filter can extract only the fuzzy entries into a separate file you can send for review.
- Fix the wrapping convention and write it down next to the files, so a re-flowed handoff never turns into a diff nobody can read.
- Explain the placeholders in an extracted comment in the source code, where re-extraction preserves it. A c-format flag lets a tool confirm that a placeholder survived, but nothing tells a translator whether it holds a key name or an item name.
- Say whether placeholders may be reordered. Several unnumbered placeholders in one printf-style string must stay in order unless the format supports positional arguments, so check the runtime's documentation before promising a translator otherwise.
PO, POT and MO, and why this loop has lasted
The three extensions are three points in one cycle, not competing formats. An extractor scans the source code and writes the POT, the template in which every msgstr is empty. Each language's PO is created from that template and later updated from it by the merge tool, and the PO is the only file a human edits or version control needs to track. msgfmt compiles each PO into an MO installed under a locale directory, and the MO is the only one the running program opens.
The merge step is what makes the cycle cheap. Unchanged entries pass through untouched, new entries arrive empty, and entries whose source string has vanished move into obsolete blocks rather than being deleted, so a string pulled for one release does not lose its translation when it comes back. The translator gets a bounded list of work after each source change instead of a full re-translation.
Every stage is a text file, so the whole pipeline diffs in version control and needs no server, no database and no account. That is why gettext has outlasted most of its contemporaries in open-source software, and why sending a PO file to a translator by mail is still a defensible thing to do.
Its weaknesses are the four above: the source text doubles as the key, fuzzy state does not show on screen until you compile, and the plural model is narrower than the Unicode one. If your work is mostly handoff between separate tools and outside translators rather than a codebase you control, a format built for that purpose fits better, which is the case for XLIFF as a handoff format and for TMX and TBX for memories and glossaries.