Process & operationsこの記事を日本語で読む

Preparing a localization kit (lockit): what to send with the text

A localization kit, shortened to lockit in the industry, is everything you send a translator alongside the text itself: the string file, a glossary, a style guide, character notes, screenshots, and the written rules for placeholders, file format and delivery. It exists because a bare list of strings tells a translator what the words are and nothing about the game they belong to.

This article is the contents list, with a real column layout, glossary rows, a character sheet row and placeholder rules you can copy. It also names the four questions translators ask on almost every project, a checklist to run before you hand anything over, and the smallest kit that still works when you are one person.

What belongs in a kit, and what breaks without each piece

The test for including something is simple: name the failure it prevents. If you cannot name one, leave it out. Each line below pairs an item with what goes wrong in its absence.

  • The string file, in the format the translator will actually work in, with stable keys. Without stable keys a re-export reorders rows, and the next batch matches the wrong translation to the wrong line.
  • Context for each string: where it appears and what is happening. Without it the translator has to guess, and a one-word string like Fire becomes a verb when it was labelling a spell.
  • A speaker for every line of dialogue. Without it Japanese cannot choose a first-person pronoun or a politeness level, so one character ends up speaking in three registers inside one scene.
  • A character limit on strings that live in a fixed box. Without it the translation is correct and still clipped in the build.
  • A screenshot reference for each interface string. Without it a button, a heading and a tooltip all read the same to the translator, and all three get the same phrasing.
  • A glossary of fixed terms: items, characters, factions, mechanics. Without it the same sword carries three names across files, or two names across translators.
  • A style guide: tone, formality, forbidden words, how numbers and dates are written — the conventions covered in the books worth reading before a first localization. Without it the text reads as a patchwork of each translator's default voice rather than one game.
  • Placeholder and tag rules: what each token expands to, what type it holds, whether it may move. Without them tokens get translated, reordered into a broken sentence, or dropped.
  • The format you expect back, the text encoding, and the deadline. Without them you receive a file you cannot import, in an encoding that renders as garbage.
  • A named person to ask. Without one, the questions the kit failed to anticipate turn into assumptions, and assumptions surface in review instead.

The string file: columns, quoting, and length units

Seven columns carry most of a kit's value. The table at the end of this section is a header plus three rows, written as comma-separated values because nearly every engine and spreadsheet can both produce and read that. Key design deserves attention of its own, covered in string key naming and design. Three details decide whether such a table survives the handoff intact.

Quoting comes first, and it is exactly where hand-built kits break. Any field containing a comma must be wrapped in double quotes, or the comma is read as a column boundary. Write the second row's context unquoted as playful, not angry and the row parses into eight fields against a seven-column header: the speaker column receives the words after the comma, and every column after that shifts by one. The row still looked fine in a text editor. Parse the file once before sending it and compare each row's field count against the header, because a shifted row is silent and a translator will translate whatever landed in the source column.

The length column needs its unit stated, because the plausible answers disagree. The word cafe written with a separate combining accent measures five UTF-16 code units, five code points and four visible characters; written with a single precomposed accented letter it measures four, four and four. One fire emoji is two UTF-16 code units and one visible character. Whether your text box counts code units, code points or visible characters decides whether a limit is a constraint or a suggestion, and Japanese adds a second question on top: whether a full-width character costs the same as a half-width one in that box.

Treat an empty cell as meaning not applicable, never as meaning not filled in yet. A blank speaker on an interface label is information. A blank speaker on a line of dialogue is a question waiting to become an email. If your pipeline already uses a structured format, prefer it over a flat table: the XLIFF file format carries context, notes and per-unit state natively, and these same columns map onto it. For engine-side plumbing, the Unity localization CSV workflow covers the import side.

key,source_en,context,speaker,max_chars,screenshot,note
UI_SHOP_CONFIRM,Confirm,Bottom-right button of the shop dialog; beside Cancel,,12,shop_dialog.png,Same width as Cancel
DLG_ARA_0031,"Don't touch that, it's mine!","Ara catches the player at her workbench; playful, not angry",Ara,,workbench_02.png,
ITEM_DESC_EMBER,"Burns for {0} damage, and sets the target alight.",Item tooltip; {0} is an integer,,80,tooltip_ember.png,"{0} can be 1, so avoid plural-only phrasing"

Glossary rows, and a character sheet that fixes the voice

A glossary row needs four things: the source term, what kind of thing it is, the approved translation, and the translations you have already rejected. That last column is the one teams leave out and the one that saves the most re-review, because it turns a preference you hold silently into a rule the translator can follow.

A character sheet does for voice what the glossary does for terms. For Japanese the load-bearing fields are the first-person pronoun, the politeness level, and how the character addresses the player. All three are grammatically required in Japanese and entirely absent from English source text, so somebody decides them. If the kit does not decide, the translator decides, line by line, and the character drifts. That gap is the practical core of English to Japanese game localization.

# glossary.csv
term,type,ja,do_not_use,note
Ember Blade,item name,エンバーブレード,燃える刃 / 炎の剣,Proper noun; tooltip and quest text
Ash Covenant,faction name,灰の盟約,アッシュ・コヴェナント,Translate the meaning; English is never shown in JA builds
Stagger,mechanic,よろけ,スタガー / ふらつき,"Status effect, not a verb; identical in UI and tutorial"

# characters.csv
name,role,first_person,speech_style,formality,addresses_player_as,note
Ara,"Blacksmith, 40s",オレ,Blunt and clipped; drops sentence endings,常体 (plain form),あんた,Never uses keigo even to nobles

Placeholder rules, and the four questions translators always ask

Every project produces the same four questions. Front-loading means each one already has an answer in the kit: where does this appear, answered by the context and screenshot columns; who says this, answered by the speaker column plus the character sheet; how long can it be, answered by the length column with its unit stated; and what is this {0}, answered by a placeholder table written once instead of explained per string.

Writing down a token's type matters because of grammar the source language does not force you to think about. A count of one is grammatically singular in English and unremarkable in Japanese: in the plural categories that Unicode's CLDR locale data defines, English resolves into one and other, while Japanese has only other. So a string such as {0} items needs two forms in English and one in Japanese, and a translator who was never told that {0} is a number cannot know which. Message syntax that handles this properly is covered in the ICU MessageFormat guide.

{0} {1}        numbered positional token; its type is in the note column
{playerName}   named token; free text the player typed, any script, any length
%s %d          legacy printf tokens; %d holds an integer, %s a string
[b] ... [/b]   markup pair; keep both halves, keep them around the same words
\n             hard line break; keep it, do not add new ones

Never translate the text inside braces or brackets.
Numbered tokens may be reordered; the number travels with its value.
A token with no note is a gap in this kit, not in your translation. Ask.

The handoff checklist, and the text teams leave out

Run this before you send, not after the first question arrives.

  • Every row parses: field count matches the header on every line, and the file opens correctly in a second tool rather than only the one that wrote it.
  • Every dialogue row has a speaker, and every interface row has a screenshot reference.
  • Every placeholder that appears in the file also appears in the placeholder table.
  • Every glossary term actually occurs in the string file, and every fixed term in the file occurs in the glossary.
  • The encoding is stated, the return format is stated, and the deadline and the contact name sit on the first page of the kit rather than in the email that delivered it.

The smallest kit that works, and what to add later

Five categories of text sit outside the string file often enough to name individually. Strings still hardcoded in scripts and scene files, which have to come out before any kit can be complete: see externalizing hardcoded text. Words baked into images, such as a signpost, a logo or a tutorial diagram, which need either a redraw or a text layer. Store page copy, including the short description and the screenshot captions. Achievement and trophy names, which live in the platform's own console and not in your project at all. And legal or safety text such as an age notice or a required credit line, where a specialist should confirm the wording rather than a translator guessing at it.

If you are one person shipping one language, the minimum is four things: a string file with context and speaker columns, a glossary covering every proper noun, a one-page style guide that fixes formality and how the player is addressed, and a folder of screenshots named after the keys they contain. That is a bounded, finite job, and it absorbs most of what would otherwise arrive as questions.

What gets added as the project grows, roughly in the order it becomes necessary: a character sheet once more than a handful of characters speak; a do-not-translate list once the glossary is too long to read end to end; a length column once the interface stops growing to fit its text; and a structured format once two or more languages are in flight at once and you need per-string status rather than just text.

Two things are worth keeping out even as the kit grows. Do not attach a build the translator has no way to reach the text in, and do not write a prose changelog of every edit since the last batch, because a difference between two versions of the string file does that job without asking anyone to read. Instead treat the glossary and style guide as living documents rather than per-handoff attachments. They are the part of the kit that gets more valuable each time it is reused, and the part a checking pass measures the translation against, as described in what localization QA is.