Localisation · runs in your browser
Build a terminology glossary from your documents
Add a set of text documents and get the recurring terms with their spelling variants and counts, ready to export as CSV or TBX. Files are read in your browser.
- Up to 50 files of 20 MB · DOCX, PPTX, XLSX, HTML, XML, MD, CSV, TXT (PDF: paste the text)
- Exports CSV and TBX (ISO 30042)
installation-guide-v4.docx 412 KB release-notes-2025-Q3.docx 96 KB help-center-export.html 1.1 MB onboarding-deck.pptx 2.4 MB … 8 more files TBX language code: en Minimum occurrences: 3
Term Variants found (count) Most frequent Flag
------------------ -------------------------- ----------------- --------------
NorthPeak Relay NorthPeak Relay (148) NorthPeak Relay 3 spellings
Northpeak Relay (31)
North Peak Relay (7)
sync token sync token (94) sync token 2 spellings
synchronisation token (12)
email address email address (203) email address 2 spellings
e-mail address (44)
SSO SSO (77) SSO 3 spellings,
single sign-on (23) abbreviation
single sign on (5) and expansion
work item work item (61) work item 2 spellings
workitem (9)
data room data room (88) data room one form
Candidates: 214 · Flagged as inconsistent: 37
Awaiting reviewer approval · nothing is applied to
your files by this tool.One clear job, from documents to term base
- 1
Add the document set
Add the files a translator would receive: manuals, release notes, help center exports, interface strings, contracts. More source material means better counts, because a term that appears in one file and nowhere else tells you very little.
- 2
Read the candidate list
Terms are grouped with every spelling variant found and a count for each. Entries where several forms compete are flagged. Sort by frequency to handle the terms that appear everywhere first, or by variant count to find the messiest ones.
- 3
Export and review
Export the list as CSV or TBX. A subject-matter expert keeps, edits or deletes each row and fills in the preferred form. The approved list then goes into your translation environment or the style guide.
Building a term list before translation starts
What you bring
You bring the documents a translator would receive: manuals, release notes, help pages, specs or contracts. The tool takes up to 50 files of 20 MB each in DOCX, PPTX, XLSX, HTML, XML, MD, CSV or TXT. For a PDF, paste the text. Files are read in your browser. More source material gives better counts, because a term that appears in only one file tells you little about how your team really writes it.
What you get back
You get a candidate list of recurring terms. Each term shows every spelling variant found, how often each one occurs and in how many files. The most frequent form is marked, and terms with competing spellings are flagged. You can export the list as CSV or as TBX, the term base format (ISO 30042) that most translation environments import. Both exports carry the same rows. Your documents are not changed.
Settings and the preferred term
You set the language code for the TBX file and the minimum number of occurrences a term needs to reach the list. The tool does not choose the preferred spelling. The most common form in an old document set is sometimes the one you want to retire. The CSV leaves an empty column for the preferred form, and someone who knows the product should fill it in. Expect to delete or reword many rows. The tool does not translate anything.
Questions before you run it
Does this tool translate my documents?
No. It builds a candidate term base that you review before translation starts. Translation is a separate step, for example with the Layout-Preserving Document Translator, the Spreadsheet Translator or the Subtitle Translator.
What is TBX and which tools can read it?
TBX (TermBase eXchange) is the ISO 30042 format for term bases. Most translation environments import it, including memoQ, Trados Studio, Phrase and Smartcat. If your team works in a spreadsheet instead, open the CSV export in Excel and skip TBX entirely.
Does it choose the preferred spelling?
No. For each term it lists every spelling it found, how often each one occurs and in how many files, and marks the most frequent form. The tool has no way of knowing which spelling your brand has standardised on, so the CSV leaves an empty column where a reviewer enters the preferred form.
Which file formats can I add?
DOCX, PPTX, XLSX, HTML, XML, Markdown, CSV, TSV and plain text, up to 50 files of 20 MB each. PDF is not read: copy the text out of the PDF and paste it into the text field instead. Files are read in your browser and are not uploaded.
How many candidate terms should I expect?
That depends on the size and the variety of the set, so there is no fixed number. Raise the minimum number of occurrences or the minimum number of files to cut the list down, though rare terms are often the ones that cause the worst translation errors.