WoluTools
← About this tool

localisation · Browser tool

Terminology Glossary Extractor

Load a set of documents and see which terms recur, which of them are written more than one way, and how the spellings are split across files.

Runs in your browserNothing is uploadedHow it works
Free toolNo account · no job limitFree, unlimitedFiles are read by the page itself and never sent anywhere.

Documents

You can also drop files onto this panel. Up to 50 documents, 20 MB each.

Read here: TXT, MD, CSV, TSV, HTML, XML, and the OOXML formats DOCX, PPTX and XLSX — those three are ZIP archives, which the page opens with the browser's own decompression and XML parser. It reads the text runs, so tracked changes, comments, headers in tables and text inside images are not included. PDF is not read here: text in a PDF is stored glyph by glyph with its own font encoding, and there is no honest way to recover it without a PDF parser. Paste that text instead.

Paste text

Extraction settings

Raising this is the fastest way to cut noise. A word repeated forty times in one file and nowhere else is usually an artefact of that file, not a term.

Candidate glossary

XLSX is not offered here. Writing a real workbook means building a ZIP of OOXML parts, and a spreadsheet library is the only sane way to do that; the page has none. The CSV opens in Excel and carries the same rows.

The column is headed most frequent form, not preferred form, and the difference matters. Counting can tell you that one spelling wins 203 to 44. It cannot know that legal settled on the other one two years ago, or that the loser is the form your competitor uses and you are avoiding on purpose. Somebody who knows the product has to decide, and the CSV leaves an empty column for that decision.

What this does

The page splits every document into sentences, drops the grammatical filler, and counts each remaining word and each run of two or three words. Two counts are kept: how often a form appears, and in how many separate files it appears. The second matters more when you are building a term base. A phrase that turns up in nine of twelve documents is part of the vocabulary; the same phrase repeated nine times inside one release note is probably one paragraph someone wrote twice.

How variants are grouped

Forms are reduced to a comparison key before counting: capitals are flattened, camel case is split, hyphens and spaces are removed, and a crude plural rule is applied. So NorthPeak Relay, Northpeak Relay and North Peak Relay land on the same key, as do e-mail address and email address, and work item and workitem. Each surface form keeps its own count, so you see the split rather than a single merged number. When the same concept is spelled one way in the manual and another way in the release notes, the row is marked differs between files — that split is usually the finding people came for, because it means two people were writing to two different assumptions.

Abbreviations

An abbreviation is joined to its expansion when the initials of a recurring phrase match it and both appear in the same document. That catches SSO next to single sign-on. It will not catch an abbreviation whose expansion never appears in your files, and it will not catch one where the letters do not line up with the words, such as an abbreviation built from a word's internal letters. Those you find by reading the list.

Where the counting is wrong

The plural rule is a set of suffix patterns, so it handles tokens and policies and fails on indices and criteria. Stop-word filtering is English; a German or French document set will produce noise until you add its function words to the ignore list. Nothing here understands meaning, so a term used in two senses looks like one term, and two terms that happen to normalise alike are merged wrongly. Read the list before you export it — it is a candidate list, which is why every row has room for a reviewer's verdict.

On the file formats

DOCX, PPTX and XLSX are ZIP archives holding XML. The page reads the archive directory itself and uses the browser's built-in decompression, then the built-in XML parser, so those three work without any library. PDF does not work that way and is refused rather than guessed at. Everything happens in this tab; no file leaves your machine, which is the point when the document set is a draft contract or an unreleased manual.