WoluTools

PDF & Documents

Split a PDF by text, heading or blank page

Upload one long PDF and say where each new file starts: a word such as Chapter, a pattern, or a blank separator page. You get the parts as named PDFs in a ZIP, plus a CSV that shows which page started each part.

or drop it here

PDF · Up to 1,000 pages · Free jobs: up to 50 pages and 10 MB
Free account needed for your 3 free jobs a day

See an example split below

Standard toolIncluded · 3 free jobs a dayFree: 3 jobs a dayPro: up to 200 jobs a day · €12.99/month or €89.99/year
  • PDF
  • Up to 1,000 pages
  • Split by text, pattern or blank page
Example splitA 48-page report cut at every chapter
Before

report.pdf48 pages · chapters start on pages 1, 12, 25 and 37

Start a new file at
Regular expression
Pattern
^Chapter \d+
File names
report-{index}-p{first_page}-{last_page}
After · split-documents.zip
  • report-001-p1-11.pdf11 pages
  • report-002-p12-24.pdf13 pages
  • report-003-p25-36.pdf12 pages
  • report-004-p37-48.pdf12 pages
boundary-map.csv
Page,Rule,Evidence
12,Regular expression,"Chapter 2 Methods …"
25,Regular expression,"Chapter 3 Results …"
37,Regular expression,"Chapter 4 Discussion …"

Also in the package: page-map.csv (first and last page of every file) and unmatched-pages.csv. A 48-page PDF under 10 MB fits one free job.

BringPDF
GetNamed PDF ZIP, boundary map, unmatched-page report and manifest
PrivacyEncrypted source · 24-hour result

How a split works

  1. 1

    Add one PDF

    Up to 1,000 pages. A free job takes up to 50 pages and 10 MB.

  2. 2

    Set the rule and the names

    Pick heading text, a regular expression or a blank separator page, then type a name pattern such as part-{index}.

  3. 3

    Check and download

    Read the boundary map to see which pages started a new file, then download the ZIP.

Splitting a long PDF by heading or blank page

What you upload and how cuts are found

You upload one PDF with up to 1,000 pages and choose one rule for where a new file starts. The rule can be a heading text, a regular expression or a blank separator page. Heading and pattern rules need text that can be extracted, so a scan without a text layer will not match. The blank page rule still works on scans. A longer file is rejected up front rather than split part of the way.

What you get back and how to check it

You get split-documents.zip with the named parts, plus boundary-map.csv, which shows which input page went into which output file. Pages that no rule claimed are listed in unmatched-pages.csv, so nothing disappears quietly. A manifest.json file records the source and the files produced. Check the boundary map first. If every page appears once and the file names look right, the ZIP holds what you expect.

Splitting by bookmarks

The splitter reads the text printed on each page, not the bookmark list (the PDF outline). In most reports and books each bookmark points to a page that starts with the same chapter heading, so a text rule cuts at the same places. Heading text matches the word anywhere on a page, not only in a title, so a word like Chapter can also hit a page that mentions it in a sentence. For exact chapter starts, use a regular expression anchored to the start of a line, for example ^Chapter \d+ or ^\d+\. for numbered sections.

Rules and names you set

You choose where a new file starts and type the heading text or pattern. The output filename pattern sets how the parts are named. It accepts {index} (001, 002 …), {first_page} and {last_page}; names are built from this pattern, not from the heading text. The ZIP also holds page-map.csv with the first and last page of every file. You see where your rule lands before any file is written. If matches are inconsistent, the tool reports the exception rather than guessing. Whether each cut sits in the right place is still yours to confirm.

Questions before you run it

What can I split on besides page numbers?

A heading text match, a regular expression, or a blank separator page. You pick one rule and see where it lands before any file is written.

How do I know no pages were lost?

Every input page is reconciled to exactly one output file in boundary-map.csv. Pages that no rule claims are listed in unmatched-pages.csv instead of disappearing quietly.

Can I control the names of the split files?

Yes, through the output filename pattern. The resulting names appear in the boundary map, so you can check them before opening the ZIP.

Can I split a PDF by its bookmarks?

Not from the bookmark list itself. The tool reads the text on each page, so type the chapter heading or a pattern such as ^Chapter \d+ and it cuts where those headings appear, which is usually where the bookmarks point.

Does it work on scanned PDFs?

Heading and regex rules need extractable text, so a scan with no text layer will not match anything. A blank separator page rule still works, since it does not depend on recognised characters.

How long a document can one job take?

Up to 1,000 pages per PDF. A longer file is rejected up front rather than split part of the way.