PDF & Documents
Split a PDF by text, heading or blank page
Upload one long PDF and say where each new file starts: a word such as Chapter, a pattern, or a blank separator page. You get the parts as named PDFs in a ZIP, plus a CSV that shows which page started each part.
or drop it here
- Up to 1,000 pages
- Split by text, pattern or blank page
report.pdf48 pages · chapters start on pages 1, 12, 25 and 37
- Start a new file at
- Regular expression
- Pattern
^Chapter \d+- File names
report-{index}-p{first_page}-{last_page}
- report-001-p1-11.pdf11 pages
- report-002-p12-24.pdf13 pages
- report-003-p25-36.pdf12 pages
- report-004-p37-48.pdf12 pages
Page,Rule,Evidence
12,Regular expression,"Chapter 2 Methods …"
25,Regular expression,"Chapter 3 Results …"
37,Regular expression,"Chapter 4 Discussion …"Also in the package: page-map.csv (first and last page of every file) and unmatched-pages.csv. A 48-page PDF under 10 MB fits one free job.
How a split works
- 1
Add one PDF
Up to 1,000 pages. A free job takes up to 50 pages and 10 MB.
- 2
Set the rule and the names
Pick heading text, a regular expression or a blank separator page, then type a name pattern such as
part-{index}. - 3
Check and download
Read the boundary map to see which pages started a new file, then download the ZIP.
Splitting a long PDF by heading or blank page
What you upload and how cuts are found
You upload one PDF with up to 1,000 pages and choose one rule for where a new file starts. The rule can be a heading text, a regular expression or a blank separator page. Heading and pattern rules need text that can be extracted, so a scan without a text layer will not match. The blank page rule still works on scans. A longer file is rejected up front rather than split part of the way.
What you get back and how to check it
You get split-documents.zip with the named parts, plus boundary-map.csv, which shows which input page went into which output file. Pages that no rule claimed are listed in unmatched-pages.csv, so nothing disappears quietly. A manifest.json file records the source and the files produced. Check the boundary map first. If every page appears once and the file names look right, the ZIP holds what you expect.
Splitting by bookmarks
The splitter reads the text printed on each page, not the bookmark list (the PDF outline). In most reports and books each bookmark points to a page that starts with the same chapter heading, so a text rule cuts at the same places. Heading text matches the word anywhere on a page, not only in a title, so a word like Chapter can also hit a page that mentions it in a sentence. For exact chapter starts, use a regular expression anchored to the start of a line, for example ^Chapter \d+ or ^\d+\. for numbered sections.
Rules and names you set
You choose where a new file starts and type the heading text or pattern. The output filename pattern sets how the parts are named. It accepts {index} (001, 002 …), {first_page} and {last_page}; names are built from this pattern, not from the heading text. The ZIP also holds page-map.csv with the first and last page of every file. You see where your rule lands before any file is written. If matches are inconsistent, the tool reports the exception rather than guessing. Whether each cut sits in the right place is still yours to confirm.
Questions before you run it
What can I split on besides page numbers?
A heading text match, a regular expression, or a blank separator page. You pick one rule and see where it lands before any file is written.
How do I know no pages were lost?
Every input page is reconciled to exactly one output file in boundary-map.csv. Pages that no rule claims are listed in unmatched-pages.csv instead of disappearing quietly.
Can I control the names of the split files?
Yes, through the output filename pattern. The resulting names appear in the boundary map, so you can check them before opening the ZIP.
Can I split a PDF by its bookmarks?
Not from the bookmark list itself. The tool reads the text on each page, so type the chapter heading or a pattern such as ^Chapter \d+ and it cuts where those headings appear, which is usually where the bookmarks point.
Does it work on scanned PDFs?
Heading and regex rules need extractable text, so a scan with no text layer will not match anything. A blank separator page rule still works, since it does not depend on recognised characters.
How long a document can one job take?
Up to 1,000 pages per PDF. A longer file is rejected up front rather than split part of the way.