PDF Toolkit — OCR, Split, Merge
Pricing
from $6.00 / 1,000 page made searchables
PDF Toolkit — OCR, Split, Merge
OCR a scanned PDF into a searchable one, extract a page range, or merge several PDFs. Pay per page OCR'd or per job, with no subscription.
Pricing
from $6.00 / 1,000 page made searchables
Rating
0.0
(0)
Developer
Richie Ellis
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
Three PDF jobs that should never need a subscription: make a scan searchable, pull out a page range, and join files into one. Pay for the pages you actually OCR, or the job you actually run. No monthly fee, no account on a third-party site, no uploading a contract to a free web tool that keeps a copy.
What it does
| Operation | What you get back |
|---|---|
| OCR | Your PDF, unchanged to look at, with a searchable, selectable, copy-pasteable text layer underneath. Pages that already have text are left alone. |
| Split | A new PDF containing only the pages you asked for — 1-3, 1,3,5, or any mix. |
| Merge | One PDF, files joined in the order you list them. |
Output files land in the run's key-value store, ready to download or pass straight into the next step of a workflow.
Input
- Operation —
ocr,splitormerge. - PDF file — upload one file, or
- PDF URLs — direct
http(s)links, up to 20 files, 64 MiB each. Merge joins them in the order given. - Page range — split only:
1-3or1,3,5.
{"operation": "ocr","pdfUrls": ["https://example.com/scanned-contract.pdf"]}
{"operation": "split","pdfUrls": ["https://example.com/report.pdf"],"pages": "4-6"}
{"operation": "merge","pdfUrls": ["https://example.com/invoice-jan.pdf","https://example.com/invoice-feb.pdf"]}
Output
One dataset item per result, and the file itself in the key-value store:
{"operation": "ocr","sourceUrl": "https://example.com/scanned-contract.pdf","pages": 12,"message": "OCR complete (12 pages).","outputs": [{ "key": "ocr-scanned-contract-00-ocr.pdf", "filename": "ocr.pdf", "bytes": 412887 }]}
An OUTPUT record summarises the run, including anything that was rejected and
why.
Pricing
| Event | Price |
|---|---|
| Page made searchable (OCR) | $0.006 per page |
| Page range extracted (split) | $0.015 per job |
| PDFs merged | $0.015 per job |
A 10-page scanned contract costs $0.06. Pulling three pages out of a 200-page report costs $0.015. You are charged only when the work succeeds — a file we reject costs you nothing beyond the run start.
Limits and guarantees
- 64 MiB (67 MB) per file, 20 files per run, 300 pages overall, 100 pages for OCR (OCR is the expensive one, so it gets the tighter cap).
- Nothing is kept. Each run gets its own isolated working directory, and working files are swept after two hours. Your output lives in your run's storage, under your account.
- Bad input is rejected, not crashed on. A file that is not really a PDF, a corrupt one, an oversized one, or a page range that is not a page range all come back as a plain message saying what was wrong.
Common uses
- Make a folder of scanned invoices or receipts searchable before filing them.
- Split a long statement or report into the section someone actually asked for.
- Merge a set of generated pages into one document at the end of a pipeline.
- Put a text layer on scans so a downstream extraction step has something to read.
Under the hood
ocrmypdf / tesseract for OCR, qpdf for split and merge, poppler for
inspection — the mature open-source tools, wrapped so you do not have to install
or babysit them. The processing core has been in production use since 7 Sep 2026
and was hardened against a written break-it checklist: non-PDF uploads, corrupt
files, oversized files, path-traversal filenames and shell-metacharacter page
ranges. subprocess is always called with an argument list, never shell=True.
Not supported yet
Password-protected PDFs, per-page image export, form filling, and languages other than English for OCR. Ask if you need one — they are small additions.