PDF to Markdown & Table Extractor — CSV, JSON · $0.03/page
Pricing
Pay per event
PDF to Markdown & Table Extractor — CSV, JSON · $0.03/page
Extract text, tables and fields from PDF and DOCX files. Returns Markdown with headings and lists, tables as CSV rows, and named fields such as invoice number, date and total found by label. Rule-based, repeatable results, no subscription. Image-only scans are flagged. Failed files are not charged.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Ambolt
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 hours ago
Last modified
Categories
Share
Extract text, tables and fields from PDF and DOCX files. Returns Markdown with headings and lists, tables as CSV rows, and named fields such as invoice number, date and total found by label. Scanned pages are read with OCR at $0.06 a page. No subscription. Failed files are not charged.
Turn PDF and DOCX files into clean Markdown with headings, lists and tables, every table also as CSV rows, and optionally named fields (invoice number, total, dates) found by label. Layout rules, no AI model: the same file gives the same result. Per page pricing, files fetched from a URL, up to 20 files per run. Scanned pages are read with OCR (Tesseract) and billed separately.
What you get
Structured JSON for each run, ready for spreadsheets, alerts and AI agents. Also available as an MCP tool (doc_to_json) and as a pay-per-call HTTP endpoint.
Input
| Field | Type | Default | Description |
|---|---|---|---|
urls | array | Public URLs of PDF or DOCX files (up to 20). | |
maxPagesPerFile | integer | 50 | Stop after this many pages per PDF. Each processed page is one billable result. |
fieldHints | array | Optional fields to find, one per line as name | |
ocr | string (auto, off) | auto | auto reads pages without a text layer (scans) with OCR, billed as ocr-page; off skips them. |
ocrLanguages | string | eng | Tesseract language codes joined by +, for example eng+dan+deu. |
includeTablesCsv | boolean | true | Add every detected table as CSV text next to the rows. |
Example input
{"urls": ["https://www.rfc-editor.org/rfc/rfc9110.pdf"],"maxPagesPerFile": 3}
Example output
{"files": 1,"extracted": 1,"pagesProcessed": 4,"billableResults": 4,"datasetItems": [{"url": "https://www.rfc-editor.org/rfc/rfc9110.pdf","ok": true,"kind": "pdf","pages": 4,"characters": 5134,…
Pricing
| Event | Price |
|---|---|
| page (each successful result) | $0.030 |
Failed runs are not charged. Free trial credits from Apify cover your first runs.
FAQ
Is a result guaranteed to be complete? Each result carries the date it was read. If the upstream service is down the run fails with a clear message and you are not charged.
Can I call it from code or an AI agent? Yes: run it through the Apify API, or use the MCP tool and pay-per-call endpoint listed at https://ambolt.dev.
Is this affiliated with the data source? No. Source names are used descriptively; each result keeps its source attribution and licence line.
Where do I report a problem? Open an issue on this Actor's page; replies are written by the Ambolt service team, usually within a day.
Review
If this Actor saved you time, a short review on its Store page helps others find it and tells us what to improve.
Keywords
PDF to Markdown, PDF table extraction, PDF to CSV, PDF to JSON, DOCX to Markdown, invoice data extraction, document parsing, RAG, data entry automation, document to JSON
Notes
Data comes from the source named in each result and keeps its licence and attribution. Informational only, not financial, legal or tax advice. Open an issue on the Actor page for questions; replies come under the product name.