PDF to Text & Markdown: Extract PDF Content by URL
Pricing
from $1.52 / 1,000 document processed (up to 20 pages)s
PDF to Text & Markdown: Extract PDF Content by URL
Extracts text or Markdown from PDFs by URL, for a list of PDF URLs from a crawl, a Sheet, or a CRM export. Rebuilds paragraphs and headings from PDF fonts, with best-effort tables, then feeds a RAG pipeline, an LLM prompt, or a row to Google Sheets. No OCR: scanned PDFs are flagged, not charged.
Pricing
from $1.52 / 1,000 document processed (up to 20 pages)s
Rating
0.0
(0)
Developer
Adrian Voss
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Give it a list of PDF URLs and get back clean text or Markdown for each one —
headings and paragraphs rebuilt from the PDF's own font sizes, a best-effort
Markdown table for simple aligned tables, and metadata (title, author, page
count). Built for feeding a RAG pipeline, an LLM prompt, or a row to Google
Sheets — anywhere you'd otherwise hand-copy text out of a PDF. No OCR yet: a
scanned, image-only PDF is flagged (isScanned) and never charged.
Who it's for
Anyone who has a list of PDF URLs — from a web crawl, a Sheet, a CRM export, or a document management system — and needs a plain PDF to Markdown conversion, not a raw byte stream. Common uses: preparing PDFs to embed for RAG or as LLM context, converting a folder of reports from PDF to Markdown for a wiki, or piping extracted text to Google Sheets or Airtable via an integration.
How to use
Run node scripts/build-schemas.js once this actor is registered in
data/portfolio.json to generate three numbered steps (Console, API curl,
schedule) built from this actor's real input field name and seed input.
Input
{"pdfUrls": ["https://www.irs.gov/pub/irs-pdf/fw4.pdf","https://arxiv.org/pdf/1706.03762"],"pages": "","outputFormat": "markdown","includePageBreaks": false}
pdfUrls is one direct PDF URL per line — for a list of PDF URLs from a
crawl, a Sheet, or a CRM export. A URL pasted without its https:// (common
when copied from a spreadsheet cell) is filled in automatically.
pages(optional) — a page range like"1-5", or a single page like"3". Leave empty to extract every page.outputFormat—"markdown"(default) rebuilds#/##/###headings from font sizes and renders any detected table as a best-effort Markdown table;"text"keeps the same paragraphs with no Markdown syntax.includePageBreaks— inserts a marker between each page's content (an invisible<!-- page N -->comment in Markdown,"--- Page N ---"in plain text) so a downstream chunker can split by page.
Output
One row per PDF, for example:
Run node scripts/build-schemas.js once this actor has a canary task run to
generate a real sample row here. A real row from this build's one live run
(see docs/eval/builds/pdf-text-extractor.md), content trimmed:
{"query": "https://www.irs.gov/pub/irs-pdf/fw4.pdf","found": true,"status": "OK","url": "https://www.irs.gov/pub/irs-pdf/fw4.pdf","title": "2026 Form W-4","author": "C:DC:TS:CAR:MP","pageCount": 5,"pagesExtracted": 5,"outputFormat": "markdown","content": "### Employee's Withholding Certificate OMB No. 1545-0074\n\n## Form W-4\n\nComplete Form W-4 so that your employer can withhold the correct federal income tax from your pay...","wordCount": 5001,"isScanned": false,"scrapedAt": "2026-09-25T20:29:42.368Z"}
A URL that isn't a PDF, doesn't exist, or is larger than the 50 MB limit
comes back as a row with "found": false and is never charged. A scanned,
image-only PDF (isScanned: true) is also never charged — see "No OCR in
this version" below.
How this works
- Download — up to 50 MB per PDF, streamed with a hard cap so one huge file can't run away with your run's memory or your bill.
- Extract — pdfjs-dist reads the text layer page by page, with each character's font size and position.
- Reconstruct — lines are grouped by vertical position, a two-column
academic layout is detected and read left-column-then-right-column
instead of line-by-line across the gutter, consecutive lines of similar
size are merged into paragraphs, and a line noticeably bigger than body
text becomes a heading (
#/##/###, biggest first). A run of 3+ consecutive lines that each split into well-separated clusters is rendered as a best-effort Markdown table. - No OCR in this version — if a PDF's text layer is almost empty (a
scanned page with no embedded text),
isScannedistrue,contentisnull, and the row is never charged.
This is a best-effort reconstruction, not a layout-perfect one:
- Two-column detection works well for the common single-gutter academic layout. An exotic multi-region magazine layout, or a page where both columns happen to align to the exact same text baseline throughout, can read out of order.
- Table detection requires at least 3 consecutive lines with clearly separated columns. A 1-2 row table, or a table with merged or empty cells in irregular places, may render as plain paragraph text instead.
- Line-wrap dehyphenation (joining a word split across two lines by a hyphen) can't always tell a soft line-break hyphen from a real one in a compound word (e.g. "English-to-German" broken exactly at a hyphen can come out as "Englishto-German"). This is a known limitation of automatic PDF text extraction generally, not specific to this actor.
Pricing
Pay-per-event. A flat per-run fee covers session/proxy warmup; you're billed
per item only when data is actually found and returned — see
.actor/pay_per_event.json for exact prices. A miss is never charged.
One document event covers up to 20 extracted pages; a longer document also
bills one page event for each page beyond 20, so a 200-page report doesn't
cost the same as a 2-page memo. A scanned PDF with no text layer bills
nothing at all, even though it still returns a row telling you so.
Use it from Clay, n8n, Make, or an AI agent
Run node scripts/build-schemas.js once this actor is priced and registered in
data/portfolio.json to generate this section — a run-sync curl example, an n8n
HTTP Request node recipe, a Clay "HTTP API" column recipe, and an MCP line, all
built from this actor's real input field name and seed input.
A common pattern: run this actor on a list of PDF URLs, then send each row's
content field to Google Sheets (one row per PDF) as a lightweight PDF-to-
Google-Sheets pipeline, or straight into a RAG pipeline's document loader —
the Markdown output chunks cleanly on its own #/## headings.
Data & privacy
This actor only reads the PDF at the URL you give it. Nothing is stored beyond the run's own dataset; no PDF content is retained after the run ends.
FAQ
Does it do OCR on scanned PDFs? Not in this version. A scanned or
image-only PDF is flagged (isScanned: true) and never charged — no
extracted text is fabricated.
What's the file size limit? 50 MB per PDF. A larger file comes back as a
free BAD_FORMAT miss.
Can I extract just a few pages from a huge PDF? Yes — set pages to a
range like "1-10". Only the requested pages count toward parsing time and
billing (pagesExtracted, not pageCount); the whole file still has to be
downloaded first, since PDF pages aren't independently fetchable over HTTP.