PDF to Text & Markdown: Extract PDF Content by URL avatar

PDF to Text & Markdown: Extract PDF Content by URL

Pricing

from $1.52 / 1,000 document processed (up to 20 pages)s

Go to Apify Store
PDF to Text & Markdown: Extract PDF Content by URL

PDF to Text & Markdown: Extract PDF Content by URL

Extracts text or Markdown from PDFs by URL, for a list of PDF URLs from a crawl, a Sheet, or a CRM export. Rebuilds paragraphs and headings from PDF fonts, with best-effort tables, then feeds a RAG pipeline, an LLM prompt, or a row to Google Sheets. No OCR: scanned PDFs are flagged, not charged.

Pricing

from $1.52 / 1,000 document processed (up to 20 pages)s

Rating

0.0

(0)

Developer

Adrian Voss

Adrian Voss

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

Give it a list of PDF URLs and get back clean text or Markdown for each one — headings and paragraphs rebuilt from the PDF's own font sizes, a best-effort Markdown table for simple aligned tables, and metadata (title, author, page count). Built for feeding a RAG pipeline, an LLM prompt, or a row to Google Sheets — anywhere you'd otherwise hand-copy text out of a PDF. No OCR yet: a scanned, image-only PDF is flagged (isScanned) and never charged.

Who it's for

Anyone who has a list of PDF URLs — from a web crawl, a Sheet, a CRM export, or a document management system — and needs a plain PDF to Markdown conversion, not a raw byte stream. Common uses: preparing PDFs to embed for RAG or as LLM context, converting a folder of reports from PDF to Markdown for a wiki, or piping extracted text to Google Sheets or Airtable via an integration.

How to use

Run node scripts/build-schemas.js once this actor is registered in data/portfolio.json to generate three numbered steps (Console, API curl, schedule) built from this actor's real input field name and seed input.

Input

{
"pdfUrls": [
"https://www.irs.gov/pub/irs-pdf/fw4.pdf",
"https://arxiv.org/pdf/1706.03762"
],
"pages": "",
"outputFormat": "markdown",
"includePageBreaks": false
}

pdfUrls is one direct PDF URL per line — for a list of PDF URLs from a crawl, a Sheet, or a CRM export. A URL pasted without its https:// (common when copied from a spreadsheet cell) is filled in automatically.

  • pages (optional) — a page range like "1-5", or a single page like "3". Leave empty to extract every page.
  • outputFormat — "markdown" (default) rebuilds #/##/### headings from font sizes and renders any detected table as a best-effort Markdown table; "text" keeps the same paragraphs with no Markdown syntax.
  • includePageBreaks — inserts a marker between each page's content (an invisible <!-- page N --> comment in Markdown, "--- Page N ---" in plain text) so a downstream chunker can split by page.

Output

One row per PDF, for example:

Run node scripts/build-schemas.js once this actor has a canary task run to generate a real sample row here. A real row from this build's one live run (see docs/eval/builds/pdf-text-extractor.md), content trimmed:

{
"query": "https://www.irs.gov/pub/irs-pdf/fw4.pdf",
"found": true,
"status": "OK",
"url": "https://www.irs.gov/pub/irs-pdf/fw4.pdf",
"title": "2026 Form W-4",
"author": "C:DC:TS:CAR:MP",
"pageCount": 5,
"pagesExtracted": 5,
"outputFormat": "markdown",
"content": "### Employee's Withholding Certificate OMB No. 1545-0074\n\n## Form W-4\n\nComplete Form W-4 so that your employer can withhold the correct federal income tax from your pay...",
"wordCount": 5001,
"isScanned": false,
"scrapedAt": "2026-09-25T20:29:42.368Z"
}

A URL that isn't a PDF, doesn't exist, or is larger than the 50 MB limit comes back as a row with "found": false and is never charged. A scanned, image-only PDF (isScanned: true) is also never charged — see "No OCR in this version" below.

How this works

  1. Download — up to 50 MB per PDF, streamed with a hard cap so one huge file can't run away with your run's memory or your bill.
  2. Extract — pdfjs-dist reads the text layer page by page, with each character's font size and position.
  3. Reconstruct — lines are grouped by vertical position, a two-column academic layout is detected and read left-column-then-right-column instead of line-by-line across the gutter, consecutive lines of similar size are merged into paragraphs, and a line noticeably bigger than body text becomes a heading (#/##/###, biggest first). A run of 3+ consecutive lines that each split into well-separated clusters is rendered as a best-effort Markdown table.
  4. No OCR in this version — if a PDF's text layer is almost empty (a scanned page with no embedded text), isScanned is true, content is null, and the row is never charged.

This is a best-effort reconstruction, not a layout-perfect one:

  • Two-column detection works well for the common single-gutter academic layout. An exotic multi-region magazine layout, or a page where both columns happen to align to the exact same text baseline throughout, can read out of order.
  • Table detection requires at least 3 consecutive lines with clearly separated columns. A 1-2 row table, or a table with merged or empty cells in irregular places, may render as plain paragraph text instead.
  • Line-wrap dehyphenation (joining a word split across two lines by a hyphen) can't always tell a soft line-break hyphen from a real one in a compound word (e.g. "English-to-German" broken exactly at a hyphen can come out as "Englishto-German"). This is a known limitation of automatic PDF text extraction generally, not specific to this actor.

Pricing

Pay-per-event. A flat per-run fee covers session/proxy warmup; you're billed per item only when data is actually found and returned — see .actor/pay_per_event.json for exact prices. A miss is never charged.

One document event covers up to 20 extracted pages; a longer document also bills one page event for each page beyond 20, so a 200-page report doesn't cost the same as a 2-page memo. A scanned PDF with no text layer bills nothing at all, even though it still returns a row telling you so.

Use it from Clay, n8n, Make, or an AI agent

Run node scripts/build-schemas.js once this actor is priced and registered in data/portfolio.json to generate this section — a run-sync curl example, an n8n HTTP Request node recipe, a Clay "HTTP API" column recipe, and an MCP line, all built from this actor's real input field name and seed input.

A common pattern: run this actor on a list of PDF URLs, then send each row's content field to Google Sheets (one row per PDF) as a lightweight PDF-to- Google-Sheets pipeline, or straight into a RAG pipeline's document loader — the Markdown output chunks cleanly on its own #/## headings.

Data & privacy

This actor only reads the PDF at the URL you give it. Nothing is stored beyond the run's own dataset; no PDF content is retained after the run ends.

FAQ

Does it do OCR on scanned PDFs? Not in this version. A scanned or image-only PDF is flagged (isScanned: true) and never charged — no extracted text is fabricated.

What's the file size limit? 50 MB per PDF. A larger file comes back as a free BAD_FORMAT miss.

Can I extract just a few pages from a huge PDF? Yes — set pages to a range like "1-10". Only the requested pages count toward parsing time and billing (pagesExtracted, not pageCount); the whole file still has to be downloaded first, since PDF pages aren't independently fetchable over HTTP.