PDF Text Extractor — Text & Metadata from URLs avatar

PDF Text Extractor — Text & Metadata from URLs

Pricing

from $2.00 / 1,000 results

Go to Apify Store
PDF Text Extractor — Text & Metadata from URLs

PDF Text Extractor — Text & Metadata from URLs

Download PDF files by URL and extract their text and metadata (page count, title, author) — whole-document or per-page. Honest per-URL status. No auth.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

alaudin burki

alaudin burki

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 hours ago

Last modified

Share

Download PDF files by URL and extract their text and metadata (page count, title, author) — whole-document or per-page. Honest per-URL status. No auth.

Built reliability-first: every row reports what was found and what was missing — you never get a silent blank, and a run summary tells you exactly what happened.

What you get

Returns one clean, structured row per result. Every row reports its own status so you never get a silent blank.

How to use it

  1. Fill in the input (see the example below).
  2. Run it once for a snapshot, or schedule it to keep the data fresh.
  3. Export to CSV/JSON/Excel, or push straight to Google Sheets, Notion, Airtable, Zapier, Make, or n8n.

Input

{
"pdfUrls": [
"https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
],
"perPage": false,
"useProxy": false,
"maxItems": 100000
}

Sample output

[
{
"result": "example",
"status": "ok",
"scrapedAt": "2026-09-03T09:00:00.000Z"
}
]

Typical uses

  • RAG / AI pipelines — turn PDFs (reports, papers, manuals) into text.
  • Document processing — extract invoices/contracts text for downstream parsing.
  • Search indexing — index PDF contents.

Pricing

$2.00 / 1,000 results ($0.002 per result), plus a near-zero start fee. You are never charged beyond your limit, and blocked or failed items are reported honestly — not billed as data.

FAQ & limitations

  • Public data only — no login walls, no cookies required.
  • Rate limits on the source may require the proxy or a retry on very large pulls.
  • Every row reports its own status, so partial results are always labeled, never faked.
  • Integrations: output works with Zapier, Make, n8n, and any webhook via Apify's integrations.
  • Formats: results export as JSON, CSV, Excel, or HTML from the dataset.
  • Portfolio Health Monitor
  • Format Converter
  • URL Metadata Extractor