PDF to Markdown Converter (Text, Headings, Metadata) avatar

PDF to Markdown Converter (Text, Headings, Metadata)

Pricing

from $1.20 / 1,000 pdfs

Go to Apify Store
PDF to Markdown Converter (Text, Headings, Metadata)

PDF to Markdown Converter (Text, Headings, Metadata)

Convert PDF files to clean Markdown for LLMs and RAG: headings detected from font sizes, paragraphs joined, page headers, footers and page numbers removed, plus title, author, dates, page count, links and token count. Give PDF URLs or pages that link to PDFs.

Pricing

from $1.20 / 1,000 pdfs

Rating

0.0

(0)

Developer

Murat Uzun

Murat Uzun

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What is PDF to Markdown Converter?

PDF to Markdown Converter turns PDF files into clean Markdown ready for ChatGPT, Claude, RAG pipelines and vector databases. Give it PDF links, or web pages that link to PDFs (report libraries, investor relations, government forms), and for each PDF it returns the Markdown with real headings (detected from font sizes), paragraphs joined across line breaks, hyphenation undone, running headers, footers and page numbers removed, plus the PDF's title, author, subject, creation date, page count, word and token counts and the links inside it. $2 per 1,000 PDFs.

Use it to feed research papers, annual reports, manuals, contracts or government documents into an LLM, to build a searchable document archive, or to extract text from many PDFs at once.

What data does PDF to Markdown Converter extract?

FieldDescription
url, finalUrl, sourcePage, fileName, sizeBytesThe PDF, where it was found, its file name and size
title, author, subject, keywords, creator, producerDocument metadata (title falls back to the first heading)
createdAt, modifiedAt, languageDates (ISO) and language from the PDF
pageCount, pagesReadPages in the file, and pages converted (see Max pages per PDF)
markdownThe document as Markdown: #/##/### headings, paragraphs, - bullets
wordCount, tokensSize of the text (tokens ≈ characters / 4)
isScannedtrue when the PDF has almost no text layer (a scan; see Limits)
linksURLs linked from inside the PDF
text, pagesOptional: raw text, and { page, text } for every page

How to use PDF to Markdown Converter

  1. Paste PDF URLs, and/or Web pages with PDF links to convert the PDFs they link to.
  2. Optional: Page markers in Markdown (<!-- page N --> before each page, handy for citations), Plain text too, Text per page.
  3. Run. PDFs are downloaded and converted three at a time.
  4. Download the rows as JSON, CSV, Excel or HTML, or read them through the API.

Example input

{
"pdfUrls": ["https://arxiv.org/pdf/1706.03762", "https://www.irs.gov/pub/irs-pdf/fw9.pdf"],
"pageUrls": ["https://www.irs.gov/forms-pubs/about-form-w-9"],
"maxPdfsPerPage": 5,
"pageMarkers": true
}

Example output

{
"url": "https://arxiv.org/pdf/1706.03762",
"fileName": "1706.03762v7.pdf",
"sizeBytes": 2215244,
"title": "Attention Is All You Need",
"createdAt": "2024-04-10T21:11:43Z",
"pageCount": 15,
"pagesRead": 15,
"wordCount": 5850,
"tokens": 9573,
"isScanned": false,
"links": ["https://github.com/tensorflow/tensor2tensor", "http://arxiv.org/abs/1607.06450"],
"markdown": "# Attention Is All You Need\n\nAshish Vaswani∗ Google Brain avaswani@google.com\n\n...\n\n## Abstract\n\nThe dominant sequence transduction models are based on complex recurrent or convolutional neural networks ...\n\n## 1 Introduction\n\n..."
}

How much does it cost?

Pay per result: $0.002 per PDF converted ($2 per 1,000), whatever its length up to Max pages per PDF, with volume discounts on paid Apify plans. Links that fail, are not PDFs, or are password-protected are listed in the ERRORS record and cost nothing. Set Maximum cost per run to cap spend.

Limits

  • No OCR. Scanned PDFs (images of pages) have no text layer; they come back with isScanned: true and little or no text.
  • Headings are inferred from font sizes, so unusual layouts can put a heading one level off, and tables come out as lines of text rather than Markdown tables.
  • Multi-column pages are read in the order the PDF stores its text; check the column order on a sample of your documents.
  • Files above 50 MB are skipped (listed in ERRORS).

Using PDF to Markdown Converter with AI agents

The Actor is pay-per-event with limited permissions, so AI agents can call it through Apify's MCP server (mcp.apify.com), for example with {"pdfUrls": ["https://arxiv.org/pdf/1706.03762"]}, and read markdown straight from the result. It also accepts url, urls, startUrls, and inputDatasetId + inputField to convert PDF links found by another Actor (for example a Google Search run with filetype:pdf).

FAQ

Does it work with PDFs behind a login? No, only PDFs anyone can download.

Can I convert a local PDF? Upload it somewhere that gives a public link (for example an Apify key-value store record) and paste that link.

Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:

Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.

Website & domain intelligence

Content for AI, LLMs and RAG

Search, video and social

Leads, jobs and company data

Developer, app and research data