PDF to Markdown & RAG Chunks: Document Parser (DOCX/PPTX/XLSX) avatar

PDF to Markdown & RAG Chunks: Document Parser (DOCX/PPTX/XLSX)

Pricing

from $2.00 / 1,000 document converteds

Go to Apify Store
PDF to Markdown & RAG Chunks: Document Parser (DOCX/PPTX/XLSX)

PDF to Markdown & RAG Chunks: Document Parser (DOCX/PPTX/XLSX)

PDF to Markdown and RAG chunks: a document parser that turns PDF, Word (DOCX), PowerPoint (PPTX) and Excel (XLSX) into clean Markdown text for RAG and LLMs. PDF table extraction inline as Markdown tables, heading-aware chunks with page numbers. Scanned pages flagged and free. No start fee.

Pricing

from $2.00 / 1,000 document converteds

Rating

0.0

(0)

Developer

Lindenwerk Data

Lindenwerk Data

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

12 hours ago

Last modified

Share

What is PDF to Markdown & RAG Chunks?

Convert PDF, Word (DOCX), PowerPoint (PPTX) and Excel (XLSX) files to clean Markdown for RAG and LLMs, built with pdfplumber, python-docx, python-pptx and openpyxl. USD 0.002 per document plus USD 0.0005 per page (USD 0.50 per 1,000 pages). No start fee; failed, empty and scanned-only files are free. No login, no API key of any other service, no LLM in the loop: the same file always gives the same Markdown.

Tables stay inline, in reading order, as Markdown tables, so a model sees "Los 1 | Wartung | 12" next to the paragraph that explains it. Headings become #/##/###, lists stay lists, running headers and footers ("Seite 3 von 46") are removed, and words hyphenated across line breaks are joined again. If you want, every document is also split into heading-aware chunks with page numbers (pageStart, pageEnd, headingPath), ready for a vector database and for answers that cite "page 12, § 4 Eignung".

It is made for German, French and other European documents: umlauts, accents, ligatures and ß come through cleanly, and the language from the PDF metadata is reported. A typical use is reading the tender documents found by our German, French and UK tender monitors (see below).

Who it's for

  • RAG and AI-agent builders who need clean Markdown and citable chunks from PDFs, Word, PowerPoint and Excel in one place, through the Apify API or MCP.
  • Bid and proposal teams who want to search, summarise or compare tender documents (Leistungsbeschreibung, cahier des charges (CCTP), statement of requirements) with an LLM.
  • Knowledge-base and search teams who index policies, manuals and reports in German, French and English.
  • Developers who want a hosted, pay-per-use converter instead of running their own PDF stack.

What you get

  • One row per document with the full Markdown, title, language, created and modified date, page count, pages with text, pages without a text layer, table count, word count, file size and a sha256 for deduplication and caching.
  • Optional chunk rows (outputMode): heading-aware, about chunkSize characters, with a small overlap inside a section (never across headings), page range and the heading path. Long tables are split by rows with the header repeated in every chunk.
  • Page markers (<!-- page: 12 -->) in the Markdown, so you can still cite pages after your own chunking.
  • Formats: PDF with a text layer, DOCX (headings from styles, numbered and bulleted lists, tables, page numbers from Word's rendered page breaks or estimated), PPTX (one section per slide, tables, speaker notes), XLSX (one Markdown table per sheet, values not formulas).
  • Honest flags instead of guesses: scanned pages are reported as pagesWithoutText and scannedPageCount, not filled with OCR noise. They are not charged.

How to convert PDFs to Markdown

  1. Click Start with the prefilled example (a one-page W3C sample PDF with a table) to see real output in a few seconds, for about USD 0.0025.
  2. Add your own files in Document URLs (public links), pick a key-value store in Your files: key-value store (see below), or send Base64 documents from your code.
  3. Choose the Output: documents only, documents and RAG chunks, or chunks only. Set Chunk size and Chunk overlap if you use chunks.
  4. Optional: a Page range such as 1-10, 15, keep or remove running headers and footers, and limits under Limits and performance.
  5. Read the results in the Output tab (views "Overview", "Markdown", "RAG chunks" and "Pages & quality") or download them as JSON, CSV or Excel. Very large documents keep their full Markdown in the run's key-value store.
  6. Call it from the Apify API, from AI agents via MCP, or chain it after our tender monitors with Apify integrations (Make, Zapier, n8n or a webhook). Save the input as a task and add a schedule to re-convert a list of documents regularly (every run converts and charges again; sha256 tells you what changed).

Input

{
"documentUrls": [
"https://dserver.bundestag.de/brd/2026/0225-26.pdf",
"https://www.fedlex.admin.ch/filestore/fedlex.data.admin.ch/eli/cc/2020/126/20250101/fr/pdf-a/fedlex-data-admin-ch-eli-cc-2020-126-20250101-fr-pdf-a.pdf"
],
"outputMode": "documents-and-chunks",
"chunkSize": 1500,
"chunkOverlap": 150,
"pageRange": "1-20"
}
  • documentUrls: public http(s) links. Apify key-value store record URLs are read with your run's permissions.
  • keyValueStoreId + keyValueStoreKeys: files you put into one of your key-value stores (all records if no keys).
  • files: API only, kept for compatibility; works like documentUrls.
  • base64Documents: [{"fileName": "angebot.docx", "contentBase64": "UEsDB..."}] for files that are not public.
  • outputMode: documents (default), documents-and-chunks or chunks.
  • Limits: maxDocuments (100), maxPagesPerDocument (500), maxFileSizeMb (30), maxRowsPerSheet (1000).

Your own files (from your computer)

This Actor runs with limited permissions: it can only read the storages you pick in its input, nothing else in your account. That is why the form has no separate upload button (a Console upload goes into a new store the run isn't allowed to read). Instead:

  1. In Apify Console open Storage > Key-value stores, open or create a store, and upload your PDF, DOCX, PPTX or XLSX files as records (the record key is the file name, e.g. angebot.docx).
  2. In this Actor's input pick that store in Your files: key-value store. Leave Record keys empty to convert every file in it, or list the keys you want.
  3. From code or an AI agent, send small files as Base64 documents instead, or pass Apify record URLs in documentUrls (they are read if the store is picked, readable by ID or the URL is pre-signed). If a record can't be read, its row says so (status: failed) and it is not charged.

Output

A document row (shortened):

{
"rowType": "document",
"source": "https://www.w3.org/WAI/WCAG21/working-examples/pdf-table/table.pdf",
"format": "pdf",
"status": "converted",
"title": "table",
"language": "EN-US",
"pageCount": 1,
"pagesConverted": 1,
"pagesWithoutText": [],
"tableCount": 1,
"wordCount": 99,
"markdown": "<!-- page: 1 -->\n\nExample table\n\nThis is an example of a data table.\n\n| Disability Category | Participants | Ballots Completed | ... |\n|---|---|---|---|\n| Blind | 5 | 1 | ... |",
"sha256": "a693998ff2a475d128c11644fbf02374249f08ac134d526a4d9b913d8b5834a5",
"chargedEvents": {"document-converted": 1, "page-converted": 1}
}

A chunk row (from the German procurement acceleration act as passed by the Bundestag, Bundesrat-Drucksache 225/26):

{
"rowType": "chunk",
"fileName": "0225-26.pdf",
"chunkId": "aacf340838a3-0007",
"pageStart": 4,
"pageEnd": 4,
"headingPath": ["Gesetzesbeschluss", "Gesetz zur Beschleunigung der Vergabe öffentlicher Aufträge", "unter Berücksichtigung einer Berichtigung in beigefügter Fassung angenommen.", "Änderung des Gesetzes gegen Wettbewerbsbeschränkungen"],
"text": "...\n5. Nach § 97 wird der folgende § 97a eingefügt:\n\n„§ 97a Losgrundsatz\n\n- (1) Leistungen sind in der Menge aufgeteilt (Teillose) ...",
"tokenEstimate": 327
}

Example source: Deutscher Bundestag/Bundesrat – DIP, BR-Drs. 225/26. Bundestag and Bundesrat documents are available free of charge in DIP (https://dip.bundestag.de). The third heading in headingPath shows that PDF heading detection is a heuristic (see the limits below).

status is one of converted, no_text (only scanned or blank pages, free), failed, encrypted, unsupported_format, too_large, robots_disallowed, refused, not_found, http_error, timeout or unreachable. Everything except converted is free. A full sample is in sample-output.json.

Pricing (pay per event)

EventPriceWhen
document-convertedUSD 0.002A document with at least one page of text was converted
page-convertedUSD 0.0005 (USD 0.50 per 1,000 pages)Each page (PDF, DOCX), slide (PPTX) or sheet (XLSX) with text
Free0Scanned or blank pages, failed, encrypted, unsupported, too large and robots.txt-blocked files, chunk rows
Actor startUSD 0.00005Apify platform minimum per run

A 10-page PDF costs USD 0.007. A 46-page tender regulation costs USD 0.025. 100 documents of 20 pages cost USD 1.20. There is no custom start fee and no minimum. Set Maximum cost per run in Apify Console: the Actor checks the budget before each document, converts only the first pages that still fit, marks that document truncated, and stops, so it never charges more than your limit. The smallest useful limit is USD 0.01.

Use with AI agents (MCP)

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com/?tools=lindenwerk/pdf-to-markdown-rag",
"headers": {"Authorization": "Bearer <APIFY_TOKEN>"}
}
}
}

Example prompts:

  • "Convert this tender PDF to Markdown and list every Eignungskriterium with its page number."
  • "Turn these three DOCX offers into chunks and tell me which one has the longest warranty."
  • "Read the price sheet in this XLSX and sum the lots."

For agents, outputMode: "chunks" with one document keeps answers short and citable. Over HTTP: POST https://api.apify.com/v2/acts/lindenwerk~pdf-to-markdown-rag/run-sync-get-dataset-items?token=<APIFY_TOKEN> with this input as the body:

{
"documentUrls": ["https://www.w3.org/WAI/WCAG21/working-examples/pdf-table/table.pdf"],
"outputMode": "chunks",
"chunkSize": 1000,
"maxDocuments": 1
}

Tender documents: German, French and UK procurement

Our tender monitors return a documentsUrl for each notice (the link to the procurement documents). Where that link is a direct, public PDF/DOCX/XLSX download, paste it into Document URLs or chain the runs with an integration:

  • German & EU Tenders Monitor: Vergabeunterlagen in German (deu), for example Leistungsbeschreibung, Preisblatt (XLSX) and Eignungskriterien. Umlauts, ß and hyphenation such as "Ver-gabe" are handled.
  • France Tenders Monitor: DCE documents in French (fra), for example the CCTP, CCAP or règlement de la consultation, with accents and French list styles intact.
  • UK Tenders Monitor: specifications and framework guidance in English.
  • SAM.gov Government Contract Opportunities Monitor: SAM.gov attachments need your own SAM.gov access, so download them and put them into a key-value store and pick it, or send them as Base64 documents.

Many tender platforms put documents behind a login or a click-through. This Actor doesn't log in or work around that: download such files yourself and add them through a key-value store or base64. Public links download in seconds. In our test run on Apify at the default 512 MB, the German procurement acceleration act (Bundesrat-Drucksache 225/26, 26 pages), the Swiss federal procurement act in French (LMP, 44 pages) and the UK framework guidance (18 pages) converted in under two minutes together, about one second per page. More memory gives the run more CPU, so large batches can finish faster.

How it works and limits

  • Text comes from the PDF's text layer (pdfplumber), tables from ruling lines in the PDF. Tables drawn only with whitespace (no lines) come out as text, not as a table. Multi-column page layouts are read line by line and can mix columns.
  • No OCR in v0.1. Pages without a text layer are flagged (scanned when an image covers the page, otherwise empty) and are free. OCR is planned for a later version as a separate, clearly priced event.
  • Headings are detected from font size and bold short lines in PDFs, and from styles in Word. Running headers and footers that repeat on at least half of the pages are removed (removeHeadersFooters).
  • DOCX has no fixed pages: page numbers come from Word's last rendered page breaks when the file has them (pageNumbering: "word-rendered"), otherwise they are estimated ("estimated"), and pages are charged on that basis.
  • Encrypted PDFs, legacy .doc/.ppt/.xls, OpenDocument, HTML pages and images are reported as free rows with a clear statusNote. Archives that unpack to huge sizes (zip bombs) are refused.
  • In our test, 5 public PDFs with 239 pages took 28 s at 512 MB, peak memory about 120 MB.

Data, privacy and responsible use

  • Honours robots.txt for the user agent token LindenwerkDocBot and sends a clear user agent. If robots.txt can't be fetched because of a server error, the file isn't downloaded. No proxies, no login, no captcha solving.
  • Safe fetching: only public http(s) hosts, every redirect checked, private and cloud-metadata addresses refused, size limits enforced while downloading, polite retries with Retry-After.
  • No personal data fields. The Actor never extracts names, email addresses or phone numbers as fields, and it never outputs the document's author metadata. The Markdown contains the document's text as it is, so only convert documents you are allowed to process.
  • Files are processed in your own Apify run and stored only in your run's storages.

FAQ

How is this different from asking an LLM to read the PDF? It's deterministic, cheap and fast, and it keeps tables and page numbers. Many teams convert first and give the Markdown or chunks to their model.

Why is my scanned PDF free but empty? It has no text layer. v0.1 doesn't run OCR, so it tells you which pages need it instead of charging for guesses.

Can it read password-protected files? No. Encrypted files are reported as encrypted and are free.

What chunk size should I use? 1,000 to 2,000 characters (about 250 to 500 tokens) works well for most embedding models. tokenEstimate is characters divided by 4.

Do you keep my files? No. They stay in your Apify account's run storage under your retention settings.

Examples

Other Lindenwerk Data Actors

Deutsch (Kurzfassung)

Dieser Actor wandelt PDF, Word (DOCX), PowerPoint (PPTX) und Excel (XLSX) in sauberes Markdown für RAG und LLMs um. Tabellen bleiben als Markdown-Tabellen an ihrer Stelle im Text, Überschriften und Listen bleiben erhalten, Kopf- und Fußzeilen werden entfernt. Optional gibt es Chunks mit Überschriftenpfad und Seitenzahlen für Zitate. Ideal für Vergabeunterlagen auf Deutsch und Französisch. Preis: USD 0,002 pro Dokument plus USD 0,0005 pro Seite mit Text, keine Startgebühr. Gescannte Seiten ohne Textebene, fehlerhafte und leere Dateien sind kostenlos. Der Actor beachtet robots.txt und erhebt keine personenbezogenen Daten als Felder.

Changelog

See the Changelog tab. Version 0.1 is the first public release.

Made by Lindenwerk Data.