# Changelog of PDF to Markdown & RAG Chunks: Document Parser (DOCX/PPTX/XLSX) (`lindenwerk/pdf-to-markdown-rag`) Actor

- **URL**: https://apify.com/lindenwerk/pdf-to-markdown-rag/changelog.md
- **Full Actor documentation**: https://apify.com/lindenwerk/pdf-to-markdown-rag.md

## Changelog

### README: cross-links and examples, no code or price change (2026-10-08, build 0.1.6)

- Product name in the README now matches the Store title ("PDF to Markdown & RAG Chunks: Document Parser (DOCX/PPTX/XLSX)").
- New Examples section linking the published example task "Convert a PDF with tables to Markdown and RAG chunks".
- "Other Lindenwerk Data Actors" lists all six sibling Actors by their current Store titles and Store URLs.
- No change to code, input schema, pricing, events, defaults or output.

All notable changes to this Actor are listed here.

### 0.1.4 (2026-10-06, before release)

- The input form no longer shows an **Upload files** button. Under limited permissions a Console upload goes into a
  new key-value store that the run can't read, so it always failed. To convert your own files, upload them to one of
  your key-value stores and pick it in **Your files: key-value store**, or send base64 from code.
- Apify record URLs that the run's token can't read get one plain retry without a token (works for stores readable
  by ID and pre-signed URLs). If that fails too, the row says clearly what to do instead. Such rows are free.
- Records that were given as URLs and also sit in the picked store are converted and charged only once.

### 0.1 (2026-10-06)

- First release.
- Converts PDF (text layer), Word (DOCX), PowerPoint (PPTX) and Excel (XLSX) to Markdown: headings, lists, tables
  inline in reading order as Markdown tables, running headers and footers removed, hyphenation across lines joined,
  page markers.
- Optional heading-aware RAG chunks with `pageStart`, `pageEnd`, `headingPath`, overlap inside a section only, long
  tables split by rows with the header repeated.
- Inputs: public URLs, uploaded files, Apify key-value store records and base64.
- Scanned and blank pages are flagged (`pagesWithoutText`, `scannedPageCount`). No OCR yet.
- Pay per event: `document-converted` USD 0.002 and `page-converted` USD 0.0005 per page with text. No start fee
  beyond the platform minimum. Failed, encrypted, unsupported, empty and scanned-only files and chunk rows are free.
  The run never charges more than the maximum cost per run.
- Honours robots.txt, honest user agent, SSRF guard on every redirect, file size cap, zip-bomb guard, polite retries.
  No personal data fields; author metadata is never output.
