Doc to Markdown — DOCX, PDF & web pages to clean markdown
Pricing
Pay per event
Doc to Markdown — DOCX, PDF & web pages to clean markdown
Convert DOCX files, PDFs, and web pages into clean markdown for LLM and agent ingestion. Main-content extraction strips navigation junk from web pages. Pay per document.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Dos
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
19 hours ago
Last modified
Categories
Share
One call turns documents into clean markdown ready for LLM context windows: DOCX with headings preserved, PDFs page-by-page, and web pages with the navigation/ads/junk stripped (main-content extraction) — plus TXT/MD passthrough.
Why agents use it
Every RAG pipeline, summarizer, and research agent needs documents as clean text. Feed URLs, receive markdown with title, character count, and per-document status. Long documents include a direct download link to the full markdown file.
Output per document
{ url, status, format, title, n_chars, markdown, truncated, full_markdown_url? }
Pricing
Pay-per-event: one small fee per converted document. No subscription.
Limits
Up to 100 documents/run, 50 MB each. Text-based PDFs only (no OCR — see our pdf-ocr-extractor for scanned documents). Failures return status: "error".
Other actors by amanatools
- PDF Text Extractor — tables and text from PDF files into clean JSON rows
- PDF OCR Extractor — scanned and image-only PDFs into searchable text
- ATS Job Scraper — open jobs straight from Greenhouse, Lever, Workday and more
- Data Cleaner — messy CSV, Excel and JSON into clean, typed data