Word, PowerPoint & EPUB to Text: DOCX, PPTX, ODT
Pricing
$3.00 / 1,000 document extracteds
Word, PowerPoint & EPUB to Text: DOCX, PPTX, ODT
Extract text, sections and metadata from Word (DOCX), PowerPoint (PPTX), OpenDocument (ODT/ODP) and EPUB files by URL, up to 50 per run. $0.003 per document, no start fee; failed or unsupported files are never charged.
Pricing
$3.00 / 1,000 document extracteds
Rating
0.0
(0)
Developer
Yodesla
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
6 days ago
Last modified
Categories
Share
Convert DOCX, PPTX, ODT, ODP and EPUB files to clean text by URL, up to 50 per run. $0.003 per successfully extracted document, no start fee, failures free.
Supported formats
- docx — Word (OOXML). Detected by the
word/document.xmlzip member. - pptx — PowerPoint (OOXML). Detected by
ppt/slides/zip members. - odt / odp — OpenDocument text / presentation. Detected by the zip
mimetypefile. - epub — EPUB 2/3. Detected by the zip
mimetypefile (application/epub+zip).
The type is detected from the file's magic bytes and zip member names, never from the URL
extension. Legacy binary .doc/.ppt (OLE2) files return unsupported_format with a
message suggesting conversion to .docx/.pptx; they are not charged.
What you get
For each URL, one dataset item:
status:ok,download_failed,unsupported_format,invalid_file,empty, orerrorfileType:docx,pptx,odt,odp,epub(null when the file could not be classified)text: full extracted text (tables rendered as pipe-separated rows whenincludeTablesis on)sections[]:- docx: paragraphs grouped by heading (
heading,paragraphs) - pptx: one entry per slide (
slide,title,text,notes) - epub: one entry per chapter (
chapter,title,text) - odt/odp: paragraphs grouped by heading
- docx: paragraphs grouped by heading (
wordCount,charCount,truncated(true whentextwas cut atmaxChars)metadata:title,author,created,modifiedwhen the file provides themfileSizeBytes,processingMs, anderrorwith a plain-language reason when something fails
Pricing
$0.003 per successfully extracted document (document-extracted event), no start fee.
Failed downloads, legacy .doc/.ppt, invalid files and empty documents are reported in the
output but never charged. The run checks your spending limit before each file and stops
cleanly when it is reached.
Limits (by design)
- Max 50 URLs per run; max 50 MB per file; 25-second processing deadline per file; parsing runs in a separate process with a 384 MiB memory limit.
- Max 500,000 characters per document (configurable via
maxChars); longer text is cut and markedtruncated. - Only public
http(s)links; local and private-network addresses are refused. - No OCR, no scanned-document support (these formats are inherently digital).
- Tables are flattened to text rows; no nested structure beyond rows and cells.
Input example
{"urls": ["https://raw.githubusercontent.com/python-openxml/python-docx/master/tests/test_files/test.docx","https://raw.githubusercontent.com/scanny/python-pptx/master/tests/test_files/test.pptx"],"includeTables": true,"maxChars": 500000}
Licences
Runtime dependencies:
- apify — Apache-2.0
- python-docx — MIT
- python-pptx — MIT
- lxml — BSD
- odfpy — used under its Apache-2.0 option (its licence also offers GPL/LGPL options; we rely on the Apache option)
EPUB parsing is our own src/epub.py built on zipfile + lxml; ebooklib (AGPL-3.0) was
removed because AGPL is not acceptable for a public network actor. httpx (BSD) is used
only by the local deploy tool, not in the actor runtime.