Word, PowerPoint & EPUB to Text: DOCX, PPTX, ODT avatar

Word, PowerPoint & EPUB to Text: DOCX, PPTX, ODT

Pricing

$3.00 / 1,000 document extracteds

Go to Apify Store
Word, PowerPoint & EPUB to Text: DOCX, PPTX, ODT

Word, PowerPoint & EPUB to Text: DOCX, PPTX, ODT

Extract text, sections and metadata from Word (DOCX), PowerPoint (PPTX), OpenDocument (ODT/ODP) and EPUB files by URL, up to 50 per run. $0.003 per document, no start fee; failed or unsupported files are never charged.

Pricing

$3.00 / 1,000 document extracteds

Rating

0.0

(0)

Developer

Yodesla

Yodesla

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

6 days ago

Last modified

Share

Convert DOCX, PPTX, ODT, ODP and EPUB files to clean text by URL, up to 50 per run. $0.003 per successfully extracted document, no start fee, failures free.

Supported formats

  • docx — Word (OOXML). Detected by the word/document.xml zip member.
  • pptx — PowerPoint (OOXML). Detected by ppt/slides/ zip members.
  • odt / odp — OpenDocument text / presentation. Detected by the zip mimetype file.
  • epub — EPUB 2/3. Detected by the zip mimetype file (application/epub+zip).

The type is detected from the file's magic bytes and zip member names, never from the URL extension. Legacy binary .doc/.ppt (OLE2) files return unsupported_format with a message suggesting conversion to .docx/.pptx; they are not charged.

What you get

For each URL, one dataset item:

  • status: ok, download_failed, unsupported_format, invalid_file, empty, or error
  • fileType: docx, pptx, odt, odp, epub (null when the file could not be classified)
  • text: full extracted text (tables rendered as pipe-separated rows when includeTables is on)
  • sections[]:
    • docx: paragraphs grouped by heading (heading, paragraphs)
    • pptx: one entry per slide (slide, title, text, notes)
    • epub: one entry per chapter (chapter, title, text)
    • odt/odp: paragraphs grouped by heading
  • wordCount, charCount, truncated (true when text was cut at maxChars)
  • metadata: title, author, created, modified when the file provides them
  • fileSizeBytes, processingMs, and error with a plain-language reason when something fails

Pricing

$0.003 per successfully extracted document (document-extracted event), no start fee. Failed downloads, legacy .doc/.ppt, invalid files and empty documents are reported in the output but never charged. The run checks your spending limit before each file and stops cleanly when it is reached.

Limits (by design)

  • Max 50 URLs per run; max 50 MB per file; 25-second processing deadline per file; parsing runs in a separate process with a 384 MiB memory limit.
  • Max 500,000 characters per document (configurable via maxChars); longer text is cut and marked truncated.
  • Only public http(s) links; local and private-network addresses are refused.
  • No OCR, no scanned-document support (these formats are inherently digital).
  • Tables are flattened to text rows; no nested structure beyond rows and cells.

Input example

{
"urls": [
"https://raw.githubusercontent.com/python-openxml/python-docx/master/tests/test_files/test.docx",
"https://raw.githubusercontent.com/scanny/python-pptx/master/tests/test_files/test.pptx"
],
"includeTables": true,
"maxChars": 500000
}

Licences

Runtime dependencies:

  • apify — Apache-2.0
  • python-docx — MIT
  • python-pptx — MIT
  • lxml — BSD
  • odfpy — used under its Apache-2.0 option (its licence also offers GPL/LGPL options; we rely on the Apache option)

EPUB parsing is our own src/epub.py built on zipfile + lxml; ebooklib (AGPL-3.0) was removed because AGPL is not acceptable for a public network actor. httpx (BSD) is used only by the local deploy tool, not in the actor runtime.