DOCX & Excel Extractor — Word & Spreadsheet Text Extraction
Pricing
from $5.00 / 1,000 file extracteds
DOCX & Excel Extractor — Word & Spreadsheet Text Extraction
Extract text, paragraphs, tables, and structured data from Word (.docx) and Excel (.xlsx) files. Supports Google Drive and Dropbox share links. Batch up to 50 files per run. $0.005 per file.
Pricing
from $5.00 / 1,000 file extracteds
Rating
0.0
(0)
Developer
Hojun Lee
Maintained by CommunityActor stats
0
Bookmarked
6
Total users
3
Monthly active users
20 hours ago
Last modified
Categories
Share
DOCX & Excel Extractor — Word, Spreadsheets, Google Drive, Dropbox
Word document extractor — DOCX to text, paragraphs, tables, and structured data via this Word file parser for .docx and Excel .xlsx. Document text extraction from direct URLs, Google Drive share links, and Dropbox URLs. Batch up to 50 files per run. $0.005/file.
⚡ Run in 30 seconds
Paste any Word or Excel URL and click Start:
# Word document (direct URL)https://example.com/report.docx# Google Drive share link (auto-resolved)https://drive.google.com/file/d/1BxiMVs0XRA5nFMdKvBdBZjgmUUqptlbs74OgVE2upms/view# Dropbox share link (auto-resolved)https://www.dropbox.com/s/abc123/data.xlsx?dl=0
Perfect companion to PDF Text Extractor — same workflow, now for Office documents.
Use cases
- LLM ingestion: Extract Word doc content and feed directly into Claude, GPT, or any RAG pipeline
- Spreadsheet → JSON: Convert Excel reports into structured JSON for databases or APIs
- Document processing pipelines: Batch-convert DOCX/XLSX archives to searchable text
- Data extraction: Pull tables from Word documents or rows from Excel sheets
- Content migration: Extract content from legacy Office files for CMS import
Output — Word (.docx)
{"index": 1,"url": "https://example.com/report.docx","file_type": "docx","ok": true,"paragraph_count": 42,"word_count": 1850,"char_count": 11200,"heading_count": 8,"table_count": 2,"full_text": "Executive Summary\n\nThis report covers...","headings": [{"text": "Executive Summary", "level": "Heading 1"}],"paragraphs": [{"text": "This report covers...", "style": "Normal"}],"tables": [{"table_index": 0, "rows": 5, "cols": 3, "data": [["Name", "Value", "Date"], ...]}],"metadata": {"title": "Q3 Report", "author": "John Smith", "created": "2026-01-15T10:00:00"},"fetched_at": "2026-08-08T07:00:00Z"}
Output — Excel (.xlsx)
{"index": 1,"url": "https://example.com/data.xlsx","file_type": "xlsx","ok": true,"sheet_count": 3,"total_rows": 2500,"sheets": [{"name": "Sales","row_count": 1001,"col_count": 8,"headers": ["Date", "Product", "Region", "Revenue", "Units", "Cost", "Margin", "Rep"],"rows": [["2026-01-01", "Widget A", "North", "5400", "36", "2160", "60%", "Alice"]]}],"fetched_at": "2026-08-08T07:00:00Z"}
Input options
| Field | Default | Description |
|---|---|---|
urls | — | List of .docx or .xlsx URLs (up to 50) |
url | — | Single file URL (shortcut when urls is empty) |
outputFormat | full | full includes all text/tables/rows; summary returns metadata + truncated text |
maxRowsPerSheet | 1000 | Max rows per sheet for Excel files (up to 100,000) |
limit | 50 | Max files per run |
Pricing
- $0.001 flat per run start
- First 1 file per run is free
- $0.005 per successfully extracted file (after the free tier)
Example: 1 Word document = free (just run start cost). 10 documents = $0.046
Supported formats
| Format | Extension | Notes |
|---|---|---|
| Word 2007+ | .docx | Full text, paragraphs, headings, tables, metadata |
| Excel 2007+ | .xlsx | All sheets, headers, rows as JSON arrays |
Legacy
.doc(pre-2007 binary format) is not supported. Convert to.docxfirst using Word or LibreOffice.
See also
- PDF Text Extractor — Same workflow for PDFs, including OCR for scanned documents
- Image OCR Extractor — Extract text from images using Tesseract OCR
Keywords: DOCX extractor, Word to text, document parser, text extraction, NLP pipeline, Word file reader, .docx converter, Google Drive, Dropbox, batch document processing
Related actors
- Excel Extractor — .xlsx and .xls to JSON, same URL + Drive/Dropbox pattern for spreadsheet extraction
- PDF Text Extractor — PDF parsing including OCR for scanned documents alongside DOCX extraction
- Image OCR Extractor — Tesseract OCR for image-only content that can't be extracted by DOCX/PDF parsers
Feedback
If this actor powers your document pipeline, a review helps others find it: Leave a review on Apify Store