PDF Layout-Preserved Text — Forms & Financials, Columns Intact
Pricing
from $3.00 / 1,000 results
PDF Layout-Preserved Text — Forms & Financials, Columns Intact
Extracts PDF text with column alignment kept intact — between plain text (scrambles forms) and structured tables (not every layout has one). Verified on IRS Form 1040: plain extraction garbled two columns together; this keeps them separate. Flags scanned pages instead of nonsense.
Pricing
from $3.00 / 1,000 results
Rating
0.0
(0)
Developer
alaudin burki
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Plain PDF text extraction reads a page left-to-right, top-to-bottom by character position, with no regard for columns. On a form, invoice, or financial statement, that scrambles unrelated fields into one nonsensical run of text.
Verified on the real target case, IRS Form 1040:
Plain mode: "...also complete spaces below. State ZIP code PresidentialElection Campaign / Check here if you, or your spouse / iffiling jointly, want $3 to go to..."Layout mode: "You Spouse" kept on its own aligned line; "Filing StatusSingle [gap] Head of household (HOH)" kept spatiallyseparate, exactly as printed.
This actor is that one capability — pdfplumber.extract_text(layout=True) —
on its own, chunked and cleaned for downstream use (an LLM prompt, a
document pipeline). It does not claim to be a structured table; that's
a different, harder, lower-confidence problem, covered by the PDF Tables
Extractor.
Where this fits
| Actor | What it does |
|---|---|
| PDF Text Extractor (#43) | Whole-document plain text — fast, but scrambles multi-column layouts |
| PDF Layout-Preserved Text (this) | Keeps visual column/row alignment intact — for forms and financial docs |
| PDF Tables Extractor (#56) | Detects actual tables with a confidence score — a different, stricter claim |
Input
{ "pdfUrls": [{ "url": "https://example.com/form.pdf" }] }
Sample output
{"sourceUrl": "https://www.irs.gov/pub/irs-pdf/f1040.pdf","page": 2,"chunkIndex": 1,"chunksOnPage": 3,"text": " Form 1040 (2025) Page 2\n Tax and 11b Amount from line 11a...","charCount": 2970,"tokensEstimate": 742,"status": "ok"}
Typical uses
- Feeding forms/financial statements to an LLM where column position carries meaning a plain-text scramble would destroy.
- Document pipelines that need a faithful text representation before further processing, without committing to a table-extraction claim.
- Anything where #43's plain text produced garbled output on a multi-column source — this is the direct fix for that specific failure.
Pricing
$3.00 / 1,000 results ($0.003 per chunk/page).
⚠️ Read before you act
- Scanned PDFs (images with no text layer) are flagged, not guessed at.
If a page looks scanned, it's skipped and reported in
QUALITY_REPORTrather than returning empty or garbled text silently. This actor does not perform OCR. - Chunks never split mid-line. A line is one row of the source's visual layout; cutting it in half would destroy the exact alignment this actor exists to preserve. Splits happen at blank-line or line boundaries only.
- Not a table extractor. If you need actual rows/columns with a confidence score, use PDF Tables Extractor instead — this actor deliberately makes no structured-table claim.
FAQ
- Why not always use plain text? Plain text is faster and fine for single-column prose. It actively scrambles forms, invoices, and anything with side-by-side fields — verified on a real IRS form above.
- Can I get one item per page instead of chunked? Yes —
oneItemPerPage: truein the input. - What happens on a scanned PDF? It's detected (both plain and layout extraction come back near-empty) and reported as skipped, not silently returned as empty text.
Related actors
- PDF Text Extractor (#43) — whole-document plain text, faster, no layout preservation.
- PDF Tables Extractor (#56) — actual table detection with a confidence score, for documents where the content really is tabular.