PDF Layout-Preserved Text — Forms & Financials, Columns Intact avatar

PDF Layout-Preserved Text — Forms & Financials, Columns Intact

Pricing

from $3.00 / 1,000 results

Go to Apify Store
PDF Layout-Preserved Text — Forms & Financials, Columns Intact

PDF Layout-Preserved Text — Forms & Financials, Columns Intact

Extracts PDF text with column alignment kept intact — between plain text (scrambles forms) and structured tables (not every layout has one). Verified on IRS Form 1040: plain extraction garbled two columns together; this keeps them separate. Flags scanned pages instead of nonsense.

Pricing

from $3.00 / 1,000 results

Rating

0.0

(0)

Developer

alaudin burki

alaudin burki

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Plain PDF text extraction reads a page left-to-right, top-to-bottom by character position, with no regard for columns. On a form, invoice, or financial statement, that scrambles unrelated fields into one nonsensical run of text.

Verified on the real target case, IRS Form 1040:

Plain mode: "...also complete spaces below. State ZIP code Presidential
Election Campaign / Check here if you, or your spouse / if
filing jointly, want $3 to go to..."
Layout mode: "You Spouse" kept on its own aligned line; "Filing Status
Single [gap] Head of household (HOH)" kept spatially
separate, exactly as printed.

This actor is that one capability — pdfplumber.extract_text(layout=True) — on its own, chunked and cleaned for downstream use (an LLM prompt, a document pipeline). It does not claim to be a structured table; that's a different, harder, lower-confidence problem, covered by the PDF Tables Extractor.

Where this fits

ActorWhat it does
PDF Text Extractor (#43)Whole-document plain text — fast, but scrambles multi-column layouts
PDF Layout-Preserved Text (this)Keeps visual column/row alignment intact — for forms and financial docs
PDF Tables Extractor (#56)Detects actual tables with a confidence score — a different, stricter claim

Input

{ "pdfUrls": [{ "url": "https://example.com/form.pdf" }] }

Sample output

{
"sourceUrl": "https://www.irs.gov/pub/irs-pdf/f1040.pdf",
"page": 2,
"chunkIndex": 1,
"chunksOnPage": 3,
"text": " Form 1040 (2025) Page 2\n Tax and 11b Amount from line 11a...",
"charCount": 2970,
"tokensEstimate": 742,
"status": "ok"
}

Typical uses

  • Feeding forms/financial statements to an LLM where column position carries meaning a plain-text scramble would destroy.
  • Document pipelines that need a faithful text representation before further processing, without committing to a table-extraction claim.
  • Anything where #43's plain text produced garbled output on a multi-column source — this is the direct fix for that specific failure.

Pricing

$3.00 / 1,000 results ($0.003 per chunk/page).

⚠️ Read before you act

  • Scanned PDFs (images with no text layer) are flagged, not guessed at. If a page looks scanned, it's skipped and reported in QUALITY_REPORT rather than returning empty or garbled text silently. This actor does not perform OCR.
  • Chunks never split mid-line. A line is one row of the source's visual layout; cutting it in half would destroy the exact alignment this actor exists to preserve. Splits happen at blank-line or line boundaries only.
  • Not a table extractor. If you need actual rows/columns with a confidence score, use PDF Tables Extractor instead — this actor deliberately makes no structured-table claim.

FAQ

  • Why not always use plain text? Plain text is faster and fine for single-column prose. It actively scrambles forms, invoices, and anything with side-by-side fields — verified on a real IRS form above.
  • Can I get one item per page instead of chunked? Yes — oneItemPerPage: true in the input.
  • What happens on a scanned PDF? It's detected (both plain and layout extraction come back near-empty) and reported as skipped, not silently returned as empty text.
  • PDF Text Extractor (#43) — whole-document plain text, faster, no layout preservation.
  • PDF Tables Extractor (#56) — actual table detection with a confidence score, for documents where the content really is tabular.