# PDF Layout-Preserved Text — Forms & Financials, Columns Intact (`alaudinburki/pdf-layout-text`) Actor

Extracts PDF text with column alignment kept intact — between plain text (scrambles forms) and structured tables (not every layout has one). Verified on IRS Form 1040: plain extraction garbled two columns together; this keeps them separate. Flags scanned pages instead of nonsense.

- **URL**: https://apify.com/alaudinburki/pdf-layout-text.md
- **Developed by:** [alaudin burki](https://apify.com/alaudinburki) (community)
- **Categories:** Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF Layout-Preserved Text — Forms & Financials, Columns Intact

Plain PDF text extraction reads a page left-to-right, top-to-bottom by
character position, with no regard for columns. On a form, invoice, or
financial statement, that scrambles unrelated fields into one nonsensical
run of text.

**Verified on the real target case, IRS Form 1040:**

```
Plain mode:   "...also complete spaces below. State ZIP code Presidential
               Election Campaign / Check here if you, or your spouse / if
               filing jointly, want $3 to go to..."

Layout mode:  "You   Spouse" kept on its own aligned line; "Filing Status
               Single [gap] Head of household (HOH)" kept spatially
               separate, exactly as printed.
```

This actor is that one capability — `pdfplumber.extract_text(layout=True)` —
on its own, chunked and cleaned for downstream use (an LLM prompt, a
document pipeline). It does **not** claim to be a structured table; that's
a different, harder, lower-confidence problem, covered by the PDF Tables
Extractor.

### Where this fits

| Actor | What it does |
|---|---|
| PDF Text Extractor (#43) | Whole-document plain text — fast, but scrambles multi-column layouts |
| **PDF Layout-Preserved Text (this)** | Keeps visual column/row alignment intact — for forms and financial docs |
| PDF Tables Extractor (#56) | Detects actual tables with a confidence score — a different, stricter claim |

### Input

```json
{ "pdfUrls": [{ "url": "https://example.com/form.pdf" }] }
```

### Sample output

```json
{
  "sourceUrl": "https://www.irs.gov/pub/irs-pdf/f1040.pdf",
  "page": 2,
  "chunkIndex": 1,
  "chunksOnPage": 3,
  "text": "     Form 1040 (2025)                                Page 2\n     Tax and 11b Amount from line 11a...",
  "charCount": 2970,
  "tokensEstimate": 742,
  "status": "ok"
}
```

### Typical uses

- **Feeding forms/financial statements to an LLM** where column position
  carries meaning a plain-text scramble would destroy.
- **Document pipelines** that need a faithful text representation before
  further processing, without committing to a table-extraction claim.
- **Anything where #43's plain text produced garbled output** on a
  multi-column source — this is the direct fix for that specific failure.

### Pricing

**$3.00 / 1,000 results** (`$0.003` per chunk/page).

### ⚠️ Read before you act

- **Scanned PDFs (images with no text layer) are flagged, not guessed at.**
  If a page looks scanned, it's skipped and reported in `QUALITY_REPORT`
  rather than returning empty or garbled text silently. This actor does
  not perform OCR.
- **Chunks never split mid-line.** A line is one row of the source's visual
  layout; cutting it in half would destroy the exact alignment this actor
  exists to preserve. Splits happen at blank-line or line boundaries only.
- **Not a table extractor.** If you need actual rows/columns with a
  confidence score, use PDF Tables Extractor instead — this actor
  deliberately makes no structured-table claim.

### FAQ

- **Why not always use plain text?** Plain text is faster and fine for
  single-column prose. It actively scrambles forms, invoices, and anything
  with side-by-side fields — verified on a real IRS form above.
- **Can I get one item per page instead of chunked?** Yes —
  `oneItemPerPage: true` in the input.
- **What happens on a scanned PDF?** It's detected (both plain and layout
  extraction come back near-empty) and reported as skipped, not silently
  returned as empty text.

### Related actors

- **PDF Text Extractor (#43)** — whole-document plain text, faster, no
  layout preservation.
- **PDF Tables Extractor (#56)** — actual table detection with a confidence
  score, for documents where the content really is tabular.

# Actor input Schema

## `pdfUrls` (type: `array`):

Public URLs of the PDFs to extract layout-preserved text from.

## `pages` (type: `string`):

Which pages to read, e.g. 1-5,8. Leave empty for every page.

## `maxPagesPerPdf` (type: `integer`):

Upper bound on pages read from each document.

## `maxChunkChars` (type: `integer`):

A page's text is split into chunks no larger than this, always at a line boundary so column alignment is never cut mid-row.

## `oneItemPerPage` (type: `boolean`):

Return one dataset item per page instead of chunking by character count. Off by default because a dense page can exceed typical LLM context limits.

## `maxItems` (type: `integer`):

Hard cap on items returned. You are never charged beyond this.

## Actor input object example

```json
{
  "pdfUrls": [
    {
      "url": "https://www.irs.gov/pub/irs-pdf/f1040.pdf"
    }
  ],
  "maxPagesPerPdf": 50,
  "maxChunkChars": 3000,
  "oneItemPerPage": false,
  "maxItems": 5000
}
```

# Actor output Schema

## `results` (type: `string`):

Layout-preserved text, chunked, one item per chunk (or per page).

## `qualityReport` (type: `string`):

Pages checked, scanned-page detection and problems.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pdfUrls": [
        {
            "url": "https://www.irs.gov/pub/irs-pdf/f1040.pdf"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("alaudinburki/pdf-layout-text").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "pdfUrls": [{ "url": "https://www.irs.gov/pub/irs-pdf/f1040.pdf" }] }

# Run the Actor and wait for it to finish
run = client.actor("alaudinburki/pdf-layout-text").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pdfUrls": [
    {
      "url": "https://www.irs.gov/pub/irs-pdf/f1040.pdf"
    }
  ]
}' |
apify call alaudinburki/pdf-layout-text --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,alaudinburki/pdf-layout-text"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vKSCbuG3ojFXGLwXS/builds/VhQ9mzUYNbnOilYrb/openapi.json
