# Find which PDFs are scans needing OCR

**Use case:** 

Audits a list of documents and keeps only the ones with no text layer, so you know exactly which files have to go through OCR before anything can read them. Documents that could not be reached at all are not charged.

## Input

```json
{
  "fileUrl": "",
  "fileFormat": "auto",
  "pdfUrls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
  ],
  "textFormat": "none",
  "extractTables": false,
  "minTableRows": 3,
  "includePageText": false,
  "maxCharsPerPdf": 200000,
  "keep": "needs_ocr",
  "keepOriginalFields": true,
  "concurrency": 3,
  "requestTimeoutSecs": 60,
  "maxDocumentMb": 50,
  "exportFormats": []
}
```

## Output

```json
{
  "pdfUrl": {
    "label": "PDF URL",
    "format": "link"
  },
  "extractStatus": {
    "label": "Status",
    "format": "text"
  },
  "hasTextLayer": {
    "label": "Has text layer",
    "format": "boolean"
  },
  "ocrRequired": {
    "label": "Needs OCR",
    "format": "boolean"
  },
  "pageCount": {
    "label": "Pages",
    "format": "number"
  },
  "wordCount": {
    "label": "Words",
    "format": "number"
  },
  "tableCount": {
    "label": "Tables",
    "format": "number"
  },
  "pdfTitle": {
    "label": "Title",
    "format": "text"
  },
  "pdfAuthor": {
    "label": "Author",
    "format": "text"
  },
  "pdfCreatedOn": {
    "label": "Created",
    "format": "text"
  },
  "fileSizeKb": {
    "label": "Size (KB)",
    "format": "number"
  },
  "text": {
    "label": "Text",
    "format": "text"
  },
  "tables": {
    "label": "Tables",
    "format": "array"
  },
  "statusDetail": {
    "label": "Detail",
    "format": "text"
  }
}
```

## About this Actor

This example demonstrates how to use [PDF Extractor: Bulk PDF to Text, Tables & Markdown from a CSV](https://apify.com/nerolabs/dataset-pdf-extract.md) with a specific input configuration. Visit the [Actor detail page](https://apify.com/nerolabs/dataset-pdf-extract.md) to learn more, explore other use cases, and run it yourself.


## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
This Task's input is already configured above — use it as-is rather than inventing a new one.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For full API examples (JavaScript, Python, CLI, MCP, OpenAPI), see this Task's Actor page: https://apify.com/nerolabs/dataset-pdf-extract.md

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).
