# PDF Accessibility Checker - Find & Test Every PDF on a Site (`that_red_bird/pdf-accessibility-checker`) Actor

How many PDFs does your site host and which fail accessibility? Finds PDFs from a sitemap or page crawl (bounded), validates each with veraPDF for PDF/UA and PDF/A, and returns pass/fail per PDF with the failed rules. Automated rules only, not a full audit.

- **URL**: https://apify.com/that_red_bird/pdf-accessibility-checker.md
- **Developed by:** [mohamed alaya](https://apify.com/that_red_bird) (community)
- **Categories:** Developer tools, Automation, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$3.00 / 1,000 pdf validateds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF Accessibility Checker - Find & Test Every PDF on a Site

Finds the PDFs on a website (sitemap plus a small page crawl) and validates each one with veraPDF against PDF/UA (accessibility) and PDF/A (archiving). You get a pass or fail per PDF with the failed rules, and a site summary.

Built for public bodies preparing for ADA Title II, EU accessibility rules (EAA), universities and agencies that host hundreds of PDFs and do not know how many fail. veraPDF is the open-source reference validator used in the PDF industry.

### Input examples

Audit a site (finds up to 20 PDFs):

```json
{ "startUrls": ["https://www.example.gov/"], "pdfUrls": [], "maxPdfs": 20, "maxPages": 50 }
```

Test a list of PDFs against PDF/UA only:

```json
{ "pdfUrls": ["https://example.org/annual-report.pdf", "https://example.org/forms/apply.pdf"], "flavours": ["ua1"] }
```

Take URLs from a crawler dataset (links ending in .pdf are tested):

```json
{ "datasetId": "WEBSITE_CONTENT_CRAWLER_DATASET_ID", "urlField": "url", "pdfUrls": [], "maxPdfs": 100 }
```

The default input validates one small public sample PDF so you can see the output.

### Input guide

| Field | What it does |
|---|---|
| `startUrls` | Websites to search. The Actor reads robots.txt sitemaps and /sitemap.xml, then crawls same-site pages. |
| `pdfUrls`, `datasetId` + `urlField`, `sitemapUrls` | More ways to supply PDFs or pages. |
| `flavours` | Standards: `ua1` (PDF/UA-1, default), `ua2`, `1b`, `2b`, `3b`, `2u`. Each standard is validated and charged per PDF. |
| `maxPdfs`, `maxPages`, `maxDepth`, `maxPdfMb` | Bounds for discovery and size. |

### Output

One row per PDF, plus a `summary` row at the end.

| Field | Meaning |
|---|---|
| `pdfUrl`, `foundOn`, `sizeKb` | Which PDF, on which page it was linked, size |
| `pdfUaStatus`, `pdfAStatus` | `pass`, `fail`, `error` or `not-checked` |
| `overall` | Follows PDF/UA when selected, otherwise PDF/A |
| `failedRuleCount`, `failedChecks` | PDF/UA rules and checks that failed |
| `topFailedRules` | Up to 8 failed rules with standard, clause, test and plain description |
| `isTagged`, `hasDocumentTitle`, `hasLanguage` | Quick indicators derived from the failed PDF/UA rules |

Sample row (shortened):

```json
{ "type": "pdf", "pdfUrl": "https://archive.ada.gov/aag_covid_statement.pdf", "foundOn": "https://www.ada.gov/resources/2021-08-25-covid-qa/",
  "sizeKb": 150, "pdfUaStatus": "fail", "pdfAStatus": "not-checked", "overall": "fail", "failedRuleCount": 3,
  "topFailedRules": [{ "standard": "ISO 14289-1:2014", "clause": "7.1", "description": "The logical structure ... StructTreeRoot ..." }] }
```

### Use it with

- Website Content Crawler -> this Actor with `datasetId`: validate every PDF link the crawl found.
- A weekly schedule + a dataset diff: track how many PDFs pass PDF/UA over time.
- Your remediation tool or vendor: send the `fail` rows and `topFailedRules` as the work list.

### FAQ

**Is passing veraPDF the same as being accessible?** No. veraPDF checks the machine-verifiable rules of PDF/UA (tagging, language, title, alt text presence, structure). Whether the alt text or reading order is meaningful needs a human.

**Why do most PDFs fail PDF/A?** Only PDFs created as archival PDF/A pass. PDF/A is off by default; add `2b` to `flavours` if you need it.

**Does it find PDFs behind logins or JavaScript menus?** No. Only links in the sitemap and plain HTML pages.

### Honest limits

- Discovery is bounded (Max pages, Max PDFs) and uses plain HTML: PDFs that only appear after JavaScript runs or behind a login are not found. robots.txt rules are not enforced for the page crawl, so keep Max pages modest.
- Automated rules only; this is not a full accessibility audit or a legal opinion.
- Scanned PDFs without text and encrypted PDFs produce errors or fail.
- Each PDF is downloaded fully (limit 30 MB by default) and validated once per standard. The run stops validating shortly before its timeout and reports the rest as errors.

### Price

Pay per event: **$0.003 per PDF per standard validated** (`pdf-validated`). With the default (PDF/UA-1 only) auditing 200 PDFs costs $0.60; adding PDF/A doubles it. PDFs that fail to download or cannot be parsed are not charged; the summary row is free. Validation takes roughly 5 to 10 seconds per PDF per standard, so a 20-PDF audit finishes in a few minutes.

# Actor input Schema

## `startUrls` (type: `array`):

Home pages or section pages of the sites to audit. The Actor reads the sitemap and crawls same-site pages (bounded by Max pages) to find PDF links. Leave empty to only test the PDF URLs below.

## `pdfUrls` (type: `array`):

Direct links to PDF files to validate. The default is one small public sample PDF so you can try the Actor; clear it when you only want to audit websites.

## `datasetId` (type: `string`):

Pick a dataset with a URL column, for example Website Content Crawler output. Links ending in .pdf are validated; other URLs are treated as pages to search for PDFs.

## `urlField` (type: `string`):

Name of the dataset column that holds the URL. Only used when a dataset is selected.

## `sitemapUrls` (type: `array`):

Sitemap XML URLs to read in addition to robots.txt and /sitemap.xml of each website. Sitemap indexes are followed.

## `flavours` (type: `array`):

ua1 is PDF/UA-1 (accessibility), ua2 is PDF/UA-2, the others are PDF/A archiving levels. Each selected standard is validated and charged separately per PDF. The overall pass/fail follows PDF/UA when a PDF/UA standard is selected.

## `maxPdfs` (type: `integer`):

Upper limit of PDFs validated per run, across all sites. Discovery stops once this many are found.

## `maxPages` (type: `integer`):

Upper limit of HTML pages fetched while looking for PDF links, shared across all websites.

## `maxDepth` (type: `integer`):

How many link clicks away from the start pages the crawl may go. 0 only reads the start pages.

## `maxPdfMb` (type: `integer`):

PDFs larger than this are skipped with an error row.

## Actor input object example

```json
{
  "startUrls": [],
  "pdfUrls": [
    "https://raw.githubusercontent.com/mozilla/pdf.js/master/test/pdfs/basicapi.pdf"
  ],
  "urlField": "url",
  "sitemapUrls": [],
  "flavours": [
    "ua1"
  ],
  "maxPdfs": 20,
  "maxPages": 50,
  "maxDepth": 2,
  "maxPdfMb": 30
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `downloadCsv` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `count` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [],
    "pdfUrls": [
        "https://raw.githubusercontent.com/mozilla/pdf.js/master/test/pdfs/basicapi.pdf"
    ],
    "urlField": "url",
    "sitemapUrls": [],
    "flavours": [
        "ua1"
    ],
    "maxPdfs": 20,
    "maxPages": 50,
    "maxDepth": 2,
    "maxPdfMb": 30
};

// Run the Actor and wait for it to finish
const run = await client.actor("that_red_bird/pdf-accessibility-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [],
    "pdfUrls": ["https://raw.githubusercontent.com/mozilla/pdf.js/master/test/pdfs/basicapi.pdf"],
    "urlField": "url",
    "sitemapUrls": [],
    "flavours": ["ua1"],
    "maxPdfs": 20,
    "maxPages": 50,
    "maxDepth": 2,
    "maxPdfMb": 30,
}

# Run the Actor and wait for it to finish
run = client.actor("that_red_bird/pdf-accessibility-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [],
  "pdfUrls": [
    "https://raw.githubusercontent.com/mozilla/pdf.js/master/test/pdfs/basicapi.pdf"
  ],
  "urlField": "url",
  "sitemapUrls": [],
  "flavours": [
    "ua1"
  ],
  "maxPdfs": 20,
  "maxPages": 50,
  "maxDepth": 2,
  "maxPdfMb": 30
}' |
apify call that_red_bird/pdf-accessibility-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,that_red_bird/pdf-accessibility-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KzQcNlxcbmRmkQ0cH/builds/V19oaX6LHZr3c6MJn/openapi.json
