# Website PDF Inventory: Tags, Title, Language, Text Layer (`madrasco/website-pdf-inventory`) Actor

Lists every PDF a website links to with the facts that decide accessibility work: tagged or not, title, language, text layer or scanned, pages, referring pages. Also lists Word files and Google Drive/Docs links. Machine checks only, not an audit. Obeys robots.txt.

- **URL**: https://apify.com/madrasco/website-pdf-inventory.md
- **Developed by:** [Jack Valmadre](https://apify.com/madrasco) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 pdf checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website PDF Inventory: Tags, Title, Language, Text Layer

Give it a website and get back a list of the PDFs the site links to, with the facts that decide how much accessibility work each one needs: whether it is **tagged**, whether it has a **title** and a **document language**, whether it has a **text layer** or is a scan that needs OCR, how many **pages** it has, and which **pages on the site link to it**. The list comes sorted so the documents most likely to need work, and linked from the most pages, come first.

It's for the people who have to start that work: web teams at cities, counties, school districts, colleges and university departments. Guides to this work, such as those written for the US ADA Title II web rule, often start with the same step: make a list of every PDF on the site. This actor makes that list for you, as a table and a spreadsheet.

It also lists, without checking them, the Word, Excel and PowerPoint files the site links to, and links to documents on other services (Google Drive, Docs, Sheets and Slides, Box, Dropbox, OneDrive, SharePoint, BoardDocs, Issuu, Scribd, Laserfiche), so they don't get forgotten.

### What it checks in each PDF

- **Tagged**: the PDF has a structure tree (tags), which screen readers rely on to read it in order. A file that has a structure tree but isn't marked as tagged (MarkInfo Marked not set) is flagged too.
- **Text layer**: whether the pages contain text (`yes`), only some do (`partial`), or none do (`none`). `none` on a file with images usually means a scan that needs OCR before it can be tagged. Up to 30 pages are sampled per file, spread across the document.
- **Title**: the document title from the PDF's metadata, and whether it looks like a file or scanner name (such as `00206B42FF8B230830145315` or `Microsoft Word - report.docx`). Also whether the title is set to show in the viewer's title bar (DisplayDocTitle).
- **Language**: the document language set on the PDF (for example `en-US`).
- **Pages**, **file size**, **form fields** (fillable forms), **last modified** (from the web server), producer and creator.
- **Optional veraPDF check**: switch on `runVeraPdf` to also validate each PDF with the open-source veraPDF validator against its PDF/UA-1 profile, and get the number of failed rules and their IDs.

Each checked PDF also has **What the checks found**: a short plain-English list such as "no text layer (likely scanned: needs OCR)", "not tagged (no structure tree)" or "no document language".

### How it finds documents

- It reads the site's sitemap, then follows links from the home page and the pages it finds, staying on the same site (add subdomains with `includeSubdomains`).
- It picks up links ending in `.pdf`, and also document links that don't end in `.pdf`, such as CivicPlus DocumentCenter and Agenda Center links (`/DocumentCenter/View/123`): it downloads each of those once and keeps it if it turns out to be a PDF, or lists it as a Word or other file if it is one. On one CivicPlus city site in our tests, all 77 PDFs it read sat behind links like these.
- PDFs the site links to on other websites are included by default (`includeOffsitePdfs`).
- It is polite: it checks robots.txt before every request (including each PDF download and each redirect), sends one request at a time, waits at least 0.5 s between requests to a site (longer if robots.txt asks for a Crawl-delay; a site asking for more than 10 s is skipped), and identifies itself with the user agent token `MadrascoDocInventory`.

### Input

- **Website** (`url`, required): the site's home page, for example `https://www.example.gov/`.
- **Most web pages to crawl** (`maxPages`, default 300).
- **Most PDFs to check** (`maxPdfs`, default 500): the PDFs downloaded and read. This is what you pay for.
- **Largest PDF to download** (`maxPdfMegabytes`, default 50): larger files are listed as `too_large` and not read.
- **Include PDFs hosted on other sites** (`includeOffsitePdfs`, default on), **Crawl subdomains too** (`includeSubdomains`, default off).
- **Also run veraPDF** (`runVeraPdf`, default off): run with 2048 MB of memory when this is on.
- **Seconds between requests** (`secondsBetweenRequests`, default 0.5).
- **Time budget** (`maxRunMinutes`, default 50): the crawl stops at half of it, and PDFs not read by the end are listed as `skipped`. Keep it below the run timeout.

Start small (say 100 pages and 50 PDFs) to see what a site looks like, then raise the limits.

### Output

- **Dataset**, one row per document, in the order to look at them (view "Documents, in the order to look at them").
- **inventory.csv**, the same rows as a spreadsheet.
- **OUTPUT**, a run summary: pages crawled, PDFs found, PDFs read, how many are not tagged, have no text layer, no title or no language, counts of each row status, and notes on anything that stopped the run early.

An example row (trimmed), from a city website:

```json
{"fixOrder": 1, "status": "ok", "documentType": "PDF",
 "pdfUrl": "https://www.cityofmesquite.com/DocumentCenter/View/24726/2023-Truth-in-Taxation",
 "whatTheChecksFound": ["no text layer (likely scanned: needs OCR)", "not tagged (no structure tree)",
   "title looks like a file or scanner name",
   "title not set to show in the window title bar (DisplayDocTitle)", "no document language"],
 "pages": 10, "tagged": false, "textLayer": "none", "title": "00206B42FF8B230830145315",
 "displayDocTitle": false, "language": null, "hasForm": false, "sizeBytes": 8873085,
 "referringPageCount": 22,
 "referringPages": ["https://www.cityofmesquite.com/129/Departments", "https://www.cityofmesquite.com/1799/City-Attorney"],
 "producer": "KONICA MINOLTA bizhub C360i"}
```

With `runVeraPdf` on, rows also carry `verapdfFailedRules`, `verapdfFailedChecks` and `verapdfRulesFailed` (rule IDs such as `7.1-3`).

**Row status** says what happened to each document:

- `ok`: downloaded and checked.
- `blocked_by_robots`: robots.txt disallows it (for a document link that doesn't end in `.pdf`, the type is shown as `document link (type unknown)`). `blocked`: not downloaded for another stated reason, for example the other site's robots.txt could not be read, or the site asked for a human check.
- `download_failed` (with the HTTP status), `too_large`, `not_a_pdf`, `unreadable` (the file could not be opened as a PDF), `failed`.
- `skipped`: not downloaded because the `maxPdfs` limit, the time budget or your maximum charge was reached (free). For a link that doesn't end in `.pdf`, the type is shown as `document link (type unknown)`, since it wasn't downloaded.
- `not_checked_other_format`: a Word, Excel, PowerPoint, OpenDocument, RTF or EPUB file, listed only.
- `external_host_not_checked`: a document on Google Drive, Box, BoardDocs or a similar service, listed only; check it by hand.
- `site_not_crawled`: the only row when the site itself couldn't be crawled (see below); the reason is in `error`.

### What it does not do

- **It is not an accessibility audit and not a compliance statement.** It reports machine checks of document properties. A PDF that is tagged, titled and has a language set can still be hard to use: reading order, headings, table structure and alt text need a person to review. veraPDF covers only the machine-checkable PDF/UA-1 rules. Nothing here tells you whether a site or document meets the ADA, Section 508, WCAG or any other law or standard.
- **It doesn't fix PDFs.** It tells you which ones to look at first.
- **When a limit is reached, the list is incomplete.** With `maxPages` reached, pages beyond it aren't read. With `maxPdfs` reached, the PDFs linked from the most pages are checked first (links ending in `.pdf` before other document links); every other document link found still gets a free `skipped` row: `.pdf` links as PDFs (counted as `pdfsSkipped` in the summary), and links that don't end in `.pdf`, such as DocumentCenter links, as `document link (type unknown)` (counted as `documentLinksNotFetched`), because they weren't downloaded to see what they are. Raise `maxPdfs` to check them.
- **Sites built with JavaScript are covered only partly.** It reads the HTML the server sends and doesn't run JavaScript, so links that only appear after scripts run are not seen, and the output can't tell you they were missed. If a site you know has many documents shows few pages crawled or no PDFs, this is the likely reason.
- **Sites that answer with a captcha or bot check are not crawled.** A captcha means the site wants a person, so the actor stops asking that site anything: if it happens on the home page, nothing is crawled: the dataset has one `site_not_crawled` row, the status message gives the reason, and the summary shows `stoppedBecause: "human-check"`; if it happens part-way, the summary shows `stoppedBecause: "human-check"` and the page where it happened, and the list covers only what was found before. Ask the site's web team for a file list instead.
- **Sites that refuse to send their robots.txt, or whose robots.txt disallows the home page, are not crawled.** If a site answers our request for robots.txt with an error such as HTTP 403, we treat that as "do not fetch". The dataset then has one `site_not_crawled` row with the reason, the status message says why, and the summary shows `stoppedBecause: "robots-refused"` (or `"robots-disallowed"`).
- **Documents on file-sharing services are listed, not opened.** Google Drive, Box, BoardDocs and similar links get a row but aren't downloaded.
- **No speed guarantee.** Time depends on the site: its size, its Crawl-delay and how fast it answers. In our tests on Apify (512 MB), crawling 120 pages and checking up to 80 PDFs took 2 to 4 minutes on two local-government sites, and a school district site asking for a 5-second Crawl-delay took about 10 minutes for 100 pages and 60 PDFs.

### Pricing

Pay per event:

- **US$0.005 per PDF checked** (event `pdf-checked`): each PDF downloaded and read (row status `ok`).
- **US$0.005 per PDF validated with veraPDF**, only when `runVeraPdf` is on (event `pdf-verapdf-checked`): each PDF veraPDF returned a result for.
- **Actor start: US$0.00005 per GB of run memory**, minimum one GB, once per run (event `apify-actor-start`).

Everything else is free: pages crawled, PDFs that could not be downloaded or read, skipped PDFs, and documents that are only listed (Word and other formats, file-sharing links). For example, checking 80 PDFs costs 80 × US$0.005 = US$0.40, plus US$0.00005 to start at the default 512 MB; with veraPDF on as well, US$0.80 plus US$0.0001 to start at 2048 MB.

If you set a maximum charge per run and it is reached, the actor stops downloading: the remaining document links are listed as `skipped`, free (links that don't end in `.pdf` as `document link (type unknown)`).

### Use with AI agents

Input fields: `url` (string, required; `startUrls` or `urls` with one entry also work), `maxPages`, `maxPdfs`, `maxPdfMegabytes` (integers), `includeOffsitePdfs`, `includeSubdomains`, `runVeraPdf` (booleans), `secondsBetweenRequests` (number), `maxRunMinutes` (integer). Minimal input: `{"url": "https://www.example.gov/", "maxPages": 100, "maxPdfs": 50}`. Results: the default dataset (one row per document, already in priority order; `status`, `pdfUrl`, `whatTheChecksFound`, `tagged`, `textLayer`, `title`, `language`, `pages`, `referringPages`), the `OUTPUT` record (summary) and `inventory.csv` in the default key-value store.

### How we tested it

This page describes build 0.2.8 (0.2.7 with a run status message that also counts the documents it did not download). We ran builds 0.2.4, 0.2.6 and 0.2.7 on Apify on 2026-09-27 on public sites we had not used while building the actor, then checked the output with other tools. The PDF checks are the same code in all three; the later builds changed how pages and links are reported (rows for skipped and robots-blocked links, a row for sites that can't be crawled, and fewer false captcha stops).

- **Sites:** the City of Mesquite, Texas (a CivicPlus site), Dakota County, Minnesota, Palo Alto Unified School District, and the University of Wisconsin–Madison Office of the Registrar. On Mesquite (0.2.4), 120 pages gave 80 PDFs (the limit we set), and 77 were read: 39 not tagged, 16 with no text layer, 38 with no title and 43 with no language. A small Mesquite run on 0.2.7 (20 pages, 5 PDFs) read 5 PDFs and listed the other 235 Agenda Center and DocumentCenter links it found as `skipped`. On Dakota County (0.2.6), 120 pages gave 182 PDF links: 72 of the 80 we allowed were read and the other 102 were listed as `skipped`; the 72 results were identical to 0.2.4's. On the school district (0.2.6), 100 pages gave 123 links to Google Drive, Docs and Slides and BoardDocs, all listed; on the registrar's site (0.2.4), 24 Box and Google Sheets links.
- **Checked against a second tool (0.2.4 rows):** we downloaded 18 of the checked PDFs again (across three of the sites, including scanned, partly scanned, untagged and untitled files) and compared them with poppler (`pdfinfo`, `pdftotext`) and `qpdf`. Pages, tagged (structure tree), title, language and text layer (yes / partial / none) agreed on all 18.
- **veraPDF (0.2.4):** on 15 PDFs from the city site at 2048 MB, veraPDF returned a failed-rule count for all 15; the whole run took 78 seconds on Apify.
- **Sites that refuse us:** two public sites that answered with a Cloudflare check were not crawled, and the run said so (0.2.4 and, for one of them, 0.2.6 with its `site_not_crawled` row); a site whose robots.txt answers HTTP 403 got a `site_not_crawled` row explaining that (0.2.6).

### Privacy

The actor reads public web pages and PDFs and keeps only document properties and URLs; it doesn't keep the text of the documents. Downloaded files are deleted once they have been checked. Results are stored only in your own Apify storage.

### Support

Open an issue on the actor's Issues tab. This actor and this page were built with AI assistance by Madrasco; a person (the owner) can be reached through the Issues tab.

veraPDF is developed by the veraPDF consortium and the Open Preservation Foundation and is used unmodified under its open-source licence (GPLv3+ / MPLv2+). This actor is independent and not affiliated with or endorsed by them, or by any of the services named above.

# Actor input Schema

## `url` (type: `string`):

The site's home page, e.g. https://www.example.gov/

## `maxPages` (type: `integer`):

Pages from the sitemap and links on the same site.

## `maxPdfs` (type: `integer`):

PDFs downloaded and read (the priced item). Other-format and file-sharing links are listed free and don't count.

## `maxPdfMegabytes` (type: `integer`):

Larger files are listed with status too\_large and not read.

## `includeOffsitePdfs` (type: `boolean`):

PDFs the site links to on other hosts (robots.txt is checked there too).

## `includeSubdomains` (type: `boolean`):

Also crawl pages on subdomains of the site (e.g. parks.example.gov).

## `runVeraPdf` (type: `boolean`):

Also counts failed PDF/UA-1 machine-checkable rules per PDF with veraPDF. Slower and charged separately; run with 2048 MB of memory.

## `secondsBetweenRequests` (type: `number`):

A larger robots.txt Crawl-delay is always honoured.

## `maxRunMinutes` (type: `integer`):

The crawl stops at half of this, and any PDFs not read by the end are listed as skipped. Keep it below the run timeout.

## Actor input object example

```json
{
  "url": "https://www.bentoncountyor.gov/",
  "maxPages": 15,
  "maxPdfs": 10,
  "maxPdfMegabytes": 50,
  "includeOffsitePdfs": true,
  "includeSubdomains": false,
  "runVeraPdf": false,
  "secondsBetweenRequests": 0.5,
  "maxRunMinutes": 50
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `spreadsheet` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "url": "https://www.bentoncountyor.gov/",
    "maxPages": 15,
    "maxPdfs": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("madrasco/website-pdf-inventory").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "url": "https://www.bentoncountyor.gov/",
    "maxPages": 15,
    "maxPdfs": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("madrasco/website-pdf-inventory").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "url": "https://www.bentoncountyor.gov/",
  "maxPages": 15,
  "maxPdfs": 10
}' |
apify call madrasco/website-pdf-inventory --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,madrasco/website-pdf-inventory"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WG6UuUL7hPtkQWBPC/builds/CcoawwUa7QtPNJ5Ej/openapi.json
