Website PDF Inventory: Tags, Title, Language, Text Layer avatar

Website PDF Inventory: Tags, Title, Language, Text Layer

Pricing

from $5.00 / 1,000 pdf checkeds

Go to Apify Store
Website PDF Inventory: Tags, Title, Language, Text Layer

Website PDF Inventory: Tags, Title, Language, Text Layer

Lists every PDF a website links to with the facts that decide accessibility work: tagged or not, title, language, text layer or scanned, pages, referring pages. Also lists Word files and Google Drive/Docs links. Machine checks only, not an audit. Obeys robots.txt.

Pricing

from $5.00 / 1,000 pdf checkeds

Rating

0.0

(0)

Developer

Jack Valmadre

Jack Valmadre

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Give it a website and get back a list of the PDFs the site links to, with the facts that decide how much accessibility work each one needs: whether it is tagged, whether it has a title and a document language, whether it has a text layer or is a scan that needs OCR, how many pages it has, and which pages on the site link to it. The list comes sorted so the documents most likely to need work, and linked from the most pages, come first.

It's for the people who have to start that work: web teams at cities, counties, school districts, colleges and university departments. Guides to this work, such as those written for the US ADA Title II web rule, often start with the same step: make a list of every PDF on the site. This actor makes that list for you, as a table and a spreadsheet.

It also lists, without checking them, the Word, Excel and PowerPoint files the site links to, and links to documents on other services (Google Drive, Docs, Sheets and Slides, Box, Dropbox, OneDrive, SharePoint, BoardDocs, Issuu, Scribd, Laserfiche), so they don't get forgotten.

What it checks in each PDF

  • Tagged: the PDF has a structure tree (tags), which screen readers rely on to read it in order. A file that has a structure tree but isn't marked as tagged (MarkInfo Marked not set) is flagged too.
  • Text layer: whether the pages contain text (yes), only some do (partial), or none do (none). none on a file with images usually means a scan that needs OCR before it can be tagged. Up to 30 pages are sampled per file, spread across the document.
  • Title: the document title from the PDF's metadata, and whether it looks like a file or scanner name (such as 00206B42FF8B230830145315 or Microsoft Word - report.docx). Also whether the title is set to show in the viewer's title bar (DisplayDocTitle).
  • Language: the document language set on the PDF (for example en-US).
  • Pages, file size, form fields (fillable forms), last modified (from the web server), producer and creator.
  • Optional veraPDF check: switch on runVeraPdf to also validate each PDF with the open-source veraPDF validator against its PDF/UA-1 profile, and get the number of failed rules and their IDs.

Each checked PDF also has What the checks found: a short plain-English list such as "no text layer (likely scanned: needs OCR)", "not tagged (no structure tree)" or "no document language".

How it finds documents

  • It reads the site's sitemap, then follows links from the home page and the pages it finds, staying on the same site (add subdomains with includeSubdomains).
  • It picks up links ending in .pdf, and also document links that don't end in .pdf, such as CivicPlus DocumentCenter and Agenda Center links (/DocumentCenter/View/123): it downloads each of those once and keeps it if it turns out to be a PDF, or lists it as a Word or other file if it is one. On one CivicPlus city site in our tests, all 77 PDFs it read sat behind links like these.
  • PDFs the site links to on other websites are included by default (includeOffsitePdfs).
  • It is polite: it checks robots.txt before every request (including each PDF download and each redirect), sends one request at a time, waits at least 0.5 s between requests to a site (longer if robots.txt asks for a Crawl-delay; a site asking for more than 10 s is skipped), and identifies itself with the user agent token MadrascoDocInventory.

Input

  • Website (url, required): the site's home page, for example https://www.example.gov/.
  • Most web pages to crawl (maxPages, default 300).
  • Most PDFs to check (maxPdfs, default 500): the PDFs downloaded and read. This is what you pay for.
  • Largest PDF to download (maxPdfMegabytes, default 50): larger files are listed as too_large and not read.
  • Include PDFs hosted on other sites (includeOffsitePdfs, default on), Crawl subdomains too (includeSubdomains, default off).
  • Also run veraPDF (runVeraPdf, default off): run with 2048 MB of memory when this is on.
  • Seconds between requests (secondsBetweenRequests, default 0.5).
  • Time budget (maxRunMinutes, default 50): the crawl stops at half of it, and PDFs not read by the end are listed as skipped. Keep it below the run timeout.

Start small (say 100 pages and 50 PDFs) to see what a site looks like, then raise the limits.

Output

  • Dataset, one row per document, in the order to look at them (view "Documents, in the order to look at them").
  • inventory.csv, the same rows as a spreadsheet.
  • OUTPUT, a run summary: pages crawled, PDFs found, PDFs read, how many are not tagged, have no text layer, no title or no language, counts of each row status, and notes on anything that stopped the run early.

An example row (trimmed), from a city website:

{"fixOrder": 1, "status": "ok", "documentType": "PDF",
"pdfUrl": "https://www.cityofmesquite.com/DocumentCenter/View/24726/2023-Truth-in-Taxation",
"whatTheChecksFound": ["no text layer (likely scanned: needs OCR)", "not tagged (no structure tree)",
"title looks like a file or scanner name",
"title not set to show in the window title bar (DisplayDocTitle)", "no document language"],
"pages": 10, "tagged": false, "textLayer": "none", "title": "00206B42FF8B230830145315",
"displayDocTitle": false, "language": null, "hasForm": false, "sizeBytes": 8873085,
"referringPageCount": 22,
"referringPages": ["https://www.cityofmesquite.com/129/Departments", "https://www.cityofmesquite.com/1799/City-Attorney"],
"producer": "KONICA MINOLTA bizhub C360i"}

With runVeraPdf on, rows also carry verapdfFailedRules, verapdfFailedChecks and verapdfRulesFailed (rule IDs such as 7.1-3).

Row status says what happened to each document:

  • ok: downloaded and checked.
  • blocked_by_robots: robots.txt disallows it (for a document link that doesn't end in .pdf, the type is shown as document link (type unknown)). blocked: not downloaded for another stated reason, for example the other site's robots.txt could not be read, or the site asked for a human check.
  • download_failed (with the HTTP status), too_large, not_a_pdf, unreadable (the file could not be opened as a PDF), failed.
  • skipped: not downloaded because the maxPdfs limit, the time budget or your maximum charge was reached (free). For a link that doesn't end in .pdf, the type is shown as document link (type unknown), since it wasn't downloaded.
  • not_checked_other_format: a Word, Excel, PowerPoint, OpenDocument, RTF or EPUB file, listed only.
  • external_host_not_checked: a document on Google Drive, Box, BoardDocs or a similar service, listed only; check it by hand.
  • site_not_crawled: the only row when the site itself couldn't be crawled (see below); the reason is in error.

What it does not do

  • It is not an accessibility audit and not a compliance statement. It reports machine checks of document properties. A PDF that is tagged, titled and has a language set can still be hard to use: reading order, headings, table structure and alt text need a person to review. veraPDF covers only the machine-checkable PDF/UA-1 rules. Nothing here tells you whether a site or document meets the ADA, Section 508, WCAG or any other law or standard.
  • It doesn't fix PDFs. It tells you which ones to look at first.
  • When a limit is reached, the list is incomplete. With maxPages reached, pages beyond it aren't read. With maxPdfs reached, the PDFs linked from the most pages are checked first (links ending in .pdf before other document links); every other document link found still gets a free skipped row: .pdf links as PDFs (counted as pdfsSkipped in the summary), and links that don't end in .pdf, such as DocumentCenter links, as document link (type unknown) (counted as documentLinksNotFetched), because they weren't downloaded to see what they are. Raise maxPdfs to check them.
  • Sites built with JavaScript are covered only partly. It reads the HTML the server sends and doesn't run JavaScript, so links that only appear after scripts run are not seen, and the output can't tell you they were missed. If a site you know has many documents shows few pages crawled or no PDFs, this is the likely reason.
  • Sites that answer with a captcha or bot check are not crawled. A captcha means the site wants a person, so the actor stops asking that site anything: if it happens on the home page, nothing is crawled: the dataset has one site_not_crawled row, the status message gives the reason, and the summary shows stoppedBecause: "human-check"; if it happens part-way, the summary shows stoppedBecause: "human-check" and the page where it happened, and the list covers only what was found before. Ask the site's web team for a file list instead.
  • Sites that refuse to send their robots.txt, or whose robots.txt disallows the home page, are not crawled. If a site answers our request for robots.txt with an error such as HTTP 403, we treat that as "do not fetch". The dataset then has one site_not_crawled row with the reason, the status message says why, and the summary shows stoppedBecause: "robots-refused" (or "robots-disallowed").
  • Documents on file-sharing services are listed, not opened. Google Drive, Box, BoardDocs and similar links get a row but aren't downloaded.
  • No speed guarantee. Time depends on the site: its size, its Crawl-delay and how fast it answers. In our tests on Apify (512 MB), crawling 120 pages and checking up to 80 PDFs took 2 to 4 minutes on two local-government sites, and a school district site asking for a 5-second Crawl-delay took about 10 minutes for 100 pages and 60 PDFs.

Pricing

Pay per event:

  • US$0.005 per PDF checked (event pdf-checked): each PDF downloaded and read (row status ok).
  • US$0.005 per PDF validated with veraPDF, only when runVeraPdf is on (event pdf-verapdf-checked): each PDF veraPDF returned a result for.
  • Actor start: US$0.00005 per GB of run memory, minimum one GB, once per run (event apify-actor-start).

Everything else is free: pages crawled, PDFs that could not be downloaded or read, skipped PDFs, and documents that are only listed (Word and other formats, file-sharing links). For example, checking 80 PDFs costs 80 × US$0.005 = US$0.40, plus US$0.00005 to start at the default 512 MB; with veraPDF on as well, US$0.80 plus US$0.0001 to start at 2048 MB.

If you set a maximum charge per run and it is reached, the actor stops downloading: the remaining document links are listed as skipped, free (links that don't end in .pdf as document link (type unknown)).

Use with AI agents

Input fields: url (string, required; startUrls or urls with one entry also work), maxPages, maxPdfs, maxPdfMegabytes (integers), includeOffsitePdfs, includeSubdomains, runVeraPdf (booleans), secondsBetweenRequests (number), maxRunMinutes (integer). Minimal input: {"url": "https://www.example.gov/", "maxPages": 100, "maxPdfs": 50}. Results: the default dataset (one row per document, already in priority order; status, pdfUrl, whatTheChecksFound, tagged, textLayer, title, language, pages, referringPages), the OUTPUT record (summary) and inventory.csv in the default key-value store.

How we tested it

This page describes build 0.2.8 (0.2.7 with a run status message that also counts the documents it did not download). We ran builds 0.2.4, 0.2.6 and 0.2.7 on Apify on 2026-09-27 on public sites we had not used while building the actor, then checked the output with other tools. The PDF checks are the same code in all three; the later builds changed how pages and links are reported (rows for skipped and robots-blocked links, a row for sites that can't be crawled, and fewer false captcha stops).

  • Sites: the City of Mesquite, Texas (a CivicPlus site), Dakota County, Minnesota, Palo Alto Unified School District, and the University of Wisconsin–Madison Office of the Registrar. On Mesquite (0.2.4), 120 pages gave 80 PDFs (the limit we set), and 77 were read: 39 not tagged, 16 with no text layer, 38 with no title and 43 with no language. A small Mesquite run on 0.2.7 (20 pages, 5 PDFs) read 5 PDFs and listed the other 235 Agenda Center and DocumentCenter links it found as skipped. On Dakota County (0.2.6), 120 pages gave 182 PDF links: 72 of the 80 we allowed were read and the other 102 were listed as skipped; the 72 results were identical to 0.2.4's. On the school district (0.2.6), 100 pages gave 123 links to Google Drive, Docs and Slides and BoardDocs, all listed; on the registrar's site (0.2.4), 24 Box and Google Sheets links.
  • Checked against a second tool (0.2.4 rows): we downloaded 18 of the checked PDFs again (across three of the sites, including scanned, partly scanned, untagged and untitled files) and compared them with poppler (pdfinfo, pdftotext) and qpdf. Pages, tagged (structure tree), title, language and text layer (yes / partial / none) agreed on all 18.
  • veraPDF (0.2.4): on 15 PDFs from the city site at 2048 MB, veraPDF returned a failed-rule count for all 15; the whole run took 78 seconds on Apify.
  • Sites that refuse us: two public sites that answered with a Cloudflare check were not crawled, and the run said so (0.2.4 and, for one of them, 0.2.6 with its site_not_crawled row); a site whose robots.txt answers HTTP 403 got a site_not_crawled row explaining that (0.2.6).

Privacy

The actor reads public web pages and PDFs and keeps only document properties and URLs; it doesn't keep the text of the documents. Downloaded files are deleted once they have been checked. Results are stored only in your own Apify storage.

Support

Open an issue on the actor's Issues tab. This actor and this page were built with AI assistance by Madrasco; a person (the owner) can be reached through the Issues tab.

veraPDF is developed by the veraPDF consortium and the Open Preservation Foundation and is used unmodified under its open-source licence (GPLv3+ / MPLv2+). This actor is independent and not affiliated with or endorsed by them, or by any of the services named above.