PDF Accessibility Checker - Find & Test Every PDF on a Site avatar

PDF Accessibility Checker - Find & Test Every PDF on a Site

Pricing

$3.00 / 1,000 pdf validateds

Go to Apify Store
PDF Accessibility Checker - Find & Test Every PDF on a Site

PDF Accessibility Checker - Find & Test Every PDF on a Site

How many PDFs does your site host and which fail accessibility? Finds PDFs from a sitemap or page crawl (bounded), validates each with veraPDF for PDF/UA and PDF/A, and returns pass/fail per PDF with the failed rules. Automated rules only, not a full audit.

Pricing

$3.00 / 1,000 pdf validateds

Rating

0.0

(0)

Developer

mohamed alaya

mohamed alaya

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Finds the PDFs on a website (sitemap plus a small page crawl) and validates each one with veraPDF against PDF/UA (accessibility) and PDF/A (archiving). You get a pass or fail per PDF with the failed rules, and a site summary.

Built for public bodies preparing for ADA Title II, EU accessibility rules (EAA), universities and agencies that host hundreds of PDFs and do not know how many fail. veraPDF is the open-source reference validator used in the PDF industry.

Input examples

Audit a site (finds up to 20 PDFs):

{ "startUrls": ["https://www.example.gov/"], "pdfUrls": [], "maxPdfs": 20, "maxPages": 50 }

Test a list of PDFs against PDF/UA only:

{ "pdfUrls": ["https://example.org/annual-report.pdf", "https://example.org/forms/apply.pdf"], "flavours": ["ua1"] }

Take URLs from a crawler dataset (links ending in .pdf are tested):

{ "datasetId": "WEBSITE_CONTENT_CRAWLER_DATASET_ID", "urlField": "url", "pdfUrls": [], "maxPdfs": 100 }

The default input validates one small public sample PDF so you can see the output.

Input guide

FieldWhat it does
startUrlsWebsites to search. The Actor reads robots.txt sitemaps and /sitemap.xml, then crawls same-site pages.
pdfUrls, datasetId + urlField, sitemapUrlsMore ways to supply PDFs or pages.
flavoursStandards: ua1 (PDF/UA-1, default), ua2, 1b, 2b, 3b, 2u. Each standard is validated and charged per PDF.
maxPdfs, maxPages, maxDepth, maxPdfMbBounds for discovery and size.

Output

One row per PDF, plus a summary row at the end.

FieldMeaning
pdfUrl, foundOn, sizeKbWhich PDF, on which page it was linked, size
pdfUaStatus, pdfAStatuspass, fail, error or not-checked
overallFollows PDF/UA when selected, otherwise PDF/A
failedRuleCount, failedChecksPDF/UA rules and checks that failed
topFailedRulesUp to 8 failed rules with standard, clause, test and plain description
isTagged, hasDocumentTitle, hasLanguageQuick indicators derived from the failed PDF/UA rules

Sample row (shortened):

{ "type": "pdf", "pdfUrl": "https://archive.ada.gov/aag_covid_statement.pdf", "foundOn": "https://www.ada.gov/resources/2021-08-25-covid-qa/",
"sizeKb": 150, "pdfUaStatus": "fail", "pdfAStatus": "not-checked", "overall": "fail", "failedRuleCount": 3,
"topFailedRules": [{ "standard": "ISO 14289-1:2014", "clause": "7.1", "description": "The logical structure ... StructTreeRoot ..." }] }

Use it with

  • Website Content Crawler -> this Actor with datasetId: validate every PDF link the crawl found.
  • A weekly schedule + a dataset diff: track how many PDFs pass PDF/UA over time.
  • Your remediation tool or vendor: send the fail rows and topFailedRules as the work list.

FAQ

Is passing veraPDF the same as being accessible? No. veraPDF checks the machine-verifiable rules of PDF/UA (tagging, language, title, alt text presence, structure). Whether the alt text or reading order is meaningful needs a human.

Why do most PDFs fail PDF/A? Only PDFs created as archival PDF/A pass. PDF/A is off by default; add 2b to flavours if you need it.

Does it find PDFs behind logins or JavaScript menus? No. Only links in the sitemap and plain HTML pages.

Honest limits

  • Discovery is bounded (Max pages, Max PDFs) and uses plain HTML: PDFs that only appear after JavaScript runs or behind a login are not found. robots.txt rules are not enforced for the page crawl, so keep Max pages modest.
  • Automated rules only; this is not a full accessibility audit or a legal opinion.
  • Scanned PDFs without text and encrypted PDFs produce errors or fail.
  • Each PDF is downloaded fully (limit 30 MB by default) and validated once per standard. The run stops validating shortly before its timeout and reports the rest as errors.

Price

Pay per event: $0.003 per PDF per standard validated (pdf-validated). With the default (PDF/UA-1 only) auditing 200 PDFs costs $0.60; adding PDF/A doubles it. PDFs that fail to download or cannot be parsed are not charged; the summary row is free. Validation takes roughly 5 to 10 seconds per PDF per standard, so a 20-PDF audit finishes in a few minutes.