# Website Accessibility WCAG Audit API (`elab/website-accessibility-wcag-audit-api`) Actor

Batch WCAG accessibility audits with a real headless browser and axe-core: real CSS selectors, screenshot evidence, mobile-viewport auditing, a flat per-violation dataset mode, a batch rollup, and an optional court-ready PDF compliance report. Never charges for a failed/blocked audit.

- **URL**: https://apify.com/elab/website-accessibility-wcag-audit-api.md
- **Developed by:** [Barak Eliov](https://apify.com/elab) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $35.00 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Accessibility WCAG Audit API

Batch WCAG accessibility audits using a real headless browser (Playwright/Chromium) and the real
[axe-core](https://github.com/dequelabs/axe-core) engine (MIT, Deque Systems) - built for the legal/
compliance-audit buyer this category actually serves, not just developers wanting raw JSON.

This is v1 (the scope approved in `SPEC.md`). See "What's not in v1 yet" below.

### What makes this different from other WCAG-audit Actors on Apify

Every incumbent Actor in this category was fetched and live-tested during this build's spec research
(see `SPEC.md` sections 0-3 for run IDs and evidence). Confirmed, working differentiators:

1. **Real, pasteable CSS selectors** - the `selector` field on every violation is axe-core's own
   `target` output, not a raw HTML snippet. Verified during this build by resolving the returned
   selector back to the exact element with Playwright (`tests/axe.test.ts`).
2. **Screenshot evidence** - an element-cropped screenshot per violation, plus one full-page
   screenshot per URL, bundled into the `page-audited` price. No incumbent offers this.
3. **Mobile-viewport auditing** - set `viewport: "mobile"` to catch mobile-relevant WCAG criteria
   (e.g. WCAG 2.2's Target Size rule) that a desktop-only audit cannot see. **Verified with a real
   fixture during this build**: the same page produces a violation at `mobile` viewport and no
   violation at `desktop` viewport (`tests/axe.test.ts`, "EDGE CASE 2"). No incumbent offers any
   viewport control.
4. **Flat per-violation dataset mode** (`outputMode: "flat"` or `"both"`) - one row per individual
   finding, spreadsheet/ticketing-friendly, instead of only a nested array per page.
5. **Batch rollup** across exactly the URLs you supply (`rowType: "summary"` row): total violations by
   impact, and the rules most commonly failing across your batch.
6. **Optional PDF compliance report** (`generatePdfReport: true`) - violations grouped by impact,
   screenshots inlined, WCAG citations, a batch rollup, and failed pages explicitly noted rather than
   hidden. No incumbent Actor offers any document/report output at all.
7. **Mandatory SSRF protection** - every URL's DNS resolution is checked against private/link-local/
   loopback/cloud-metadata IP ranges before any browser navigation. Always on, not a toggle.
8. **Never charges for a failed/blocked audit.** This directly fixes a bug this build's own spec
   research caught live in the category leader: auditing a stable, public W3C demo page returned a
   403-blocked, zero-data result that was still written to its priced dataset. Verified here with a
   real local-server 403 response end-to-end (`tests/edge-case-blocked.test.ts`).

### Input example

```json
{
  "urls": ["https://example.com", "https://info.cern.ch/hypertext/WWW/TheProject.html"],
  "wcagLevel": "wcag21aa",
  "viewport": "desktop",
  "outputMode": "both",
  "generatePdfReport": false
}
```

Full input fields (all have defaults; see `.actor/input_schema.json` for the authoritative list):
`urls`, `wcagLevel`, `viewport`, `viewportWidth`, `viewportHeight`, `includeScreenshots`,
`includeIncomplete`, `includePasses`, `outputMode`, `generatePdfReport`, `waitForSelector`,
`timeoutSecs`, `proxyConfiguration`, `maxConcurrency`.

### Output example

This dataset has three row shapes, distinguished by `rowType`.

**Page row** (one per audited URL):

```json
{
  "rowType": "page",
  "url": "https://info.cern.ch/hypertext/WWW/TheProject.html",
  "finalUrl": "https://info.cern.ch/hypertext/WWW/TheProject.html",
  "status": "success",
  "wcagLevel": "wcag21aa",
  "viewport": "desktop",
  "score": 88,
  "violationCount": 2,
  "incompleteCount": 0,
  "passCount": 0,
  "violationsByImpact": { "critical": 0, "serious": 0, "moderate": 2, "minor": 0 },
  "fullPageScreenshotUrl": "https://api.apify.com/v2/key-value-stores/.../records/fullpage-....png",
  "waitForSelectorTimedOut": false,
  "errorCategory": null,
  "errorMessage": null,
  "auditedAt": "2026-10-01T18:00:00.000Z",
  "engineVersion": "4.13.0"
}
```

**Violation row** (one per individual finding, when `outputMode` is `flat`/`both`):

```json
{
  "rowType": "violation",
  "pageUrl": "https://info.cern.ch/hypertext/WWW/TheProject.html",
  "ruleId": "region",
  "impact": "moderate",
  "status": "violation",
  "description": "Ensure all page content is contained by landmarks",
  "helpText": "All page content must be contained by landmarks",
  "helpUrl": "https://dequeuniversity.com/rules/axe/4.13/region",
  "wcagCriteria": null,
  "selector": "body > h1",
  "htmlSnippet": "<h1>World Wide Web</h1>",
  "screenshotUrl": "https://api.apify.com/v2/key-value-stores/.../records/violation-....png",
  "nodeCount": 1
}
```

Note `wcagCriteria: null` above: axe-core's `region` rule is a "best-practice" rule with no numbered
WCAG success-criterion mapping. This is never guessed or hidden - see "WCAG citation accuracy" below.

**Summary row** (exactly one per run, also saved to the key-value store as `batchSummary`):

```json
{
  "rowType": "summary",
  "totalPagesAudited": 2,
  "totalPagesFailed": 0,
  "totalViolations": 2,
  "violationsByImpact": { "critical": 0, "serious": 0, "moderate": 2, "minor": 0 },
  "topFailingRules": [{ "ruleId": "region", "pageCount": 1 }],
  "pdfReportUrl": null
}
```

A failed URL's page row has `status: "error"`, `errorCategory` set to one of `blocked` / `not_found` /
`timeout` / `render_error` / `invalid_url` / `upstream_changed`, and every audit-specific field left
`null`. **Failed page rows are never charged.**

### Errors (never silent, never charged)

| Category | When |
|---|---|
| `invalid_url` | Malformed URL, non-http(s) URL, or the URL resolves to a private/link-local/loopback/cloud-metadata IP address (SSRF protection, always on) |
| `not_found` | DNS failure, connection refused, 404/410 |
| `blocked` | 401/403/429 or a bot-challenge page - **the exact failure mode this spec exists to fix vs. the category leader** |
| `timeout` | Navigation, the axe-core audit, or screenshot capture exceeded `timeoutSecs` combined |
| `render_error` | Browser crash or an unclassified render failure |
| `upstream_changed` | 5xx upstream server error |

A blocked/not-found/upstream-error page **never reaches axe-core at all** - it is categorized from the
HTTP response status alone, before any audit work happens, and is never charged.

A `waitForSelector` that never appears is **not** an error: the page is still audited and
`waitForSelectorTimedOut: true` is reported (the successful audit is charged normally - a real audit
was in fact delivered).

### WCAG citation accuracy (read this before citing a report)

Not every axe-core rule maps to a single numbered WCAG success criterion. axe-core also ships
"best-practice" rules (e.g. `region`, `landmark-one-main`) that catch real accessibility problems but
are not themselves a WCAG requirement. For those, `wcagCriteria` is explicitly `null` - never guessed,
never silently omitted. If you need every row to cite a numbered criterion, filter on
`wcagCriteria !== null`.

`score` is **this Actor's own derived heuristic** (100 minus a weighted penalty per violation impact -
critical hurts most, minor least). It is not an official W3C/WCAG conformance score; no incumbent
Actor provides one either. Use `violationsByImpact`/`violationCount` for anything that must be
defensible on its own.

`status: "incomplete"` violation rows are axe-core checks that could not be resolved automatically and
need manual human review (e.g. some `color-contrast` cases). They are never counted in `score` or
`violationCount`, and are kept distinguishable from confirmed violations via the `status` field.

### WCAG level and the mobile-viewport differentiator

`wcagLevel` values are cumulative: `wcag21aa` includes WCAG 2.0 A, 2.0 AA, 2.1 A, and 2.1 AA.
`wcag22aa` additionally enables axe-core's `target-size` rule (WCAG 2.2 AA, success criterion 2.5.8,
Target Size Minimum) - the rule actually exercised by this build's mobile-viewport verification test.
**Select `wcagLevel: "wcag22aa"` if you want the mobile-viewport differentiator to have a rule that can
fire on it.** At `wcag21aa` and below, `target-size` is correctly never run (verified in
`tests/axe.test.ts`).

### SSRF protection (always on)

Identical approach and identical module (`src/ssrf.ts`) as `actors/universal-screenshot-api`: before
any browser navigation, the URL's hostname is DNS-resolved and every resolved IP is checked against
RFC1918/link-local/loopback/CGNAT/cloud-metadata ranges (IPv4 and IPv6). A hit is rejected as
`invalid_url` with zero navigation attempted.

**Known limitation:** this is a pre-navigation check, not a pinned connection - a target that changes
DNS or redirects to a private address *after* the check passed would not be caught by this v1 (shared
with most browser-automation tools; closing it fully needs IP-pinned egress, out of scope for v1).

### Screenshots and the PDF compliance report

`includeScreenshots` (default on) captures an element-cropped PNG per violation plus one full-page PNG
per URL, bundled into the `page-audited` price.

`generatePdfReport: true` produces one PDF per run: violations grouped by impact (critical first),
WCAG citations, inlined screenshots, and the batch rollup. Failed pages are listed with their error
category - **never silently dropped from the report.** Charged as a separate `report-generated` event,
and **only when it actually succeeds** - if every page in the batch failed, PDF generation is skipped
and nothing is charged.

### Output modes

- `perPage` - one row per URL, violations summarized in `violationsByImpact`/`violationCount`.
- `flat` - one row per individual finding (`rowType: "violation"`), no page rows.
- `both` (default) - both row types in the same dataset, filter by `rowType`.

Regardless of `outputMode`, the chargeable unit is always one audited page - `outputMode` only changes
dataset shape, never price.

### Limits

- Max 100 URLs per run (internal safety cap; no `maxItems` input field in v1).
- Max 120s `timeoutSecs`, minimum 5s. Bounds navigation, the axe-core audit, and screenshot capture
  combined - exceeding it reports `errorCategory: "timeout"` and is not charged.
- Full-page screenshots are capped at 20,000px height (cost/reliability safety limit, not a Chromium
  limit - same rationale as `actors/universal-screenshot-api`).
- This Actor audits exactly the URLs you list. It does not crawl or discover additional pages.
- `buildSelector` returns axe-core's own selector for plain-document elements (the common case,
  verified to resolve back to the correct element). For shadow-DOM or cross-origin-iframe targets,
  axe-core's `target` is a multi-entry traversal chain that is not valid plain CSS; this Actor falls
  back to a best-effort, human-readable join of those entries, which will not always paste directly
  into DevTools in that specific (rare) case.
- axe-core rules tagged `experimental` or `deprecated` are not run, even if disabled-by-default and
  otherwise WCAG-tagged (`css-orientation-lock`, `label-content-name-mismatch`, `p-as-heading`,
  `table-fake-caption`, `td-has-header`, `aria-roledescription`, `audio-caption`) - a deliberate choice
  favoring defensibility over maximum coverage for a legal/compliance audience. `target-size` (WCAG 2.2
  AA) is the one disabled-by-default rule force-enabled, and only when `wcagLevel` is `wcag22aa`.

### Use with AI agents / MCP

Call the Actor with `urls` as an array (never a comma-separated string). Check `rowType` first, then
`status`/`errorCategory` on page rows. Only `status: "success"` page rows have populated
`violationsByImpact`/`fullPageScreenshotUrl`. Filter `rowType: "violation"` rows by `status ===
"violation"` if you need only confirmed findings, not manual-review items. The single `rowType:
"summary"` row gives you a batch-level answer without having to aggregate page rows yourself.

### FAQ

**Am I charged for a blocked or failed audit?** No. This is the single most important guarantee in
this Actor, verified end-to-end in `tests/edge-case-blocked.test.ts` against a real HTTP 403 response.

**Is the `score` an official WCAG conformance score?** No - see "WCAG citation accuracy" above.

**Does `wcagCriteria: null` mean the finding is wrong or unimportant?** No - it means axe-core
classifies that specific rule as a "best-practice" check rather than a numbered WCAG requirement. The
finding itself is still real and still worth fixing; just don't cite a WCAG clause number for it.

**Can this bypass CAPTCHAs or logins?** No. It audits publicly-rendered pages only, same operating
pattern as `actors/universal-screenshot-api`.

### What's not in v1 yet (explicitly deferred, not silently dropped)

Site-crawl/auto-discovery mode, CI/CD webhook callbacks, scheduled recurring audits, cookie-consent-
gated page auditing, and custom rule-set authoring. Each exists in at least one incumbent but adds real
scope beyond this spec's 8 approved differentiators.

### Local development

```bash
npm install
npx playwright install chromium   # one-time, downloads a browser to the shared Playwright cache
npm run typecheck
npm test
npm run build
apify run --input-file smoke-input.json --purge
```

# Actor input Schema

## `urls` (type: `array`):

Public page URLs to audit. Each URL produces one audit. This Actor audits exactly the pages you list - it does not crawl or discover additional pages on its own. Max 100 URLs per run.

## `wcagLevel` (type: `string`):

WCAG version/level to test against (axe-core rule tags, cumulative - e.g. 'wcag21aa' includes WCAG 2.0 A/AA and WCAG 2.1 A/AA). wcag21aa covers the level most current ADA/DOJ guidance references. wcag22aa is required to exercise WCAG 2.2's mobile-relevant Target Size rule.

## `viewport` (type: `string`):

Viewport to render and audit against. Use 'mobile' to catch mobile-relevant violations (e.g. WCAG 2.2's Target Size rule) that a desktop-only audit cannot see - verified during this build to produce a genuinely different, correct result from 'desktop' on the same page. No incumbent Actor we reviewed offers any viewport control at all.

## `viewportWidth` (type: `integer`):

Used only when Viewport is 'custom'.

## `viewportHeight` (type: `integer`):

Used only when Viewport is 'custom'.

## `includeScreenshots` (type: `boolean`):

Capture an element-cropped screenshot of the first affected node for every violation, plus one full-page screenshot per URL. Bundled into the 'page audited' price - no extra charge.

## `includeIncomplete` (type: `boolean`):

Also include axe-core 'incomplete' checks that need manual human review, alongside confirmed violations. Each row's 'status' field distinguishes 'violation' (confirmed) from 'incomplete' (needs review) - they are never merged indistinguishably.

## `includePasses` (type: `boolean`):

Also return the count of rules that passed. Off by default to keep output focused on actionable findings.

## `outputMode` (type: `string`):

'perPage' returns one row per URL with a violationsByImpact summary. 'flat' returns one row per individual violation (spreadsheet/ticketing-friendly). 'both' returns both row types in the same dataset, tagged by a 'rowType' field. No incumbent Actor we reviewed offers a flat mode at all.

## `generatePdfReport` (type: `boolean`):

Generate one PDF compliance report per run covering all audited URLs (violations grouped by impact, screenshots inlined, WCAG citations, batch rollup; failed pages are noted, not hidden). Charged separately as a 'report generated' event, and only when it actually succeeds - never charged if zero pages audited successfully. No incumbent Actor we reviewed offers any report/document output.

## `waitForSelector` (type: `string`):

Optional CSS selector to wait for before auditing (for SPAs/client-rendered pages). If it never appears within the timeout, the page is still audited and 'waitForSelectorTimedOut' is set true rather than failing.

## `waitUntil` (type: `string`):

When to consider the page ready to audit. 'load' (default, recommended) is reliable even on complex real-world sites with continuous ad/analytics network activity. 'networkidle' waits for network activity to stop - more thorough for slow-loading content, but never fires on many real sites and will time out. 'domcontentloaded' is fastest but may run before late-rendered content appears.

## `timeoutSecs` (type: `integer`):

Max seconds for navigation, the axe-core audit, and screenshot capture combined, before this URL is reported as a timeout error (not charged).

## `proxyConfiguration` (type: `object`):

Apify Proxy configuration for the audit request.

## `maxConcurrency` (type: `integer`):

Maximum pages audited in parallel within one run.

## Actor input object example

```json
{
  "urls": [
    "https://example.com"
  ],
  "wcagLevel": "wcag21aa",
  "viewport": "desktop",
  "viewportWidth": 1280,
  "viewportHeight": 800,
  "includeScreenshots": true,
  "includeIncomplete": true,
  "includePasses": false,
  "outputMode": "both",
  "generatePdfReport": false,
  "waitForSelector": "",
  "waitUntil": "load",
  "timeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "maxConcurrency": 5
}
```

# Actor output Schema

## `results` (type: `string`):

One row per audited URL (rowType "page"), optionally one row per individual finding (rowType "violation", when outputMode is "flat"/"both"), and exactly one batch rollup row (rowType "summary").

## `files` (type: `string`):

Full-page and element-cropped PNG screenshots referenced by fullPageScreenshotUrl/screenshotUrl in dataset rows, the 'batchSummary' record, and the PDF compliance report (when generatePdfReport is true) referenced by pdfReportUrl.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("elab/website-accessibility-wcag-audit-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://example.com"],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("elab/website-accessibility-wcag-audit-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call elab/website-accessibility-wcag-audit-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,elab/website-accessibility-wcag-audit-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pXXq2NbbqTYXz4oQ6/builds/ViQ8fCI8yBi2Xs5fn/openapi.json
