# Static HTML Accessibility Audit (`dr.skywalker/accessibility-audit`) Actor

Fast first-pass static HTML accessibility checks mapped to WCAG 2.2. Not a replacement for manual testing or a full axe-core run.

- **URL**: https://apify.com/dr.skywalker/accessibility-audit.md
- **Developed by:** [Luqin Wang](https://apify.com/dr.skywalker) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Accessibility Audit

A new Apify actor that runs static HTML accessibility checks and returns prioritized findings mapped to WCAG 2.2.

**This is a fast first-pass scan, not a replacement for manual testing or a full axe-core run.** There is no headless browser, JavaScript execution, or browser accessibility tree. Every page and summary includes `audit_method: "static-html"` and the disclaimer. Findings require confirmation; they are not a conformance verdict.

### Input

```json
{"urls": ["https://example.com", "https://example.com/contact"], "maxPages": 200, "timeoutSec": 20}
```

Alternatively use `urlsText`, one URL per line. Supply exactly one URL source. Exact duplicate URLs are audited once, preserving order. The first `maxPages` unique URLs are selected; unselected URLs appear in `SUMMARY.skipped_urls`.

| Bound | Value |
| --- | --- |
| Supplied URLs / maxPages | 1–200; maxPages defaults to 200 |
| timeoutSec | Integer 5–60; default 20; one total fetch deadline |
| URL length | 8,000 characters; absolute HTTP/HTTPS only |
| Per-page HTML | 8 MiB in bytes; no silent truncation |
| Aggregate admitted HTML | 8 MiB across the run, including replayed snapshots; each page accounts for the larger of wire bytes and decoded UTF-8 bytes |
| DOM admission | 100,000 nodes including elements, comments, text, and declarations; depth 128 |
| Per-page displayed lists | 2,000 issues; 2,000 not-checkable entries |
| Per-page finding bytes | 2 MiB each for issues and not-checkable entries |
| Storage/dataset record | At most 5 MiB serialized JSON |

Excess input is rejected. Oversized responses/DOMs and pages exceeding the remaining aggregate allowance become `unreachable` without a page charge. Fetching receives the remaining allowance before reading a body. Severity and per-check counts cover every observed issue; `issues_truncated` and `not_checkable_truncated` count entries omitted by either list or byte limits. Only one large page result is loaded at a time.

Fetching uses bounded chunks, at most five redirects, and one deadline covering DNS, connections, headers, redirects, and reads. Header deadline cancellation shuts down the active connection socket. It validates public destinations on each hop, pins IP connections, and preserves TLS identity. Credentials, private addresses, environment proxies, non-HTML content, and missing Content-Type are rejected. Identity encoding is requested; compressed responses and premature EOF in declared bodies are rejected. Charset decoding supports UTF-8, ASCII, Latin-1, Windows-1252, UTF-16, and UTF-32. Unsupported labels and broken encodings fall back to UTF-8 replacement decoding. Redirect failures retain the responding URL/status and history. A bad page never stops other URLs.

### Checks

| Check IDs | Severity | WCAG |
| --- | --- | --- |
| image-alt | critical | 1.1.1 |
| image-alt-quality — filename or over 125 chars | warning | 1.1.1 |
| heading-order, heading-empty | high | 1.3.1 |
| heading-multiple-h1 | warning | 1.3.1 |
| form-label | critical | 1.3.1, 3.3.2 |
| button-name | critical | 4.1.2 |
| landmark-missing, landmark-label | medium | 1.3.1 |
| link-name — empty/generic text | medium | 2.4.4 |
| color-contrast — known ratio below 4.5:1 | high | 1.4.3 |
| document-language | medium | 3.1.1 |
| document-title | medium | 2.4.2 |
| document-title-duplicate | warning | 2.4.2 |
| aria-role, aria-attribute | high | 4.1.2 |
| aria-hidden-focus | critical | 4.1.2 |
| duplicate-id | high | 1.3.1, 4.1.2 |
| focus-outline | high | 2.4.7 |
| custom-button-keyboard | high | 2.1.1, 4.1.2 |
| table-headers, table-scope | medium | 1.3.1 |
| viewport-zoom | medium | 1.4.4 |
| media-captions | medium | 1.2.2 |

Form labels must be nonempty matching/wrapping labels, aria-label, or resolvable aria-labelledby. Placeholders and title-only inputs do not satisfy this check. Button names also consider text, image alternatives, title, and native submit/reset defaults. Accessible-name calculation is approximate, capped at 512 characters, and shared label aggregates are cached per ID.

Duplicate IDs are flagged because references become ambiguous. Inline removal of a focus outline is flagged for review; a rendered replacement indicator can satisfy WCAG. Custom buttons must be keyboard reachable; JavaScript activation and unsupported or stylesheet focus styling require browser confirmation and produce `not-checkable` results.

Contrast handles all opaque CSS named colors, hex, RGB/RGBA and HSL/HSLA, determinable ancestor backgrounds, and font color. Missing colors, transparency, images, unsupported CSS, and pages containing stylesheets produce explicit `status: "not-checkable"` entries. Stylesheets are not fetched or interpreted; their cascade can override inline colors.

ARIA validation covers WAI-ARIA 1.2 vocabulary and basic value types, including invalid/abstract roles. It does not validate every role/property relationship, ownership rule, referenced ID, or browser behavior. Newer draft vocabulary may need manual confirmation.

Requested heuristics can flag conforming HTML: decorative empty alt, long alternatives, multiple h1 headings, absent optional landmarks, links with adequate surrounding context, and implicit table headers. Large text can qualify for 3:1 contrast. Presentation/none tables are excluded; other tables are treated as potential data tables. Both video and audio are checked for a captions track as requested; presence does not establish usable captions. Audio-only media commonly needs a transcript under 1.2.1. Remediation hints explain these limits.

### Output

The default dataset contains one item per audited or unreachable page in input order. Each has `url`, `final_url`, `http_status`, `status`, `recordId`, `audit_complete`, `title`, `issues`, severity `counts`, `not_checkable`, `check_results`, omission counts, `redirect_chain`, and `error`.

`check_results` lists every declared check ID with `status: "pass"`, `"fail"`, or `"not-checkable"`, plus complete `issue_count` and `not_checkable_count` totals. A failure takes precedence over unknown cases; a pass means no failure or unknown was observed by the static heuristic, including when there are no applicable elements. Duplicate-title results are resolved across admitted pages before publication. Unreachable or interrupted pages report all checks as not-checkable.

```json
{
  "check_id": "image-alt",
  "severity": "critical",
  "wcag": ["1.1.1"],
  "selector": "html:nth-of-type(1) > body:nth-of-type(1) > img:nth-of-type(1)",
  "message": "Image alt is missing or empty (requested first-pass heuristic).",
  "remediation": "Add a meaningful alternative for informative images. Empty alt can be correct for decorative images; confirm intent manually."
}
```

Issues sort by critical/high/medium/warning, check ID, then document order. Unknown checks are separate from severity counts. Bounded selectors identify relevant HTML paths; deep paths may match multiple elements and need contextual inspection. `.actor/results_schema.json` describes page and summary JSON; `.actor/output_schema.json` exposes Apify output links.

`SUMMARY` in the key-value store contains severity totals, audited/unreachable/published counts, unprocessed/skipped URLs, and duplicate titles. Full normalized title hashes prevent display truncation from causing false matches. Redirect aliases of one final page are not distinct-page duplicates. Counts cover confirmed dataset publications, including the confirmed prefix on publication failure. Duplicate-title groups describe admitted audited pages.

### Billing and recovery

The declared price is **$0.001 per page audited**, event `page-audited`, with no job fee. Unreachable pages are free. Publishing must configure the event from `.actor/pay_per_event.json` in Apify pricing; PPE runs reject missing events.

Finding lists are trimmed to their byte budgets while preserving complete counts, so ordinary output overflow still delivers an audited result. If a defensive page processing/output limit interrupts checks after charging, the accepted charge and source snapshot are retained. That page publishes `status: "audited"`, `audit_complete: false`, an explicit error, and not-checkable check results; later pages continue. Replays reuse that checkpoint without recharging. Storage and billing failures still fail closed.

1. Preserve admitted HTML in bounded `SOURCE-...-PART-...` records and an integrity-checked manifest.
2. Write immutable `BILLING-JOURNAL-...` intent before charging; durably save a separate `-ACCEPTANCE` receipt before running checks.
3. Save `FITTED-PAGE-...` audit checkpoints before charging another page.
4. Add cross-page findings and preserve exact deliverables under `RESULT-...`.
5. Write `DATASET-PUBLICATION-...` intent, publish, verify the exact dataset row, and write `-COMPLETE`.

Cloud result POST retries are disabled. The offline test uses a client stub to verify `max_retries=0`; actual SDK retry behavior has not been integration-tested. Lost responses are reconciled against exact dataset rows. Uncertain charges and missing rows after ambiguous writes are never blindly replayed. Restarts reuse snapshots/checkpoints without refetching or recharging. Changed input or billing mode requires fresh storage.

Budget exhaustion audits/publishes only the accepted prefix. `RUN-PLAN` preserves all URLs; `PRESERVATION-REPORT` identifies unpaid work or reconciliation needs. Billing/storage failures fail the run closed while retaining available source/result records. Storage outages may prevent writing the preservation report too. Reconcile unresolved intents against the platform before deleting a journal or resending an ambiguous POST.

Ordinary non-PPE runs perform audits without charges and record `status: "not-charged"`, `chargedCount: 0` receipts. Local PPE simulation uses `ACTOR_TEST_PAY_PER_EVENT=true`.

### Development

```bash
python3.11 -m venv .venv
.venv/bin/python -m pip install -r requirements-dev.txt
.venv/bin/python -m pytest -q
mkdir -p storage/key_value_stores/default
## Save input JSON as storage/key_value_stores/default/INPUT.json
APIFY_LOCAL_STORAGE_DIR=./storage .venv/bin/python -m src.main
```

Docker targets `apify/actor-python:3.11`. Dependencies are pinned Apify and BeautifulSoup; parsing uses Python's built-in HTML parser. Tests use fixture HTML, mocked network connections, and an SDK-shaped actor without real charges.

This workspace uses Python 3.12.12, cached BeautifulSoup 4.15.0 and pytest 8.4.2 because sandbox networking blocked Python 3.11 and SDK downloads. No Docker build, cloud run, or real SDK integration is claimed. `BUILD_DONE.md` records the initial build verification; `REPAIR_DONE.md` records the repair verification.

### Reference adaptations

Layout, admission bounds, error classes, charge-first journals, source preservation, and publication checkpoints follow the references. Adaptations: one per-page event instead of a job plus item event; requested 8 MiB HTML cap instead of the SEO reference's 5 MiB; URL arrays/pasted lists instead of uploads; sequential page checkpoints instead of conversion batches. Separate immutable receipts prevent delayed intent writes from overwriting acceptance. Ambiguous missing rows require reconciliation instead of automatic missing-suffix replay. Free runs record admission explicitly.

Sources: [WCAG 2.2](https://www.w3.org/TR/WCAG22/), [WAI-ARIA 1.2](https://www.w3.org/TR/wai-aria-1.2/), [CSS named colors](https://www.w3.org/TR/css-color-4/#named-colors), and [Apify SDK 4.0.2 source](https://github.com/apify/apify-sdk-python/tree/v4.0.2/src/apify).

# Actor input Schema

## `urls` (type: `array`):

HTTP or HTTPS URLs to audit. At most 200; exact duplicate URLs are audited once. Use either this field or urlsText.

## `urlsText` (type: `string`):

One HTTP or HTTPS URL per line, at most 200. Use either this field or urls.

## `maxPages` (type: `integer`):

Audit the first N unique supplied URLs. Unselected URLs are listed in SUMMARY.

## `timeoutSec` (type: `integer`):

Total fetch deadline including DNS, connections, redirects, and body reads.

## Actor input object example

```json
{
  "urlsText": "https://example.com",
  "maxPages": 200,
  "timeoutSec": 20
}
```

# Actor output Schema

## `results` (type: `string`):

Each page has audit\_method static-html, URL and final URL, prioritized WCAG issues, severity counts, and explicit not-checkable findings.

## `summary` (type: `string`):

Totals by severity, unreachable and unprocessed pages, and duplicate titles.

## `files` (type: `string`):

Source snapshots, accepted audit checkpoints, BILLING-JOURNAL, DATASET-PUBLICATION, and preservation records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urlsText": "https://example.com"
};

// Run the Actor and wait for it to finish
const run = await client.actor("dr.skywalker/accessibility-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urlsText": "https://example.com" }

# Run the Actor and wait for it to finish
run = client.actor("dr.skywalker/accessibility-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urlsText": "https://example.com"
}' |
apify call dr.skywalker/accessibility-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dr.skywalker/accessibility-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/8bHfFsFikDnkc5HjI/builds/gib32DeWN4SZdV8wi/openapi.json
