Static HTML Accessibility Audit
Pricing
Pay per usage
Static HTML Accessibility Audit
Fast first-pass static HTML accessibility checks mapped to WCAG 2.2. Not a replacement for manual testing or a full axe-core run.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Luqin Wang
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Accessibility Audit
A new Apify actor that runs static HTML accessibility checks and returns prioritized findings mapped to WCAG 2.2.
This is a fast first-pass scan, not a replacement for manual testing or a full axe-core run. There is no headless browser, JavaScript execution, or browser accessibility tree. Every page and summary includes audit_method: "static-html" and the disclaimer. Findings require confirmation; they are not a conformance verdict.
Input
{"urls": ["https://example.com", "https://example.com/contact"], "maxPages": 200, "timeoutSec": 20}
Alternatively use urlsText, one URL per line. Supply exactly one URL source. Exact duplicate URLs are audited once, preserving order. The first maxPages unique URLs are selected; unselected URLs appear in SUMMARY.skipped_urls.
| Bound | Value |
|---|---|
| Supplied URLs / maxPages | 1–200; maxPages defaults to 200 |
| timeoutSec | Integer 5–60; default 20; one total fetch deadline |
| URL length | 8,000 characters; absolute HTTP/HTTPS only |
| Per-page HTML | 8 MiB in bytes; no silent truncation |
| Aggregate admitted HTML | 8 MiB across the run, including replayed snapshots; each page accounts for the larger of wire bytes and decoded UTF-8 bytes |
| DOM admission | 100,000 nodes including elements, comments, text, and declarations; depth 128 |
| Per-page displayed lists | 2,000 issues; 2,000 not-checkable entries |
| Per-page finding bytes | 2 MiB each for issues and not-checkable entries |
| Storage/dataset record | At most 5 MiB serialized JSON |
Excess input is rejected. Oversized responses/DOMs and pages exceeding the remaining aggregate allowance become unreachable without a page charge. Fetching receives the remaining allowance before reading a body. Severity and per-check counts cover every observed issue; issues_truncated and not_checkable_truncated count entries omitted by either list or byte limits. Only one large page result is loaded at a time.
Fetching uses bounded chunks, at most five redirects, and one deadline covering DNS, connections, headers, redirects, and reads. Header deadline cancellation shuts down the active connection socket. It validates public destinations on each hop, pins IP connections, and preserves TLS identity. Credentials, private addresses, environment proxies, non-HTML content, and missing Content-Type are rejected. Identity encoding is requested; compressed responses and premature EOF in declared bodies are rejected. Charset decoding supports UTF-8, ASCII, Latin-1, Windows-1252, UTF-16, and UTF-32. Unsupported labels and broken encodings fall back to UTF-8 replacement decoding. Redirect failures retain the responding URL/status and history. A bad page never stops other URLs.
Checks
| Check IDs | Severity | WCAG |
|---|---|---|
| image-alt | critical | 1.1.1 |
| image-alt-quality — filename or over 125 chars | warning | 1.1.1 |
| heading-order, heading-empty | high | 1.3.1 |
| heading-multiple-h1 | warning | 1.3.1 |
| form-label | critical | 1.3.1, 3.3.2 |
| button-name | critical | 4.1.2 |
| landmark-missing, landmark-label | medium | 1.3.1 |
| link-name — empty/generic text | medium | 2.4.4 |
| color-contrast — known ratio below 4.5:1 | high | 1.4.3 |
| document-language | medium | 3.1.1 |
| document-title | medium | 2.4.2 |
| document-title-duplicate | warning | 2.4.2 |
| aria-role, aria-attribute | high | 4.1.2 |
| aria-hidden-focus | critical | 4.1.2 |
| duplicate-id | high | 1.3.1, 4.1.2 |
| focus-outline | high | 2.4.7 |
| custom-button-keyboard | high | 2.1.1, 4.1.2 |
| table-headers, table-scope | medium | 1.3.1 |
| viewport-zoom | medium | 1.4.4 |
| media-captions | medium | 1.2.2 |
Form labels must be nonempty matching/wrapping labels, aria-label, or resolvable aria-labelledby. Placeholders and title-only inputs do not satisfy this check. Button names also consider text, image alternatives, title, and native submit/reset defaults. Accessible-name calculation is approximate, capped at 512 characters, and shared label aggregates are cached per ID.
Duplicate IDs are flagged because references become ambiguous. Inline removal of a focus outline is flagged for review; a rendered replacement indicator can satisfy WCAG. Custom buttons must be keyboard reachable; JavaScript activation and unsupported or stylesheet focus styling require browser confirmation and produce not-checkable results.
Contrast handles all opaque CSS named colors, hex, RGB/RGBA and HSL/HSLA, determinable ancestor backgrounds, and font color. Missing colors, transparency, images, unsupported CSS, and pages containing stylesheets produce explicit status: "not-checkable" entries. Stylesheets are not fetched or interpreted; their cascade can override inline colors.
ARIA validation covers WAI-ARIA 1.2 vocabulary and basic value types, including invalid/abstract roles. It does not validate every role/property relationship, ownership rule, referenced ID, or browser behavior. Newer draft vocabulary may need manual confirmation.
Requested heuristics can flag conforming HTML: decorative empty alt, long alternatives, multiple h1 headings, absent optional landmarks, links with adequate surrounding context, and implicit table headers. Large text can qualify for 3:1 contrast. Presentation/none tables are excluded; other tables are treated as potential data tables. Both video and audio are checked for a captions track as requested; presence does not establish usable captions. Audio-only media commonly needs a transcript under 1.2.1. Remediation hints explain these limits.
Output
The default dataset contains one item per audited or unreachable page in input order. Each has url, final_url, http_status, status, recordId, audit_complete, title, issues, severity counts, not_checkable, check_results, omission counts, redirect_chain, and error.
check_results lists every declared check ID with status: "pass", "fail", or "not-checkable", plus complete issue_count and not_checkable_count totals. A failure takes precedence over unknown cases; a pass means no failure or unknown was observed by the static heuristic, including when there are no applicable elements. Duplicate-title results are resolved across admitted pages before publication. Unreachable or interrupted pages report all checks as not-checkable.
{"check_id": "image-alt","severity": "critical","wcag": ["1.1.1"],"selector": "html:nth-of-type(1) > body:nth-of-type(1) > img:nth-of-type(1)","message": "Image alt is missing or empty (requested first-pass heuristic).","remediation": "Add a meaningful alternative for informative images. Empty alt can be correct for decorative images; confirm intent manually."}
Issues sort by critical/high/medium/warning, check ID, then document order. Unknown checks are separate from severity counts. Bounded selectors identify relevant HTML paths; deep paths may match multiple elements and need contextual inspection. .actor/results_schema.json describes page and summary JSON; .actor/output_schema.json exposes Apify output links.
SUMMARY in the key-value store contains severity totals, audited/unreachable/published counts, unprocessed/skipped URLs, and duplicate titles. Full normalized title hashes prevent display truncation from causing false matches. Redirect aliases of one final page are not distinct-page duplicates. Counts cover confirmed dataset publications, including the confirmed prefix on publication failure. Duplicate-title groups describe admitted audited pages.
Billing and recovery
The declared price is $0.001 per page audited, event page-audited, with no job fee. Unreachable pages are free. Publishing must configure the event from .actor/pay_per_event.json in Apify pricing; PPE runs reject missing events.
Finding lists are trimmed to their byte budgets while preserving complete counts, so ordinary output overflow still delivers an audited result. If a defensive page processing/output limit interrupts checks after charging, the accepted charge and source snapshot are retained. That page publishes status: "audited", audit_complete: false, an explicit error, and not-checkable check results; later pages continue. Replays reuse that checkpoint without recharging. Storage and billing failures still fail closed.
- Preserve admitted HTML in bounded
SOURCE-...-PART-...records and an integrity-checked manifest. - Write immutable
BILLING-JOURNAL-...intent before charging; durably save a separate-ACCEPTANCEreceipt before running checks. - Save
FITTED-PAGE-...audit checkpoints before charging another page. - Add cross-page findings and preserve exact deliverables under
RESULT-.... - Write
DATASET-PUBLICATION-...intent, publish, verify the exact dataset row, and write-COMPLETE.
Cloud result POST retries are disabled. The offline test uses a client stub to verify max_retries=0; actual SDK retry behavior has not been integration-tested. Lost responses are reconciled against exact dataset rows. Uncertain charges and missing rows after ambiguous writes are never blindly replayed. Restarts reuse snapshots/checkpoints without refetching or recharging. Changed input or billing mode requires fresh storage.
Budget exhaustion audits/publishes only the accepted prefix. RUN-PLAN preserves all URLs; PRESERVATION-REPORT identifies unpaid work or reconciliation needs. Billing/storage failures fail the run closed while retaining available source/result records. Storage outages may prevent writing the preservation report too. Reconcile unresolved intents against the platform before deleting a journal or resending an ambiguous POST.
Ordinary non-PPE runs perform audits without charges and record status: "not-charged", chargedCount: 0 receipts. Local PPE simulation uses ACTOR_TEST_PAY_PER_EVENT=true.
Development
python3.11 -m venv .venv.venv/bin/python -m pip install -r requirements-dev.txt.venv/bin/python -m pytest -qmkdir -p storage/key_value_stores/default# Save input JSON as storage/key_value_stores/default/INPUT.jsonAPIFY_LOCAL_STORAGE_DIR=./storage .venv/bin/python -m src.main
Docker targets apify/actor-python:3.11. Dependencies are pinned Apify and BeautifulSoup; parsing uses Python's built-in HTML parser. Tests use fixture HTML, mocked network connections, and an SDK-shaped actor without real charges.
This workspace uses Python 3.12.12, cached BeautifulSoup 4.15.0 and pytest 8.4.2 because sandbox networking blocked Python 3.11 and SDK downloads. No Docker build, cloud run, or real SDK integration is claimed. BUILD_DONE.md records the initial build verification; REPAIR_DONE.md records the repair verification.
Reference adaptations
Layout, admission bounds, error classes, charge-first journals, source preservation, and publication checkpoints follow the references. Adaptations: one per-page event instead of a job plus item event; requested 8 MiB HTML cap instead of the SEO reference's 5 MiB; URL arrays/pasted lists instead of uploads; sequential page checkpoints instead of conversion batches. Separate immutable receipts prevent delayed intent writes from overwriting acceptance. Ambiguous missing rows require reconciliation instead of automatic missing-suffix replay. Free runs record admission explicitly.
Sources: WCAG 2.2, WAI-ARIA 1.2, CSS named colors, and Apify SDK 4.0.2 source.