HTML Selector Fixture Regression Report avatar

HTML Selector Fixture Regression Report

Pricing

$0.25 / completed report

Go to Apify Store
HTML Selector Fixture Regression Report

HTML Selector Fixture Regression Report

Replay CSS extraction rules on supplied baseline/current HTML. Flag missing nodes, extra matches and lost attributes with exact assertions, hashes, paired evidence, JSON, CSV and HTML.

Pricing

$0.25 / completed report

Rating

0.0

(0)

Developer

Gilad Ronen

Gilad Ronen

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

What does HTML Selector Fixture Regression Report do?

Replay CSS extraction rules against saved HTML fixtures before changing an extractor or delivering another scraping batch. The report shows selectors that lost nodes, matched extra occurrences, lost an attribute, or breached a literal value assertion. Baseline and current evidence sit together with source hashes and JSON pointers so a developer can reproduce each result.

This Actor processes HTML that you supply. Its parser is Cheerio. It does not fetch pages, execute JavaScript, render a browser, repair selectors, or determine whether a live website is correct. Natural text changes are informational unless a declared assertion fails.

When is a selector fixture report useful?

Use the same sanitized fixture pack after a CSS configuration change, before a scheduled customer delivery, or after a source page redesign. Scraping operators can keep known examples of product prices, titles, identifiers, and optional fields as a small regression corpus. An API call from GitHub Actions or n8n can produce a review packet whenever that corpus or extractor configuration changes.

The sample contains four distinct cases: a renamed price class, a missing data-price attribute, duplicate price nodes, and a harmless title change. Three pairs fail declared checks; the title change remains a passing natural delta. Ordinary local tests can perform this job too; this Actor provides a hosted batch interface and consistent evidence downloads.

How to run a CSS selector regression test

  1. Save current HTML and, when available, a baseline for each fixture. Remove confidential content that is unnecessary for testing.
  2. Give each fixture and rule a unique string id. Values such as 001 stay distinct from 1.
  3. Define a selector, extraction mode, explicit whitespace policy, and inclusive minMatches / maxMatches.
  4. For attribute mode, give an exact lower-case HTML attribute name. A missing attribute is a failure, while a present empty attribute remains an empty value.
  5. Optionally supply expectedValues or a literal fixtureIds applicability list. Omitting the list applies the rule to all supplied fixtures.
  6. Run the Actor and inspect each pair's outcome, delta, baseline/current findings, and retained occurrence evidence.

See the Input tab for the full configuration. Example:

{
"fixtures": [{"id":"001", "baselineHtml":"<p class=\"price\">$10</p>", "currentHtml":"<p class=\"price-new\">$10</p>"}],
"rules": [{"id":"price", "selector":".price", "mode":"text", "whitespace":"trim", "minMatches":1, "maxMatches":1}],
"maxRetainedValues":25
}

Exact extraction and comparison policies

Values are individual matched nodes in DOM order. Duplicate values stay separate. Entity references are decoded by HTML parsing. preserve leaves decoded text untouched; trim removes edge JavaScript whitespace; collapse replaces JavaScript whitespace runs with one space and trims the edges. Expected literals use the same policy and must match the full list in DOM order, including its length. These assertions apply independently to every supplied snapshot, including the baseline.

All matched nodes contribute to exact count, missing-attribute, and empty-value totals. Only the first maxRetainedValues occurrences are retained. An omitted or clipped value makes baseline value comparison incomplete; it never hides a count violation. Literal assertions compare complete normalized values rather than a retained prefix. An absent baseline is explicitly unavailable, with delta: not_checked.

Supported selectors include tags, IDs, classes, attribute comparisons, group unions, descendant/child/sibling combinators, :empty, :root, first/last/only child and type pseudo-classes, and bounded :nth-* integer or an+b formulas. Nested or relational pseudo-classes such as :has, :is, :not, text filters such as :contains, jQuery positional extensions, namespaces, and pseudo-elements are unsupported. Invalid or unsupported syntax produces incomplete evidence instead of a clean result.

Capacity limits

One input may contain at most 50 fixtures and 50 rules, 1,000 applicable fixture/rule pairs, 4 MB JSON, and 3 MB combined baseline/current HTML. Each snapshot is at most 500,000 UTF-8 bytes. Weighted work sums each applicable rule's current/baseline HTML bytes and is limited to 30 MB. Parsed snapshots above 20,000 nodes or depth 100 are incomplete and have unknown match counts.

Selectors are at most 300 characters, five groups, forty tokens, and eight combinators; nth formula numeric coefficients are at most 1,000. Retention defaults to 25 and may be 1–100 occurrences per rule/snapshot. Each displayed value retains at most 8,192 UTF-8 bytes, with a 2 MB total retained-value budget across the batch. Original byte lengths and omissions are reported. Split larger batches or narrow fixtureIds. Full JSON must be below 8 MB and every export below 9 MB before a report can be charged.

Output and downloads

FieldMeaning
summaryCounts of assertion failures, incomplete pairs, natural deltas, and unavailable baselines
rowsOne paired evidence row per applicable fixture/rule tuple
current / baselineExact match totals, retained values, assertion states, source pointers, hashes, and findings
outcomefailure, incomplete, or pass for the pair
deltachanged, unchanged, not_checked, or incomplete

The primary dataset contains one complete JSON report. Download OUTPUT for full JSON, review.csv for flat paired evidence, and report.html for a standalone readable report. The HTML download should be opened locally; it escapes all content and contains no scripts or external assets. CSV cells neutralize spreadsheet formula prefixes. Use the evidence view for pair rows and the summary view for the batch overview.

Pricing and recovery

A completed report costs $0.25, including platform usage, through one report-completed event. There is no startup or dataset-item fee. Malformed overall input and insufficient event budget produce no report event. A report with assertion failures or explicitly incomplete supported coverage is still a completed report.

The dataset is the primary paid deliverable. Downloads are constructed and size-checked before publication or charging, then saved as convenience exports. If an export save is interrupted after the dataset charge, some downloads can remain absent until the same run is resurrected. Recovery checks the original input and saved report integrity, keeps one dataset item, and does not charge the report event again. A fresh run is a new billable report. Local tests simulate billing and do not establish platform billing configuration.

Interpretation and support

Static snapshots cannot prove live freshness, access success, pagination, JavaScript rendering, or source truth. Cheerio may recover malformed HTML; this is not an HTML conformance validator. Input and report hashes identify bytes and accidental changes, not signed evidence. Use authorized fixtures and supply the assertions that matter to your extraction contract. For reproducible questions, share a sanitized fixture/rule pair through the Issues tab. The API tab contains programmatic integration examples.