Web Scraper — CSS Selector & Data Extractor avatar

Web Scraper — CSS Selector & Data Extractor

Pricing

from $1.20 / 1,000 pages

Go to Apify Store
Web Scraper — CSS Selector & Data Extractor

Web Scraper — CSS Selector & Data Extractor

Web Scraper pulls any CSS selector off any page, returning text, HTML, attributes or match counts — one row per URL, no browser required.

Pricing

from $1.20 / 1,000 pages

Rating

0.0

(0)

Developer

Murat Uzun

Murat Uzun

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

15 hours ago

Last modified

Share

What is CSS Selector Extractor (any page, any field)?

CSS Selector Extractor (any page, any field) is an Apify Actor that turns a list of URLs and a list of CSS selectors into structured data — no scraper to write, no browser to configure. Give it a page and tell it "get the text at .price" or "get the href at a.download-link" and it returns one clean row per URL with exactly the values those selectors matched. It is the generic "scrape this bit of this page" tool people reach for before writing a custom Cheerio or Playwright script, and it works on any static HTML page — product pages, blog posts, directories, documentation, internal tools — using plain HTTP requests, not a headless browser. Runs on Apify's schedule, webhook and API infrastructure, so checking the same selectors across a list of pages on a timer is a five-minute setup.

What data does CSS Selector Extractor extract?

For every URL, the Actor runs each configured selector against the page and returns:

FieldTypeDescription
url, finalUrlstringThe input URL and the URL after following redirects
statusCode, contentTypenumber, stringHTTP status code and Content-Type header of the fetched page
fieldsobjectOne key per selector name — a string, an array of strings, or a number for count
matchCountsobjectOne key per selector name — how many DOM nodes that selector actually matched
htmlSizeBytesnumberSize of the fetched HTML in bytes
usedProxybooleanWhether this URL needed Apify Proxy to load
errorstringSet when the URL could not be fetched or parsed; every other field but url and scrapedAt stays null
scrapedAtstringTimestamp of the fetch (ISO 8601 UTC)

How to use CSS Selector Extractor

  1. Fill URLs with one page per line — product pages, articles, directory listings, anything with static HTML.
  2. Fill Selectors with a JSON array of {"name", "selector", "type"} objects. type is one of text (element text), html (element's outer HTML), attr (an attribute value — add "attribute":"href" or similar) or count (how many nodes matched, no value extraction).
  3. Leave First match only on to get one value per selector, or turn it off to get every match as an array (capped by Max matches per selector).
  4. Click Start and export the dataset as JSON, CSV, Excel or HTML from the Output tab — one row per URL, ready to join, chart or feed into another system.

Example input

{
"urls": ["https://example.com", "https://apify.com"],
"selectors": [
{ "name": "title", "selector": "h1", "type": "text" },
{ "name": "canonicalUrl", "selector": "link[rel='canonical']", "type": "attr", "attribute": "href" }
],
"firstMatchOnly": true,
"trimWhitespace": true
}

Example output

{
"url": "https://news.ycombinator.com/",
"finalUrl": "https://news.ycombinator.com/",
"statusCode": 200,
"contentType": "text/html; charset=utf-8",
"fields": {
"firstTitle": "Show HN: A CSS selector scraper",
"firstLink": "https://example.com/show-hn",
"storyCount": 30
},
"matchCounts": {
"firstTitle": 30,
"firstLink": 30,
"storyCount": 30
},
"htmlSizeBytes": 34611,
"usedProxy": false,
"error": null,
"scrapedAt": "2026-09-14T01:57:46.921Z"
}

A URL that fails to load or parse still produces exactly one row, with error set and every other field null — a bad URL or a typo'd selector never crashes the run or leaves a URL missing from the dataset.

Input parameters

ParameterTypeDefaultDescription
urlsarray["https://example.com"]One URL per line to fetch and run the selectors against
selectorsarray (JSON)[{"name":"title","selector":"h1","type":"text"}]{name, selector, type, attribute?} objects; type is text, html, attr or count; capped at 25 entries
firstMatchOnlybooleantrueReturn only the first match per selector; off returns every match as an array
maxMatchesPerSelectorinteger20Most values returned per selector when firstMatchOnly is off (1-500)
trimWhitespacebooleantrueCollapse whitespace runs and trim ends of extracted text/attr values
maxConcurrencyinteger5Parallel page fetches (1-20)
proxyModestringautoauto tries every URL directly first and only switches to Apify Proxy after a block is detected; always and never force one path
proxyConfigurationobjectResidential proxyWhich Apify Proxy group proxyMode uses when a request needs the proxy

Pricing

CSS Selector Extractor uses pay-per-event pricing: $0.001 per result row, plus a negligible actor-start fee. Extracting five selectors from 100 URLs costs about ten cents, whether one selector matches or five do — the price is per URL, not per field. Set Maximum cost per run and the Actor trims the URL list to what the budget covers instead of overspending.

CSS Selector Extractor vs. writing your own scraper

A one-off Node or Python script with Cheerio or BeautifulSoup can do the same extraction, but it means writing fetch/retry/timeout logic, a proxy fallback, whitespace cleanup and a dataset writer from scratch for every new site. This Actor already handles retries, direct-first proxy escalation on a block, and a guaranteed one-row-per-URL output shape — describe the fields once as a small JSON list and get a typed dataset back, schedulable and API-callable without touching code.

Using CSS Selector Extractor with AI agents and MCP

CSS Selector Extractor is pay-per-event with limited permissions — the two requirements for an Actor to be callable through the Apify MCP server at mcp.apify.com. An agent that needs one specific value off a page (a price, a title, a status, a download link) can call this Actor directly with a urls list and a selectors array instead of fetching and parsing HTML itself, and gets back a typed row with the value already extracted, trimmed, and counted.

FAQ

Does this use a headless browser? No. It fetches HTML with plain HTTP requests and parses it with Cheerio, which is far faster and cheaper than a browser but means it cannot extract content that is only added by client-side JavaScript after the initial page load.

What happens if my selector matches nothing? That is not an error — the field's value is null and its matchCounts entry is 0. Only a fetch failure (bad URL, timeout, non-2xx status) or an unparseable response sets the row's error.

What does type: "count" do? It returns how many elements the selector matched as a number, ignoring firstMatchOnly — useful for "how many products/comments/results are on this page" without extracting any text.

Why is my attr field null even though the element exists? Either attribute was left out of that selector's definition, or the matched element simply doesn't have that attribute set — both return null rather than throwing.

Why did a site return blocked or empty content? Some sites answer bot traffic with a challenge page even at HTTP 200. With proxyMode set to auto (the default), the Actor detects that content and retries the rest of the run through Apify Proxy automatically — no proxy spend unless a block is actually detected.

Is this legal to run? Extracting publicly visible page content for your own use is common practice, but you are responsible for complying with the target site's Terms of Service and applicable law (e.g. robots.txt, copyright, personal-data rules) for your specific use case.

What are the limitations? No JavaScript rendering, up to 25 selectors per run, and up to 500 matches returned per selector even when firstMatchOnly is off. A failing URL still produces a row, with the reason in error.

Support and feedback

Found a page shape or selector edge case it should handle better? Open an issue on the Issues tab.

Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:

Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.

Website & domain intelligence

Content for AI, LLMs and RAG

Search, video and social

Leads, jobs and company data

Developer, app and research data