Web Scraper — CSS Selector & Data Extractor
Pricing
from $1.20 / 1,000 pages
Web Scraper — CSS Selector & Data Extractor
Web Scraper pulls any CSS selector off any page, returning text, HTML, attributes or match counts — one row per URL, no browser required.
Pricing
from $1.20 / 1,000 pages
Rating
0.0
(0)
Developer
Murat Uzun
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
15 hours ago
Last modified
Categories
Share
What is CSS Selector Extractor (any page, any field)?
CSS Selector Extractor (any page, any field) is an Apify Actor that turns a list of URLs and a list of CSS selectors into structured data — no scraper to write, no browser to configure. Give it a page and tell it "get the text at .price" or "get the href at a.download-link" and it returns one clean row per URL with exactly the values those selectors matched. It is the generic "scrape this bit of this page" tool people reach for before writing a custom Cheerio or Playwright script, and it works on any static HTML page — product pages, blog posts, directories, documentation, internal tools — using plain HTTP requests, not a headless browser. Runs on Apify's schedule, webhook and API infrastructure, so checking the same selectors across a list of pages on a timer is a five-minute setup.
What data does CSS Selector Extractor extract?
For every URL, the Actor runs each configured selector against the page and returns:
| Field | Type | Description |
|---|---|---|
url, finalUrl | string | The input URL and the URL after following redirects |
statusCode, contentType | number, string | HTTP status code and Content-Type header of the fetched page |
fields | object | One key per selector name — a string, an array of strings, or a number for count |
matchCounts | object | One key per selector name — how many DOM nodes that selector actually matched |
htmlSizeBytes | number | Size of the fetched HTML in bytes |
usedProxy | boolean | Whether this URL needed Apify Proxy to load |
error | string | Set when the URL could not be fetched or parsed; every other field but url and scrapedAt stays null |
scrapedAt | string | Timestamp of the fetch (ISO 8601 UTC) |
How to use CSS Selector Extractor
- Fill URLs with one page per line — product pages, articles, directory listings, anything with static HTML.
- Fill Selectors with a JSON array of
{"name", "selector", "type"}objects.typeis one oftext(element text),html(element's outer HTML),attr(an attribute value — add"attribute":"href"or similar) orcount(how many nodes matched, no value extraction). - Leave First match only on to get one value per selector, or turn it off to get every match as an array (capped by Max matches per selector).
- Click Start and export the dataset as JSON, CSV, Excel or HTML from the Output tab — one row per URL, ready to join, chart or feed into another system.
Example input
{"urls": ["https://example.com", "https://apify.com"],"selectors": [{ "name": "title", "selector": "h1", "type": "text" },{ "name": "canonicalUrl", "selector": "link[rel='canonical']", "type": "attr", "attribute": "href" }],"firstMatchOnly": true,"trimWhitespace": true}
Example output
{"url": "https://news.ycombinator.com/","finalUrl": "https://news.ycombinator.com/","statusCode": 200,"contentType": "text/html; charset=utf-8","fields": {"firstTitle": "Show HN: A CSS selector scraper","firstLink": "https://example.com/show-hn","storyCount": 30},"matchCounts": {"firstTitle": 30,"firstLink": 30,"storyCount": 30},"htmlSizeBytes": 34611,"usedProxy": false,"error": null,"scrapedAt": "2026-09-14T01:57:46.921Z"}
A URL that fails to load or parse still produces exactly one row, with error set and every other field null — a bad URL or a typo'd selector never crashes the run or leaves a URL missing from the dataset.
Input parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
urls | array | ["https://example.com"] | One URL per line to fetch and run the selectors against |
selectors | array (JSON) | [{"name":"title","selector":"h1","type":"text"}] | {name, selector, type, attribute?} objects; type is text, html, attr or count; capped at 25 entries |
firstMatchOnly | boolean | true | Return only the first match per selector; off returns every match as an array |
maxMatchesPerSelector | integer | 20 | Most values returned per selector when firstMatchOnly is off (1-500) |
trimWhitespace | boolean | true | Collapse whitespace runs and trim ends of extracted text/attr values |
maxConcurrency | integer | 5 | Parallel page fetches (1-20) |
proxyMode | string | auto | auto tries every URL directly first and only switches to Apify Proxy after a block is detected; always and never force one path |
proxyConfiguration | object | Residential proxy | Which Apify Proxy group proxyMode uses when a request needs the proxy |
Pricing
CSS Selector Extractor uses pay-per-event pricing: $0.001 per result row, plus a negligible actor-start fee. Extracting five selectors from 100 URLs costs about ten cents, whether one selector matches or five do — the price is per URL, not per field. Set Maximum cost per run and the Actor trims the URL list to what the budget covers instead of overspending.
CSS Selector Extractor vs. writing your own scraper
A one-off Node or Python script with Cheerio or BeautifulSoup can do the same extraction, but it means writing fetch/retry/timeout logic, a proxy fallback, whitespace cleanup and a dataset writer from scratch for every new site. This Actor already handles retries, direct-first proxy escalation on a block, and a guaranteed one-row-per-URL output shape — describe the fields once as a small JSON list and get a typed dataset back, schedulable and API-callable without touching code.
Using CSS Selector Extractor with AI agents and MCP
CSS Selector Extractor is pay-per-event with limited permissions — the two requirements for an Actor to be callable through the Apify MCP server at mcp.apify.com. An agent that needs one specific value off a page (a price, a title, a status, a download link) can call this Actor directly with a urls list and a selectors array instead of fetching and parsing HTML itself, and gets back a typed row with the value already extracted, trimmed, and counted.
FAQ
Does this use a headless browser? No. It fetches HTML with plain HTTP requests and parses it with Cheerio, which is far faster and cheaper than a browser but means it cannot extract content that is only added by client-side JavaScript after the initial page load.
What happens if my selector matches nothing? That is not an error — the field's value is null and its matchCounts entry is 0. Only a fetch failure (bad URL, timeout, non-2xx status) or an unparseable response sets the row's error.
What does type: "count" do? It returns how many elements the selector matched as a number, ignoring firstMatchOnly — useful for "how many products/comments/results are on this page" without extracting any text.
Why is my attr field null even though the element exists? Either attribute was left out of that selector's definition, or the matched element simply doesn't have that attribute set — both return null rather than throwing.
Why did a site return blocked or empty content? Some sites answer bot traffic with a challenge page even at HTTP 200. With proxyMode set to auto (the default), the Actor detects that content and retries the rest of the run through Apify Proxy automatically — no proxy spend unless a block is actually detected.
Is this legal to run? Extracting publicly visible page content for your own use is common practice, but you are responsible for complying with the target site's Terms of Service and applicable law (e.g. robots.txt, copyright, personal-data rules) for your specific use case.
What are the limitations? No JavaScript rendering, up to 25 selectors per run, and up to 500 matches returned per selector even when firstMatchOnly is off. A failing URL still produces a row, with the reason in error.
Support and feedback
Found a page shape or selector edge case it should handle better? Open an issue on the Issues tab.
Related Actors
Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:
Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.
Website & domain intelligence
- Email Extractor — Website Contact & Social Finder — e-mails, phones and social profiles per domain
- Tech Stack Detector — Wappalyzer & BuiltWith Alternative — CMS, e-commerce, analytics, pixels and payments per domain
- Domain DNS & Email Security Checker — SPF, DKIM, DMARC, MX provider, registrar and domain age
- Domain Security Audit (TLS, HTTP headers, redirects, robots) — TLS expiry, security headers, redirect chain, robots and llms.txt
- Subdomain Finder (Certificate Transparency) — every subdomain seen in CT logs, with a live DNS check
- Bulk Core Web Vitals & PageSpeed Audit — Lighthouse scores, LCP, CLS, INP and top fixes per URL
- On-Page SEO Audit — title, meta, headings, links, images and schema issues per page
- Sitemap URL Extractor & Change Monitor — every sitemap URL, or new and removed pages between runs
- Wayback Machine Snapshot & Page Change Tracker — how a page changed over time, or every archived snapshot
- Bulk Domain WHOIS & RDAP Lookup — registrar, dates, status and nameservers per domain
- Website Screenshot Generator — full-page or viewport PNG/JPEG screenshots of any URL
Content for AI, LLMs and RAG
- AI Web Search & Read: Google results as clean Markdown — a query turned into clean Markdown from the top search results
- Website to Markdown — Content Crawler for LLM & RAG — any site as clean Markdown per page, no browser
- Article & News Extractor (clean text, author, date, markdown) — clean article text, author, date and Markdown per URL
- Structured Data & JSON-LD Extractor (Schema.org, Open Graph) — Schema.org and Open Graph data from any page
- Google News Scraper (RSS search by keyword, topic, site) — news results by keyword, topic or site
- Press Release Monitor: PR Newswire, BusinessWire, GlobeNewswire — PR Newswire, Business Wire and GlobeNewswire releases
Search, video and social
- YouTube Shorts Scraper — Shorts from channels, hashtags and searches with view counts
- Pinterest Pins Scraper — latest pins of public Pinterest profiles and boards
- YouTube Transcript Scraper — captions and subtitles as text + timed segments, per video or channel
- Google Search Results Scraper — SERP API — organic SERP results per keyword and country
- YouTube Comments Scraper — Comments & Replies — comments and replies with likes, no API key
- YouTube Channel Latest Videos (RSS, no API key) — the latest 15 videos of any channel from RSS
- YouTube Channel Scraper (videos, shorts, live) — a channel's full video, shorts and stream list
- YouTube Search Results Scraper (videos, channels, no API key) — videos, channels and playlists per query
- YouTube Video Details Scraper (views, likes, description, tags) — views, likes, description, tags and chapters per video
- Apple Podcasts Lookup & Episodes Scraper — podcast metadata and episodes from iTunes and RSS
- Bluesky Post, Search & Profile Scraper — posts, profiles, followers and threads from the AT Protocol API
- Telegram Channel Posts Scraper — posts, views and media flags from any public channel
- Substack Publication & Posts Scraper — archive, authors and paywall status per publication
- Google Play Reviews Scraper — reviews, ratings, replies and app versions per app
- App Store Reviews Scraper — iOS reviews and ratings per app and country
- Google Trends Scraper — interest over time, by region, and related queries per keyword
- Google Ads Transparency Scraper — ads any advertiser runs on Google, with format and dates
- Keyword Suggestions Scraper (Google, YouTube, Amazon, Bing) — autocomplete keyword ideas from four search engines
- Bilibili Scraper (Videos, Search, Popular) — Chinese video platform: views, likes, coins, danmaku, uploader
- Mastodon Scraper (Hashtags, Accounts, Trending) — public fediverse posts by hashtag, account or trending
- Meetup Events Scraper (Search by Keyword & City) — upcoming events with RSVPs, fees, venues and groups
- Eventbrite Scraper (Events by Keyword & City) — events by keyword and city with venue, dates and organizer
Leads, jobs and company data
- Career Site Jobs API (Greenhouse, Lever, Ashby, Workday +1) — company domains in, their open jobs out, ATS detected automatically
- Workday Jobs Scraper — jobs with full descriptions from any Workday career site
- Google Maps Scraper — businesses with phone, website, address, rating and coordinates per search
- LinkedIn Jobs Scraper — job titles, companies, locations and full descriptions from LinkedIn job search
- Company 360: full company profile from a domain — one row per domain: contacts, tech, security, hiring and company facts
- Hiring Signals Scraper (Greenhouse, Lever, Ashby, Workable) — open jobs and hiring velocity from 10 public ATS boards
- Y Combinator Companies & Founders Scraper — YC startups by batch, industry and hiring status
- Wikidata Entity & Company Enrichment (facts, IDs, links) — HQ, founders, employees, revenue and social IDs per company
- Email Validator & Verifier — Bulk Email Check — syntax, MX, disposable, role and free-provider checks
- OpenStreetMap POI Extractor (Overpass API: shops, amenities) — shops and amenities by radius, bbox or area
- Stock, Crypto & FX Quotes — one row per symbol from Yahoo, Binance and ECB rates
- Remote Jobs Aggregator (RemoteOK, WWR, Hacker News) — one clean row per remote job, de-duplicated across feeds
- Greenhouse Jobs Scraper — jobs with descriptions from any Greenhouse job board
- Lever Jobs Scraper — jobs with descriptions from any Lever careers page
- Ashby Jobs Scraper — jobs, salaries and descriptions from any Ashby job board
- SmartRecruiters Jobs Scraper — jobs with descriptions from any SmartRecruiters company
- Seek Jobs Scraper (Australia & New Zealand) — Seek job ads with salary, work type and location
- Dice Jobs Scraper — US tech jobs from Dice with salary and remote flag
- AutoScout24 Scraper — European car listings with price, mileage and seller
- Rightmove Scraper — UK property for sale or rent with price and agent
- Wellfound Jobs Scraper (AngelList Startup Jobs) — startup jobs with salary and equity ranges, company size and stage
- Yandex Maps Scraper (Places, Ratings, Phones) — businesses in Russia, Türkiye and the CIS with phones, ratings, hours
- Craigslist Scraper (Listings, Prices, Locations) — listings in any area and category with price, date and coordinates
- JobStreet Scraper (Malaysia, Singapore, PH, ID + JobsDB) — JobStreet and JobsDB jobs in 6 Asian countries with parsed salaries
- InfoJobs Scraper (Spain Jobs, Salaries, Companies) — Spanish jobs with salary range, contract type and full description
- Redfin Scraper (Homes for Sale, Prices, Details) — US homes for sale or sold from any Redfin search, with price and details
- Kleinanzeigen Scraper (Ads, Prices, Locations) — German classifieds with price, VB flag, ZIP, city and seller type
Developer, app and research data
- npm, PyPI & Crates.io Package Health Checker — releases, downloads, deprecation and a health score
- GitHub Repository Health & Activity Report — stars, commits, contributors and risk flags per repo
- VS Code Marketplace Extension Scraper (installs, ratings) — installs, ratings and versions per extension
- Chrome Web Store Extension Scraper (installs, ratings) — users, rating, version and developer per extension
- Google Play Scraper — apps, ratings, installs, developer contact and reviews
- App Store (iOS) App Metadata, Ratings & Top Charts Lookup — ratings, price, version and charts per app
- CrossRef DOI & Citation Metadata Lookup — papers, authors, journals and citation counts
- FDA Recalls & Adverse Events Monitor (openFDA) — food, drug and device recalls from openFDA
- iCal / ICS Calendar Feed to Events Extractor — any public calendar feed as event rows
- Shopify Store Products Scraper — catalog, prices, variants and stock per store
- Hacker News Search & Front Page Scraper — stories, comments and points by query or front page
- GitHub Trending Repositories Scraper — trending repos and developers by language and period
- Stack Overflow & Stack Exchange Q&A Scraper — questions, answers and scores by query, tag or site
- Bulk Image Downloader — download image URLs to storage with size, dimensions and a ZIP
- Google Flights Scraper (Prices, Airlines, Stops) — flight prices, airlines, times, stops and CO2 by route and date
- Google Hotels Scraper (Prices, Ratings, Reviews) — hotel prices per night, stars, rating and reviews by city and dates
- AliExpress Scraper (Search Products & Prices) — AliExpress search results with USD price, discount and rank
- Lazada Scraper (Products, Prices, Sold, Ratings) — Lazada products in 6 countries with price, rating, units sold and seller