SimilarWeb Website Scraper: Bulk Traffic, WHOIS & Rankings
Pricing
from $1.00 / 1,000 results
SimilarWeb Website Scraper: Bulk Traffic, WHOIS & Rankings
Bulk scrape SimilarWeb traffic analytics for any domain: global, country and category rankings, 3 months of visits, bounce rate, the full traffic source split, top keywords with search volume and CPC, and AI chatbot referrals. Plus WHOIS and 1-to-5-word keyword density. $1 per 1,000 results.
Pricing
from $1.00 / 1,000 results
Rating
5.0
(1)
Developer
Sourabh Kumar
Maintained by CommunityActor stats
5
Bookmarked
269
Total users
31
Monthly active users
13 days ago
Last modified
Categories
Share
SimilarWeb scraper: traffic, rankings, AI referrals, WHOIS, keyword density
Look up any website's traffic, global and country rank, engagement, traffic sources and top keywords. Add WHOIS and on-page keyword density for the same domain in one run.
$1 per 1,000 results. No login, no API key, no SimilarWeb account.
Built for bulk: paste one domain or a few thousand, and get one row back for each.
Why this scraper
- π° $1 per 1,000 results on every paid plan. Both modes bill the same rate, plus a small start fee per run.
- π¦ Bulk lookups in one run. Hand it a whole list of domains instead of querying them one at a time, and export the lot as JSON, CSV or Excel.
- π€ AI traffic, with the actual prompts. Visits from ChatGPT, Perplexity, Gemini and the rest, ranked by share, with three months of history and the real prompt text people typed.
- π Keywords with volume and CPC. Not just the keyword string: monthly search volume, cost per click, and estimated traffic value.
- π The full traffic split. Ten sources, each with its own percentage, adding up to 1.0. Direct, organic and paid search, organic and paid social, referrals, mail, display ads, affiliate, and AI.
- π WHOIS without an API key. Registrar, creation and expiry dates, and name servers for most TLDs.
- π Keyword density on any homepage. One to five word phrase frequency, so you can audit content without a crawler.
- π« You are not billed for misses. A typo, an untracked domain, or a homepage behind a bot wall is delivered as a row with a reason, and charged nothing.
- π Long traffic runs survive restarts. Traffic mode remembers which domains are already done, so a server restart mid-run never re-fetches or double-bills them.
Two modes
π Traffic mode
Everything SimilarWeb publishes for a domain: global, country and category rank, three months of visits, bounce rate, pages per visit, average visit duration, the ten-way traffic source split, top countries, top keywords with volume and CPC, AI chatbot referrals with prompts, and a screenshot.
π Domain analysis mode
WHOIS plus on-page keyword density. Registrar, creation and expiry dates, name servers, and the most frequent one to five word phrases on the site's homepage, with counts and frequencies.
Top use cases
- Competitor research. Compare rank, visits, engagement and traffic mix across a list of rivals in one run.
- AI visibility tracking. Measure what share of a site's traffic now comes from AI assistants, which ones, and what people are asking them.
- SEO and content audits. Pull a competitor's top keywords with search volume and CPC, then check what phrases their homepage actually leans on.
- Lead scoring and CRM enrichment. Add traffic volume, category and domain age to company records so sales can sort by real size.
- Domain investing. Pair expiry dates with traffic trends to spot domains worth buying before they drop.
- Budget allocation. See how much of a market leader's traffic is paid versus organic before you copy their channel mix.
How much does SimilarWeb Scraper cost?
Pay per result. $1 per 1,000 results on every paid Apify plan, $2 per 1,000 on the free plan. Plus a start fee of $0.005 per GB of memory, which is $0.005 at the default setting.
- Apify Free plan, $5 monthly credit: about 2,500 results per month.
- Apify Starter plan, $29 monthly credit: about 29,000 results per month.
Failed and empty results are free. A domain SimilarWeb does not track, a typo, or a homepage that refuses the fetch still arrives as a row explaining why, and it is not billed. Proxy and compute are included, not added to your bill.
Input examples
Traffic mode
{"mode": "traffic","domains": ["washingtonpost.com", "nytimes.com", "theguardian.com"],"maxItems": 100}
Domain analysis mode
{"mode": "domainAnalysis","domains": ["cnn.com", "github.com"],"includeWhois": true,"includeKeywordDensity": true,"keywordDensityNGrams": [1, 2, 3],"keywordDensityTopN": 50}
Bring your own proxies
{"mode": "traffic","domains": ["stripe.com"],"proxyConfiguration": {"useApifyProxy": false,"proxyUrls": ["http://user:pass@proxy.example.com:8080"]}}
When you supply proxy URLs the run uses only those addresses, so you are never billed for our bandwidth instead of yours.
All input fields
| Field | Type | Default | What it does |
|---|---|---|---|
mode | enum | traffic | traffic or domainAnalysis. |
domains | string[] | sample list | Paste domains or full URLs. Schemes, www. and paths are stripped, and accented domains are converted for you. |
maxItems | int | none | Caps how many domains are processed and charged. |
maxConcurrency | int | 2 traffic, 8 domain analysis | Domains handled at once. Traffic caps at 4, domain analysis at 15. |
proxyCountry | string | US | Country the lookup is made from. |
forceFresh | bool | false | Traffic mode. Skips the monthly cache and fetches live. |
includeWhois | bool | true | Domain analysis mode. |
includeKeywordDensity | bool | true | Domain analysis mode. |
keywordDensityNGrams | int[] | [1,2,3,4,5] | Phrase lengths to count, 1 to 8. |
keywordDensityTopN | int | 50 | How many phrases to return per length. |
proxyConfiguration | object | Apify Proxy | Set useApifyProxy: false and pass proxyUrls to use your own. |
About the monthly cache
SimilarWeb refreshes its public numbers once a month, so two lookups of the same domain in the same month return the same figures. Repeat lookups are served from a shared cache in about a second, and marked with _servedFromCache. When SimilarWeb publishes a new month, the cache steps aside on its own. Set forceFresh to always go live.
Output examples
Real rows from live runs, trimmed only where an array repeats.
Traffic mode, one row per domain:
{"domain": "washingtonpost.com","siteName": "washingtonpost.com","title": "The Washington Post - Breaking news and latest headlines, U.S. news, world news, and video","globalRank": 863,"countryRank": { "country": "US", "countryId": 840, "rank": 225 },"categoryRank": { "rank": 19, "category": "News_and_Media" },"category": "news_and_media","totalVisits": 64868363,"estimatedMonthlyVisits": {"2026-06-01": 64811912,"2026-07-01": 71645024,"2026-08-01": 64868363},"bounceRate": 0.563250001913124,"pagesPerVisit": 2.6753718981830037,"avgVisitDuration": 196.71870534656807,"engagementMonth": "2026-08","trafficSourcesRanked": [{ "source": "Direct", "percentage": 0.5206999367487237, "rank": 1 },{ "source": "SearchOrganic", "percentage": 0.22444199153612904, "rank": 2 },{ "source": "Mail", "percentage": 0.0687842533107792, "rank": 3 }],"topCountries": [{ "countryCode": "US", "countryId": 840, "share": 0.8652628431885249 },{ "countryCode": "CA", "countryId": 124, "share": 0.03174341288281115 }],"topKeywords": [{ "keyword": "weather", "volume": 18362370, "cpc": 0.47, "estimatedValue": 614180 },{ "keyword": "washington post", "volume": 328460, "cpc": 2.36, "estimatedValue": 208216 }],"aiTrafficDetails": {"totalAiVisits": 283658.3730000001,"aiReferralShare": 0.004372830762449376,"aiTrafficTier": "<1M","topChatbots": [{ "name": "chatgpt.com", "share": 0.8064 }],"topPrompts": ["What are the key facts and background on the recent major political decision affecting national policy?"],"aiPromptsStatus": { "code": 0, "error": null }},"dataSource": "estimated","largeScreenshot": "https://site-images.similarcdn.com/image?url=washingtonpost.com&t=1&s=1&h=390af69...","snapshotDate": "2026-08-01T00:00:00+00:00","isSmall": false}
Domain analysis mode, one row per domain:
{"domain": "cnn.com","whois": {"registrar": "Nom-iq Ltd. dba COM LAUDE","createdDate": "1993-09-22T04:00:00Z","updatedDate": "2025-04-22T17:03:07Z","expiresDate": "2027-09-21T04:00:00Z","registrantOrg": null,"registrantCountry": null,"nameServers": ["ns-1242.awsdns-27.org","ns-1652.awsdns-14.co.uk","ns-378.awsdns-47.com","ns-587.awsdns-09.net"]},"whoisError": null,"keywordDensity": {"1": [{ "ngram": "subscribers", "count": 97, "frequency": 0.0401656314699793 },{ "ngram": "cnn", "count": 90, "frequency": 0.037267080745341616 },{ "ngram": "video", "count": 62, "frequency": 0.02567287784679089 }]},"keywordDensityError": null,"htmlFetchProxyTier": "apify-datacenter-US"}
Export the dataset as JSON, CSV, Excel, XML or HTML table, or pull it through the API.
Reading a short row
Rows that worked carry no _error key at all. When something goes wrong the key appears with the reason, and that row is not billed:
| Common value | Meaning |
|---|---|
no_data | SimilarWeb tracks no data for this domain. Often a typo or a very small site. |
forbidden_by_similarweb | SimilarWeb refuses this specific domain, whatever address asks. |
partial_keyword_density_failed | WHOIS came back, the homepage did not. |
all_subtasks_failed | Neither half returned anything. |
nothing_requested | Both domain analysis options were switched off, so there was nothing to collect. |
Domain analysis rows also carry whoisError and keywordDensityError separately, so you can see
which half failed and why. A homepage that answers with an interstitial instead of content shows up
there as thin_body.
Limitations
- In
aiPromptsStatus,code: 0means the prompts came back fine. It is SimilarWeb's own status value, not an error code. A non-zero code, ornull, means no prompts were published for that domain. - Chatbots and prompts are capped at 3 each by the public data, the same way countries and keywords are capped at 5.
- The numbers are SimilarWeb's estimates, refreshed monthly, not live analytics. Treat them as a market-size signal, not a billing record.
- Top countries and top keywords are capped at 5 by the public data itself, checked across 40 domains.
- No competitor list, company profile, or rank-change deltas. You get three months of absolute visits instead, which is what most people compute a trend from anyway.
- Use
trafficSourcesRankedfor percentages that add up. The shortertrafficSourcesobject double counts paid traffic inside search and social, so its six numbers do not sum to 1.0. The ranked list does. - Keyword density reads the homepage only, and many large retail and news sites render their content with JavaScript, so those return navigation and menu text rather than article copy.
- Some homepages sit behind a bot wall. Keyword density comes back empty with a reason and is not charged; the WHOIS half still works.
- Some TLDs publish no WHOIS at all, including
.de,.jpand.ru. Most registrars also redact the registrant name and country, so those fields are usuallynull. - The free plan has no second proxy pool. Paid plans fall back to a second pool when the first is blocked; on the free plan, where Apify covers the proxy, the run finishes with what it collected.
- A short run tells you it is short. The status message says how many domains came back and why the rest did not.
Frequently asked questions
How much does this SimilarWeb scraper cost?
$1 per 1,000 results on every paid plan, $2 per 1,000 on the free plan, plus a $0.005 start fee per run at the default memory. Failed and empty lookups are free. No subscription, pause whenever.
Do I need a SimilarWeb account or API key?
No. The scraper reads public data only. You never provide a login, a cookie, or a key.
Is it legal to scrape SimilarWeb?
Scraping publicly available pages is generally allowed in the US and most of the EU, provided you do not collect personal data covered by GDPR or CCPA without a lawful basis. This actor touches public endpoints only, but what you do with the output is your responsibility. Apify's breakdown: Is web scraping legal?
How accurate is the traffic data?
It is SimilarWeb's own published estimate, the same figure their public page shows. Sites with a verified analytics integration are marked dataSource: "ga-verified"; everything else is "estimated". Expect differences from what a site owner sees in their own analytics.
Can I bulk scrape hundreds of domains in one run?
Yes, that is what it is built for. Pass the whole list in domains. Runs of several hundred domains are routine, and repeat lookups inside the same month come from cache in about a second each.
Can I integrate it with other tools?
Yes. Push results to Make, Zapier, Slack, Airbyte, GitHub, Google Sheets and more. Every run is a webhook source. Full list: Apify integrations
Can I run it through the Apify API?
Yes:
curl -X POST "https://api.apify.com/v2/acts/sourabhbgp~similarweb-scraper/runs?token=APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"mode":"traffic","domains":["washingtonpost.com","nytimes.com"]}'
Docs: Apify API reference
Can I use it through an MCP server?
Yes. Apify's MCP server exposes every actor as a tool, so Claude, Cursor and any other MCP client can call this scraper. Setup: Apify MCP docs
Your feedback
Found a bug, want a field, or seeing something odd? Open the Issues tab. Reports reach a human, and fixes usually ship the same week.