SimilarWeb Website Scraper: Bulk Traffic, WHOIS & Rankings avatar

SimilarWeb Website Scraper: Bulk Traffic, WHOIS & Rankings

Pricing

from $1.00 / 1,000 results

Go to Apify Store
SimilarWeb Website Scraper: Bulk Traffic, WHOIS & Rankings

SimilarWeb Website Scraper: Bulk Traffic, WHOIS & Rankings

Bulk scrape SimilarWeb traffic analytics for any domain: global, country and category rankings, 3 months of visits, bounce rate, the full traffic source split, top keywords with search volume and CPC, and AI chatbot referrals. Plus WHOIS and 1-to-5-word keyword density. $1 per 1,000 results.

Pricing

from $1.00 / 1,000 results

Rating

5.0

(1)

Developer

Sourabh Kumar

Sourabh Kumar

Maintained by Community

Actor stats

5

Bookmarked

269

Total users

31

Monthly active users

13 days ago

Last modified

Share

SimilarWeb scraper: traffic, rankings, AI referrals, WHOIS, keyword density

Look up any website's traffic, global and country rank, engagement, traffic sources and top keywords. Add WHOIS and on-page keyword density for the same domain in one run.

$1 per 1,000 results. No login, no API key, no SimilarWeb account.

Built for bulk: paste one domain or a few thousand, and get one row back for each.

Why this scraper

  • πŸ’° $1 per 1,000 results on every paid plan. Both modes bill the same rate, plus a small start fee per run.
  • πŸ“¦ Bulk lookups in one run. Hand it a whole list of domains instead of querying them one at a time, and export the lot as JSON, CSV or Excel.
  • πŸ€– AI traffic, with the actual prompts. Visits from ChatGPT, Perplexity, Gemini and the rest, ranked by share, with three months of history and the real prompt text people typed.
  • πŸ”‘ Keywords with volume and CPC. Not just the keyword string: monthly search volume, cost per click, and estimated traffic value.
  • πŸ“Š The full traffic split. Ten sources, each with its own percentage, adding up to 1.0. Direct, organic and paid search, organic and paid social, referrals, mail, display ads, affiliate, and AI.
  • πŸ“ WHOIS without an API key. Registrar, creation and expiry dates, and name servers for most TLDs.
  • πŸ” Keyword density on any homepage. One to five word phrase frequency, so you can audit content without a crawler.
  • 🚫 You are not billed for misses. A typo, an untracked domain, or a homepage behind a bot wall is delivered as a row with a reason, and charged nothing.
  • πŸ”„ Long traffic runs survive restarts. Traffic mode remembers which domains are already done, so a server restart mid-run never re-fetches or double-bills them.

Two modes

πŸ“Š Traffic mode

Everything SimilarWeb publishes for a domain: global, country and category rank, three months of visits, bounce rate, pages per visit, average visit duration, the ten-way traffic source split, top countries, top keywords with volume and CPC, AI chatbot referrals with prompts, and a screenshot.

πŸ”Ž Domain analysis mode

WHOIS plus on-page keyword density. Registrar, creation and expiry dates, name servers, and the most frequent one to five word phrases on the site's homepage, with counts and frequencies.

Top use cases

  • Competitor research. Compare rank, visits, engagement and traffic mix across a list of rivals in one run.
  • AI visibility tracking. Measure what share of a site's traffic now comes from AI assistants, which ones, and what people are asking them.
  • SEO and content audits. Pull a competitor's top keywords with search volume and CPC, then check what phrases their homepage actually leans on.
  • Lead scoring and CRM enrichment. Add traffic volume, category and domain age to company records so sales can sort by real size.
  • Domain investing. Pair expiry dates with traffic trends to spot domains worth buying before they drop.
  • Budget allocation. See how much of a market leader's traffic is paid versus organic before you copy their channel mix.

How much does SimilarWeb Scraper cost?

Pay per result. $1 per 1,000 results on every paid Apify plan, $2 per 1,000 on the free plan. Plus a start fee of $0.005 per GB of memory, which is $0.005 at the default setting.

  • Apify Free plan, $5 monthly credit: about 2,500 results per month.
  • Apify Starter plan, $29 monthly credit: about 29,000 results per month.

Failed and empty results are free. A domain SimilarWeb does not track, a typo, or a homepage that refuses the fetch still arrives as a row explaining why, and it is not billed. Proxy and compute are included, not added to your bill.

Input examples

Traffic mode

{
"mode": "traffic",
"domains": ["washingtonpost.com", "nytimes.com", "theguardian.com"],
"maxItems": 100
}

Domain analysis mode

{
"mode": "domainAnalysis",
"domains": ["cnn.com", "github.com"],
"includeWhois": true,
"includeKeywordDensity": true,
"keywordDensityNGrams": [1, 2, 3],
"keywordDensityTopN": 50
}

Bring your own proxies

{
"mode": "traffic",
"domains": ["stripe.com"],
"proxyConfiguration": {
"useApifyProxy": false,
"proxyUrls": ["http://user:pass@proxy.example.com:8080"]
}
}

When you supply proxy URLs the run uses only those addresses, so you are never billed for our bandwidth instead of yours.

All input fields

FieldTypeDefaultWhat it does
modeenumtraffictraffic or domainAnalysis.
domainsstring[]sample listPaste domains or full URLs. Schemes, www. and paths are stripped, and accented domains are converted for you.
maxItemsintnoneCaps how many domains are processed and charged.
maxConcurrencyint2 traffic, 8 domain analysisDomains handled at once. Traffic caps at 4, domain analysis at 15.
proxyCountrystringUSCountry the lookup is made from.
forceFreshboolfalseTraffic mode. Skips the monthly cache and fetches live.
includeWhoisbooltrueDomain analysis mode.
includeKeywordDensitybooltrueDomain analysis mode.
keywordDensityNGramsint[][1,2,3,4,5]Phrase lengths to count, 1 to 8.
keywordDensityTopNint50How many phrases to return per length.
proxyConfigurationobjectApify ProxySet useApifyProxy: false and pass proxyUrls to use your own.

About the monthly cache

SimilarWeb refreshes its public numbers once a month, so two lookups of the same domain in the same month return the same figures. Repeat lookups are served from a shared cache in about a second, and marked with _servedFromCache. When SimilarWeb publishes a new month, the cache steps aside on its own. Set forceFresh to always go live.

Output examples

Real rows from live runs, trimmed only where an array repeats.

Traffic mode, one row per domain:

{
"domain": "washingtonpost.com",
"siteName": "washingtonpost.com",
"title": "The Washington Post - Breaking news and latest headlines, U.S. news, world news, and video",
"globalRank": 863,
"countryRank": { "country": "US", "countryId": 840, "rank": 225 },
"categoryRank": { "rank": 19, "category": "News_and_Media" },
"category": "news_and_media",
"totalVisits": 64868363,
"estimatedMonthlyVisits": {
"2026-06-01": 64811912,
"2026-07-01": 71645024,
"2026-08-01": 64868363
},
"bounceRate": 0.563250001913124,
"pagesPerVisit": 2.6753718981830037,
"avgVisitDuration": 196.71870534656807,
"engagementMonth": "2026-08",
"trafficSourcesRanked": [
{ "source": "Direct", "percentage": 0.5206999367487237, "rank": 1 },
{ "source": "SearchOrganic", "percentage": 0.22444199153612904, "rank": 2 },
{ "source": "Mail", "percentage": 0.0687842533107792, "rank": 3 }
],
"topCountries": [
{ "countryCode": "US", "countryId": 840, "share": 0.8652628431885249 },
{ "countryCode": "CA", "countryId": 124, "share": 0.03174341288281115 }
],
"topKeywords": [
{ "keyword": "weather", "volume": 18362370, "cpc": 0.47, "estimatedValue": 614180 },
{ "keyword": "washington post", "volume": 328460, "cpc": 2.36, "estimatedValue": 208216 }
],
"aiTrafficDetails": {
"totalAiVisits": 283658.3730000001,
"aiReferralShare": 0.004372830762449376,
"aiTrafficTier": "<1M",
"topChatbots": [{ "name": "chatgpt.com", "share": 0.8064 }],
"topPrompts": [
"What are the key facts and background on the recent major political decision affecting national policy?"
],
"aiPromptsStatus": { "code": 0, "error": null }
},
"dataSource": "estimated",
"largeScreenshot": "https://site-images.similarcdn.com/image?url=washingtonpost.com&t=1&s=1&h=390af69...",
"snapshotDate": "2026-08-01T00:00:00+00:00",
"isSmall": false
}

Domain analysis mode, one row per domain:

{
"domain": "cnn.com",
"whois": {
"registrar": "Nom-iq Ltd. dba COM LAUDE",
"createdDate": "1993-09-22T04:00:00Z",
"updatedDate": "2025-04-22T17:03:07Z",
"expiresDate": "2027-09-21T04:00:00Z",
"registrantOrg": null,
"registrantCountry": null,
"nameServers": [
"ns-1242.awsdns-27.org",
"ns-1652.awsdns-14.co.uk",
"ns-378.awsdns-47.com",
"ns-587.awsdns-09.net"
]
},
"whoisError": null,
"keywordDensity": {
"1": [
{ "ngram": "subscribers", "count": 97, "frequency": 0.0401656314699793 },
{ "ngram": "cnn", "count": 90, "frequency": 0.037267080745341616 },
{ "ngram": "video", "count": 62, "frequency": 0.02567287784679089 }
]
},
"keywordDensityError": null,
"htmlFetchProxyTier": "apify-datacenter-US"
}

Export the dataset as JSON, CSV, Excel, XML or HTML table, or pull it through the API.

Reading a short row

Rows that worked carry no _error key at all. When something goes wrong the key appears with the reason, and that row is not billed:

Common valueMeaning
no_dataSimilarWeb tracks no data for this domain. Often a typo or a very small site.
forbidden_by_similarwebSimilarWeb refuses this specific domain, whatever address asks.
partial_keyword_density_failedWHOIS came back, the homepage did not.
all_subtasks_failedNeither half returned anything.
nothing_requestedBoth domain analysis options were switched off, so there was nothing to collect.

Domain analysis rows also carry whoisError and keywordDensityError separately, so you can see which half failed and why. A homepage that answers with an interstitial instead of content shows up there as thin_body.

Limitations

  • In aiPromptsStatus, code: 0 means the prompts came back fine. It is SimilarWeb's own status value, not an error code. A non-zero code, or null, means no prompts were published for that domain.
  • Chatbots and prompts are capped at 3 each by the public data, the same way countries and keywords are capped at 5.
  • The numbers are SimilarWeb's estimates, refreshed monthly, not live analytics. Treat them as a market-size signal, not a billing record.
  • Top countries and top keywords are capped at 5 by the public data itself, checked across 40 domains.
  • No competitor list, company profile, or rank-change deltas. You get three months of absolute visits instead, which is what most people compute a trend from anyway.
  • Use trafficSourcesRanked for percentages that add up. The shorter trafficSources object double counts paid traffic inside search and social, so its six numbers do not sum to 1.0. The ranked list does.
  • Keyword density reads the homepage only, and many large retail and news sites render their content with JavaScript, so those return navigation and menu text rather than article copy.
  • Some homepages sit behind a bot wall. Keyword density comes back empty with a reason and is not charged; the WHOIS half still works.
  • Some TLDs publish no WHOIS at all, including .de, .jp and .ru. Most registrars also redact the registrant name and country, so those fields are usually null.
  • The free plan has no second proxy pool. Paid plans fall back to a second pool when the first is blocked; on the free plan, where Apify covers the proxy, the run finishes with what it collected.
  • A short run tells you it is short. The status message says how many domains came back and why the rest did not.

Frequently asked questions

How much does this SimilarWeb scraper cost?

$1 per 1,000 results on every paid plan, $2 per 1,000 on the free plan, plus a $0.005 start fee per run at the default memory. Failed and empty lookups are free. No subscription, pause whenever.

Do I need a SimilarWeb account or API key?

No. The scraper reads public data only. You never provide a login, a cookie, or a key.

Scraping publicly available pages is generally allowed in the US and most of the EU, provided you do not collect personal data covered by GDPR or CCPA without a lawful basis. This actor touches public endpoints only, but what you do with the output is your responsibility. Apify's breakdown: Is web scraping legal?

How accurate is the traffic data?

It is SimilarWeb's own published estimate, the same figure their public page shows. Sites with a verified analytics integration are marked dataSource: "ga-verified"; everything else is "estimated". Expect differences from what a site owner sees in their own analytics.

Can I bulk scrape hundreds of domains in one run?

Yes, that is what it is built for. Pass the whole list in domains. Runs of several hundred domains are routine, and repeat lookups inside the same month come from cache in about a second each.

Can I integrate it with other tools?

Yes. Push results to Make, Zapier, Slack, Airbyte, GitHub, Google Sheets and more. Every run is a webhook source. Full list: Apify integrations

Can I run it through the Apify API?

Yes:

curl -X POST "https://api.apify.com/v2/acts/sourabhbgp~similarweb-scraper/runs?token=APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"mode":"traffic","domains":["washingtonpost.com","nytimes.com"]}'

Docs: Apify API reference

Can I use it through an MCP server?

Yes. Apify's MCP server exposes every actor as a tool, so Claude, Cursor and any other MCP client can call this scraper. Setup: Apify MCP docs

Your feedback

Found a bug, want a field, or seeing something odd? Open the Issues tab. Reports reach a human, and fixes usually ship the same week.