Fast Similarweb Scraper Light avatar

Fast Similarweb Scraper Light

Pricing

from $0.70 / 1,000 results

Go to Apify Store
Fast Similarweb Scraper Light

Fast Similarweb Scraper Light

Fast Similarweb scraper for website traffic, rankings, engagement, and ai traffic sources.

Pricing

from $0.70 / 1,000 results

Rating

0.0

(0)

Developer

karamelo

karamelo

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

12 hours ago

Last modified

Categories

Share

Overview: Comprehensive Website Traffic and Ranking Intelligence

Similarweb Traffic & Rank Scraper Light is a high-performance web data automation designed to collect public website traffic estimates, ranking metrics, audience engagement analytics, acquisition channel breakdowns, geographic distributions, and top organic keywords for any domain or list of websites. Built for marketing teams, SEO professionals, digital agencies, data scientists, and investment analysts, this Actor simplifies the process of extracting digital competitive intelligence in bulk without requiring an expensive enterprise seat or manual domain-by-domain navigation.

Whether you need to benchmark three major industry competitors or analyze a list of five hundred prospective merchant websites, this Actor automates the complete data extraction lifecycle. Users submit bare domain names, standard website URLs, or existing profile links, and the Actor normalizes, deduplicates, and evaluates each website against public traffic records. The resulting dataset provides clean, structured, spreadsheet-ready rows containing essential performance indicators such as total estimated monthly visits, monthly traffic trends, global and country rankings, category standings, bounce rates, visit durations, pages per visit, traffic sources, top audience countries, and top driving keywords.

By executing requests through an optimized parallel architecture, the Actor delivers rapid results while maintaining strict data integrity. Every run produces a structured tabular dataset viewable directly in the Apify Console and downloadable in JSON, CSV, XML, or Excel formats.

What does Similarweb Traffic & Rank Scraper Light do?

The Actor serves as an automated bridge between raw web traffic metrics and your business workflows. When provided with a list of domains, it performs the following core functions:

  • Hostname Normalization: Cleans and standardizes varied input formats, transforming protocols, subdirectories, search queries, tracking parameters, and profile links into uniform root domain identifiers.
  • Traffic and Ranking Retrieval: Captures estimated total monthly visits, 3-month historical visit trends, worldwide global ranks, primary country ranks with ISO codes, and industry category ranks.
  • Engagement Quality Assessment: Gathers critical user behavior metrics including bounce rate percentages, average page views per session, and average time spent on site in seconds.
  • Acquisition Channel Decomposition: Details the proportions of incoming traffic generated by direct navigation, organic search, paid search, organic social networks, paid social campaigns, referral links, email campaigns, and display advertising.
  • Audience Geographic Footprint: Extracts the leading countries sending traffic to each target domain, along with their respective percentage shares of total traffic.
  • Organic Keyword Discovery: Collects top performing search keywords driving organic traffic, accompanied by search volumes, estimated traffic value, and cost-per-click estimates where published.
  • Structured Storage & Export: Pushes clean, normalized records directly to the default Apify Dataset for immediate analysis or programmatic downstream consumption.

Why use Similarweb Traffic & Rank Scraper Light?

  • High-Throughput Parallel Processing — Analyze dozens or hundreds of target websites in minutes using configurable concurrency that scales to match your project deadlines and platform allowances.
  • Fast Turnaround — Direct data extraction delivers rapid results, minimizing compute memory consumption and keeping operational platform expenses low.
  • Robust Failure Resilience — Transient network interruptions and rate restrictions are resolved automatically through intelligent exponential backoff and retry routines that protect batch completion.
  • Input Flexibility — Supports mixed input formats including bare hostnames, fully qualified HTTPS URLs, tracking links, and profile URLs without manual data cleaning.
  • Clean Denormalized Records — Delivers flat, fully structured dataset rows ready for immediate import into Google Sheets, Microsoft Excel, PostgreSQL, BigQuery, or Power BI.
  • Zero-Bloat Dataset Storage — Delivers compact, standardized datasets directly to the default dataset without unnecessary storage overhead.
  • Transparent Execution Feedback — Provides clear progress updates and concise completion tallies in the Apify Console run log.

Key Data and Metrics Extracted

The Actor extracts a rich collection of commercial, technical, and marketing metrics organized into logical groups:

Data GroupPrimary Fields IncludedPractical Business Value
Identity & Categorydomain, siteName, title, description, category, similarwebUrlIdentify the website, its official title, classification, and public profile link.
RankingsglobalRank, countryRank, countryRankCountryCode, categoryRank, categoryRankCategoryUnderstand where the domain stands globally, within its primary national market, and among category peers.
Volume & TrendmonthlyVisits, monthlyVisitsHistory, engagementMonthMeasure total monthly visit volume and observe 3-month historical trajectory direction.
User EngagementbounceRate, pagesPerVisit, avgVisitDurationSecEvaluate audience retention, content stickiness, and overall visit quality.
Traffic AcquisitiontrafficDirect, trafficSearchOrganic, trafficSocialOrganic, trafficReferrals, trafficMail, trafficDisplayAdsPinpoint primary acquisition channels and identify marketing channel dependencies.
Geographic MarketstopCountries (country code, country name, traffic share)Identify top regional customer bases and evaluate international expansion opportunities.
Search KeywordstopKeywords (keyword name, search volume, CPC, estimated value)Uncover high-intent organic search terms and inform keyword research strategies.
Execution DiagnosticshasData, scrapedAtTrack data availability and timestamp extraction.

Who is this Actor for?

  • Competitive Intelligence Analysts — Monitor rival market leaders, track monthly traffic growth or contraction, and detect emerging disruptors across digital verticals.
  • SEO & Growth Strategists — Analyze organic search shares, uncover competitor keyword strategies, and benchmark channel acquisition splits against industry averages.
  • B2B Sales & Lead Generation Teams — Enrich prospective customer lists by scoring corporate domains according to estimated traffic volume and engagement before conducting outbound outreach.
  • E-Commerce & Marketplace Operators — Evaluate merchant reach, benchmark niche retail competitors, and determine vendor digital viability.
  • Venture Capital & Private Equity Investors — Conduct preliminary due diligence on potential portfolio companies by verifying claimed traffic scale, historical momentum, and user retention.
  • Data Engineers & Automated Pipeline Builders — Feed live digital analytics directly into data warehouses, dashboards, and automated alerting webhooks.

Use Cases and Real-World Applications

Competitive Benchmarking and Digital Intelligence

Gain an objective, third-party perspective on market share and digital presence across your competitive set. By passing twenty leading industry players into the Actor, market intelligence teams can compare global rankings, monthly visits, and historical trend trajectories over a three-month horizon. Identifying competitors with expanding monthly visit volumes highlights organizations executing successful marketing initiatives, while examining traffic channel splits reveals whether growth stems from brand awareness (direct traffic), search dominance, paid media, or referral partnerships.

SEO Strategy, Organic Keyword Valuation, and Content Planning

Content and SEO directors frequently use this Actor to audit how competitors capture organic demand. The scraper extracts the percentage of incoming traffic attributable to organic search alongside top organic keywords, search volumes, and estimated keyword monetary values. These findings expose high-value topics and commercial terms where competitors enjoy organic visibility, helping your editorial team prioritize content clusters that capture targeted search traffic without inflated pay-per-click advertising expenses.

B2B Lead Qualification and Merchant Scoring

Outbound sales development teams often waste valuable time reaching out to inactive, stagnant, or low-traffic domains. By integrating this Actor into your CRM ingestion pipeline, you can automatically score prospective domains upon signup or list import. Domains with healthy monthly visits and steady engagement metrics can be flagged for high-priority executive follow-up, while inactive or zero-data domains can be routed to lower-touch automated email nurturing sequences.

Investment Screening, M&A Due Diligence, and Market Sizing

Private equity analysts and corporate development officers require rapid verification of company performance claims. When reviewing an acquisition candidate, analysts run the target domain alongside known competitors to validate whether historical traffic numbers are stable, growing, or declining. The presence of diversified traffic sources (healthy direct and organic search shares) indicates sustainable customer interest, whereas heavy dependence on paid advertising signals high customer acquisition costs.

Step-by-Step Getting Started Guide

Quick Start via Apify Console

  1. Navigate to the Actor page on the Apify Store and click Try for free or Run.
  2. In the Domains input field, enter or paste the domain names or website URLs you wish to analyze (e.g. google.com, github.com, openai.com).
  3. Set your desired Max Concurrency (default is 10 simultaneous requests) and Max Retries (default is 3).
  4. Keep the default Proxy Configuration enabled to ensure reliable access.
  5. Click Save & Start to begin collection.
  6. Monitor the live log for progress updates. Once completed, inspect the resulting records in the Dataset tab or export them as JSON, CSV, or Excel.

Example 1: Minimal Single Domain Run

To quickly check a single company website with default settings:

{
"domains": [
"stripe.com"
]
}

Example 2: Batch Analysis with High Concurrency

To process a multi-domain competitive list with optimized throughput:

{
"domains": [
"stripe.com",
"paypal.com",
"square.com",
"adyen.com",
"checkout.com"
],
"maxConcurrency": 15,
"maxRetries": 3
}

Example 3: High-Throughput Batch Processing

To analyze multiple domains at higher concurrency:

{
"domains": [
"https://github.com/pricing",
"https://gitlab.com",
"https://bitbucket.org"
],
"maxConcurrency": 20,
"maxRetries": 3
}

Input Reference

The Actor accepts a clean JSON input object configuring the scope, performance, and proxy options for the run:

ParameterTypeDefaultRequiredDescription
domainsarray of strings["google.com", "github.com", "openai.com"]YesList of domain names or URLs to analyze. Accepts bare hostnames (stripe.com), URLs with protocols (https://stripe.com), subpaths, and profile links. Automatically normalized.
maxConcurrencyinteger10NoMaximum number of domain requests executed simultaneously. Allowed range: 1 to 50. Recommended range for standard runs: 5 to 20.
maxRetriesinteger3NoMaximum retry attempts per domain upon transient network or rate-limiting errors before moving to the next item. Range: 0 to 10.
proxyConfigurationobjectDefault ProxyNoStandard Apify Proxy configuration object to maintain consistent access.

Output Dataset Structure

Each processed domain generates one standardized dataset record. If Similarweb does not publish data for a domain (such as obscure or newly registered websites), the Actor emits a record with hasData: false and null metric fields so your downstream analysis retains full reconciliation of all submitted domains.

FieldTypeDescription
domainstringThe cleaned, lowercase target root domain (e.g. github.com).
siteNamestring | nullThe website brand name or reported site identifier.
titlestring | nullPage title associated with the website profile.
descriptionstring | nullMetadata description summarizing the website content or business.
categorystring | nullPrimary industry or topic category classification.
categoryRankinteger | nullRank position of the website within its assigned category.
categoryRankCategorystring | nullHuman-readable category title corresponding to categoryRank.
globalRankinteger | nullWorldwide website traffic rank (1 represents the most visited domain globally).
countryRankinteger | nullRanking of the website within its primary national geographic market.
countryRankCountryCodestring | nullTwo-letter ISO country code for the primary geographic ranking market.
monthlyVisitsinteger | nullEstimated total visits received during the most recent snapshot month.
bounceRatenumber | nullProportion of single-page sessions without further interaction (0.0 to 1.0).
pagesPerVisitnumber | nullAverage number of web pages viewed per user session.
avgVisitDurationSecnumber | nullAverage duration of a user visit in seconds.
engagementMonthstring | nullCalendar month of the extracted traffic metrics formatted as YYYY-MM.
monthlyVisitsHistoryarray | nullChronological list of historical monthly visit counts formatted as { "date": "YYYY-MM-DD", "visits": number }.
trafficDirectnumber | nullFraction of traffic originating from direct address bar entries or bookmarks (0.0 to 1.0).
trafficSearchOrganicnumber | nullFraction of traffic originating from unpaid organic search engine results.
trafficSearchPaidnumber | nullFraction of traffic originating from paid search engine advertising.
trafficSocialOrganicnumber | nullFraction of traffic originating from unpaid organic social media links.
trafficSocialPaidnumber | nullFraction of traffic originating from paid social media advertisements.
trafficReferralsnumber | nullFraction of traffic originating from referral hyperlinks on other websites.
trafficMailnumber | nullFraction of traffic originating from direct email messages and campaigns.
trafficDisplayAdsnumber | nullFraction of traffic originating from display and banner ad networks.
trafficGenAinumber | nullFraction of traffic originating from Generative AI search and referral platforms.
trafficAffiliatenumber | nullFraction of traffic originating from affiliate partnerships and publisher links.
topCountriesarray | nullList of leading visitor countries with ISO codes, country names, and traffic shares.
topKeywordsarray | nullTop search terms driving organic visits, with search volume and cost-per-click values.
competitorsarray | nullList of top similarity competitor domains identified by Similarweb.
aiTrafficDetailsobject | nullIn-depth AI traffic metrics including chatbot split, distribution, and top prompts.
largeScreenshotstring | nullDirect image URL to the website preview screenshot on Similarweb CDN.
similarwebUrlstringDirect link to the public Similarweb website report page.
hasDatabooleanIndicates whether the source platform maintained and returned traffic data for this domain.
scrapedAtstringISO 8601 timestamp marking the moment the data was captured.
apiDataobject | nullComplete unparsed raw Similarweb API JSON payload (as provided in reference code).
rawobject | nullComplete unparsed raw Similarweb API JSON payload.

Sample Dataset Record

Below is an authentic sample record generated by the Actor for github.com:

{
"domain": "github.com",
"siteName": "github.com",
"title": "GitHub · Change is constant. GitHub keeps you ahead.",
"description": "Join the world's most widely adopted, AI-powered developer platform where millions of developers build software.",
"category": "computers_electronics_and_technology/programming_and_developer_software",
"categoryRank": 4,
"categoryRankCategory": "Computers_Electronics_and_Technology/Programming_and_Developer_Software",
"globalRank": 50,
"countryRank": 81,
"countryRankCountryCode": "US",
"monthlyVisits": 649321442,
"bounceRate": 0.3665,
"pagesPerVisit": 5.77,
"avgVisitDurationSec": 384.66,
"engagementMonth": "2026-08",
"monthlyVisitsHistory": [
{ "date": "2026-06-01", "visits": 615239605 },
{ "date": "2026-07-01", "visits": 637885711 },
{ "date": "2026-08-01", "visits": 649321442 }
],
"trafficDirect": 0.526,
"trafficSearchOrganic": null,
"trafficSearchPaid": null,
"trafficSocialOrganic": null,
"trafficSocialPaid": null,
"trafficReferrals": 0.0956,
"trafficMail": 0.0115,
"trafficDisplayAds": 0.0024,
"topCountries": [
{ "countryCode": "US", "country": "United States", "share": 0.1861 },
{ "countryCode": "CN", "country": "China", "share": 0.1166 },
{ "countryCode": "IN", "country": "India", "share": 0.1047 },
{ "countryCode": "RU", "country": "Russia", "share": 0.0833 },
{ "countryCode": "DE", "country": "Germany", "share": 0.0411 }
],
"topKeywords": [
{ "keyword": "github", "volume": 10019460, "estimatedValue": 11327530, "cpc": 1.63 },
{ "keyword": "yt-dlp", "volume": 638970, "estimatedValue": 478520, "cpc": 4.63 }
],
"similarwebUrl": "https://www.similarweb.com/website/github.com/",
"hasData": true,
"scrapedAt": "2026-09-12T21:15:03.648Z"
}

Actionable Pillar Workflows

Workflow 1: Benchmarking Direct Competitor Digital Footprints

  1. Define the Scope: Compile a focused list of 5 to 15 key direct competitors within your market vertical.
  2. Execute Collection: Run the Actor with domains populated with your competitor URLs. Keep maxConcurrency at 10 and maxRetries at 3.
  3. Filter and Compare Key Ratios:
    • Sort the dataset by monthlyVisits to understand relative scale.
    • Calculate engagement index scores using pagesPerVisit and avgVisitDurationSec to identify which competitors boast the most engaged audience.
    • Compare trafficDirect against trafficReferrals to see who relies on established brand recognition versus external syndication.
  4. Identify Gaps: Flag competitors with high bounceRate figures as opportunities to win over dissatisfied users through better landing page experiences.
  5. Report to Leadership: Export the results to a Google Sheet or executive presentation highlighting month-over-month traffic trajectory changes.

Workflow 2: Evaluating Inbound Channel Vulnerability for SEO Audits

  1. Target Identification: Enter the prospective client domain and three primary organic competitors.
  2. Extract Channel Splits: Review the trafficDirect, trafficSearchOrganic, and trafficReferrals shares.
  3. Evaluate Search Reliance: A site with over 60% dependence on organic search is highly vulnerable to algorithmic volatility. Use the topKeywords list to evaluate whether traffic concentrates on brand keywords (e.g. company name) or commercial non-branded queries.
  4. Deliver Actionable SEO Recommendations: Present a diversification roadmap recommending content expansion or referral partnerships to balance acquisition channels.

Workflow 3: Automated Lead Scoring for SaaS Sales Pipelines

  1. Trigger via Webhook: Configure your CRM or inbound signup form to trigger an Apify run whenever a new enterprise lead registers with a corporate email.
  2. Extract Company Domain: Parse the domain from the email address (company.com) and dispatch it to the Actor via the Apify REST API.
  3. Automate Pipeline Routing:
    • If globalRank is below 50,000 and monthlyVisits exceeds 250,000: tag the lead as Enterprise Tier 1 and assign an Account Executive for immediate phone outreach.
    • If monthlyVisits is between 10,000 and 250,000: assign to mid-market SDR queues.
    • If hasData: false: assign to self-serve automated onboarding.
  4. Enrich CRM Record: Write the returned category, topCountries, and estimated monthlyVisits directly into the CRM account profile.

Programmatic Automation and API Usage

You can launch runs, monitor status, and download datasets programmatically using the official Apify API clients or direct HTTP REST calls.

Node.js / JavaScript Client Example

Install the official client using npm install apify-client:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({
token: process.env.APIFY_TOKEN,
});
async function runSimilarwebScraper() {
const input = {
domains: ['stripe.com', 'openai.com', 'shopify.com'],
maxConcurrency: 10,
maxRetries: 3,
};
console.log('Starting Similarweb Light scraper run...');
const run = await client.actor('karamelo/similarweb-scraper-light').call(input);
console.log(`Run completed with status: ${run.status}`);
console.log(`Dataset ID: ${run.defaultDatasetId}`);
// Fetch and display extracted records
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const item of items) {
console.log(`- ${item.domain}: ${item.monthlyVisits?.toLocaleString() ?? 'No data'} visits (Rank: #${item.globalRank ?? 'N/A'})`);
}
}
runSimilarwebScraper().catch(console.error);

Python Client Example

Install the official client using pip install apify-client:

import os
from apify_client import ApifyClient
client = ApifyClient(os.getenv('APIFY_TOKEN'))
actor_input = {
'domains': ['stripe.com', 'openai.com', 'shopify.com'],
'maxConcurrency': 10,
'maxRetries': 3,
}
print('Starting Similarweb Light Actor run...')
run = client.actor('karamelo/similarweb-scraper-light').call(run_input=actor_input)
print(f"Run finished with status: {run['status']}")
dataset_items = client.dataset(run['defaultDatasetId']).list_items().items
for record in dataset_items:
domain = record.get('domain')
visits = record.get('monthlyVisits')
rank = record.get('globalRank')
print(f"Domain: {domain} | Visits: {visits} | Global Rank: #{rank}")

cURL Webhook / REST Lifecycle

To trigger an asynchronous run via direct REST endpoint:

curl -X POST "https://api.apify.com/v2/acts/karamelo~similarweb-scraper-light/runs?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"domains": ["airbnb.com", "booking.com", "expedia.com"],
"maxConcurrency": 10
}'

To retrieve the resulting dataset items in CSV format once finished:

curl -G "https://api.apify.com/v2/datasets/{DATASET_ID}/items" \
-d "token=$APIFY_TOKEN" \
-d "format=csv" \
-o similarweb_traffic_results.csv

Data Export Formats and Platform Integrations

Apify automatically exposes multiple structured formats for every dataset generated by this Actor:

  • JSON / JSONL — Full hierarchical objects including nested history trends, top countries, and organic keywords.
  • CSV — Flat spreadsheet format where array fields are neatly serialized for easy parsing in spreadsheet software.
  • Microsoft Excel (XLSX) — Native workbook ready for financial analysis, charts, and stakeholder reporting.
  • XML / HTML Table — For syndication pipelines, CMS integrations, and internal report rendering.

Platform Integrations

  • Google Sheets & Microsoft Excel — Use Apify's native Google Drive integration or the dataset CSV URL formula (=IMPORTDATA("...")) to refresh competitive rankings on a recurring schedule.
  • Airtable & Notion — Sync newly extracted domain profiles automatically via Apify webhooks and Make or Zapier.
  • Slack & Discord Webhooks — Dispatch run completion summaries and alerting notifications directly to team communication channels.
  • Cloud Warehouses (BigQuery, Snowflake, AWS S3) — Pipe high-volume domain traffic records directly into corporate enterprise data warehouses for longitudinal trend modeling.

Operational Best Practices and Concurrency Optimization

  • Batch Size Guidelines: For optimal performance and monitoring, we recommend submitting domains in batches of 50 to 500 items per run. Processing multi-domain batches is substantially more economical than launching hundreds of single-domain Actor runs because system startup and proxy initialization overhead is shared.
  • Concurrency Calibration: Use the default concurrency of 10 for batches under 100 domains. For larger batches of 200 to 500 domains, scaling concurrency up to 25 or 30 accelerates completion without degrading network stability.
  • Proxy Configuration: Always utilize the default Apify Proxy configuration to prevent rate limiting and ensure uninterrupted collection.
  • Direct Dataset Architecture: Extracted records are piped directly to the default Apify Dataset, eliminating redundant file storage overhead and keeping run executions fast and clean.
  • Handling No-Data Domains: Obscure, local, or newly launched domains may lack public Similarweb records. The Actor marks these records with hasData: false, allowing you to cleanly filter active from inactive domains in spreadsheet or database queries.

Pricing and Resource Consumption Guide

Similarweb Traffic & Rank Scraper Light operates on standard Apify platform compute consumption without monthly subscription locks or expensive seat licenses.

Because this Actor utilizes an optimized lightweight design, its memory footprint is minimal—typically consuming only 256 MB to 512 MB of memory throughout execution. A typical batch of 100 domains requires approximately 1 to 2 minutes of run time, consuming mere pennies of standard Apify platform compute credits.

Batch SizeTypical ConcurrencyEstimated Run DurationEstimated Compute Usage
10 domains515–30 seconds~0.003 CU
50 domains1045–90 seconds~0.012 CU
200 domains202–4 minutes~0.045 CU
500 domains255–8 minutes~0.110 CU

Note: Estimates reflect standard network conditions and may vary depending on target website responsiveness, proxy latency, and selected retry allowances.

Known Limitations and Platform Boundaries

  • Public Estimate Nature: All traffic counts, ranks, and channel distributions represent Similarweb's statistical modeling and estimates based on desktop and mobile web panels, ISP measurements, and public datasets. They do not reflect internal server log files or Google Analytics telemetry.
  • Low-Traffic Domain Thresholds: Domains with very low monthly visitor counts (typically fewer than 5,000 monthly visits) may show hasData: false or lack granular keyword and country breakdowns.
  • Historical Depth: The free public profile endpoint surfaces a rolling 3-month historical visit trend. Multi-year historical lookbacks require enterprise Similarweb access.
  • Subdomain Aggregation: Traffic figures generally represent root domain aggregation unless a specific major subdomain maintains an independent public ranking profile.

Diagnostic Troubleshooting Matrix

SymptomLikely CauseRecommended Resolution
Domain returns hasData: false with null ranksThe domain has low traffic volume or is not indexed in Similarweb's public database.Confirm the domain has established web traffic by visiting its public profile URL directly on Similarweb.
Run times out or experiences high retry countsConcurrency is configured higher than network bandwidth can sustain.Lower maxConcurrency to 5 or 10 and ensure proxyConfiguration is enabled.
Input domain is missing from datasetInput format was invalid or contained unparseable characters.Check the run log for normalization warnings; ensure domain inputs follow standard hostname conventions (e.g. domain.com).
Incomplete keyword or country listsThe target website does not attract enough regional traffic to populate secondary data tables.Review the monthlyVisits field; low-volume sites naturally publish fewer keyword and geographic details.
Dataset export shows blank columns for some fieldsFields like paid advertising or mail traffic are zero or not applicable for that website.Use null coalescing (IFNULL, COALESCE) in downstream SQL or formulas to treat empty metrics as 0.

Frequently Asked Questions (FAQ)

Do I need a Similarweb account or API subscription to use this Actor?

No. You do not need a Similarweb subscription, API key, or registered account. The Actor retrieves publicly accessible web analytics data using standard platform requests.

How fresh is the traffic data?

Similarweb publishes its web analytics on a monthly cadence. The engagementMonth field (e.g. 2026-08) in each dataset item clearly indicates the specific calendar month reflected by the data.

Can I input full URLs with paths instead of bare domains?

Yes. The Actor automatically cleans and standardizes URLs, extracting the root domain name whether you provide stripe.com, https://www.stripe.com/docs, or https://www.similarweb.com/website/stripe.com/.

What happens if a domain has no traffic data?

If Similarweb has no data for a submitted domain, the Actor outputs a row with hasData: false and null metric values. This ensures that every submitted domain is accounted for in your output dataset without failing the run.

Web analytics platforms employ automated rate monitoring against rapid burst traffic. Apify Proxy distributes requests across exit nodes, ensuring uninterrupted and reliable data extraction across large batches.

Can I run this Actor on a schedule?

Yes. You can use Apify Schedules to run the Actor automatically on a daily, weekly, or monthly cadence to track competitor ranking shifts and visit trends over time.

How many domains can I scrape in a single run?

There is no hard limit imposed by the Actor. We recommend processing batches of 50 to 500 domains per run for easy management and rapid turnaround.

Is historical traffic data included?

Yes. Each successful domain record contains the monthlyVisitsHistory array with the last three recorded monthly visit estimates, allowing you to plot trajectory trends.

Where are the extracted dataset records stored?

Results are automatically stored in the run's default Apify Dataset. You can view them in tabular format in the Apify Console or export them directly in JSON, CSV, XML, or Excel format.

Can I integrate the output into Google Sheets automatically?

Yes. You can copy the dataset's live CSV link from the Apify Console and use the =IMPORTDATA("url") function in Google Sheets, or configure the official Apify Google Sheets integration.

Ethical and Responsible Data Collection

Similarweb Traffic & Rank Scraper Light is intended for legitimate competitive analysis, market research, business intelligence, and digital strategy development. Users must respect the terms of service of target platforms, implement sensible concurrency levels to prevent undue load on public infrastructure, and adhere to applicable data privacy and intellectual property laws within their operational jurisdictions.