Fast Similarweb Scraper Light
Pricing
from $0.70 / 1,000 results
Fast Similarweb Scraper Light
Fast Similarweb scraper for website traffic, rankings, engagement, and ai traffic sources.
Pricing
from $0.70 / 1,000 results
Rating
0.0
(0)
Developer
karamelo
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
12 hours ago
Last modified
Categories
Share
Overview: Comprehensive Website Traffic and Ranking Intelligence
Similarweb Traffic & Rank Scraper Light is a high-performance web data automation designed to collect public website traffic estimates, ranking metrics, audience engagement analytics, acquisition channel breakdowns, geographic distributions, and top organic keywords for any domain or list of websites. Built for marketing teams, SEO professionals, digital agencies, data scientists, and investment analysts, this Actor simplifies the process of extracting digital competitive intelligence in bulk without requiring an expensive enterprise seat or manual domain-by-domain navigation.
Whether you need to benchmark three major industry competitors or analyze a list of five hundred prospective merchant websites, this Actor automates the complete data extraction lifecycle. Users submit bare domain names, standard website URLs, or existing profile links, and the Actor normalizes, deduplicates, and evaluates each website against public traffic records. The resulting dataset provides clean, structured, spreadsheet-ready rows containing essential performance indicators such as total estimated monthly visits, monthly traffic trends, global and country rankings, category standings, bounce rates, visit durations, pages per visit, traffic sources, top audience countries, and top driving keywords.
By executing requests through an optimized parallel architecture, the Actor delivers rapid results while maintaining strict data integrity. Every run produces a structured tabular dataset viewable directly in the Apify Console and downloadable in JSON, CSV, XML, or Excel formats.
What does Similarweb Traffic & Rank Scraper Light do?
The Actor serves as an automated bridge between raw web traffic metrics and your business workflows. When provided with a list of domains, it performs the following core functions:
- Hostname Normalization: Cleans and standardizes varied input formats, transforming protocols, subdirectories, search queries, tracking parameters, and profile links into uniform root domain identifiers.
- Traffic and Ranking Retrieval: Captures estimated total monthly visits, 3-month historical visit trends, worldwide global ranks, primary country ranks with ISO codes, and industry category ranks.
- Engagement Quality Assessment: Gathers critical user behavior metrics including bounce rate percentages, average page views per session, and average time spent on site in seconds.
- Acquisition Channel Decomposition: Details the proportions of incoming traffic generated by direct navigation, organic search, paid search, organic social networks, paid social campaigns, referral links, email campaigns, and display advertising.
- Audience Geographic Footprint: Extracts the leading countries sending traffic to each target domain, along with their respective percentage shares of total traffic.
- Organic Keyword Discovery: Collects top performing search keywords driving organic traffic, accompanied by search volumes, estimated traffic value, and cost-per-click estimates where published.
- Structured Storage & Export: Pushes clean, normalized records directly to the default Apify Dataset for immediate analysis or programmatic downstream consumption.
Why use Similarweb Traffic & Rank Scraper Light?
- High-Throughput Parallel Processing — Analyze dozens or hundreds of target websites in minutes using configurable concurrency that scales to match your project deadlines and platform allowances.
- Fast Turnaround — Direct data extraction delivers rapid results, minimizing compute memory consumption and keeping operational platform expenses low.
- Robust Failure Resilience — Transient network interruptions and rate restrictions are resolved automatically through intelligent exponential backoff and retry routines that protect batch completion.
- Input Flexibility — Supports mixed input formats including bare hostnames, fully qualified HTTPS URLs, tracking links, and profile URLs without manual data cleaning.
- Clean Denormalized Records — Delivers flat, fully structured dataset rows ready for immediate import into Google Sheets, Microsoft Excel, PostgreSQL, BigQuery, or Power BI.
- Zero-Bloat Dataset Storage — Delivers compact, standardized datasets directly to the default dataset without unnecessary storage overhead.
- Transparent Execution Feedback — Provides clear progress updates and concise completion tallies in the Apify Console run log.
Key Data and Metrics Extracted
The Actor extracts a rich collection of commercial, technical, and marketing metrics organized into logical groups:
| Data Group | Primary Fields Included | Practical Business Value |
|---|---|---|
| Identity & Category | domain, siteName, title, description, category, similarwebUrl | Identify the website, its official title, classification, and public profile link. |
| Rankings | globalRank, countryRank, countryRankCountryCode, categoryRank, categoryRankCategory | Understand where the domain stands globally, within its primary national market, and among category peers. |
| Volume & Trend | monthlyVisits, monthlyVisitsHistory, engagementMonth | Measure total monthly visit volume and observe 3-month historical trajectory direction. |
| User Engagement | bounceRate, pagesPerVisit, avgVisitDurationSec | Evaluate audience retention, content stickiness, and overall visit quality. |
| Traffic Acquisition | trafficDirect, trafficSearchOrganic, trafficSocialOrganic, trafficReferrals, trafficMail, trafficDisplayAds | Pinpoint primary acquisition channels and identify marketing channel dependencies. |
| Geographic Markets | topCountries (country code, country name, traffic share) | Identify top regional customer bases and evaluate international expansion opportunities. |
| Search Keywords | topKeywords (keyword name, search volume, CPC, estimated value) | Uncover high-intent organic search terms and inform keyword research strategies. |
| Execution Diagnostics | hasData, scrapedAt | Track data availability and timestamp extraction. |
Who is this Actor for?
- Competitive Intelligence Analysts — Monitor rival market leaders, track monthly traffic growth or contraction, and detect emerging disruptors across digital verticals.
- SEO & Growth Strategists — Analyze organic search shares, uncover competitor keyword strategies, and benchmark channel acquisition splits against industry averages.
- B2B Sales & Lead Generation Teams — Enrich prospective customer lists by scoring corporate domains according to estimated traffic volume and engagement before conducting outbound outreach.
- E-Commerce & Marketplace Operators — Evaluate merchant reach, benchmark niche retail competitors, and determine vendor digital viability.
- Venture Capital & Private Equity Investors — Conduct preliminary due diligence on potential portfolio companies by verifying claimed traffic scale, historical momentum, and user retention.
- Data Engineers & Automated Pipeline Builders — Feed live digital analytics directly into data warehouses, dashboards, and automated alerting webhooks.
Use Cases and Real-World Applications
Competitive Benchmarking and Digital Intelligence
Gain an objective, third-party perspective on market share and digital presence across your competitive set. By passing twenty leading industry players into the Actor, market intelligence teams can compare global rankings, monthly visits, and historical trend trajectories over a three-month horizon. Identifying competitors with expanding monthly visit volumes highlights organizations executing successful marketing initiatives, while examining traffic channel splits reveals whether growth stems from brand awareness (direct traffic), search dominance, paid media, or referral partnerships.
SEO Strategy, Organic Keyword Valuation, and Content Planning
Content and SEO directors frequently use this Actor to audit how competitors capture organic demand. The scraper extracts the percentage of incoming traffic attributable to organic search alongside top organic keywords, search volumes, and estimated keyword monetary values. These findings expose high-value topics and commercial terms where competitors enjoy organic visibility, helping your editorial team prioritize content clusters that capture targeted search traffic without inflated pay-per-click advertising expenses.
B2B Lead Qualification and Merchant Scoring
Outbound sales development teams often waste valuable time reaching out to inactive, stagnant, or low-traffic domains. By integrating this Actor into your CRM ingestion pipeline, you can automatically score prospective domains upon signup or list import. Domains with healthy monthly visits and steady engagement metrics can be flagged for high-priority executive follow-up, while inactive or zero-data domains can be routed to lower-touch automated email nurturing sequences.
Investment Screening, M&A Due Diligence, and Market Sizing
Private equity analysts and corporate development officers require rapid verification of company performance claims. When reviewing an acquisition candidate, analysts run the target domain alongside known competitors to validate whether historical traffic numbers are stable, growing, or declining. The presence of diversified traffic sources (healthy direct and organic search shares) indicates sustainable customer interest, whereas heavy dependence on paid advertising signals high customer acquisition costs.
Step-by-Step Getting Started Guide
Quick Start via Apify Console
- Navigate to the Actor page on the Apify Store and click Try for free or Run.
- In the Domains input field, enter or paste the domain names or website URLs you wish to analyze (e.g.
google.com,github.com,openai.com). - Set your desired Max Concurrency (default is
10simultaneous requests) and Max Retries (default is3). - Keep the default Proxy Configuration enabled to ensure reliable access.
- Click Save & Start to begin collection.
- Monitor the live log for progress updates. Once completed, inspect the resulting records in the Dataset tab or export them as JSON, CSV, or Excel.
Example 1: Minimal Single Domain Run
To quickly check a single company website with default settings:
{"domains": ["stripe.com"]}
Example 2: Batch Analysis with High Concurrency
To process a multi-domain competitive list with optimized throughput:
{"domains": ["stripe.com","paypal.com","square.com","adyen.com","checkout.com"],"maxConcurrency": 15,"maxRetries": 3}
Example 3: High-Throughput Batch Processing
To analyze multiple domains at higher concurrency:
{"domains": ["https://github.com/pricing","https://gitlab.com","https://bitbucket.org"],"maxConcurrency": 20,"maxRetries": 3}
Input Reference
The Actor accepts a clean JSON input object configuring the scope, performance, and proxy options for the run:
| Parameter | Type | Default | Required | Description |
|---|---|---|---|---|
domains | array of strings | ["google.com", "github.com", "openai.com"] | Yes | List of domain names or URLs to analyze. Accepts bare hostnames (stripe.com), URLs with protocols (https://stripe.com), subpaths, and profile links. Automatically normalized. |
maxConcurrency | integer | 10 | No | Maximum number of domain requests executed simultaneously. Allowed range: 1 to 50. Recommended range for standard runs: 5 to 20. |
maxRetries | integer | 3 | No | Maximum retry attempts per domain upon transient network or rate-limiting errors before moving to the next item. Range: 0 to 10. |
proxyConfiguration | object | Default Proxy | No | Standard Apify Proxy configuration object to maintain consistent access. |
Output Dataset Structure
Each processed domain generates one standardized dataset record. If Similarweb does not publish data for a domain (such as obscure or newly registered websites), the Actor emits a record with hasData: false and null metric fields so your downstream analysis retains full reconciliation of all submitted domains.
| Field | Type | Description |
|---|---|---|
domain | string | The cleaned, lowercase target root domain (e.g. github.com). |
siteName | string | null | The website brand name or reported site identifier. |
title | string | null | Page title associated with the website profile. |
description | string | null | Metadata description summarizing the website content or business. |
category | string | null | Primary industry or topic category classification. |
categoryRank | integer | null | Rank position of the website within its assigned category. |
categoryRankCategory | string | null | Human-readable category title corresponding to categoryRank. |
globalRank | integer | null | Worldwide website traffic rank (1 represents the most visited domain globally). |
countryRank | integer | null | Ranking of the website within its primary national geographic market. |
countryRankCountryCode | string | null | Two-letter ISO country code for the primary geographic ranking market. |
monthlyVisits | integer | null | Estimated total visits received during the most recent snapshot month. |
bounceRate | number | null | Proportion of single-page sessions without further interaction (0.0 to 1.0). |
pagesPerVisit | number | null | Average number of web pages viewed per user session. |
avgVisitDurationSec | number | null | Average duration of a user visit in seconds. |
engagementMonth | string | null | Calendar month of the extracted traffic metrics formatted as YYYY-MM. |
monthlyVisitsHistory | array | null | Chronological list of historical monthly visit counts formatted as { "date": "YYYY-MM-DD", "visits": number }. |
trafficDirect | number | null | Fraction of traffic originating from direct address bar entries or bookmarks (0.0 to 1.0). |
trafficSearchOrganic | number | null | Fraction of traffic originating from unpaid organic search engine results. |
trafficSearchPaid | number | null | Fraction of traffic originating from paid search engine advertising. |
trafficSocialOrganic | number | null | Fraction of traffic originating from unpaid organic social media links. |
trafficSocialPaid | number | null | Fraction of traffic originating from paid social media advertisements. |
trafficReferrals | number | null | Fraction of traffic originating from referral hyperlinks on other websites. |
trafficMail | number | null | Fraction of traffic originating from direct email messages and campaigns. |
trafficDisplayAds | number | null | Fraction of traffic originating from display and banner ad networks. |
trafficGenAi | number | null | Fraction of traffic originating from Generative AI search and referral platforms. |
trafficAffiliate | number | null | Fraction of traffic originating from affiliate partnerships and publisher links. |
topCountries | array | null | List of leading visitor countries with ISO codes, country names, and traffic shares. |
topKeywords | array | null | Top search terms driving organic visits, with search volume and cost-per-click values. |
competitors | array | null | List of top similarity competitor domains identified by Similarweb. |
aiTrafficDetails | object | null | In-depth AI traffic metrics including chatbot split, distribution, and top prompts. |
largeScreenshot | string | null | Direct image URL to the website preview screenshot on Similarweb CDN. |
similarwebUrl | string | Direct link to the public Similarweb website report page. |
hasData | boolean | Indicates whether the source platform maintained and returned traffic data for this domain. |
scrapedAt | string | ISO 8601 timestamp marking the moment the data was captured. |
apiData | object | null | Complete unparsed raw Similarweb API JSON payload (as provided in reference code). |
raw | object | null | Complete unparsed raw Similarweb API JSON payload. |
Sample Dataset Record
Below is an authentic sample record generated by the Actor for github.com:
{"domain": "github.com","siteName": "github.com","title": "GitHub · Change is constant. GitHub keeps you ahead.","description": "Join the world's most widely adopted, AI-powered developer platform where millions of developers build software.","category": "computers_electronics_and_technology/programming_and_developer_software","categoryRank": 4,"categoryRankCategory": "Computers_Electronics_and_Technology/Programming_and_Developer_Software","globalRank": 50,"countryRank": 81,"countryRankCountryCode": "US","monthlyVisits": 649321442,"bounceRate": 0.3665,"pagesPerVisit": 5.77,"avgVisitDurationSec": 384.66,"engagementMonth": "2026-08","monthlyVisitsHistory": [{ "date": "2026-06-01", "visits": 615239605 },{ "date": "2026-07-01", "visits": 637885711 },{ "date": "2026-08-01", "visits": 649321442 }],"trafficDirect": 0.526,"trafficSearchOrganic": null,"trafficSearchPaid": null,"trafficSocialOrganic": null,"trafficSocialPaid": null,"trafficReferrals": 0.0956,"trafficMail": 0.0115,"trafficDisplayAds": 0.0024,"topCountries": [{ "countryCode": "US", "country": "United States", "share": 0.1861 },{ "countryCode": "CN", "country": "China", "share": 0.1166 },{ "countryCode": "IN", "country": "India", "share": 0.1047 },{ "countryCode": "RU", "country": "Russia", "share": 0.0833 },{ "countryCode": "DE", "country": "Germany", "share": 0.0411 }],"topKeywords": [{ "keyword": "github", "volume": 10019460, "estimatedValue": 11327530, "cpc": 1.63 },{ "keyword": "yt-dlp", "volume": 638970, "estimatedValue": 478520, "cpc": 4.63 }],"similarwebUrl": "https://www.similarweb.com/website/github.com/","hasData": true,"scrapedAt": "2026-09-12T21:15:03.648Z"}
Actionable Pillar Workflows
Workflow 1: Benchmarking Direct Competitor Digital Footprints
- Define the Scope: Compile a focused list of 5 to 15 key direct competitors within your market vertical.
- Execute Collection: Run the Actor with
domainspopulated with your competitor URLs. KeepmaxConcurrencyat10andmaxRetriesat3. - Filter and Compare Key Ratios:
- Sort the dataset by
monthlyVisitsto understand relative scale. - Calculate engagement index scores using
pagesPerVisitandavgVisitDurationSecto identify which competitors boast the most engaged audience. - Compare
trafficDirectagainsttrafficReferralsto see who relies on established brand recognition versus external syndication.
- Sort the dataset by
- Identify Gaps: Flag competitors with high
bounceRatefigures as opportunities to win over dissatisfied users through better landing page experiences. - Report to Leadership: Export the results to a Google Sheet or executive presentation highlighting month-over-month traffic trajectory changes.
Workflow 2: Evaluating Inbound Channel Vulnerability for SEO Audits
- Target Identification: Enter the prospective client domain and three primary organic competitors.
- Extract Channel Splits: Review the
trafficDirect,trafficSearchOrganic, andtrafficReferralsshares. - Evaluate Search Reliance: A site with over 60% dependence on organic search is highly vulnerable to algorithmic volatility. Use the
topKeywordslist to evaluate whether traffic concentrates on brand keywords (e.g.company name) or commercial non-branded queries. - Deliver Actionable SEO Recommendations: Present a diversification roadmap recommending content expansion or referral partnerships to balance acquisition channels.
Workflow 3: Automated Lead Scoring for SaaS Sales Pipelines
- Trigger via Webhook: Configure your CRM or inbound signup form to trigger an Apify run whenever a new enterprise lead registers with a corporate email.
- Extract Company Domain: Parse the domain from the email address (
company.com) and dispatch it to the Actor via the Apify REST API. - Automate Pipeline Routing:
- If
globalRankis below 50,000 andmonthlyVisitsexceeds 250,000: tag the lead as Enterprise Tier 1 and assign an Account Executive for immediate phone outreach. - If
monthlyVisitsis between 10,000 and 250,000: assign to mid-market SDR queues. - If
hasData: false: assign to self-serve automated onboarding.
- If
- Enrich CRM Record: Write the returned
category,topCountries, and estimatedmonthlyVisitsdirectly into the CRM account profile.
Programmatic Automation and API Usage
You can launch runs, monitor status, and download datasets programmatically using the official Apify API clients or direct HTTP REST calls.
Node.js / JavaScript Client Example
Install the official client using npm install apify-client:
import { ApifyClient } from 'apify-client';const client = new ApifyClient({token: process.env.APIFY_TOKEN,});async function runSimilarwebScraper() {const input = {domains: ['stripe.com', 'openai.com', 'shopify.com'],maxConcurrency: 10,maxRetries: 3,};console.log('Starting Similarweb Light scraper run...');const run = await client.actor('karamelo/similarweb-scraper-light').call(input);console.log(`Run completed with status: ${run.status}`);console.log(`Dataset ID: ${run.defaultDatasetId}`);// Fetch and display extracted recordsconst { items } = await client.dataset(run.defaultDatasetId).listItems();for (const item of items) {console.log(`- ${item.domain}: ${item.monthlyVisits?.toLocaleString() ?? 'No data'} visits (Rank: #${item.globalRank ?? 'N/A'})`);}}runSimilarwebScraper().catch(console.error);
Python Client Example
Install the official client using pip install apify-client:
import osfrom apify_client import ApifyClientclient = ApifyClient(os.getenv('APIFY_TOKEN'))actor_input = {'domains': ['stripe.com', 'openai.com', 'shopify.com'],'maxConcurrency': 10,'maxRetries': 3,}print('Starting Similarweb Light Actor run...')run = client.actor('karamelo/similarweb-scraper-light').call(run_input=actor_input)print(f"Run finished with status: {run['status']}")dataset_items = client.dataset(run['defaultDatasetId']).list_items().itemsfor record in dataset_items:domain = record.get('domain')visits = record.get('monthlyVisits')rank = record.get('globalRank')print(f"Domain: {domain} | Visits: {visits} | Global Rank: #{rank}")
cURL Webhook / REST Lifecycle
To trigger an asynchronous run via direct REST endpoint:
curl -X POST "https://api.apify.com/v2/acts/karamelo~similarweb-scraper-light/runs?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"domains": ["airbnb.com", "booking.com", "expedia.com"],"maxConcurrency": 10}'
To retrieve the resulting dataset items in CSV format once finished:
curl -G "https://api.apify.com/v2/datasets/{DATASET_ID}/items" \-d "token=$APIFY_TOKEN" \-d "format=csv" \-o similarweb_traffic_results.csv
Data Export Formats and Platform Integrations
Apify automatically exposes multiple structured formats for every dataset generated by this Actor:
- JSON / JSONL — Full hierarchical objects including nested history trends, top countries, and organic keywords.
- CSV — Flat spreadsheet format where array fields are neatly serialized for easy parsing in spreadsheet software.
- Microsoft Excel (XLSX) — Native workbook ready for financial analysis, charts, and stakeholder reporting.
- XML / HTML Table — For syndication pipelines, CMS integrations, and internal report rendering.
Platform Integrations
- Google Sheets & Microsoft Excel — Use Apify's native Google Drive integration or the dataset CSV URL formula (
=IMPORTDATA("...")) to refresh competitive rankings on a recurring schedule. - Airtable & Notion — Sync newly extracted domain profiles automatically via Apify webhooks and Make or Zapier.
- Slack & Discord Webhooks — Dispatch run completion summaries and alerting notifications directly to team communication channels.
- Cloud Warehouses (BigQuery, Snowflake, AWS S3) — Pipe high-volume domain traffic records directly into corporate enterprise data warehouses for longitudinal trend modeling.
Operational Best Practices and Concurrency Optimization
- Batch Size Guidelines: For optimal performance and monitoring, we recommend submitting domains in batches of 50 to 500 items per run. Processing multi-domain batches is substantially more economical than launching hundreds of single-domain Actor runs because system startup and proxy initialization overhead is shared.
- Concurrency Calibration: Use the default concurrency of
10for batches under 100 domains. For larger batches of 200 to 500 domains, scaling concurrency up to25or30accelerates completion without degrading network stability. - Proxy Configuration: Always utilize the default Apify Proxy configuration to prevent rate limiting and ensure uninterrupted collection.
- Direct Dataset Architecture: Extracted records are piped directly to the default Apify Dataset, eliminating redundant file storage overhead and keeping run executions fast and clean.
- Handling No-Data Domains: Obscure, local, or newly launched domains may lack public Similarweb records. The Actor marks these records with
hasData: false, allowing you to cleanly filter active from inactive domains in spreadsheet or database queries.
Pricing and Resource Consumption Guide
Similarweb Traffic & Rank Scraper Light operates on standard Apify platform compute consumption without monthly subscription locks or expensive seat licenses.
Because this Actor utilizes an optimized lightweight design, its memory footprint is minimal—typically consuming only 256 MB to 512 MB of memory throughout execution. A typical batch of 100 domains requires approximately 1 to 2 minutes of run time, consuming mere pennies of standard Apify platform compute credits.
| Batch Size | Typical Concurrency | Estimated Run Duration | Estimated Compute Usage |
|---|---|---|---|
| 10 domains | 5 | 15–30 seconds | ~0.003 CU |
| 50 domains | 10 | 45–90 seconds | ~0.012 CU |
| 200 domains | 20 | 2–4 minutes | ~0.045 CU |
| 500 domains | 25 | 5–8 minutes | ~0.110 CU |
Note: Estimates reflect standard network conditions and may vary depending on target website responsiveness, proxy latency, and selected retry allowances.
Known Limitations and Platform Boundaries
- Public Estimate Nature: All traffic counts, ranks, and channel distributions represent Similarweb's statistical modeling and estimates based on desktop and mobile web panels, ISP measurements, and public datasets. They do not reflect internal server log files or Google Analytics telemetry.
- Low-Traffic Domain Thresholds: Domains with very low monthly visitor counts (typically fewer than 5,000 monthly visits) may show
hasData: falseor lack granular keyword and country breakdowns. - Historical Depth: The free public profile endpoint surfaces a rolling 3-month historical visit trend. Multi-year historical lookbacks require enterprise Similarweb access.
- Subdomain Aggregation: Traffic figures generally represent root domain aggregation unless a specific major subdomain maintains an independent public ranking profile.
Diagnostic Troubleshooting Matrix
| Symptom | Likely Cause | Recommended Resolution |
|---|---|---|
Domain returns hasData: false with null ranks | The domain has low traffic volume or is not indexed in Similarweb's public database. | Confirm the domain has established web traffic by visiting its public profile URL directly on Similarweb. |
| Run times out or experiences high retry counts | Concurrency is configured higher than network bandwidth can sustain. | Lower maxConcurrency to 5 or 10 and ensure proxyConfiguration is enabled. |
| Input domain is missing from dataset | Input format was invalid or contained unparseable characters. | Check the run log for normalization warnings; ensure domain inputs follow standard hostname conventions (e.g. domain.com). |
| Incomplete keyword or country lists | The target website does not attract enough regional traffic to populate secondary data tables. | Review the monthlyVisits field; low-volume sites naturally publish fewer keyword and geographic details. |
| Dataset export shows blank columns for some fields | Fields like paid advertising or mail traffic are zero or not applicable for that website. | Use null coalescing (IFNULL, COALESCE) in downstream SQL or formulas to treat empty metrics as 0. |
Frequently Asked Questions (FAQ)
Do I need a Similarweb account or API subscription to use this Actor?
No. You do not need a Similarweb subscription, API key, or registered account. The Actor retrieves publicly accessible web analytics data using standard platform requests.
How fresh is the traffic data?
Similarweb publishes its web analytics on a monthly cadence. The engagementMonth field (e.g. 2026-08) in each dataset item clearly indicates the specific calendar month reflected by the data.
Can I input full URLs with paths instead of bare domains?
Yes. The Actor automatically cleans and standardizes URLs, extracting the root domain name whether you provide stripe.com, https://www.stripe.com/docs, or https://www.similarweb.com/website/stripe.com/.
What happens if a domain has no traffic data?
If Similarweb has no data for a submitted domain, the Actor outputs a row with hasData: false and null metric values. This ensures that every submitted domain is accounted for in your output dataset without failing the run.
Why is proxy configuration recommended?
Web analytics platforms employ automated rate monitoring against rapid burst traffic. Apify Proxy distributes requests across exit nodes, ensuring uninterrupted and reliable data extraction across large batches.
Can I run this Actor on a schedule?
Yes. You can use Apify Schedules to run the Actor automatically on a daily, weekly, or monthly cadence to track competitor ranking shifts and visit trends over time.
How many domains can I scrape in a single run?
There is no hard limit imposed by the Actor. We recommend processing batches of 50 to 500 domains per run for easy management and rapid turnaround.
Is historical traffic data included?
Yes. Each successful domain record contains the monthlyVisitsHistory array with the last three recorded monthly visit estimates, allowing you to plot trajectory trends.
Where are the extracted dataset records stored?
Results are automatically stored in the run's default Apify Dataset. You can view them in tabular format in the Apify Console or export them directly in JSON, CSV, XML, or Excel format.
Can I integrate the output into Google Sheets automatically?
Yes. You can copy the dataset's live CSV link from the Apify Console and use the =IMPORTDATA("url") function in Google Sheets, or configure the official Apify Google Sheets integration.
Ethical and Responsible Data Collection
Similarweb Traffic & Rank Scraper Light is intended for legitimate competitive analysis, market research, business intelligence, and digital strategy development. Users must respect the terms of service of target platforms, implement sensible concurrency levels to prevent undue load on public infrastructure, and adhere to applicable data privacy and intellectual property laws within their operational jurisdictions.