# Similarweb Scraper — Website Traffic, Rankings & Analytics (`crawloop/similarweb-scraper`) Actor

Scrape Similarweb website traffic into JSON: monthly visits, ranks, bounce rate, traffic sources, keywords, and AI referrals. Pull Top Websites lists by country or category. A Similarweb API alternative for Python, Node.js, and MCP.

- **URL**: https://apify.com/crawloop/similarweb-scraper.md
- **Developed by:** [Andrej Kiva](https://apify.com/crawloop) (community)
- **Categories:** SEO tools, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.29 / 1,000 website snapshots

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Similarweb Scraper — Website Traffic, Rankings & Analytics

> Unofficial tool for publicly accessible Similarweb website and Top Websites pages. Similarweb and related names are trademarks of their respective owners. Not affiliated with, sponsored by, or endorsed by Similarweb Ltd. For informational, research, and competitive-intelligence use only. Respect applicable terms of use and law.

**Similarweb Scraper** ◄── you are here

Scrape **Similarweb website traffic** into structured JSON: **monthly visits**, **global / country / category rank**, **bounce rate**, **pages per visit**, **visit duration**, **traffic sources** (including **GenAI**), **top countries**, **top keywords**, **AI chatbot referrals**, and public **Top Websites** ranking lists. Paste **domains**, **URLs**, or **Similarweb website URLs**. A practical **Similarweb API alternative** — call from **Python**, **Node.js**, **cURL**, or **Apify MCP**.

**Best for:** Similarweb scraper jobs, website traffic checks, SEO channel mix, competitor ranking, lead scoring, Top Websites market maps, and JSON/CSV datasets for BI or AI assistants.

> **Crawloop digital intelligence** — website traffic here, then product / app / review siblings.

| Similarweb (website traffic) | Product Hunt (launches) | Google Play (ASO) | Trustpilot (reviews) |
| :--- | :--- | :--- | :--- |
| **Similarweb Scraper** ◄── you are here | [Product Hunt Scraper](https://apify.com/crawloop/producthunt-scraper) | [Google Play Scraper](https://apify.com/crawloop/google-play-scraper) | [Trustpilot Scraper](https://apify.com/crawloop/trustpilot-scraper) |

### When to use this Actor

- You need a **Similarweb scraper** / **website traffic scraper** for one or many domains as JSON
- You want **visits, ranks, sources, keywords, and AI traffic** in a **single website row** (no extra run per data type)
- You pull **Top Websites rankings** by **category** or **country** (50 public sites per list) for market mapping
- You want a **Similarweb API alternative** from Python, Node.js, or an MCP / AI assistant
- You enrich CRM / SEO sheets without a Similarweb login or official API key

### When not to use this Actor

- **Official Similarweb REST / Batch API** multi-year enterprise exports — this Actor reads the **public monthly snapshot**
- **Paywalled Pro-only HTML** (some referral hostnames, company firmographics, demographics)
- **App store ASO** — use [Google Play Scraper](https://apify.com/crawloop/google-play-scraper) or [App Stores Scraper](https://apify.com/crawloop/app-stores-scraper)
- **Consumer reviews** — use [Trustpilot Scraper](https://apify.com/crawloop/trustpilot-scraper)

### Modes

| Mode | What it does |
| :--- | :--- |
| `website` | One Similarweb traffic snapshot per domain (**default**) |
| `ranking` | Public **Top Websites** lists (global / country / category) |
| `both` | Website rows and ranking rows in the same run |

### Key features

- **Full public website snapshot in one row** — visits, 3-month history, ranks, granular traffic sources, keywords, AI traffic split (ChatGPT / Claude / Gemini / Perplexity / Copilot)
- **Top Websites ranking mode** — 50 sites per list with rank change, bounce, pages/visit, duration
- **No Similarweb login or API key** — public website-data snapshot + public ranking pages
- **Bulk parallel domains** — no small per-run domain cap
- **Typed dataset rows** — `website`, `ranking`, `error`
- **HTTP (no browser)** for the core snapshot — fast enough for bulk SEO and lead-gen lists
- **Export-ready** — download the default dataset as JSON, CSV, Excel, or JSONL

### Input

| Parameter | Description |
| :--- | :--- |
| `mode` | `website` / `ranking` / `both` |
| `websites` | Domains, website URLs, or Similarweb website URLs |
| `categories` / `countries` | Ranking list slugs (`finance`, `united-states`) |
| `includeGlobal` | Worldwide Top 50 |
| `scrapeAllCategories` | Discover and scrape every public category list |
| `maxItems` / `concurrency` | Caps and parallelism |
| `proxyConfiguration` | Residential recommended for high volume |

#### Website input example

```json
{
  "mode": "website",
  "websites": ["amazon.com", "apify.com", "https://www.similarweb.com/website/shopify.com/"],
  "concurrency": 8
}
```

#### Ranking input example

```json
{
  "mode": "ranking",
  "categories": ["e-commerce-and-shopping", "ai-chatbots-and-tools"],
  "countries": ["united-states"],
  "includeGlobal": true
}
```

### Output

#### Website traffic fields

| Field | Description |
| :--- | :--- |
| `domain` / `title` / `description` / `category` | Site identity |
| `globalRank` / `countryRank` / `categoryRank` | Popularity ranks |
| `visits` / `bounceRate` / `pagesPerVisit` / `avgVisitDuration` | Engagement |
| `estimatedMonthlyVisits` | Last three snapshot months |
| `trafficSources` | Direct, organic/paid search, organic/paid social, referrals, mail, display, affiliate, **genAi** |
| `topCountries` | Traffic share with country names |
| `topKeywords` | Keyword, volume, CPC, estimated value |
| `aiTraffic` | Chatbot split, GenAI visit estimate, referral share |
| `competitors` | Similar sites when the public payload includes them |
| `snapshotDate` / `scrapedAt` | Similarweb month vs scrape time |

#### Website output example

```json
{
  "type": "website",
  "domain": "apify.com",
  "title": "Apify: Full-stack web scraping and data extraction platform",
  "category": "computers_electronics_and_technology/computers_electronics_and_technology",
  "globalRank": 9105,
  "countryRank": 3547,
  "countryRankCountryCode": "IN",
  "categoryRank": 117,
  "visits": 4490242,
  "bounceRate": 0.3587,
  "pagesPerVisit": 7.47,
  "avgVisitDurationFormatted": "00:05:10",
  "trafficSources": {
    "direct": 0.4288,
    "searchOrganic": 0.3495,
    "searchPaid": 0.0563,
    "genAi": 0.0316
  },
  "topCountries": [
    { "countryCode": "US", "countryName": "United States", "visitsShare": 0.2063 }
  ],
  "topKeywords": [
    { "name": "apify", "volume": 955500, "cpc": 0.98 }
  ],
  "aiTraffic": {
    "totalVisits": 125827,
    "referralShare": 0.028,
    "chatbots": [{ "name": "chatgpt.com", "share": 59.03 }]
  },
  "snapshotDate": "2026-07-01T00:00:00+00:00",
  "hasData": true
}
```

#### Top Websites ranking fields

| Field | Description |
| :--- | :--- |
| `rank` / `rankChange` / `isNewRank` | Position in the public list |
| `domain` / `categoryId` | Ranked site |
| `pagesPerVisit` / `bounceRate` / `avgVisitDurationFormatted` | Engagement on the list |
| `listCountry` / `listCategory` / `listUrl` | Which Top Websites list was scraped |

```json
{
  "type": "ranking",
  "rank": 1,
  "domain": "google.com",
  "categoryId": "computers_electronics_and_technology/search_engines",
  "rankChange": 0,
  "pagesPerVisit": 8.68,
  "bounceRate": 0.285,
  "avgVisitDurationFormatted": "00:09:58",
  "listUrl": "https://www.similarweb.com/top-websites/"
}
```

### Use cases

- **Competitor traffic analysis** — compare visits, sources, and ranks across a domain list
- **SEO channel mix** — organic vs paid vs social vs GenAI referral share
- **Website traffic checker** — bulk-enrich domains before outreach
- **Lead scoring** — attach global rank and monthly visits to CRM accounts
- **Market mapping** — scrape Top Websites for a category or country, then re-run `website` mode on the shortlist
- **AI search monitoring** — track ChatGPT / Claude / Perplexity share in `aiTraffic`

### Integration examples

#### Node.js

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('crawloop/similarweb-scraper').call({
  mode: 'website',
  websites: ['amazon.com', 'apify.com'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

#### Python

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("crawloop/similarweb-scraper").call(
    run_input={"mode": "website", "websites": ["amazon.com", "apify.com"]}
)
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item["domain"], item.get("visits"), item.get("globalRank"))
```

#### cURL

```bash
curl "https://api.apify.com/v2/acts/crawloop~similarweb-scraper/runs?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d "{\"mode\":\"website\",\"websites\":[\"amazon.com\",\"apify.com\"]}"
```

### MCP and AI assistants

Use this Actor from AI tools via [Apify MCP](https://docs.apify.com/platform/integrations/mcp).
Connect your Apify account, then call this Actor by its Store ID / name.

Example prompts:

- "Run Similarweb Scraper for amazon.com and apify.com and return visits, global rank, traffic sources, and AI referrals as JSON"
- "Scrape Similarweb Top Websites for e-commerce-and-shopping in the United States and list the top 20 domains"
- "Chain Similarweb Scraper then [Google Play Scraper](https://apify.com/crawloop/google-play-scraper) for the same brand's web vs app traffic"

### Suite next step

After website traffic, run [Google Play Scraper](https://apify.com/crawloop/google-play-scraper) or [App Stores Scraper](https://apify.com/crawloop/app-stores-scraper) for mobile ASO, or [Product Hunt Scraper](https://apify.com/crawloop/producthunt-scraper) for launch-day positioning.

### FAQ

**Is this a Similarweb API alternative?**
Yes for the public monthly snapshot — visits, ranks, sources, keywords, AI traffic, and Top Websites lists — without a Similarweb API key. It is not a replacement for paid multi-year REST/Batch exports.

**How do I scrape Similarweb with Python or Node.js?**
Call `crawloop/similarweb-scraper` with the Apify client (examples above) or from an MCP assistant. Results land in the default dataset; export JSON, CSV, or Excel from the run.

**Website mode vs ranking mode?**
`website` enriches the domains you pass. `ranking` scrapes public Top Websites lists (50 sites each). Use `both` to map a market, then enrich the shortlist.

**Do I need a Similarweb login?**
No.

**Why are some competitor / referral hostnames empty?**
Part of the HTML overview is paywalled. The public snapshot still returns visits, ranks, sources, keywords, and AI traffic.

**Does a zero-traffic domain fail the run?**
No. Other domains continue. Empty snapshots are marked `hasData: false`.

**Can I pass Similarweb URLs?**
Yes — `https://www.similarweb.com/website/example.com/` is normalized to `example.com`.

**How current is the data?**
Public snapshots are monthly. See `snapshotDate` on each website row.

### Related Actors

- [Product Hunt Scraper](https://apify.com/crawloop/producthunt-scraper)
- [Google Play Scraper](https://apify.com/crawloop/google-play-scraper)
- [App Stores Scraper](https://apify.com/crawloop/app-stores-scraper)
- [Trustpilot Scraper](https://apify.com/crawloop/trustpilot-scraper)

# Actor input Schema

## `mode` (type: `string`):

website = per-domain traffic snapshot (visits, ranks, sources, keywords, AI traffic). ranking = public Top Websites lists. both = run both in one job.

## `websites` (type: `array`):

Domains, website URLs, or Similarweb website URLs. Protocol and www are stripped. Used in website and both modes.

## `categories` (type: `array`):

Category slugs for ranking mode (e-commerce-and-shopping, finance, news-and-media, ai-chatbots-and-tools). 50 sites per list.

## `countries` (type: `array`):

Country slugs for ranking mode (united-states, germany, united-kingdom, japan). 50 sites per list. Combined with categories when both are set.

## `includeGlobal` (type: `boolean`):

Also scrape the worldwide Top Websites list. Default on when ranking mode has no categories or countries.

## `scrapeAllCategories` (type: `boolean`):

Ranking mode: discover all public category slugs from the Top Websites catalog and scrape each list (~50 sites each).

## `maxItems` (type: `integer`):

Hard cap on successful website + ranking rows. 0 = unlimited. Failed lookups do not count.

## `concurrency` (type: `integer`):

Parallel domain / list-page fetches.

## `proxyConfiguration` (type: `object`):

Residential proxies recommended for high-volume runs. Datacenter often works for the public data API.

## Actor input object example

```json
{
  "mode": "website",
  "websites": [
    "amazon.com",
    "apify.com",
    "https://www.similarweb.com/website/shopify.com/"
  ],
  "includeGlobal": false,
  "scrapeAllCategories": false,
  "maxItems": 0,
  "concurrency": 8,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Default dataset items.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "amazon.com",
        "apify.com",
        "https://www.similarweb.com/website/shopify.com/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("crawloop/similarweb-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": [
        "amazon.com",
        "apify.com",
        "https://www.similarweb.com/website/shopify.com/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("crawloop/similarweb-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "amazon.com",
    "apify.com",
    "https://www.similarweb.com/website/shopify.com/"
  ]
}' |
apify call crawloop/similarweb-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,crawloop/similarweb-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/hFK9GyE0nbRMblEdZ/builds/rErb1AZVCQaIQ0ois/openapi.json
