# Similar Sites Finder - Competitors from the Link Graph (`crawlplant/similar-sites-finder`) Actor

Sites like any domain, found from who links to them: sites the same websites link to are usually in the same niche. Competitor and alternative discovery for any list of domains, from the Common Crawl web graph (133M domains). Works as an MCP tool for AI agents.

- **URL**: https://apify.com/crawlplant/similar-sites-finder.md
- **Developed by:** [Piotr Zimniak](https://apify.com/crawlplant) (community)
- **Categories:** SEO tools, Lead generation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.50 / 1,000 similar sites

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Similar Sites Finder - Competitors from the Link Graph

*Independent tool, not affiliated with Common Crawl, Similarweb, Semrush or Ahrefs. It reads the public Common Crawl web
graph and never visits the domains you check.*

**Sites like any domain, found from who links to them.** When the same websites link to two sites, the two are usually in
the same niche: reviews, "best tools" lists, directories and articles cite competitors together. Paste one domain or a
list and get the **most similar sites for each, most similar first**, with their **Open Authority 0-100**. Built on the
Common Crawl web graph: **133 million domains and 2.1 billion links**.

### Real results (September 2026 graph)

The first eight results, in order:

| Domain | Most similar sites |
|---|---|
| apify.com | brightdata.com, zyte.com, scrapingbee.com, scraperapi.com, octoparse.com, firecrawl.dev, parsehub.com, serpapi.com |
| ahrefs.com | semrush.com, g2.com, backlinko.com, web.dev, moz.com, screamingfrog.co.uk, seroundtable.com, searchenginejournal.com |
| scrapy.org | crummy.com, toscrape.com, parsehub.com, scrapinghub.com, selenium.dev, python-requests.org, octoparse.com, zyte.com |
| booking.com | avisworld.com, dragonpass.com, loungekey.com, expediapartnercentral.com, rentalcars.com, priceless.com, agoda.com, sixt.com |
| zalando.de | oo34.net, flaconi.de, tracdelight.com, douglas.de, aboutyou.de, fashionid.de, otto.de, tracdelight.io |

### Why this one

- **Competitors the way the web sees them.** Similarity comes from real links between sites, not from keywords or
  categories, so it finds alternatives, direct competitors and close neighbours in any language or country.
- **Big general sites don't crowd the list.** Search engines, social networks and link directories link to everything; they
  are weighted down or left out, so the list stays about your niche.
- **Fast and in bulk.** About a tenth of a second per domain: the graph is indexed on our side, nothing is scraped.
- **Nothing to block.** An open dataset: no captchas, no failed runs.
- **See it live.** Market maps around apify.com, ahrefs.com, duolingo.com and glossier.com: [crawlplant.com/similar-sites-finder](https://crawlplant.com/similar-sites-finder/).

### What can you use it for?

- **Competitor research**: find the players in a market you don't know yet, then compare them with the other Actors below.
- **Link building**: feed the similar sites into [Link Gap Finder](https://apify.com/crawlplant/link-gap-finder) to see who
  links to them and not to you.
- **Lead generation and market maps**: expand a seed list of companies into their whole niche.
- **AI agents**: "what are the alternatives to X?" answered from link data.

### Quick start

1. Click **Try for free** with the default input (sites similar to apify.com).
2. Replace **Domains** with your own list.
3. Download the table as CSV, Excel or JSON.

#### Copy to your AI assistant

Paste this into ChatGPT, Claude or any agent so it knows how to use the Actor:

```
crawlplant/similar-sites-finder on Apify: the sites most similar to each domain in a list (competitors, alternatives),
found from the websites that link to both, in the Common Crawl web graph (133M domains, 2.1B links).
Input: domains (domains or URLs), maxReferrersPerDomain (default 50: similar sites per domain, up to 1,000),
minAuthority (0-100), maxItems (default 500).
Row: domain, input, similarDomain, similarity (higher = more alike), sharedReferrers, referrersCompared,
similarDomainAuthority (Open Authority 0-100), similarDomainRank, position (1 = most similar), ccRelease.
Price: $5.00 per 1,000 rows on the Free plan; the default run (50 rows) about $0.25.
```

### Ready-to-use examples

**1. The 50 sites most similar to a domain**

```json
{ "domains": ["apify.com"] }
```

**2. Competitor lists for several companies at once**

```json
{ "domains": ["ahrefs.com", "booking.com", "zalando.de"], "maxReferrersPerDomain": 20 }
```

**3. Only established similar sites (Open Authority 50 or more)**

```json
{ "domains": ["scrapy.org"], "minAuthority": 50 }
```

**4. A long list for a market map**

```json
{ "domains": ["https://www.zalando.de/"], "maxReferrersPerDomain": 500, "maxItems": 500 }
```

**5. Alternatives to a SaaS tool, top 10**

```json
{ "domains": ["ahrefs.com"], "maxReferrersPerDomain": 10 }
```

**6. Expand a seed list of companies into their niche**

```json
{ "domains": ["booking.com", "expedia.com", "airbnb.com"], "maxReferrersPerDomain": 30, "minAuthority": 40 }
```

### How to…

#### Find competitors of a website

Put the site in `domains` (example 1). The top rows are the sites most often linked from the same pages: usually direct
competitors and close alternatives.

#### Find alternatives to a tool or service

Run the tool's domain (example 5): review sites and "best tools" lists link to the alternatives together, so they come
out on top.

#### Build a market map from a few seed companies

Paste several companies of one market (example 6). Each gets its own list; merge them and count how often a site appears.

#### Turn competitors into link-building prospects

Paste the similar sites into [Link Gap Finder](https://apify.com/crawlplant/link-gap-finder) as competitors to get the
sites that link to them and not to you.

### Input options

| Option | Default | Description |
|---|---|---|
| `domains` | `apify.com` | Domains, URLs or e-mail addresses; reduced to the registrable domain |
| `maxReferrersPerDomain` | `50` | Similar sites per domain, most similar first (up to 1,000) |
| `minAuthority` | `0` | Keep only similar sites with at least this Open Authority (0-100) |
| `maxItems` | `500` | Maximum rows in total |

### Example output

A real row from September 2026:

```json
{
  "type": "similar",
  "domain": "apify.com",
  "input": "apify.com",
  "similarDomain": "brightdata.com",
  "similarDomainAuthority": 71,
  "similarDomainRank": 22413,
  "position": 1,
  "sharedReferrers": 61,
  "referrersCompared": 300,
  "similarity": 1.313,
  "ccRelease": "cc-main-2026-jul-aug-sep",
  "ccReleaseDate": "2026-09-21",
  "scrapedAt": "2026-09-29T20:51:31.478Z",
  "source": "live"
}
```

### Output fields

| Field | Description |
|---|---|
| `domain`, `input` | The domain looked up and what you typed |
| `similarDomain` | A site similar to it |
| `similarity` | Similarity score: higher = more of its links come from the same sites (comparable within one domain's list) |
| `sharedReferrers`, `referrersCompared` | How many of the compared sites linking to `domain` also link to the similar site |
| `similarDomainAuthority`, `similarDomainRank` | Its Open Authority (0-100) and harmonic-centrality rank among all 133M domains |
| `position` | Its place in the list, 1 = most similar |
| `ccRelease`, `ccReleaseDate` | The Common Crawl web graph release and its approximate date |
| `scrapedAt`, `source` | When the run read the data; always `live` |

### How does it work?

The Actor takes up to 300 of the strongest sites linking to your domain that link to at most 300 domains each (so
directories and portals that link to everything don't count), collects the sites they link to, and ranks those by how
often they are cited together with yours relative to how many sites link to them in total (a cosine similarity over
links). A domain needs a few dozen linking sites for a good list; for very small sites the list can be short or empty.

### Pricing

Pay per result: platform usage is included. One price per row, and you pay only for the rows you get.

| Event | No discount (Free plan) | Bronze (Starter) | Silver (Scale) | Gold (Business) |
|---|---|---|---|---|
| Similar site (per 1,000) | $5.00 | $4.50 | $4.00 | $3.50 |
| Actor start (per run) | $0.00005 | $0.00005 | $0.00005 | $0.00005 |

| Example on the Free plan | Rows | Cost |
|---|---|---|
| Default run: 50 similar sites of one domain | 50 | ~$0.25 |
| Top 10 alternatives of a tool | 10 | ~$0.05 |
| 30 similar sites for each of 100 companies | 3,000 | ~$15 |

Set a **maximum cost per run** in the run options to stop a large run at your budget.

### Reliability

- There's nothing to block: the data comes from our index of the Common Crawl web graph, rebuilt when a new release comes
  out (about monthly). No captchas, no logins, no browser, no requests to the sites you check.
- A domain that isn't in the graph gets a warning in the run log and no rows, and you don't pay for it.
- Set a **maximum cost per run** in the run options to stop a large run at your budget.

### Run it through the API

JavaScript:

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('crawlplant/similar-sites-finder').call({ domains: ['apify.com'], maxReferrersPerDomain: 20 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((r) => `${r.similarDomain} (${r.similarity})`));
```

Python:

```python
from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("crawlplant/similar-sites-finder").call(run_input={"domains": ["apify.com"], "maxReferrersPerDomain": 20})
for r in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(r["similarDomain"], r["similarity"], r["sharedReferrers"])
```

### Use with AI agents

Works as a tool in the [Apify MCP server](https://mcp.apify.com):

```
https://mcp.apify.com?tools=crawlplant/similar-sites-finder
```

Try *"what are the main competitors of example.com?"*.

### More from the same data

- [Link Gap Finder](https://apify.com/crawlplant/link-gap-finder): sites that link to your competitors but not to you.
- [Referring Domains Checker](https://apify.com/crawlplant/referring-domains-checker): every domain that links to a site.
- [Bulk Domain Authority Checker](https://apify.com/crawlplant/commoncrawl-domain-metrics): authority, referring-domain
  counts and rank history for any list of domains.

### FAQ

#### How fresh is the data?

The newest Common Crawl web graph release (about monthly). `ccRelease` says which one.

#### What about the biggest platforms?

For sites that almost everyone links to (facebook.com, google.com, youtube.com), sharing linking sites says little, so
their similar sites are generic. The Actor works best for companies, products, publications and niche sites.

#### Why is a known competitor missing?

Similarity needs sites that link to both. Two sites in the same market that nobody cites together, or a very new site
with few links, won't show up. Try a larger `maxReferrersPerDomain` or run the competitor itself.

#### Is it legal?

It reads the Common Crawl web graph under the [Common Crawl terms of use](https://commoncrawl.org/terms-of-use): link
statistics between domains, no page content and no personal data.

### Sources and credits

Common Crawl web graph, [commoncrawl.org](https://commoncrawl.org/web-graphs), used under the Common Crawl terms of use.

### Privacy

Only public link statistics about domains are read; no personal data is collected. Each run sends the developer anonymous
feature-usage statistics (the options used, never the domains); your Apify account id is replaced by a one-way hash.

# Actor input Schema

## `domains` (type: `array`):

Domains, URLs or e-mail addresses, one per line. You get the sites most like each of them.

## `maxReferrersPerDomain` (type: `integer`):

How many similar sites to return for each domain, most similar first (up to 1,000).

## `minAuthority` (type: `integer`):

Return only similar sites scoring at least this (0-100). 0 = all of them.

## `maxItems` (type: `integer`):

Maximum rows to return (and pay for).

## Actor input object example

```json
{
  "domains": [
    "apify.com"
  ],
  "maxReferrersPerDomain": 50,
  "minAuthority": 0,
  "maxItems": 500
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("crawlplant/similar-sites-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("crawlplant/similar-sites-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call crawlplant/similar-sites-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,crawlplant/similar-sites-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JsnGhQVMK6lPQ9Rdp/builds/yPifojhwc6RHAgPvH/openapi.json
