# Referring Domains Checker - Full Lists from Open Data (`crawlplant/referring-domains-checker`) Actor

Every domain that links to a site, strongest first, with its Open Authority 0-100: full referring-domain lists for any list of domains, from the Common Crawl web graph (133M domains, 2.1B links). No captchas, no blocked runs. Works as an MCP tool for AI agents.

- **URL**: https://apify.com/crawlplant/referring-domains-checker.md
- **Developed by:** [Piotr Zimniak](https://apify.com/crawlplant) (community)
- **Categories:** SEO tools, Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 referring domains

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Referring Domains Checker - Full Lists from Open Data

*Independent tool, not affiliated with Common Crawl, Ahrefs, Moz, Majestic or Semrush. It reads the public Common Crawl web
graph and never visits the domains you check.*

**Who links to a site?** Paste one domain or thousands and get **every domain that links to each of them, strongest
first**, each with its **Open Authority score from 0 to 100**. The lists come from the Common Crawl web graph: **133 million
domains and 2.1 billion links**, rebuilt every month. No captchas, no blocked runs, no fee per domain looked up.

### Why this one

- **Every referring domain in the graph.** apify.com has 3,556 referring domains in the September 2026 graph and you can
  get all of them; bbc.co.uk has 249,224. Set how many you want per domain.
- **Strongest first.** Rows are sorted by the linking domain's authority, so the top of the table is the links that matter.
  `minAuthority` drops the long tail.
- **Bulk.** Paste thousands of domains in one run: the graph is indexed on our side, so each domain's list comes back in
  milliseconds and nothing is scraped.
- **Nothing to block.** The data is an open dataset, not a scraped SEO tool: no captchas, no "try again later".
- **Any input format.** URLs, hostnames and e-mail addresses are reduced to the registrable domain
  (`https://www.bbc.co.uk/news` → `bbc.co.uk`, `info@scrapy.org` → `scrapy.org`).
- **See it live.** Every referring domain of apify.com and scrapy.org plotted by Open Authority, with the top 20 of each: [crawlplant.com/referring-domains-checker](https://crawlplant.com/referring-domains-checker/).

### What can you use it for?

- **Link building**: see who links to a competitor and pitch the same sites.
- **SEO audits**: list a site's strongest referring domains before a migration or a pitch.
- **Due diligence**: check whether a domain you're about to buy or partner with has real, established sites linking to it.
- **Research and AI agents**: referring-domain graphs for any list of domains, as JSON.

### Quick start

1. Click **Try for free** with the default input (apify.com, 100 strongest referring domains).
2. Replace **Domains** with your list.
3. Download the table as CSV, Excel or JSON, or save the input as a **task** and schedule it monthly: the graph is rebuilt
   about once a month.

#### Copy to your AI assistant

Paste this into ChatGPT, Claude or any agent so it knows how to use the Actor:

```
crawlplant/referring-domains-checker on Apify: the domains that link to each domain in a list, strongest first, from the
Common Crawl web graph (133M domains, 2.1B links). No scraping of SEO tools.
Input: domains (domains, URLs or e-mails; reduced to the registrable domain), maxReferrersPerDomain (default 100,
strongest first), minAuthority (0-100: keep only referring domains scoring at least this), maxItems (default 1000).
Row: domain, input, referringDomain, referringDomainAuthority (Open Authority 0-100), referringDomainRank
(harmonic-centrality rank, 1 = best linked), position (1 = strongest), referringDomains (total for the domain), ccRelease,
ccReleaseDate. Price: $2.00 per 1,000 rows on the Free plan; the default run (100 rows) about $0.20.
```

### Ready-to-use examples

**1. The 100 strongest referring domains of a site**

```json
{ "domains": ["apify.com"] }
```

**2. Only strong referring domains, for several sites**

```json
{ "domains": ["apify.com", "scrapy.org", "zyte.com"], "minAuthority": 60, "maxReferrersPerDomain": 1000 }
```

**3. Every referring domain of a site**

```json
{ "domains": ["scrapy.org"], "maxReferrersPerDomain": 100000, "maxItems": 100000 }
```

**4. Leads list: URLs and e-mail addresses work as input**

```json
{ "domains": ["https://www.bbc.co.uk/news", "info@scrapy.org"], "maxReferrersPerDomain": 20 }
```

**5. Outreach list: the 200 strongest referring domains of three competitors**

```json
{ "domains": ["zyte.com", "scrapingbee.com", "brightdata.com"], "maxReferrersPerDomain": 200 }
```

**6. Only top-tier referring domains (Open Authority 80 or more)**

```json
{ "domains": ["apify.com"], "minAuthority": 80, "maxReferrersPerDomain": 10000 }
```

### How to…

#### Find a competitor's backlinks (referring domains)

Put the competitor in `domains` (example 5). Every row is a site that links to it, strongest first: the top of the table is
where a link is worth the most. For the sites that link to competitors and not to you, use
[Link Gap Finder](https://apify.com/crawlplant/link-gap-finder).

#### Check referring domains in bulk

Paste the whole list (example 2). `maxReferrersPerDomain` caps the rows per domain and `maxItems` the whole run, so the
cost is known before you start.

#### Keep only strong referring domains

`minAuthority` drops everything below a score (example 6). 60 keeps roughly the top 150,000 domains of the web, 80 roughly the
top 5,000.

#### Get every referring domain of a site

Set `maxReferrersPerDomain` and `maxItems` above the site's total (example 3). Each row carries `referringDomains`, the
total, so you can see how far the list goes.

### Input options

| Option | Default | Description |
|---|---|---|
| `domains` | `apify.com` | Domains, URLs or e-mail addresses; reduced to the registrable domain, duplicates merged |
| `maxReferrersPerDomain` | `100` | Referring domains per domain, strongest first |
| `minAuthority` | `0` | Keep only referring domains with at least this Open Authority (0-100) |
| `maxItems` | `1000` | Maximum rows in total |

### Example output

A real row from September 2026:

```json
{
  "type": "referrer",
  "domain": "scrapy.org",
  "input": "scrapy.org",
  "referringDomain": "wikipedia.org",
  "referringDomainAuthority": 98,
  "referringDomainRank": 14,
  "position": 2,
  "referringDomains": 1556,
  "ccRelease": "cc-main-2026-jul-aug-sep",
  "ccReleaseDate": "2026-09-21",
  "scrapedAt": "2026-09-29T19:24:48.595Z",
  "source": "live"
}
```

### Output fields

| Field | Description |
|---|---|
| `domain`, `input` | The domain looked up (registrable, punycode for international names) and what you typed |
| `referringDomain` | A domain with at least one link to it |
| `referringDomainAuthority` | The referring domain's Open Authority, 0-100 |
| `referringDomainRank` | The referring domain's harmonic-centrality rank among all 133M domains (1 = best linked) |
| `position` | Its place in the list, 1 = strongest |
| `referringDomains` | How many domains link to `domain` in total |
| `ccRelease`, `ccReleaseDate` | The Common Crawl web graph release and its approximate date |
| `scrapedAt`, `source` | When the run read the data; always `live` |

### What is a referring domain here?

A domain with at least one link to yours in the pages of Common Crawl's last three monthly crawls. Common Crawl collects
billions of pages a month but not the whole web, so totals are lower than in commercial backlink indexes, and the same
method applies to every domain, so they compare fairly. Links are counted per domain: no individual URLs, anchors or
nofollow flags.

### How is Open Authority calculated?

`openAuthority = round(100 × (1 − (ln r / ln N)²))`, where `r` is the domain's harmonic-centrality rank in the Common Crawl
web graph and `N` the number of domains in it (133 million). Rank 100 scores 94, rank 5,000 scores 79, the top million 45.
A domain nobody links to scores 0.

### Pricing

Pay per result: platform usage is included. One price per row, and you pay only for the rows you get.

| Event | No discount (Free plan) | Bronze (Starter) | Silver (Scale) | Gold (Business) |
|---|---|---|---|---|
| Referring domain (per 1,000) | $2.00 | $1.80 | $1.60 | $1.40 |
| Actor start (per run) | $0.00005 | $0.00005 | $0.00005 | $0.00005 |

| Example on the Free plan | Rows | Cost |
|---|---|---|
| Default run: 100 referring domains of one site | 100 | ~$0.20 |
| 200 strongest referring domains of 3 competitors | 600 | ~$1.20 |
| Every referring domain of scrapy.org (1,556) | 1,556 | ~$3.11 |
| 20 strongest referring domains of 1,000 domains | 20,000 | ~$40 |

Set a **maximum cost per run** in the run options to stop a large run at your budget.

### Reliability

- There's nothing to block: the data comes from our index of the Common Crawl web graph, rebuilt when a new release comes
  out (about monthly). No captchas, no logins, no browser, no requests to the sites you check.
- A domain that isn't in the graph gets a warning in the run log and no rows, and you don't pay for it.
- Set a **maximum cost per run** in the run options to stop a large run at your budget.

### Run it through the API

JavaScript:

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('crawlplant/referring-domains-checker').call({ domains: ['apify.com'], maxReferrersPerDomain: 50 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((r) => `${r.referringDomain} (${r.referringDomainAuthority})`));
```

Python:

```python
from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("crawlplant/referring-domains-checker").call(run_input={"domains": ["apify.com"], "maxReferrersPerDomain": 50})
for r in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(r["referringDomain"], r["referringDomainAuthority"])
```

### Use with AI agents

The Actor works as a tool in the [Apify MCP server](https://mcp.apify.com) for Claude, ChatGPT, Cursor and others:

```
https://mcp.apify.com?tools=crawlplant/referring-domains-checker
```

Try *"which strong sites link to example.com?"*.

### More from the same data

- [Bulk Domain Authority Checker](https://apify.com/crawlplant/commoncrawl-domain-metrics): authority, referring-domain
  counts, Chrome traffic bucket and rank history for any list of domains.
- [Link Gap Finder](https://apify.com/crawlplant/link-gap-finder): sites that link to your competitors but not to you.
- [Similar Sites Finder](https://apify.com/crawlplant/similar-sites-finder): competitors and alternatives from the link graph.

### FAQ

#### How fresh is the data?

Each run reads the newest Common Crawl web graph release (about monthly, each covering three monthly crawls). `ccRelease`
says which one.

#### Does it visit my site?

No. It only reads the public web graph; the sites you check never see a request.

#### Is it legal?

It reads the Common Crawl web graph under the [Common Crawl terms of use](https://commoncrawl.org/terms-of-use): link
statistics between domains, no page content and no personal data. Check the terms that apply to your own use of the results.

### Sources and credits

Common Crawl web graph, [commoncrawl.org](https://commoncrawl.org/web-graphs), used under the Common Crawl terms of use.

### Privacy

Only public link statistics about domains are read; no personal data is collected. Each run sends the developer anonymous
feature-usage statistics (the options used, never the domains); your Apify account id is replaced by a one-way hash.

# Actor input Schema

## `domains` (type: `array`):

Domains, URLs or e-mail addresses, one per line. Each is reduced to its registrable domain (https://www.bbc.co.uk/news -> bbc.co.uk). You get the domains that link to each of them.

## `maxReferrersPerDomain` (type: `integer`):

How many referring domains to return for each of your domains, strongest (highest Open Authority) first. A well-known site can have millions; most small sites have a handful.

## `minAuthority` (type: `integer`):

Return only referring domains scoring at least this (0-100). 0 = every referring domain.

## `maxItems` (type: `integer`):

Maximum rows to return (and pay for).

## Actor input object example

```json
{
  "domains": [
    "apify.com"
  ],
  "maxReferrersPerDomain": 100,
  "minAuthority": 0,
  "maxItems": 1000
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("crawlplant/referring-domains-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("crawlplant/referring-domains-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call crawlplant/referring-domains-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,crawlplant/referring-domains-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XkhaxODTzGnsX5bhm/builds/gVCJbEeahA6cZMK3Z/openapi.json
