# Bulk Domain Authority & Referring Domains - Open Data (`crawlplant/commoncrawl-domain-metrics`) Actor

Open Authority 0-100 and referring domains for any list of domains, plus referrer lists, link gap, similar sites and rank history to 2018. From the Common Crawl web graph (133M domains), Majestic and Chrome UX Report. No blocked runs; works as an MCP tool for AI agents.

- **URL**: https://apify.com/crawlplant/commoncrawl-domain-metrics.md
- **Developed by:** [Piotr Zimniak](https://apify.com/crawlplant) (community)
- **Categories:** SEO tools, Lead generation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 domain profiles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk Domain Authority & Referring Domains - Open Data

*Independent tool, not affiliated with Common Crawl, Majestic, Google, Tranco, Ahrefs, Moz or Semrush. It reads only public open
datasets and never visits the domains you check.*

Authority, referring domains and real-traffic rank for **any list of domains**, from the open web graph instead of scraped
SEO tools. Every domain gets an **Open Authority score from 0 to 100**, computed from its rank among the **133 million domains
and 2.1 billion links** of the Common Crawl web graph, its **referring-domain count** and how strong those referring domains
are. Other modes give **the domains that link to it** (`referrers`), **the sites that link to your competitors but not to
you** (`linkGap`), **sites like it** found from who links to them (`similar`), **the domains it links to** (`outlinks`)
and **referring domains newly seen or no longer seen since the previous monthly graph** (`changes`). Next to the score you
get **Majestic Million** referring subnets and IPs and the **Chrome UX Report** traffic bucket, a popularity rank based on
real Chrome visits. Paste 3 domains or 10,000.
**Rank history goes back to 2018.**

### Why this one

- **No captchas, no scraping.** The data comes from public datasets, not from Ahrefs, Moz or Semrush pages. Nothing to solve,
  nothing to retry, no "try again later".
- **Who links to you.** `referrers` mode lists the referring domains of each domain in your list, strongest first, each with
  its own Open Authority: 3,556 for apify.com, 249,224 for bbc.co.uk (September 2026 graph).
- **Link gap, from full lists.** `linkGap` compares the referring domains of up to 20 competitors with yours and returns
  the ones you're missing, the domains linking to the most competitors first. Up to 100,000 of each competitor's
  strongest referring domains are compared, so the list is as deep as you need.
- **Find competitors you didn't know.** `similar` finds the sites that the same websites link to: usually the same niche.
  Feed them straight into `linkGap`.
- **Trend at a glance.** `openAuthorityChange3m` and `openAuthorityChange12m` show whether a domain is gaining or losing
  authority, from our monthly history of the graph.
- **How good the links are, not just how many.** `ccReferringAuthoritySpread` splits a domain's referring domains into Open
  Authority bands (90-100 ... 0-9), so a profile of a few strong sites and one of thousands of weak ones look different.
- **Links and real traffic in one row.** You see the link-graph rank (Common Crawl), links from distinct networks (Majestic) and
  a Chrome-traffic bucket (CrUX) side by side. That shows sites that have links but no visitors, and sites that have visitors
  but few links. In a real run over 1,000 domains sampled across the Tranco top million, 399 were in the Majestic Million
  and 459 in the Chrome top million (September 2026).
- **Fast in bulk.** Ranks and referring-domain counts come from an index of the whole graph: every domain gets an exact
  rank, however small, and 1,000 domains take about a second on our side.
- **History back to 2018.** History mode returns the domain's rank in every Common Crawl web graph release you ask for, with a
  top-percent figure that stays comparable when the graph grows or shrinks.
- **A score you can check.** Open Authority is one published formula on a public rank ([below](#how-is-open-authority-calculated)).
- **Handles any input format.** URLs, hostnames, e-mail addresses and international domain names are all reduced to the registrable
  domain (`https://www.bbc.co.uk/news` → `bbc.co.uk`, `info@scrapy.org` → `scrapy.org`).
- **See it live.** Where well-known domains sit on the web, the biggest 12-month climbers and a referring-domain profile, from this month's graph: [crawlplant.com/commoncrawl-domain-metrics](https://crawlplant.com/commoncrawl-domain-metrics/).

### What can you use it for?

- **Link building**: rank a prospect list by authority and drop the long tail with `minAuthority`; get an outreach list of
  strong sites that link to your competitors and not to you (`linkGap`).
- **Lead qualification**: check whether a company's website has real traffic and links before sales spends time on it.
- **SEO audits and competitor comparisons**: authority, link popularity and traffic bucket for your site and its rivals.
- **Research**: popularity of domain lists, web graph studies (harmonic centrality, PageRank), longitudinal rank data.
- **AI agents and RAG**: weight or filter sources by domain authority before citing them.

### Quick start

1. Click **Try for free** (or **Start**) with the default input: three well-known domains, done in about 10 seconds, cost under
   $0.01 on the Free plan.
2. Replace **Domains** with your list (one per line; URLs and e-mail addresses work).
3. Download the table as CSV, Excel or JSON, or save the input as a **task** to rerun it.

#### Copy to your AI assistant

Paste this into ChatGPT, Claude or any agent so it knows how to use the Actor:

```
crawlplant/commoncrawl-domain-metrics on Apify: bulk domain authority from open data (no scraping of SEO tools).
Input: domains (list of domains, URLs or e-mails; reduced to the registrable domain), mode ("profiles" default: one row per
domain; "history": one row per domain per Common Crawl web graph release, back to 2018; "referrers": one row per domain
linking to each of your domains, strongest first, up to maxReferrersPerDomain (default 100); "linkGap": domains linking to
the competitors in `domains` but not to `yourDomain`, most competitors first, comparing each competitor's strongest
maxReferrersPerDomain referring domains, hubs left out unless excludeHubs is false; "similar": sites like each domain by shared linking sites; "outlinks": domains
each domain links to; "changes": referring domains gained and lost since the previous monthly graph), sources (default
["commoncrawl","majestic","crux"]; "tranco" adds the Tranco rank and, in history mode, its ~30-day daily history),
graphScanDepth (only when the ranks file is read directly: history mode or an older graphRelease), minAuthority (0-100 filter;
in referrers mode it filters the referring domains),
sortBy ("input" | "authority"), historyReleases (default 6), graphRelease (release id, empty = newest), maxItems.
Profile row: domain, openAuthority (0-100, from the Common Crawl harmonic-centrality rank r among N domains:
round(100*(1-(ln r/ln N)^2))), openAuthorityMax (upper bound when the domain is below the searched depth), authorityTier
(top-100 ... top-100m, long-tail, below-scan-depth, not-in-graph), ccHarmonicRank, ccHarmonicCentrality, ccPageRankRank,
ccPageRank, ccTopPercent, ccHosts, ccReferringDomains (distinct domains linking to it in the graph),
ccReferringAuthoritySpread (referring domains per Open Authority band, "90-100" ... "0-9"), openAuthorityChange3m and
openAuthorityChange12m (score now minus 3 / 12 monthly releases ago), ccRelease, majesticRank, majesticRefSubnets, majesticRefIps (+ previous day), cruxRankBucket
(1000 = among the 1,000 most visited sites in Chrome ... 1000000), cruxOrigin, trancoRank. History row: domain, dataset
(commoncrawl | tranco), date, rank, topPercent, openAuthority. Referrer row: domain, referringDomain, referringDomainAuthority,
referringDomainRank, position (1 = strongest), referringDomains (total), ccRelease. Link-gap row: domain (yours),
referringDomain, referringDomainAuthority, competitorsLinked, linksTo (competitor domains), ccRelease. Similar row: domain,
similarDomain, similarity, sharedReferrers, similarDomainAuthority. Outlink row: domain, linkedDomain, linkedDomainAuthority,
position. Change row: domain, change (gained | lost), referringDomain, referringDomainAuthority, gainedTotal, lostTotal,
stableTotal, previousRelease. Default run ~5 s.
```

### Modes

| Mode | What you get |
|---|---|
| `profiles` (default) | One row per domain: Open Authority, Common Crawl ranks, Majestic links, Chrome traffic bucket (and Tranco if chosen) |
| `history` | One row per domain per Common Crawl web graph release (newest first, back to 2018), plus Tranco daily ranks when `tranco` is a source |
| `referrers` | One row per domain linking to each of your domains, strongest first, with its Open Authority |
| `linkGap` | One row per domain that links to your competitors but not to you, the ones linking to the most competitors first |
| `similar` | One row per site like each of your domains, most similar first (sites the same websites link to) |
| `outlinks` | One row per domain each of your domains links to, strongest first |
| `changes` | One row per referring domain gained or lost since the previous monthly graph |

### Ready-to-use examples

**1. Check domain authority for a list of domains**

```json
{ "domains": ["apify.com", "scrapy.org", "commoncrawl.org", "bbc.co.uk"] }
```

**2. Qualify a lead list of e-mail addresses and website URLs**

```json
{ "domains": ["info@scrapy.org", "https://www.bbc.co.uk/news", "sales@apify.com"] }
```

**3. Keep only link-building prospects with Open Authority 40 or more, strongest first**

```json
{ "domains": ["apify.com", "scrapy.org", "commoncrawl.org", "example.com"], "minAuthority": 40, "sortBy": "authority" }
```

**4. Who links to a site: its 100 strongest referring domains**

```json
{ "mode": "referrers", "domains": ["apify.com"] }
```

**5. Referring domains of competitors, only those with Open Authority 60 or more**

```json
{ "mode": "referrers", "domains": ["ahrefs.com", "semrush.com"], "minAuthority": 60, "maxReferrersPerDomain": 1000 }
```

**6. Link popularity only: Majestic referring subnets and IPs**

```json
{ "domains": ["apify.com", "scrapy.org"], "sources": ["majestic"] }
```

**7. Does this site get real visitors? Chrome traffic bucket only**

```json
{ "domains": ["apify.com", "scrapy.org"], "sources": ["crux"] }
```

**8. Every source, including the Tranco research ranking**

```json
{ "domains": ["apify.com", "github.com"], "sources": ["commoncrawl", "majestic", "crux", "tranco"] }
```

**9. Authority history over the last 12 graph releases**

```json
{ "mode": "history", "domains": ["apify.com", "scrapy.org"], "historyReleases": 12 }
```

**10. Authority history since 2018**

```json
{ "mode": "history", "domains": ["bbc.co.uk"], "historyReleases": 60 }
```

**11. Daily Tranco rank for the last month**

```json
{ "mode": "history", "domains": ["apify.com"], "sources": ["tranco"] }
```

**12. Profiles from an older graph release (read from the ranks file, searched to `graphScanDepth`)**

```json
{ "domains": ["apify.com", "scrapy.org"], "graphRelease": "cc-main-2026-jun-jul-aug" }
```

**13. Compare competitors, sorted by authority**

```json
{ "domains": ["ahrefs.com", "semrush.com", "moz.com", "majestic.com", "similarweb.com"], "sortBy": "authority" }
```

**14. Link gap: sites that link to your competitors and not to you**

```json
{ "mode": "linkGap", "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com", "scrapingbee.com"] }
```

**15. Link gap, strong sites only, deeper comparison**

```json
{ "mode": "linkGap", "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com"], "minAuthority": 50, "maxReferrersPerDomain": 20000 }
```

**16. Find sites like yours (competitors and alternatives)**

```json
{ "mode": "similar", "domains": ["apify.com"], "maxReferrersPerDomain": 30 }
```

**17. Which domains does a site link to?**

```json
{ "mode": "outlinks", "domains": ["scrapy.org"], "maxReferrersPerDomain": 200 }
```

**18. Referring domains gained and lost since last month**

```json
{ "mode": "changes", "domains": ["apify.com", "scrapy.org"], "minAuthority": 30 }
```

### How to…

#### Check domain authority in bulk

Paste the list into `domains` (example 1). Every domain in the Common Crawl graph gets its exact rank, score and
referring-domain count; a domain the graph doesn't have gets `ccStatus: "not-in-graph"` and a score of 0.

#### Find the sites that link to a domain

Use `mode: "referrers"` (examples 4 and 5). You get one row per referring domain, strongest first, with its Open Authority
and link rank. `maxReferrersPerDomain` sets how many per domain (default 100) and `minAuthority` keeps only the strong ones.
A referring domain is a domain with at least one link to yours in the pages of Common Crawl's last three monthly crawls.

#### Filter link-building prospects by authority

Set `minAuthority` (for example 40) and `sortBy: "authority"` (example 3). Domains under the minimum are left out, and you don't
pay for them. `minAuthority: 1` drops only the domains without a score.

#### Find link-building opportunities (link gap)

Use `mode: "linkGap"` with your site in `yourDomain` and up to 20 competitors in `domains` (examples 14 and 15). Each row is
a domain that links to at least one competitor and not to you; `competitorsLinked` and `linksTo` say which. Rows come
sorted by the number of competitors, then by authority, so the top of the table is the best outreach list. Hubs are left
out by default (`excludeHubs`): infrastructure such as CDNs and link shorteners, and platforms that link to more than 30,000
domains; they link to everyone and are rarely a realistic pitch. Each
competitor's strongest `maxReferrersPerDomain` referring domains are compared (default 100; raise it for a deeper list,
up to 100,000), and `minAuthority` keeps only strong sites.

#### Find sites like a domain (competitor discovery)

Use `mode: "similar"` (example 16). Two sites that the same websites link to are usually in the same niche: the Actor
takes the strongest sites linking to your domain, looks at what else they link to and ranks the domains they share by
`similarity`. `sharedReferrers` says how many of the compared linking sites (`referrersCompared`) link to both. Big
general sites (search engines, social networks) are weighted down and infrastructure is left out. The result is a ready
competitor list for `linkGap`. It works best for companies and niche sites; for the biggest platforms (facebook.com,
google.com) almost every site links to them, so their similar sites are generic.

#### See which domains a site links to

Use `mode: "outlinks"` (example 17): the domains a site links to, strongest first, with their Open Authority.

#### See referring domains gained and lost since last month

Use `mode: "changes"` (example 18): one row per referring domain that appears in the newest monthly graph and not in the
previous one (`gained`), or the other way round (`lost`), strongest first, with the domain's totals (`gainedTotal`,
`lostTotal`, `stableTotal`). Each monthly graph is built from a different sample of pages, so a typical list turns over by
15-20% from one month to the next without any link being added or removed: read a change as "seen / no longer seen in
this month's crawl". The strong domains at the top of the lists and the totals over several months are the useful
signal; `minAuthority` keeps only those. Lists of up to 1,000,000 referring domains are compared; a bigger site (facebook.com has
17 million) gets a note in the run log instead of rows, and nothing is charged for it.

#### Judge the quality of a link profile

`ccReferringAuthoritySpread` counts a domain's referring domains per Open Authority band. Two sites with 5,000 referring
domains each can differ a lot: one with hundreds in the 60-100 bands has links from established sites, one with nearly
all of them in 0-9 has mostly a long tail of small or unlinked sites.

#### Check if a website gets real traffic

`cruxRankBucket` comes from the Chrome UX Report, based on real Chrome users. 1000 means the site is among the 1,000 most
visited sites in Chrome worldwide, then 5000, 10000, 50000, 100000, 500000 and 1000000. Empty means it is not in the top
million. Compare it with the link metrics: many links but no traffic bucket often means a link network, not a real audience.

#### Track a domain's authority over time

Profiles carry `openAuthorityChange3m` and `openAuthorityChange12m`: the score now minus 3 and 12 monthly releases ago.
For the full series use `mode: "history"` (examples 9 and 10): about a year of releases comes from our index for every
domain in seconds, and older releases back to 2018 are read from Common Crawl's files. A new web graph comes out about
every month, each covering the last three monthly crawls. Compare `topPercent` across releases rather than `rank`: the
number of domains in the graph changes between releases (from 70 to 209 million since 2018).

#### Weight RAG sources by domain authority

Run the domains of your retrieved sources and use `openAuthority` or `authorityTier` as a trust weight. The rows come back in
seconds, so the Actor works inside an agent loop (see [Use with AI agents](#use-with-ai-agents)).

#### Get the Common Crawl harmonic centrality and PageRank of a domain

`ccHarmonicRank` / `ccHarmonicCentrality` and `ccPageRankRank` / `ccPageRank` are Common Crawl's own values for the domain in
the chosen release, the same numbers as in its published ranks file, with no extra processing.

### Input options

| Option | Default | Description |
|---|---|---|
| `mode` | `profiles` | `profiles`, `history`, `referrers`, `linkGap`, `similar`, `outlinks` or `changes` |
| `domains` | 3 examples | Domains, URLs or e-mail addresses; reduced to the registrable domain, duplicates merged. linkGap: your competitors (up to 20) |
| `yourDomain` | - | linkGap: your site; its referring domains are left out |
| `excludeHubs` | `true` | linkGap: leave out infrastructure (CDNs, hosts of user content, link shorteners) and sites linking to more than 30,000 domains |
| `sources` | `commoncrawl`, `majestic`, `crux` | Add `tranco` for the Tranco rank (and daily history in history mode) |
| `graphScanDepth` | `10000000` | History mode and older releases: how many of the best-ranked domains to search in the ranks file; `0` = all 133M |
| `graphRelease` | newest | A release id such as `cc-main-2026-jun-jul-aug` ([list](https://index.commoncrawl.org/graphinfo.json)) |
| `minAuthority` | `0` | Profiles: keep domains scoring at least this; `1` drops domains without a score. Link modes: keep the listed domains scoring at least this |
| `maxReferrersPerDomain` | `100` | Rows per domain in referrers, similar, outlinks and changes (changes: per list), strongest first. linkGap: strongest referring domains compared per competitor (up to 100,000) |
| `sortBy` | `input` | `input` or `authority` (highest first) |
| `historyReleases` | `6` | History: how many graph releases, newest first |
| `maxItems` | `10000` | Maximum rows |

### Example output

A real row from September 2026:

```json
{
  "type": "profile",
  "domain": "apify.com",
  "input": "apify.com",
  "openAuthority": 79,
  "openAuthorityMax": null,
  "authorityTier": "top-10k",
  "ccStatus": "found",
  "ccHarmonicRank": 5287,
  "ccHarmonicCentrality": 17206752,
  "ccPageRankRank": 19402,
  "ccPageRank": 0.00000163412,
  "ccTopPercent": 0.003968,
  "ccHosts": 29,
  "ccReferringDomains": 3556,
  "ccScannedTo": null,
  "ccGraphDomains": 133241980,
  "ccRelease": "cc-main-2026-jul-aug-sep",
  "ccReleaseDate": "2026-09-21",
  "majesticRank": 16349,
  "majesticTldRank": 7895,
  "majesticRefSubnets": 2582,
  "majesticRefIps": 5805,
  "majesticPrevRank": 16436,
  "majesticPrevRefSubnets": 2571,
  "majesticPrevRefIps": 5754,
  "cruxRankBucket": 100000,
  "cruxOrigin": "https://apify.com",
  "scrapedAt": "2026-09-29T16:51:35.073Z",
  "source": "live"
}
```

### Output fields

| Field | Description |
|---|---|
| `type` | `profile` or `history` |
| `domain`, `input` | Registrable domain looked up (punycode for international names) and what you typed |
| `openAuthority` | 0-100 score from the Common Crawl link-graph rank; `null` when the domain is below the searched depth |
| `openAuthorityMax` | For domains below the searched depth: the highest score they could have |
| `authorityTier` | `top-100`, `top-1k`, `top-10k`, `top-100k`, `top-1m`, `top-10m`, `top-100m`, `long-tail`, `no-inlinks`, `below-scan-depth`, `not-in-graph`, `unavailable` |
| `ccStatus` | `found`, `below-scan-depth` (not in the searched top N), `not-in-graph` (whole graph searched), `unavailable` |
| `ccHarmonicRank`, `ccHarmonicCentrality` | Rank and value of the domain's harmonic centrality: how close the rest of the web is to it through links |
| `ccPageRankRank`, `ccPageRank` | Rank and value of its PageRank in the same graph |
| `ccTopPercent` | Rank as a share of all domains in the graph (0.004 = top 0.004%) |
| `ccHosts` | Hostnames of the domain seen in the graph |
| `ccReferringDomains` | Distinct domains linking to it in the graph release |
| `ccReferringAuthoritySpread` | Its referring domains per Open Authority band: `90-100`, `80-89`, ... `0-9` |
| `openAuthorityChange3m`, `openAuthorityChange12m` | Open Authority now minus 3 / 12 monthly graph releases ago; null when the domain wasn't in that graph |
| `ccScannedTo` | For domains below the depth: how deep the search went |
| `ccGraphDomains`, `ccRelease`, `ccReleaseDate` | Size, id and approximate date (last crawl week) of the graph release |
| `majesticRank`, `majesticTldRank` | Position in the Majestic Million, overall and within its top-level domain |
| `majesticRefSubnets`, `majesticRefIps` | Distinct IP subnets (class C) and IP addresses of the sites linking to it |
| `majesticPrevRank`, `majesticPrevRefSubnets`, `majesticPrevRefIps` | The same the day before |
| `cruxRankBucket`, `cruxOrigin` | Chrome UX Report popularity bucket and the domain's best-ranked origin |
| `trancoRank`, `trancoListId`, `trancoListDate` | With `tranco`: rank in the latest daily Tranco list (top 1M) |
| `dataset`, `period`, `date`, `rank`, `topPercent`, `graphDomains` | History rows: source, release id or day, date, rank and top percent |
| `referringDomain`, `referringDomainAuthority`, `referringDomainRank`, `position`, `referringDomains` | Referrer rows: the linking domain, its Open Authority and link rank, its place in the list (1 = strongest) and the total count |
| `competitorsLinked`, `linksTo`, `competitorsCompared` | Link-gap rows: how many and which of your competitors the domain links to, out of how many compared (`domain` is yours) |
| `similarDomain`, `similarity`, `sharedReferrers`, `referrersCompared`, `similarDomainAuthority` | Similar rows: the similar site, its similarity score, how many of the compared linking sites link to both, and its Open Authority |
| `linkedDomain`, `linkedDomainAuthority`, `linkedDomainRank`, `linkedDomains` | Outlink rows: a domain it links to, its score and rank, and how many domains it links to in total |
| `change`, `stillInGraph`, `gainedTotal`, `lostTotal`, `stableTotal`, `previousRelease` | Change rows: `gained` or `lost` since `previousRelease`, whether a lost domain is still in the graph, and the domain's totals |
| `scrapedAt`, `source` | When the run read the data; always `live` |

Every field is `null` when the domain is not in that source's list (Majestic, CrUX and Tranco list the top million).

### How is Open Authority calculated?

`openAuthority = round(100 × (1 − (ln r / ln N)²))`, where `r` is the domain's harmonic-centrality rank in the Common Crawl
web graph and `N` the number of domains in it. A domain nobody links to scores 0.

| Rank in the graph | Open Authority |
|---|---|
| 1 | 100 |
| 100 | 94 |
| 5,000 | 79 |
| 100,000 | 62 |
| 1,000,000 | 45 |
| 10,000,000 | 26 |
| 50,000,000 | 10 |

Harmonic centrality measures how close the rest of the web is to a domain through links, so a link from a well-connected site
counts for more than many links from isolated ones. Common Crawl sorts its published domain ranks by it. Beyond the top ~20
million domains PageRank values are nearly all equal, which is why the score uses the harmonic rank.

### Run it through the API

JavaScript:

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('crawlplant/commoncrawl-domain-metrics').call({ domains: ['apify.com', 'scrapy.org'], sortBy: 'authority' });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((r) => `${r.domain}: ${r.openAuthority}`));
```

Python:

```python
from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("crawlplant/commoncrawl-domain-metrics").call(run_input={"domains": ["apify.com", "bbc.co.uk"]})
for r in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(r["domain"], r["openAuthority"], r["majesticRefSubnets"], r["cruxRankBucket"])
```

### Use with AI agents

The Actor works as a tool in the [Apify MCP server](https://mcp.apify.com), for Claude, ChatGPT, Cursor, VS Code, LangChain and
others. Every field is described. Server URL with only this tool:

```
https://mcp.apify.com?tools=crawlplant/commoncrawl-domain-metrics
```

Try *"how authoritative are these five domains?"*, *"which of my leads have websites with real traffic?"* or *"has
example.com gained or lost authority since last year?"* or *"which strong sites link to my competitor?"*.

### Pricing

Pay per result: platform usage is included. Each row type has one price, and you pay only for the rows you get: domains
left out by `minAuthority` are free.

| Event | No discount (Free plan) | Bronze (Starter) | Silver (Scale) | Gold (Business) |
|---|---|---|---|---|
| Domain profile (per 1,000) | $1.00 | $0.90 | $0.80 | $0.70 |
| Rank history row (per 1,000) | $1.50 | $1.35 | $1.20 | $1.05 |
| Link row: referring domain, outlink or change (per 1,000) | $2.00 | $1.80 | $1.60 | $1.40 |
| Link gap prospect (per 1,000) | $4.00 | $3.60 | $3.20 | $2.80 |
| Similar site (per 1,000) | $5.00 | $4.50 | $4.00 | $3.50 |
| Actor start (per run) | $0.00005 | $0.00005 | $0.00005 | $0.00005 |

| Example on the Free plan | Rows | Cost |
|---|---|---|
| Default run: 3 well-known domains | 3 | ~$0.003 |
| 1,000 domains across the web | 1,000 | ~$1.00 |
| 100 strongest referring domains of one site | 100 | ~$0.20 |
| Link gap, 500 outreach prospects | 500 | ~$2.00 |
| 50 sites similar to one domain | 50 | ~$0.25 |
| History: 10 domains × 6 releases | 60 | ~$0.09 |

Set a **maximum cost per run** in the run options to stop a large run at your budget.

### Reliability

- There's nothing to block: the link data comes from our index of the Common Crawl web graph, rebuilt when a new release
  comes out, and the other sources are public dataset files (Majestic, GitHub, Tranco). No captchas, no logins, no browser.
- If the index can't be reached, profiles read Common Crawl's ranks file directly (down to `graphScanDepth`, without
  referring-domain counts) and the run summary says so.
- Each file is streamed from the top and the download stops as soon as every domain on your list is found. A dropped
  connection resumes where it stopped.
- If one source can't be read, the run still finishes with the other sources, and the run summary (`OUTPUT` in the key-value
  store) says which source was missing and why.
- The run respects its time limit: a very deep search stops in time and returns what it found, with a note.

### Data sources and freshness

| Source | What it measures | Coverage | Updated |
|---|---|---|---|
| [Common Crawl web graph](https://commoncrawl.org/web-graphs) | Links between 133M domains (harmonic centrality, PageRank, referring domains) | every linked domain | a new release about monthly, each covering 3 crawls; archive back to 2018 |
| [Majestic Million](https://majestic.com/reports/majestic-million) | Referring subnets and IPs | top 1M | daily |
| [Chrome UX Report top lists](https://github.com/zakird/crux-top-lists) | Real Chrome visits, as a rank bucket | top 1M sites | monthly |
| [Tranco](https://tranco-list.eu/) (optional) | Research ranking combining several popularity lists | top 1M | daily; the API keeps about 30 days per domain |

Every run reads the newest data available at that moment. `ccRelease`, `ccReleaseDate` and `trancoListDate` say which
version you got.

### Troubleshooting

- **`below-scan-depth` for a domain I know**: only when the ranks file is read directly (history mode, an older
  `graphRelease`): the domain ranks below the searched part (default: top 10 million). Raise `graphScanDepth`, or set it to `0`.
- **A subdomain shows its parent domain**: the graph ranks registrable domains, so `blog.example.com` is `example.com`.
  Hosting platforms count as one domain too (`name.github.io` is `github.io`), as they do in Common Crawl.
- **Majestic, CrUX or Tranco fields are empty**: those lists cover the top million sites only. An empty field means the domain
  isn't in that list, not an error.
- **An input was skipped**: IP addresses and text that isn't a domain are listed in the run summary.
- **History has fewer releases than asked**: a domain has rows only for releases where it ranked within `graphScanDepth`, and
  history searches older releases only for domains found in the newest one.

### FAQ

#### What is a good domain authority score?

On Open Authority, 45 is the top 1 million domains of 133 million, 62 the top 100,000 and 79 the top 5,000. Half of all
domains in the graph score 7 or less.

#### Is Open Authority the same as Domain Rating or Domain Authority?

No. It is our own score from Common Crawl's public link graph, not Ahrefs DR or Moz DA, and the numbers are not meant to match.
All of them grow with links from well-linked sites. Open Authority is free of any vendor's crawler and can be reproduced from
public data.

#### How much does it cost to check 10,000 domains?

About $10 on the Free plan ($1.00 per 1,000 domain profiles). Link rows cost $2.00 per 1,000 (referring domains, outlinks,
changes), link gap prospects $4.00 and similar sites $5.00 per 1,000. See [Pricing](#pricing).

#### Can it count backlinks or referring domains?

Yes, at the domain level: `ccReferringDomains` counts the distinct domains linking to each domain, and `referrers` mode lists
them, strongest first. The count comes from the pages of Common Crawl's last three monthly crawls, so it compares domains
fairly within one release; commercial backlink indexes crawl more pages and report higher totals. Majestic's referring
subnets and IPs are there too for the top million domains. It doesn't list individual backlink URLs or anchor texts.

#### How fresh is the data?

Each run reads the newest files: the Majestic Million from today, the newest Common Crawl web graph (about monthly) and the
monthly Chrome UX Report list.

#### Does it visit or crawl my domains?

No. It only reads public datasets, so the sites you check never see a request.

#### Is it legal?

The Actor reads public datasets under their terms: Common Crawl's terms of use, Majestic Million under CC BY 3.0 and the Chrome
UX Report under Creative Commons Attribution. The optional Tranco source combines several lists, including Cloudflare Radar
(CC BY-NC 4.0). The sources are credited below. Check the terms that apply to your own use of the results.

### Limits

- Referring domains are counted per registrable domain, from the newest graph release only (not in history mode).
- Majestic, CrUX and Tranco cover the top million domains or sites each.
- Tranco daily history is read for up to 50 domains per run (its API allows one request per second).
- One row per registrable domain: subdomains are merged into their domain.

### Sources and credits

- Common Crawl web graph, [commoncrawl.org](https://commoncrawl.org/), used under the
  [Common Crawl terms of use](https://commoncrawl.org/terms-of-use).
- [The Majestic Million](https://majestic.com/reports/majestic-million) by Majestic, licensed under CC BY 3.0.
- Chrome UX Report data by Google, via [zakird/crux-top-lists](https://github.com/zakird/crux-top-lists).
- Tranco (optional source, off by default): V. Le Pochat et al., "Tranco: A Research-Oriented Top Sites Ranking Hardened
  Against Manipulation", NDSS 2019, [tranco-list.eu](https://tranco-list.eu/). Its default list combines Chrome UX Report,
  Cloudflare Radar (CC BY-NC 4.0), Farsight, Majestic and Cisco Umbrella rankings.

### Privacy

Only public datasets about domains are read; no personal data is collected. Each run sends the developer anonymous feature-usage
statistics (the mode and which options were used, never the domains); your Apify account id is replaced by a one-way hash on
arrival.

# Actor input Schema

## `mode` (type: `string`):

profiles = one row per domain with every metric (Open Authority score, Common Crawl ranks, referring-domain count, Majestic links, Chrome traffic bucket); history = one row per domain per Common Crawl graph release (back to 2018), plus Tranco daily ranks when Tranco is a source; referrers = the domains that link to each of your domains, strongest first, one row per linking domain. linkGap = domains that link to your competitors (in "domains") but not to your site ("yourDomain"), most competitors first. similar = sites like each domain (the same websites link to them); outlinks = the domains each domain links to; changes = referring domains gained and lost since the previous monthly graph.

## `domains` (type: `array`):

Domains, URLs or e-mail addresses, one per line. Each is reduced to its registrable domain: https://www.bbc.co.uk/news -> bbc.co.uk, info@scrapy.org -> scrapy.org. Duplicates are merged. In linkGap mode: your competitors (up to 20).

## `yourDomain` (type: `string`):

linkGap mode: your site. Its referring domains are left out, so you get the sites that link to your competitors and not to you.

## `excludeHubs` (type: `boolean`):

linkGap mode: leave out infrastructure (CDNs, hosts of user content, link shorteners) and sites that link to more than 30,000 domains, such as the biggest platforms. They link to everyone, so they are rarely outreach prospects.

## `sources` (type: `array`):

commoncrawl = link-graph ranks of 133M domains and the Open Authority score; majestic = Majestic Million rank and referring subnets / IPs (top 1M); crux = Chrome UX Report traffic bucket (top 1M sites by real Chrome visits); tranco = Tranco research ranking (top 1M) and its daily history in history mode.

## `graphScanDepth` (type: `integer`):

How many of the best-ranked domains of the Common Crawl ranks file to search when it is read directly: in history mode, for an older graphRelease, or if our link index can't be reached (profiles normally get an exact rank for every domain from it). 10,000,000 (default) covers every domain with an Open Authority of 26 or more; domains below it get openAuthorityMax instead of a score. 0 = the whole graph (133M domains, 2.5 GB, a few minutes).

## `graphRelease` (type: `string`):

A web graph release id, e.g. cc-main-2026-jun-jul-aug (list: https://index.commoncrawl.org/graphinfo.json). Empty = the newest. In history mode, the history starts here and goes back.

## `minAuthority` (type: `integer`):

Profiles: return only domains scoring at least this (0-100); 1 drops domains without a score, so you don't pay for them. Link modes (referrers, linkGap, similar, outlinks, changes): return only the listed domains scoring at least this. 0 = everything.

## `sortBy` (type: `string`):

Profiles mode: input = your order; authority = highest Open Authority first (domains without a score last).

## `historyReleases` (type: `integer`):

History mode: how many Common Crawl graph releases to read, newest first (a new release covers the last three monthly crawls). 6 is about half a year of recent releases; the archive goes back to early 2018 (about 50 releases).

## `maxReferrersPerDomain` (type: `integer`):

How many rows per domain in referrers, similar, outlinks and changes (changes: gained and lost each), strongest first. A well-known site can have millions of referring domains; most small sites have a handful. linkGap mode: how many of each competitor's strongest referring domains to compare (up to 100,000).

## `maxItems` (type: `integer`):

Maximum rows to return (and pay for).

## Actor input object example

```json
{
  "mode": "profiles",
  "domains": [
    "apify.com",
    "github.com",
    "wikipedia.org"
  ],
  "excludeHubs": true,
  "sources": [
    "commoncrawl",
    "majestic",
    "crux"
  ],
  "graphScanDepth": 10000000,
  "minAuthority": 0,
  "sortBy": "input",
  "historyReleases": 6,
  "maxReferrersPerDomain": 100,
  "maxItems": 10000
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("crawlplant/commoncrawl-domain-metrics").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("crawlplant/commoncrawl-domain-metrics").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call crawlplant/commoncrawl-domain-metrics --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,crawlplant/commoncrawl-domain-metrics"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/HjB8e45EfLodlpPCh/builds/Q0ZZG8ldN9Bd8UF7v/openapi.json
