# Link Gap Finder - Sites Linking to Competitors, Not You (`crawlplant/link-gap-finder`) Actor

Link intersect from open data: sites that link to your competitors but not to you, the ones linking to the most competitors first, with Open Authority 0-100. Up to 100,000 referring domains per competitor from the Common Crawl web graph. Works as an MCP tool for AI agents.

- **URL**: https://apify.com/crawlplant/link-gap-finder.md
- **Developed by:** [Piotr Zimniak](https://apify.com/crawlplant) (community)
- **Categories:** SEO tools, Lead generation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.80 / 1,000 link gap prospects

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Link Gap Finder - Sites Linking to Competitors, Not You

*Independent tool, not affiliated with Common Crawl, Ahrefs, Moz, Majestic or Semrush. It reads the public Common Crawl web
graph and never visits the domains you check.*

**Your outreach list in one run.** Give your site and up to 20 competitors: you get **every site that links to a
competitor but not to you**, the ones linking to the most competitors first, each with its **Open Authority score from 0
to 100**. It works like a link intersect tool, on full referring-domain lists from the Common Crawl web graph (**133
million domains, 2.1 billion links**).

### Why this one

- **Deep lists on both sides.** Up to 100,000 of each competitor's strongest referring domains are compared with all of
  yours, so you choose how deep the outreach list goes.
- **Best prospects on top.** A site that links to three of your competitors already writes about your niche: rows are
  sorted by how many competitors a site links to, then by its authority.
- **No noise from hubs.** CDNs, link shorteners, hosts of user content and platforms that link to more than 30,000 domains
  are left out by default: they link to everyone and are not a realistic pitch.
- **Seconds, not hours.** The graph is indexed on our side; a gap over three competitors is computed in milliseconds.
- **Nothing to block.** An open dataset, not a scraped SEO tool: no captchas, no failed runs.
- **See it live.** Real link gaps, prospect by prospect, with the competitors each one links to: [crawlplant.com/link-gap-finder](https://crawlplant.com/link-gap-finder/).

### Real example (September 2026 graph)

scrapy.org against apify.com, zyte.com and scrapingbee.com, comparing their full lists: **3,935 sites** link to at least
one of them and not to scrapy.org. At the top, zapier.com, techradar.com and n8n.io link to all three, then linktr.ee, substack.com and
herokuapp.com to two.

### What can you use it for?

- **Link-building outreach**: a ready prospect list of sites that already link to your niche.
- **Competitor analysis**: see where your competitors get links that you don't.
- **SEO agencies**: a link gap report per client in one run, as CSV or JSON.
- **AI agents**: an outreach agent calls it through MCP, gets the prospects and drafts the pitches.

### Quick start

1. Click **Try for free** with the default input (scrapy.org against three competitors).
2. Put your site in **Your domain** and your competitors in **Competitors**.
3. Download the table as CSV or Excel and start outreach from the top.

#### Copy to your AI assistant

Paste this into ChatGPT, Claude or any agent so it knows how to use the Actor:

```
crawlplant/link-gap-finder on Apify: the sites that link to your competitors but not to your site (link intersect), from
the Common Crawl web graph (133M domains, 2.1B links). No scraping of SEO tools.
Input: yourDomain (your site), domains (up to 20 competitors), maxReferrersPerDomain (default 1000: strongest referring
domains compared per competitor, up to 100,000), minAuthority (0-100), excludeHubs (default true: leave out CDNs,
shorteners and sites linking to 30,000+ domains), maxItems (default 500).
Row: domain (yours), referringDomain, referringDomainAuthority (Open Authority 0-100), referringDomainRank,
competitorsLinked, linksTo (competitor domains), competitorsCompared, ccRelease. Sorted by competitorsLinked, then
authority. Price: $4.00 per 1,000 rows on the Free plan; the default run (500 rows) about $2.00.
```

### Ready-to-use examples

**1. Link gap against three competitors**

```json
{ "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com", "scrapingbee.com"] }
```

**2. Strong sites only (Open Authority 50 or more)**

```json
{ "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com"], "minAuthority": 50 }
```

**3. Deep comparison of big competitors**

```json
{ "yourDomain": "my-shop.com", "domains": ["competitor-one.com", "competitor-two.com"], "maxReferrersPerDomain": 100000, "maxItems": 5000 }
```

**4. URLs work as input**

```json
{ "yourDomain": "https://www.my-site.com/", "domains": ["https://www.competitor.com/pricing", "https://blog.other.io/"] }
```

**5. The best prospects only: the top 100 against five competitors**

```json
{ "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com", "scrapingbee.com", "brightdata.com", "octoparse.com"], "maxItems": 100 }
```

**6. A small outreach list to start with**

```json
{ "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com"], "minAuthority": 60, "maxItems": 50 }
```

### How to…

#### Find link-building opportunities from competitors

Put your site in `yourDomain` and your competitors in `domains` (example 1). The top rows link to the most competitors:
they already write about your niche and are the easiest pitch.

#### Find sites that link to several competitors

Rows are sorted by `competitorsLinked`, then by authority: with five competitors (example 5) the first rows link to the most of them;
`maxItems` keeps only the top of the list.

#### Keep only strong sites

`minAuthority` drops every site below a score (example 6). 60 keeps roughly the top 150,000 domains of the web.

#### Why are big brands' results full of big sites?

Link gap works best against competitors of your own size. Against the largest brands (Shopify, Notion, HubSpot), the sites
linking to several of them are mostly large media, universities and platforms, because those link to every big brand.
Pick competitors that play in your league, or use `minAuthority` together with a lower `maxItems` to keep a short list.

#### Don't know your competitors yet?

[Similar Sites Finder](https://apify.com/crawlplant/similar-sites-finder) lists the sites most like yours from the same
link data; paste its results here.

### Input options

| Option | Default | Description |
|---|---|---|
| `yourDomain` | `scrapy.org` | Your site: its referring domains are left out |
| `domains` | 3 competitors | Up to 20 competitor domains or URLs |
| `maxReferrersPerDomain` | `1000` | Strongest referring domains compared per competitor (up to 100,000) |
| `minAuthority` | `0` | Keep only sites with at least this Open Authority (0-100) |
| `excludeHubs` | `true` | Leave out infrastructure (CDNs, hosts of user content, link shorteners) and sites linking to more than 30,000 domains |
| `maxItems` | `500` | Maximum rows |

### Example output

A real row from September 2026:

```json
{
  "type": "linkGap",
  "domain": "scrapy.org",
  "referringDomain": "zapier.com",
  "referringDomainAuthority": 88,
  "referringDomainRank": 586,
  "competitorsLinked": 3,
  "linksTo": ["apify.com", "zyte.com", "scrapingbee.com"],
  "competitorsCompared": 3,
  "ccRelease": "cc-main-2026-jul-aug-sep",
  "ccReleaseDate": "2026-09-21",
  "scrapedAt": "2026-09-29T20:51:33.977Z",
  "source": "live"
}
```

### Output fields

| Field | Description |
|---|---|
| `domain` | Your domain |
| `referringDomain` | A site that links to at least one competitor and not to you |
| `referringDomainAuthority`, `referringDomainRank` | Its Open Authority (0-100) and harmonic-centrality rank among all 133M domains |
| `competitorsLinked`, `linksTo` | How many and which of your competitors it links to |
| `competitorsCompared` | How many competitors were compared |
| `ccRelease`, `ccReleaseDate` | The Common Crawl web graph release and its approximate date |
| `scrapedAt`, `source` | When the run read the data; always `live` |

### What counts as a link here?

A link from any page of the site to the domain in Common Crawl's last three monthly crawls, counted once per pair of
domains. Common Crawl collects billions of pages a month but not the whole web, so a missing link here can exist on a
page it didn't crawl. The same method applies to every domain, so the comparison is fair.

### Pricing

Pay per result: platform usage is included. One price per row, and you pay only for the rows you get.

| Event | No discount (Free plan) | Bronze (Starter) | Silver (Scale) | Gold (Business) |
|---|---|---|---|---|
| Link gap prospect (per 1,000) | $4.00 | $3.60 | $3.20 | $2.80 |
| Actor start (per run) | $0.00005 | $0.00005 | $0.00005 | $0.00005 |

| Example on the Free plan | Rows | Cost |
|---|---|---|
| Default run: 500 prospects against 3 competitors | 500 | ~$2.00 |
| Top 100 prospects against 5 competitors | 100 | ~$0.40 |
| 5,000 prospects, deep comparison | 5,000 | ~$20 |

You pay per row returned, not per competitor compared. Set a **maximum cost per run** to stop at your budget.

### Reliability

- There's nothing to block: the data comes from our index of the Common Crawl web graph, rebuilt when a new release comes
  out (about monthly). No captchas, no logins, no browser, no requests to the sites you check.
- A domain that isn't in the graph gets a warning in the run log and no rows, and you don't pay for it.
- Set a **maximum cost per run** in the run options to stop a large run at your budget.

### Run it through the API

JavaScript:

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('crawlplant/link-gap-finder').call({ yourDomain: 'scrapy.org', domains: ['apify.com', 'zyte.com'], maxItems: 100 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((r) => `${r.referringDomain} -> ${r.linksTo.join(", ")}`));
```

Python:

```python
from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("crawlplant/link-gap-finder").call(run_input={"yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com"], "maxItems": 100})
for r in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(r["referringDomain"], r["competitorsLinked"], r["linksTo"])
```

### Use with AI agents

Works as a tool in the [Apify MCP server](https://mcp.apify.com):

```
https://mcp.apify.com?tools=crawlplant/link-gap-finder
```

Try *"which sites link to my three competitors but not to me?"*.

### More from the same data

- [Referring Domains Checker](https://apify.com/crawlplant/referring-domains-checker): every domain that links to a site.
- [Similar Sites Finder](https://apify.com/crawlplant/similar-sites-finder): find your competitors from the link graph.
- [Bulk Domain Authority Checker](https://apify.com/crawlplant/commoncrawl-domain-metrics): authority, referring-domain
  counts and rank history for any list of domains.

### FAQ

#### How is Open Authority calculated?

`round(100 × (1 − (ln r / ln N)²))` from the domain's harmonic-centrality rank `r` among the `N` domains of the Common Crawl
web graph: rank 100 scores 94, the top million 45.

#### How fresh is the data?

The newest Common Crawl web graph release (about monthly). `ccRelease` says which one.

#### Is it legal?

It reads the Common Crawl web graph under the [Common Crawl terms of use](https://commoncrawl.org/terms-of-use): link
statistics between domains, no page content and no personal data.

### Sources and credits

Common Crawl web graph, [commoncrawl.org](https://commoncrawl.org/web-graphs), used under the Common Crawl terms of use.

### Privacy

Only public link statistics about domains are read; no personal data is collected. Each run sends the developer anonymous
feature-usage statistics (the options used, never the domains); your Apify account id is replaced by a one-way hash.

# Actor input Schema

## `yourDomain` (type: `string`):

Your site. Its referring domains are left out, so you get the sites that link to your competitors and not to you.

## `domains` (type: `array`):

Up to 20 competitor domains or URLs, one per line.

## `maxReferrersPerDomain` (type: `integer`):

How many of each competitor's strongest referring domains to compare with yours (up to 100,000). More = a longer, deeper list.

## `minAuthority` (type: `integer`):

Return only sites scoring at least this (0-100). 0 = every site.

## `excludeHubs` (type: `boolean`):

Leave out infrastructure (CDNs, hosts of user content, link shorteners) and sites that link to more than 30,000 domains, such as the biggest platforms. They link to everyone, so they are rarely outreach prospects.

## `maxItems` (type: `integer`):

Maximum rows to return (and pay for).

## Actor input object example

```json
{
  "yourDomain": "scrapy.org",
  "domains": [
    "apify.com",
    "zyte.com",
    "scrapingbee.com"
  ],
  "maxReferrersPerDomain": 1000,
  "minAuthority": 0,
  "excludeHubs": true,
  "maxItems": 500
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("crawlplant/link-gap-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("crawlplant/link-gap-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call crawlplant/link-gap-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,crawlplant/link-gap-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/6Gedf0pEZ6Ep8fgLT/builds/4Msayt4IogMz7cCgf/openapi.json
