Link Gap Finder - Sites Linking to Competitors, Not You avatar

Link Gap Finder - Sites Linking to Competitors, Not You

Pricing

from $2.80 / 1,000 link gap prospects

Go to Apify Store
Link Gap Finder - Sites Linking to Competitors, Not You

Link Gap Finder - Sites Linking to Competitors, Not You

Link intersect from open data: sites that link to your competitors but not to you, the ones linking to the most competitors first, with Open Authority 0-100. Up to 100,000 referring domains per competitor from the Common Crawl web graph. Works as an MCP tool for AI agents.

Pricing

from $2.80 / 1,000 link gap prospects

Rating

0.0

(0)

Developer

Piotr Zimniak

Piotr Zimniak

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

Independent tool, not affiliated with Common Crawl, Ahrefs, Moz, Majestic or Semrush. It reads the public Common Crawl web graph and never visits the domains you check.

Your outreach list in one run. Give your site and up to 20 competitors: you get every site that links to a competitor but not to you, the ones linking to the most competitors first, each with its Open Authority score from 0 to 100. It works like a link intersect tool, on full referring-domain lists from the Common Crawl web graph (133 million domains, 2.1 billion links).

Why this one

  • Deep lists on both sides. Up to 100,000 of each competitor's strongest referring domains are compared with all of yours, so you choose how deep the outreach list goes.
  • Best prospects on top. A site that links to three of your competitors already writes about your niche: rows are sorted by how many competitors a site links to, then by its authority.
  • No noise from hubs. CDNs, link shorteners, hosts of user content and platforms that link to more than 30,000 domains are left out by default: they link to everyone and are not a realistic pitch.
  • Seconds, not hours. The graph is indexed on our side; a gap over three competitors is computed in milliseconds.
  • Nothing to block. An open dataset, not a scraped SEO tool: no captchas, no failed runs.
  • See it live. Real link gaps, prospect by prospect, with the competitors each one links to: crawlplant.com/link-gap-finder.

Real example (September 2026 graph)

scrapy.org against apify.com, zyte.com and scrapingbee.com, comparing their full lists: 3,935 sites link to at least one of them and not to scrapy.org. At the top, zapier.com, techradar.com and n8n.io link to all three, then linktr.ee, substack.com and herokuapp.com to two.

What can you use it for?

  • Link-building outreach: a ready prospect list of sites that already link to your niche.
  • Competitor analysis: see where your competitors get links that you don't.
  • SEO agencies: a link gap report per client in one run, as CSV or JSON.
  • AI agents: an outreach agent calls it through MCP, gets the prospects and drafts the pitches.

Quick start

  1. Click Try for free with the default input (scrapy.org against three competitors).
  2. Put your site in Your domain and your competitors in Competitors.
  3. Download the table as CSV or Excel and start outreach from the top.

Copy to your AI assistant

Paste this into ChatGPT, Claude or any agent so it knows how to use the Actor:

crawlplant/link-gap-finder on Apify: the sites that link to your competitors but not to your site (link intersect), from
the Common Crawl web graph (133M domains, 2.1B links). No scraping of SEO tools.
Input: yourDomain (your site), domains (up to 20 competitors), maxReferrersPerDomain (default 1000: strongest referring
domains compared per competitor, up to 100,000), minAuthority (0-100), excludeHubs (default true: leave out CDNs,
shorteners and sites linking to 30,000+ domains), maxItems (default 500).
Row: domain (yours), referringDomain, referringDomainAuthority (Open Authority 0-100), referringDomainRank,
competitorsLinked, linksTo (competitor domains), competitorsCompared, ccRelease. Sorted by competitorsLinked, then
authority. Price: $4.00 per 1,000 rows on the Free plan; the default run (500 rows) about $2.00.

Ready-to-use examples

1. Link gap against three competitors

{ "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com", "scrapingbee.com"] }

2. Strong sites only (Open Authority 50 or more)

{ "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com"], "minAuthority": 50 }

3. Deep comparison of big competitors

{ "yourDomain": "my-shop.com", "domains": ["competitor-one.com", "competitor-two.com"], "maxReferrersPerDomain": 100000, "maxItems": 5000 }

4. URLs work as input

{ "yourDomain": "https://www.my-site.com/", "domains": ["https://www.competitor.com/pricing", "https://blog.other.io/"] }

5. The best prospects only: the top 100 against five competitors

{ "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com", "scrapingbee.com", "brightdata.com", "octoparse.com"], "maxItems": 100 }

6. A small outreach list to start with

{ "yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com"], "minAuthority": 60, "maxItems": 50 }

How to…

Put your site in yourDomain and your competitors in domains (example 1). The top rows link to the most competitors: they already write about your niche and are the easiest pitch.

Rows are sorted by competitorsLinked, then by authority: with five competitors (example 5) the first rows link to the most of them; maxItems keeps only the top of the list.

Keep only strong sites

minAuthority drops every site below a score (example 6). 60 keeps roughly the top 150,000 domains of the web.

Why are big brands' results full of big sites?

Link gap works best against competitors of your own size. Against the largest brands (Shopify, Notion, HubSpot), the sites linking to several of them are mostly large media, universities and platforms, because those link to every big brand. Pick competitors that play in your league, or use minAuthority together with a lower maxItems to keep a short list.

Don't know your competitors yet?

Similar Sites Finder lists the sites most like yours from the same link data; paste its results here.

Input options

OptionDefaultDescription
yourDomainscrapy.orgYour site: its referring domains are left out
domains3 competitorsUp to 20 competitor domains or URLs
maxReferrersPerDomain1000Strongest referring domains compared per competitor (up to 100,000)
minAuthority0Keep only sites with at least this Open Authority (0-100)
excludeHubstrueLeave out infrastructure (CDNs, hosts of user content, link shorteners) and sites linking to more than 30,000 domains
maxItems500Maximum rows

Example output

A real row from September 2026:

{
"type": "linkGap",
"domain": "scrapy.org",
"referringDomain": "zapier.com",
"referringDomainAuthority": 88,
"referringDomainRank": 586,
"competitorsLinked": 3,
"linksTo": ["apify.com", "zyte.com", "scrapingbee.com"],
"competitorsCompared": 3,
"ccRelease": "cc-main-2026-jul-aug-sep",
"ccReleaseDate": "2026-09-21",
"scrapedAt": "2026-09-29T20:51:33.977Z",
"source": "live"
}

Output fields

FieldDescription
domainYour domain
referringDomainA site that links to at least one competitor and not to you
referringDomainAuthority, referringDomainRankIts Open Authority (0-100) and harmonic-centrality rank among all 133M domains
competitorsLinked, linksToHow many and which of your competitors it links to
competitorsComparedHow many competitors were compared
ccRelease, ccReleaseDateThe Common Crawl web graph release and its approximate date
scrapedAt, sourceWhen the run read the data; always live

A link from any page of the site to the domain in Common Crawl's last three monthly crawls, counted once per pair of domains. Common Crawl collects billions of pages a month but not the whole web, so a missing link here can exist on a page it didn't crawl. The same method applies to every domain, so the comparison is fair.

Pricing

Pay per result: platform usage is included. One price per row, and you pay only for the rows you get.

EventNo discount (Free plan)Bronze (Starter)Silver (Scale)Gold (Business)
Link gap prospect (per 1,000)$4.00$3.60$3.20$2.80
Actor start (per run)$0.00005$0.00005$0.00005$0.00005
Example on the Free planRowsCost
Default run: 500 prospects against 3 competitors500~$2.00
Top 100 prospects against 5 competitors100~$0.40
5,000 prospects, deep comparison5,000~$20

You pay per row returned, not per competitor compared. Set a maximum cost per run to stop at your budget.

Reliability

  • There's nothing to block: the data comes from our index of the Common Crawl web graph, rebuilt when a new release comes out (about monthly). No captchas, no logins, no browser, no requests to the sites you check.
  • A domain that isn't in the graph gets a warning in the run log and no rows, and you don't pay for it.
  • Set a maximum cost per run in the run options to stop a large run at your budget.

Run it through the API

JavaScript:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('crawlplant/link-gap-finder').call({ yourDomain: 'scrapy.org', domains: ['apify.com', 'zyte.com'], maxItems: 100 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((r) => `${r.referringDomain} -> ${r.linksTo.join(", ")}`));

Python:

from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("crawlplant/link-gap-finder").call(run_input={"yourDomain": "scrapy.org", "domains": ["apify.com", "zyte.com"], "maxItems": 100})
for r in client.dataset(run["defaultDatasetId"]).iterate_items():
print(r["referringDomain"], r["competitorsLinked"], r["linksTo"])

Use with AI agents

Works as a tool in the Apify MCP server:

https://mcp.apify.com?tools=crawlplant/link-gap-finder

Try "which sites link to my three competitors but not to me?".

More from the same data

FAQ

How is Open Authority calculated?

round(100 × (1 − (ln r / ln N)²)) from the domain's harmonic-centrality rank r among the N domains of the Common Crawl web graph: rank 100 scores 94, the top million 45.

How fresh is the data?

The newest Common Crawl web graph release (about monthly). ccRelease says which one.

It reads the Common Crawl web graph under the Common Crawl terms of use: link statistics between domains, no page content and no personal data.

Sources and credits

Common Crawl web graph, commoncrawl.org, used under the Common Crawl terms of use.

Privacy

Only public link statistics about domains are read; no personal data is collected. Each run sends the developer anonymous feature-usage statistics (the options used, never the domains); your Apify account id is replaced by a one-way hash.