# Meta & Facebook Ad Library Lead Finder + Website Audit (`berkaydev/meta-ad-library-lead-finder`) Actor

Find businesses paying for Facebook and Instagram ads, then audit the landing page they pay to send traffic to. Returns e-mail, phone, CMS, SSL, mobile, site age and a lead score per advertiser. Budget already proven - not a cold list. Large brands filtered out automatically.

- **URL**: https://apify.com/berkaydev/meta-ad-library-lead-finder.md
- **Developed by:** [Gezgin Data](https://apify.com/berkaydev) (community)
- **Categories:** Lead generation, Marketing, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $30.00 / 1,000 website auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Meta Ad Library Lead Finder

**Find businesses that are paying for ads and sending that traffic to a website working against
them.**

Every other Ad Library scraper answers the same question: *what are my competitors advertising?*
This one answers a different one: *which advertisers have a problem I can fix?*

The difference is the second half. Search the Meta Ad Library for what you sell into — `zahnarzt`,
`roofing`, `fitness studio` — and for every business currently running ads, this Actor opens the
landing page they are paying to send traffic to and reads it: contact e-mail, phone, CMS, SSL,
mobile layout, how old the footer is, whether any analytics are installed. Then it scores the lead.

### Why an advertiser is a better lead

The hardest part of qualifying a small business is finding out whether they have money and whether
they are willing to spend it on getting customers. An advertiser has already answered both. They
are paying Meta today.

What they often have not done is look at where that traffic lands. A dental practice buying ads
into a site with no SSL, or a salon pointing paid clicks at a page last touched in 2019, is not a
cold lead — it is a business with budget and a visible, checkable problem.

### What comes back

One row per advertiser:

| Field | Meaning |
|---|---|
| `adDomain`, `rootDomain` | The advertiser's host, and the registrable domain used to keep one row per business |
| `landingUrl`, `website`, `finalUrl` | The URL the ad pointed to, the page that was audited, and where it ended after redirects |
| `foundVia`, `countries` | Which of your search terms and countries surfaced them |
| `adsFromThisDomain` | How many separate ads this domain ran — a rough spend signal |
| `email`, `emailSource`, `emailType`, `emailConfidence` | Address, where it was found, who is behind it, how much to trust it |
| `phone`, `phoneSource`, `contactPageUrl`, `contactPageChecked` | Contact details, where the number came from, and what was actually opened |
| `emailDomainMatchesSite` | `false` when the address sits on someone else's domain — a privacy vendor, an agency, a parent brand. Those still appear, at reduced confidence, with `email_domain_differs_from_site` in `consistencyFlags` |
| `likelyJsRendered` | The page turned out to be an empty shell. Findings that depend on having seen it are suppressed |
| `trackingTools`, `opportunityCount`, `scrapedAt` | Which analytics were detected, how many opportunity tags fired, and when the row was produced |
| `cms`, `hasSsl`, `mobileFriendly`, `copyrightYear`, `hasTracking`, `socialLinks`, `hasContactForm` | The technical state of the landing page |
| `contactabilityScore`, `opportunityScore`, `opportunityTags`, `leadPriority`, `scoreBreakdown` | Can you reach them, do they visibly need help, and how each number was reached |
| `siteGrade` | A to F for the landing page, in the usual direction: **A is a healthy site, F is a broken one — so F is the best prospect.** Empty when the page could not be read, because a grade for a page nobody saw is a guess |
| `outreachHook`, `outreachBasis` | One sentence you can open with, built only from what was verified — *"Your ads currently point at a page that does not respond."* — and which finding produced it. German for DE, AT and CH, English elsewhere |
| `emailDomainHasMx` | Whether the address's domain publishes a mail server, checked by DNS. `false` means almost certainly undeliverable and lowers the confidence; empty means the lookup itself failed, which is not the same as a negative answer |
| `statusNote` | Plain sentence explaining any row that looks empty |
| `siteReachable`, `fetchStatus` | `ok`, `http_403`, `timeout`, `dns_error`. A 403 means the host refused *us* — it is never scored as a fault of the business |

**Large advertisers are dropped automatically.** A domain running many separate ads is a brand with
a media budget, not a local business you can sell a website to. Counting ads per domain filters them
far more reliably than any list of brand names, and the threshold is yours to set.

### Input

| Field | Default | What it does |
|---|---|---|
| `queries` | — | What they advertise: `zahnarzt`, `dentist`, `roofing`. **Required.** |
| `countries` | `["DE"]` | Two-letter codes. Ads are indexed per country, so each one is a separate search and adds advertisers |
| `activeOnly` | `true` | Only businesses running ads right now |
| `relevanceFilter` | `strict` | `strict` keeps an advertiser only if the trade appears in their domain, page title or H1, and drops obvious retailers. `loose` also accepts a mention anywhere on the page. `off` returns everything |
| `skipBigAdvertisers` | `true` | Drop domains running many separate ads: those are brands with a media budget |
| `bigAdvertiserThreshold` | `4` | How many ads count as "large" |
| `auditWebsite` | `true` | Open and profile the landing page. This is the point of the Actor |
| `maxContactPages` | `5` | How hard to look for contact details when the landing page yields none |
| `maxAdvertisersPerSearch`, `maxConcurrency`, `maxRunTimeSecs`, `proxyConfiguration` | | Ordinary controls |

### Use it when, and don't when

**Use it** to build a list of local businesses that are already spending on ads and whose website
you could improve — web design, hosting, SEO, CRO, marketing services.

**Do not use it** for competitor ad research. It returns advertisers and their websites, not ad
creatives, copy, spend or reach. Several Ad Library scrapers do that far better and cheaper.

### Honest limits

Read this before you buy.

**One search yields roughly 20 advertisers.** This is measured, not estimated: across three search
terms the page stopped producing new advertisers after about twenty, and two different scrolling
strategies changed nothing. Breadth therefore comes from **running several search terms and several
countries**, not from depth in one. If you need thousands of raw ads rather than a short list of
qualified leads with audited websites, a plain Ad Library scraper will serve you better and cost
less.

**The search matches ad wording, not industry — the filter fixes most of that, not all.**
Unfiltered, `physiotherapie` in Germany and Austria returned 31 advertisers of which about seven
were physiotherapists; the rest were a pet-supply shop, a pet insurer, an insole brand, a hotel and
a school, each of whom mentions the word in their ads. `relevanceFilter: strict` judges an
advertiser by what it calls *itself* — domain, page title, H1 — rather than by anything it happens
to mention, and drops obvious retailers.

Measured after that change: **`zahnarzt` returned six advertisers and all six were dental
practices. `physiotherapie` returned ten, of which seven were physiotherapists.** Across both,
roughly four rows in five are on target, against roughly one in four unfiltered.

What still slips through are **suppliers to the trade** — an anatomy-poster shop, a practice-equipment
site — because they name the trade as prominently as its members do. Telling "I am one" from "I
sell to them" is a question of meaning, not of words, and this Actor does not attempt it. Prefer
narrow search terms, and expect to discard the occasional supplier.

**Precision costs volume.** Strict filtering takes a search from roughly twenty or thirty rows down
to six or ten. That is the intended trade: a short right list beats a long mixed one. If you would
rather sort it yourself, set `relevanceFilter` to `loose` or `off`.

**Placeholder addresses are filtered, including the German ones.** `max@mustermann.de` and
`email@adresse.de` appeared as contacts on a live run before that was fixed. If you see one, report
it — handing a salesperson a placeholder is worse than handing them an empty cell.

**The Ad Library is public, and that is the whole basis of this Actor.** EU and US ad-transparency
rules require it. No login, no cookies, no proxy — but it also means only what Meta publishes is
available, and Meta decides that.

**No browser for the landing pages.** They are read over plain HTTP, which is what makes this fast
and cheap. A site that builds itself in the browser hands us a shell; those rows are flagged
`likelyJsRendered` and the "no contact form" and "no social presence" findings are suppressed,
because we never actually saw the page.

**The scoring weights are reasoned, not calibrated.** Nobody has measured which lead converts.
Every row carries a `scoreBreakdown` so you can see each contribution and disagree with it.

**This Actor is new.** The extraction and scoring engine behind the audit is the same one used by
our Website Contact Scraper, where it finds e-mail addresses on roughly three quarters of the sites
in German and US test batches. The Ad Library half has no track record yet. Treat your first run as
a sample, not a promise.

### Pricing

You are billed only for rows you actually receive. An advertiser dropped by the relevance filter, a
search the Ad Library refused to load, and a landing page that refused us all cost nothing —
charging for our own misses is the fastest way to lose the trust this Actor needs to be picked at
all.

Pay per event: `advertiser-found` when a business is returned, `website-audited` when its landing
page was actually fetched and profiled. A search that finds nothing, and a landing page that refuses
us, cost you nothing.

### Legal

Only public data is collected — the Ad Library is published by Meta under advertising-transparency
law, and landing pages are read as any visitor reads them. Under GDPR, contacting businesses whose
details you obtain this way needs a lawful basis; for B2B outreach in the EU that is usually
legitimate interest, and in Germany §7 UWG additionally restricts unsolicited advertising. Every row
records where its address came from. This is not legal advice.

# Actor input Schema

## `queries` (type: `array`):

Search terms, one per line — the same words you would type into the Meta Ad Library. `zahnarzt`, `dentist`, `fitness studio`, `roofing`. Every business currently running ads for that term becomes a lead.

## `countries` (type: `array`):

Two-letter country codes. Ads are indexed per country, so each one is a separate search — and each one adds advertisers.

## `activeOnly` (type: `boolean`):

On by default: a business paying for ads today is a warmer lead than one that stopped last year. Turn it off to reach further back and find more advertisers.

## `skipBigAdvertisers` (type: `boolean`):

A domain running many separate ads is a brand with a media budget, not a local business you can sell a website to. This drops them automatically — far more reliable than any list of brand names.

## `bigAdvertiserThreshold` (type: `integer`):

A domain with at least this many ads in the run is treated as a large advertiser. Raise it if your niche is dominated by chains and you are losing real leads.

## `relevanceFilter` (type: `string`):

The Ad Library matches ad wording, not industry — searching `physiotherapie` also returns pet shops, insurers and hotels whose ads happen to mention the word. `strict` keeps an advertiser only if the trade appears in their domain, page title, an H1 or the meta description — what a business says it *is*. `loose` also accepts a mention anywhere on the page, which is what let the hotels through. `off` returns everything and lets you filter yourself.

## `auditWebsite` (type: `boolean`):

This is the point of the Actor. Opens the page the advertiser is paying to send traffic to and reads e-mail, phone, CMS, SSL, mobile layout, copyright year, analytics, socials and contact form — then scores the lead.

## `maxContactPages` (type: `integer`):

If the landing page has no contact details, the Actor follows the contact link and probes the usual paths. Set to 0 to read the landing page only.

## `maxAdvertisersPerSearch` (type: `integer`):

Upper bound per search term and country.

## `maxScrolls` (type: `integer`):

The library stops producing new advertisers after a while; the Actor detects that and moves on. Raising this rarely helps.

## `maxConcurrency` (type: `integer`):

Higher is faster, lower is gentler on the sites you visit.

## `proxyConfiguration` (type: `object`):

Not needed for the Ad Library itself, and not needed for most landing pages. Some hosts — more often in the US — refuse datacenter addresses and answer 403. Those rows say so and carry no findings.

## `maxRunTimeSecs` (type: `integer`):

The run stops cleanly before this limit and keeps everything found so far.

## Actor input object example

```json
{
  "queries": [
    "zahnarzt"
  ],
  "countries": [
    "DE"
  ],
  "activeOnly": true,
  "skipBigAdvertisers": true,
  "bigAdvertiserThreshold": 4,
  "relevanceFilter": "strict",
  "auditWebsite": true,
  "maxContactPages": 5,
  "maxAdvertisersPerSearch": 50,
  "maxScrolls": 15,
  "maxConcurrency": 6,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "maxRunTimeSecs": 3000
}
```

# Actor output Schema

## `leads` (type: `string`):

No description

## `runStatus` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "zahnarzt"
    ],
    "countries": [
        "DE"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("berkaydev/meta-ad-library-lead-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["zahnarzt"],
    "countries": ["DE"],
}

# Run the Actor and wait for it to finish
run = client.actor("berkaydev/meta-ad-library-lead-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "zahnarzt"
  ],
  "countries": [
    "DE"
  ]
}' |
apify call berkaydev/meta-ad-library-lead-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,berkaydev/meta-ad-library-lead-finder"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/INz3EgWnpRoIG6D3g/builds/hGL9Ok7EaHp3ClCbc/openapi.json
