Yell.com Scraper — UK Business Directory | No Login avatar

Yell.com Scraper — UK Business Directory | No Login

Pricing

from $3.49 / 1,000 yell.com scraper — uk business directory | no logins

Go to Apify Store
Yell.com Scraper — UK Business Directory | No Login

Yell.com Scraper — UK Business Directory | No Login

Scrape UK business listings from Yell.com: name, phone, website, address, category, rating. Apify Residential proxy auto-configured (buyer-paid). Pay per business.

Pricing

from $3.49 / 1,000 yell.com scraper — uk business directory | no logins

Rating

0.0

(0)

Developer

Vitalii Bondarev

Vitalii Bondarev

Maintained by Community

Actor stats

1

Bookmarked

6

Total users

3

Monthly active users

3 days ago

Last modified

Share

Yell.com Scraper — UK Business Directory

Pay-per-result: $1.50/1K businesses. Phone contacts included. No monthly fee, no minimum.

Yell.com has 5 million+ UK business listings. Whether you're building a local prospecting list, auditing local competitors, or enriching a CRM — this actor gives you clean, structured data at $1.50 per 1,000 records.

Scrape UK business listings from Yell.com — the UK's leading online business directory. Ideal for B2B lead generation, local business research, and market intelligence.

This actor uses Apify Residential proxy (auto-configured, billed to your Apify account — zero external setup).

What you get

FieldDescription
business_idYell internal numeric ID
nameBusiness display name
categoryBusiness category (e.g. "Plumbers")
schema_typeschema.org type (e.g. "Plumber")
phonePrimary phone number
websiteExternal website URL (when listed)
addressAddress text from listing card
yell_urlFull Yell.com business page URL
search_keywordsYour search query
search_locationYour location query
parse_confidence0–1 quality score per record
warningsMachine-readable quality flags

Input

{
"keywords": "plumber",
"location": "London",
"maxResults": 100,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": ["RESIDENTIAL"]
}
}
  • keywords — business type or service (e.g. "dentist", "electrician")
  • location — UK city or area (e.g. "Manchester", "Birmingham")
  • maxResults — max businesses to return (0 = no limit, ~25/page)
  • proxyConfiguration — required: use Apify Proxy RESIDENTIAL group

Pricing example

Run sizeCost
100 businesses$0.15
500 businesses$0.75
1,000 businesses$1.50
5,000 businesses$7.50

Apify RESIDENTIAL proxy is billed separately to your Apify account at standard proxy rates — no extra subscription.

Output sample

{
"business_id": "90123456",
"name": "City Plumbers Ltd",
"category": "Plumbers",
"schema_type": "Plumber",
"phone": "020 7946 0123",
"website": "https://cityplumbers.co.uk",
"address": "14 Baker Street, London W1U 3BW",
"yell_url": "https://www.yell.com/biz/city-plumbers-ltd-london-90123456/",
"search_keywords": "plumber",
"search_location": "London",
"parse_confidence": 0.9,
"warnings": []
}

FAQ

Do I need a proxy or API key? No external proxy subscription needed. Apify RESIDENTIAL proxy is pre-configured in the input — Yell.com requires residential IPs and this is handled automatically. Proxy costs are billed to your Apify account at standard rates.

What export formats are available? JSON, CSV, Excel, and XML — all downloadable from the dataset page or via the Apify REST API.

Can I schedule regular runs? Yes. Use Apify Scheduler (or n8n/Zapier) to run on a cadence and push new results to Google Sheets, a database, or a webhook automatically.

What if the actor returns no results? Verify that keywords and location match Yell's search terms (e.g. "London" not "London, UK"). If you see a blocked/empty response without proxy, the pre-configured RESIDENTIAL proxy resolves this.

Why this beats the competition

This actorTypical Yell scrapers
Parse methodschema.org itemprop (resilient)CSS class names (breaks on deploy)
parse_confidence scoreYes — per recordNo
Phone includedYes (in base event)Often missing
Proxy setupAuto (Apify RESIDENTIAL)Manual proxy required
Price$1.50/1K$3-5/1K or flat fee

Every record ships a parse_confidence score (0.0–1.0). Below 0.7 is a machine-readable signal that the page structure has drifted — your data pipeline can filter automatically.

Use with AI agents (MCP)

This scraper is callable as a tool by AI agents (Claude Desktop, Cursor, VS Code, n8n, LangGraph, CrewAI, or any MCP-compatible client) via Apify's hosted Model Context Protocol server. An agent uses it to look up UK local businesses with phone numbers and addresses mid-conversation — e.g. "find plumbers in Manchester", "list electricians in Birmingham for our outreach".

Point your MCP client at this tool:

{
"mcpServers": {
"apify": {
"command": "npx",
"args": [
"mcp-remote",
"https://mcp.apify.com/?tools=bovi/yell-businesses",
"--header",
"Authorization: Bearer <YOUR_APIFY_TOKEN>"
]
}
}
}

Technical notes

  • Data parsed from schema.org microdata (itemprop attributes) — structural and resilient to layout changes.
  • 25 businesses per page, up to 5,000+ results per query.
  • Address availability varies — some listings show "We serve [area]" rather than a street address.
  • Not affiliated with Yell.com.

Integrations

Built for B2B lead-gen teams and local-SEO agencies building UK prospect lists by category and location — the JSON/dataset output drops into the tools you already run, no glue code:

  • n8n / Make / Zapier — trigger a run or pipe every new dataset item into 500+ apps (Google Sheets, Airtable, Slack, HubSpot, your database) with no code: n8n, Make, Zapier.
  • Webhooks — fire your own endpoint the moment a run finishes, to push results straight into your pipeline (docs).
  • MCP server — expose this actor as a tool to Claude, Cursor, or any MCP client so an AI agent can pull this data mid-conversation (guide).
  • API & SDKs — fetch the dataset as JSON, CSV, or Excel through the Apify REST API or the Python / JS SDKs.

See all Apify integrations.

Usage statistics

This Actor creates a small, content-free summary at the end of each run. It is used only to monitor reliability and improve this Actor. A copy is saved as USAGE_STATS in your own Apify key-value store, so you can see the exact record created for your run.

Set disableUsageStats to true in the input to opt out. Nothing is sent then; your USAGE_STATS record only says that statistics were disabled.

Only these fields are recorded:

  • schema version, Actor name and build number;
  • UTC start and finish hour (not a precise timestamp);
  • run duration, number of results and time to the first result, each as a coarse range;
  • whether the result was empty, the end status, and an error type from a fixed list;
  • memory setting and counts of charged events;
  • names of the input fields you set, never their values;
  • the selected option for input fields that offer a fixed list of choices (for example a sort order).

We do not collect input text, search terms, URLs, domains, usernames, email addresses, names, proxy credentials, tokens, scraped records, output items, raw error messages, stack traces, or hashes of any of those values. Records are kept for no longer than 13 months, used only as aggregated operational statistics, and never sold or shared.

Additional fields (Phase 2)

This Actor also records your Apify user ID, whether Apify marks the account as paying, the size range of list inputs, the selected country when the input offers a fixed list of countries, and one category from a fixed Actor taxonomy. We use these fields only for aggregate reliability, repeat-use and cross-Actor analysis; reports suppress any cell with fewer than five distinct users.

The same disableUsageStats: true input flag turns these fields off too. The user ID is removed after 13 months; we do not export, sell, share, or attempt to re-identify this data.

Run-outcome signals (v2)

To learn whether a run did what it was asked to do, the record also holds a few more coarse ranges and yes/no flags. None of them contains content:

  • the result limit you asked for (a range, when the input has one) and what share of it was delivered;
  • results delivered per input item you listed (a range);
  • output quality as ranges: how fully the result fields were filled, the share of rows that look like errors, the share of duplicate rows, and how many different fields appeared. These are counted in memory while results are saved; no result content is kept;
  • how the run was started (console, API, schedule, webhook, another Actor);
  • how it ended: stopped by you, timed out, reached the requested limit, stopped by the charge limit, and how many times the platform moved the run;
  • if this Actor reports it: how many items to process worked or failed (ranges) and one failure reason from a fixed list;
  • a short code made from the names of the input fields you set, never their values.

Repeat-run fingerprint (v2)

When your Apify user ID is recorded (see above), the record also holds an 8-character one-way code made from your input (proxy settings left out) and this Actor's name. It only lets us see that the same account ran the same input again soon after an unsatisfying run; we never see the input itself. It is stored only in the database, never published, and reports use it in aggregate with the same five-user minimum. It is the one exception to the statement above that no hashes are collected, and disableUsageStats: true turns it off.