HomeAdvisor Scraper — Contractor Leads & Data
Pricing
from $5.00 / 1,000 results
HomeAdvisor Scraper — Contractor Leads & Data
Collect home-service pros from HomeAdvisor & Angi by trade and city. Every business enriched with verified websites, emails, phones, and socials — streamed to your dataset in real time as clean JSON. Perfect for contractor lead gen, market mapping, and CRM enrichment.
Pricing
from $5.00 / 1,000 results
Rating
0.0
(0)
Developer
Emmanuel
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
6 days ago
Last modified
Categories
Share
HomeAdvisor Real-Time Data
Collect home-service pro leads from the HomeAdvisor / Angi directory at scale. Search multiple trades and cities in one run, enrich every business with lead details (verified website, emails, phones, socials), and get clean, structured JSON streamed to your Apify dataset in real time.
Built for home-improvement marketers, contractor sales teams, agencies, and data engineers who need reliable contractor lead data without slow, expensive tools.
⚠️ Paid only / free tier
Free (Apify free-plan) accounts are limited: a free run exports a maximum of 2 results and then stops gracefully with an upgrade message. Upgrade to a paid Apify plan for full, unlimited output. Paying users are never capped by this Actor. See the FAQ for details.
Why this Actor
| HomeAdvisor Real-Time Data | Typical scrapers | |
|---|---|---|
| Speed | Fast per-result collection | Often 5–15 s per business |
| Memory | 512 MB default | 2–4 GB+ |
| Setup | Organized input UI, run immediately | Fragile & high-maintenance |
| Cost | Low compute, efficient connections | High |
| Output | Structured JSON, LLM-ready | Messy HTML |
| Predictable | Every business is exported — enrichment never filters rows | Output count varies with data quality |
| Multi-market | Many trade + city pairs per run | Usually one query at a time |
Predictable output = predictable pricing. The Enable lead details option enriches every row; businesses without public details are still exported (with empty fields). That keeps runtime per 1,000 businesses stable so you always know what a run costs.
What you get — 30+ data points per business
Every record includes featureType and scrapedAt so you can filter, join, and pipe into any workflow.
| Field | Description |
|---|---|
name | Business name |
profileUrl | Directory profile link |
proId | Listing ID (stable key for dedup/joins) |
phone, emails[], phones[] | Contact channels |
website, websiteDomain | Verified business website + domain |
websiteConfidence | high / medium / low — how strongly the site matched the business |
email | Best contact email |
street, city, state, zipCode, location | Full address breakdown |
areaServed | Service area when listed |
category, services[] | Trade and services |
rating, reviewCount | Reputation signals |
description | Short description/snippet |
imageUrl | Business photo |
licensed, insured, backgroundChecked | Trust badges when displayed |
yearsInBusiness, hireCount | Track record signals |
socials, facebookUrl, instagramUrl, linkedinUrl | Social profiles |
searchCategory, searchLocation, searchTaskIndex, searchTaskLabel | Traceability for multi-task runs |
leadDetails | true when lead-detail enrichment ran on this row |
scrapedAt | ISO timestamp |
Search results (featureType: "search")
One row per business from a trade + city search, with all fields above when available.
Scrape By URL (featureType: "scrape_by_url")
Paste directory profile URLs to enrich a list you already have — same output shape.
Who it is for
- Home-improvement contractors — find partners, subcontractors, or out-of-area competitors
- Roofing / HVAC / plumbing / electrical suppliers — build targeted B2B prospect lists by trade and city
- Marketing agencies — prospect businesses with thin digital presence (no website = prime web-design leads)
- SaaS & CRM teams — feed vertical CRM data for the home-services vertical
- Lead-gen & outbound sales teams — phone + email lists ready for sequences
- Franchise developers — map service-market density across cities
- Insurance & property-tech — verify contractor coverage areas and reputation
- Data engineers & AI builders — structured JSON for pipelines, scoring, and LLM workflows
Use cases
- Local lead generation — pull roofers, plumbers, electricians, and remodelers with phone + email + website for outbound campaigns
- "No website" prospecting — businesses with ratings but no website are perfect web-design / digital-marketing leads
- Multi-city prospecting — one run across dozens of trade + city pairs
- Market mapping — compare competitor density, ratings, and review counts across metros
- Territory planning — measure how many pros serve each city before expanding
- CRM enrichment — backfill missing websites/emails on contractor records you already have
- Reputation snapshots — capture rating + review counts over time
- Trade-specific campaigns — target one trade (e.g. HVAC) across 20 states
- AI & LLM pipelines — JSON for RAG, scoring models, outreach drafts, and territory summaries
- Zapier / Make / n8n automations — real-time webhook pushes each lead into your stack as it is collected
Features
Search terms & locations — on by default
Run multiple trade + location pairs in one actor run — e.g. roofing in Austin, TX and plumbing in Dallas, TX side by side. Each task has its own max results. Category aliases are understood (roofer → roofing, plumber → plumbing, contractor → general contractors, and many more).
Location must include the state (e.g. Austin, TX or New York, NY).
Enable lead details — the lead-gen engine
When on, every business is additionally enriched with:
- Verified website — the business's own site is matched and verified against its name and geography before it is ever attached to a row (high/medium/low confidence is reported). Aggregators and social networks are never presented as the business's site.
- Emails — contact emails published on the business's own pages (contact/about pages included)
- Phones — normalized US phone numbers
- Social profiles — Facebook, Instagram, LinkedIn, X/Twitter, YouTube, TikTok, Pinterest
- Trust badges — licensed / insured / background-checked when displayed
This adds a little extra time per business but never removes a row: every business goes to the dataset even if it has no lead details. Turn it off for the fastest possible directory-only runs.
Scrape By URL
Paste directory profile URLs (one per line) to enrich a list you already have.
Webhooks
Real-time push of every record to your CRM, Slack, Zapier, Make, or custom endpoint — see Webhook delivery.
Input reference
| Input | Type | Default | Description |
|---|---|---|---|
| Search terms & locations | |||
enableSearch | boolean | true | Trade + location search |
searchTasks | object[] | 2 example tasks | Primary input: { category, location, zip?, maxResults? } per row |
enableSimpleSearch | boolean | false | Turn on the Simple search section (used only when the task table is empty) |
simpleCategory | string | — | Simple search: one trade. Legacy key category also accepted. |
simpleLocations | string | — | Simple search: one or more US cities + state, semicolon-separated (e.g. Austin, TX; Dallas, TX). Legacy key location also accepted. |
maxResultsPerTask | integer | 10 | Default max businesses per task |
maxPagesPerTask | integer | 5 | Maximum depth per task (safety cap, 1–50) |
| Scrape By URL | |||
enableScrapeByUrl | boolean | false | Scrape profile URLs |
scrapeUrls | string[] | — | Profile URL list |
| Lead details | |||
enableLeadDetails | boolean | true | Enrich every business with website/email/phone/socials. Never filters rows. |
| Output & limits | |||
maxItems | integer | 10000 | Global cap on total rows (set higher for large runs) |
concurrency | integer | 4 | Parallel detail lookups (1–8) |
delayBetweenRequestsMs | integer | 1200 | Polite pacing between fetches (0–5000 ms) |
webhookUrl | string | — | Optional real-time POST URL |
webhookFormat | enum | json | json (full record) or slack (Slack message) |
proxyConfiguration | object | residential US | Apify proxy settings |
Full schema: see .actor/input_schema.json or the Input tab on Apify Console.
Output reference
Each dataset row is one business. Filter by featureType:
featureType | Description |
|---|---|
search | Business from trade + city search |
scrape_by_url | Business from a profile URL |
Field-by-field documentation: see the tables above and .actor/dataset_schema.json.
Run summary: every run also writes an OUTPUT record to the key-value store with totals, lead-details flag, spending-limit status, and the paywall object:
{"totalPushed": 4,"tasksRequested": 1,"tasksCompleted": 1,"leadDetailsEnabled": true,"errors": [],"spendingLimitReached": false,"paywall": {"detected": true,"isPaying": true,"pricingTier": "GOLD","limited": false,"blocked": false,"freeTierMaxItems": null}}
Export formats: JSON, CSV, Excel, RSS, or via API.
Webhook delivery (optional)
Every record is always saved to the Apify dataset first. If you set webhookUrl in the Output & limits section, each new record is also POSTed in real time to your endpoint — useful for CRMs, Slack, Zapier, Make, or custom pipelines.
| Setting | Description |
|---|---|
webhookUrl | Your endpoint (https recommended). Leave empty to use dataset only. |
webhookFormat | json — the full record object. slack — compact Slack incoming-webhook message. |
Webhook delivery is best-effort: a failed webhook never stops the run or prevents dataset writes.
Payload (json format) — the exact record object documented in Output reference, including leadDetails, websiteConfidence, and social fields when present. No internal or debug data is ever included.
Payload (slack format):
{"text": ":hammer_and_wrench: *Quick Roofing*\n*Category:* roofing • *Rating:* 4.8 (214 reviews)\n*Phone:* (512) 555-0142\n*Location:* Austin, TX\n*Website:* https://quickroofing.com\n<View profile>"}
Example — search with Slack alerts
{"enableSearch": true,"searchTasks": [{ "category": "roofing", "location": "Austin, TX", "maxResults": 10 }],"webhookUrl": "https://hooks.slack.com/services/YOUR/WEBHOOK/URL","webhookFormat": "slack"}
Example — lead gen with JSON webhook
{"enableSearch": true,"searchTasks": [{ "category": "plumbing", "location": "Phoenix, AZ", "maxResults": 25 }],"enableLeadDetails": true,"webhookUrl": "https://your-crm.example.com/api/leads","webhookFormat": "json"}
Quick start examples
Multi-city lead gen (default pattern)
{"enableSearch": true,"searchTasks": [{ "category": "roofing", "location": "Phoenix, AZ", "maxResults": 25 },{ "category": "roofing", "location": "Tucson, AZ", "maxResults": 25 }],"enableLeadDetails": true}
Fast directory-only scan (no enrichment)
{"enableSearch": true,"searchTasks": [{ "category": "hvac", "location": "Miami, FL", "maxResults": 50 }],"enableLeadDetails": false,"maxItems": 50}
Enrich a URL list
{"enableSearch": false,"enableScrapeByUrl": true,"scrapeUrls": ["https://www.angi.com/companylist/us/tx/austin/quick-roofing-reviews-10944876.htm"],"enableLeadDetails": true}
One trade, many cities (simple search)
{"enableSearch": true,"searchTasks": [],"enableSimpleSearch": true,"simpleCategory": "electricians","simpleLocations": "Austin, TX; Dallas, TX; Houston, TX","maxResultsPerTask": 20}
API quick start
curl -X POST "https://api.apify.com/v2/acts/YOUR_ACTOR_ID/runs?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"enableSearch": true,"searchTasks": [{ "category": "roofing", "location": "Austin, TX", "maxResults": 25 },{ "category": "plumbing", "location": "Dallas, TX", "maxResults": 25 }],"enableLeadDetails": true}'
Dataset items: GET https://api.apify.com/v2/datasets/{datasetId}/items?format=json
LLM & MCP integration
Output is JSON Lines–friendly structured data — ideal for ChatGPT, Claude, Gemini, LangChain, LlamaIndex, and custom agents.
Recommended workflow
- Run the Actor with the tasks you need.
- Fetch dataset items via the Apify API or export JSON/CSV.
- Pass records to your LLM with a system prompt, or index into a vector store.
Example record for an LLM prompt
{"name": "Quick Roofing","category": "roofing","city": "Austin","state": "TX","phone": "(512) 555-0142","email": "office@quickroofing.com","website": "https://quickroofing.com","websiteConfidence": "high","rating": 4.8,"reviewCount": 214,"searchTaskLabel": "roofing | Austin, TX"}
Apify MCP (Model Context Protocol)
Use the Apify MCP server so AI assistants can:
- Run this Actor with natural-language instructions ("find 20 roofers in Austin with emails")
- Read dataset results directly in the chat
- Chain with other Actors (enrich → score → CRM)
Typical MCP tool flow:
User: "Find 15 highly rated plumbers in Dallas and draft outreach intros"→ MCP runs the Actor with searchTasks=[{ category: "plumbing", location: "Dallas, TX" }], enableLeadDetails=true→ MCP reads dataset items→ LLM summarizes each business and drafts the messages
Proxy & performance
- Apify residential proxy (US) is enabled by default — no extra setup required on Apify. Change it in the Connection section of the input if you need a different country or custom proxy URLs.
- Default memory: 512 MB — enough headroom for most runs.
- Results are streamed to the dataset as they are collected; long runs do not pile up data in memory.
Reliability notes
- Runs use US-based rotating residential connections by default (Apify residential proxy) — no user setup required.
- Temporary rate limits are handled automatically: the run switches connections and retries, and logs a plain-language message. It never exposes technical details in logs or output.
- If the directory has no page for a trade/city combination, the task logs a warning and continues with the next task.
Free tier (Apify free plan)
This Actor requires a paid Apify plan for full output.
| Your Apify plan | What happens |
|---|---|
| Apify free plan | Runs are capped: a maximum of 2 results per run, then the run finishes gracefully with a message to upgrade. |
| Any paid Apify plan (Bronze → Diamond) | ✅ Full, uncapped output up to your maxItems — plus your run's max-total-charge limit is always respected. |
When a free-tier run hits the cap, the Actor finishes gracefully with a clear status message ("Free tier limit reached — results were capped. Upgrade to a paid Apify plan for full, unlimited data."). Your collected sample stays in the dataset — nothing is lost and nothing looks like an error.
The paywall status is also written to the run summary as a transparent paywall object (detected, isPaying, pricingTier, limited, blocked) — see Output reference.
FAQ
Is this really limited on the Apify free plan? Yes. Free-plan runs export a maximum of 2 results. This is a policy restriction, not a bug — upgrade to any paid Apify plan for full output.
Does lead-details filtering change how many results I get? No — and that's deliberate. Enable lead details enriches every business; it never removes rows. Businesses without public contact details are still exported (with those fields empty). This keeps your output count — and therefore your per-run cost — predictable.
Why do I need to include the state in the location? The directory is organized by city + state (e.g. "Austin, TX"). Locations without a state cannot be resolved — the run will warn you and skip that task.
Which trades are supported? Roofing, plumbing, HVAC, electricians, painters, general contractors, landscaping, flooring, house cleaning, handyman, remodeling, kitchen/bathroom remodeling, windows, siding, concrete, fencing, tree service, pest control, movers — plus common aliases (roofer → roofing, plumber → plumbing, contractor → general contractors).
How accurate are the websites and emails?
Websites are only attached after they are verified against the business's name and geography, with a reported confidence level (high, medium, low). Aggregator and social-media pages are never presented as the business's own site. Emails come from the verified business's own pages. When nothing reliable is found, fields stay empty — the row is still exported.
Do you support every city? Any US city that the directory covers. Very small towns may have no dedicated page — the Actor warns and continues.
How fast is a run?
A directory page (~10–20 businesses) is collected in seconds. With Enable lead details on, add a little extra time per business for website verification and contact discovery. 4 parallel lookups by default; raise concurrency to 8 for the fastest runs.
Can I get an alert for every lead as it arrives?
Yes — set webhookUrl (and webhookFormat: slack for Slack). Each record is POSTed to your endpoint in real time as it is exported.
Does the run respect my max total charge? Yes. The Actor tracks the user's maximum charge for the run and stops collecting — gracefully, with a clear message — as soon as that limit is reached. You are never charged beyond what you allowed.
Where is the run summary?
In the run's key-value store under the OUTPUT record — totals, lead-details flag, spending-limit status, and the transparent paywall object. Linked in the run's Output tab.
Is this affiliated with HomeAdvisor or Angi? No. This Actor is an independent data tool. Use responsibly and comply with applicable laws and the sites' terms.
Limitations & compliance
- Data availability depends on what each business's directory listing and its own public pages expose.
- Not affiliated with HomeAdvisor or Angi. Use responsibly and comply with applicable laws and their Terms of Service.
- Always respect rate limits and local regulations when collecting business data.
Contact & custom work
Need something beyond this Actor? I build custom scrapers, data pipelines, and full-stack web applications for startups and enterprises.
- Email: dubem115@gmail.com
- GitHub: github.com/DrunkCodes
Reach out for:
- Custom Apify Actors (any website or data source)
- Local lead-gen data projects at scale
- LLM & MCP integrations with your data stack
- Web apps, dashboards, and automation tools
HomeAdvisor Real-Time Data · by DrunkCodes