Avvo Scraper — Lawyer & Attorney Leads
Pricing
from $3.00 / 1,000 results
Avvo Scraper — Lawyer & Attorney Leads
Scrape Avvo lawyer leads at scale. Search any practice area and location, then enrich every result with phone numbers, emails, firm websites, ratings, reviews and license details. Clean structured JSON or CSV, streamed live — ready for legal marketing, lead-gen agencies and CRM enrichment.
Pricing
from $3.00 / 1,000 results
Rating
0.0
(0)
Developer
Emmanuel
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Avvo Real-Time Data
Collect lawyer directory data from Avvo at scale. Search multiple practice areas and cities in one run, enrich every lawyer with profile details and firm contacts, and stream clean, structured JSON to your Apify dataset in real time.
⚠️ Paid-only note (please read): Free Apify accounts are limited to a small sample of results per run (2 by default). To get full, unlimited data, upgrade to a paid Apify plan. This restriction is stated here, in the input form, and in the run output — it is a deliberate policy, not a bug.
Built for legal marketers, lead-gen agencies, law-firm growth teams, recruiters, and data engineers who need fresh lawyer contact data without slow, expensive, low-yield tooling.
Why this Actor
| Avvo Real-Time Data | Typical scraper | |
|---|---|---|
| Speed | Fast per-result collection | Often 5–15 s per profile |
| Memory | 512 MB default | 2–4 GB+ |
| Setup | Organized input UI, run immediately | Fragile & high-maintenance |
| Output | Structured JSON, LLM-ready | Often messy HTML |
| Predictability | Every lawyer is exported — output is never filtered by lead completeness | Filtering makes run cost unpredictable |
| Multi-market | Many practice-area + location pairs per run | Usually one query at a time |
What it does
- Search Avvo's public lawyer directory by practice area + location (city, state). Run many pairs in a single run.
- Lead details (enabled, not filtered): for each lawyer, pull the public profile details — phone, firm name, practice areas, bar-admission year when listed — and, when the lawyer's firm site is published, scan that site for emails and social profiles.
- Profile URLs: paste profile links you already have and enrich them directly.
- Streaming output: every record is pushed to the dataset the moment it is final, so long runs stay memory-light and results arrive in real time.
- Webhooks: push each record to your CRM, Slack, Zapier, Make, or custom URL as it is collected.
Every lawyer found goes to the dataset — even if no lead details were published for them. Each row carries a leadDetails status object (complete / partial / none) so you can segment downstream without the run silently dropping rows. Enrichment adds a little extra time per lawyer; output volume stays fully predictable, which keeps per-run pricing predictable too.
Who it is for
- Legal marketing agencies — build and refresh lawyer prospect lists across states and practice areas.
- Lead-generation teams — SDR teams prospecting attorneys with phone numbers and firm sites.
- Law-firm growth / BD — map competitors by city and practice area; monitor who is listing where.
- Legal recruiters & staffing — candidate discovery by practice area and market.
- Legal-tech & SaaS — feed CRM and product data with structured lawyer profiles.
- Consultants & analysts — market mapping: how many injury lawyers operate in Austin vs Dallas.
- AI / LLM builders — clean JSON records for enrichment agents, scoring, and RAG pipelines.
Use cases
- Attorney lead lists — export lawyers by practice area + city with phone and firm details for outreach.
- Email discovery for B2B outreach — with lead details enabled, emails are added when a firm publishes them on its own website.
- Multi-market prospecting — one run across dozens of practice-area + location pairs.
- Competitive mapping — who practices personal injury in Phoenix? Family law in Chicago?
- Bar-admission intelligence —
licensedSinceandyearsExperiencewhen the directory lists them. - CRM enrichment — start from profile URLs you already have and backfill structured fields.
- Reputation snapshots — rating and review counts per lawyer for market research.
- Territory planning — density analysis by practice area across states.
- CRM & warehouse feeds — dataset via Apify API, optional real-time webhook, or scheduled runs.
- AI pipelines — JSON Lines-friendly output for LLM workflows.
Features
🔎 Search (on by default)
Pair one or more practice areas with locations. The directory organizes listings by state and city; your location is parsed into state + city automatically (e.g. Austin, TX). Each task collects until it reaches its lawyersPerTask cap, walking up to maxDepthPerTask for very large markets.
Every result includes name, profile link, and the data published on the listing (rating, review count, firm name when shown).
🎯 Lead details (enabled by default)
Enriches every lawyer with whatever the public profile publishes:
- Phone (when listed)
- Firm name, practice areas, specialties
- Years of experience / licensed since (when listed)
- Firm website (when published)
- Emails + social profiles — when the firm's own website is published, that site is scanned for contact emails and social accounts
- A
leadDetailsstatus object:complete,partial(with a short reason), ornone
Nothing is filtered out. If a lawyer publishes no contact details, their row still lands in the dataset with leadDetails.status: "none" — so 1,000 lawyers in ≈ 1,000 rows out, always.
Set crawlFirmSites: false to skip the firm-website email scan and finish runs faster.
🔗 Lawyer Profiles by URL
Paste Avvo profile URLs (one per line) to enrich specific lawyers — same rich output as search enrichment.
🌐 Connection
A residential US proxy connection is enabled by default for reliability. Change it in the input only if you need a different setup.
Input reference
Full schema: see .actor/input_schema.json or the Input tab on Apify Console.
| Input | Type | Default | Description |
|---|---|---|---|
| Search terms & locations | |||
enableSearch | boolean | true | Practice-area + location search |
searchTasks | object[] | 2 example tasks | { practiceArea, location, lawyersPerTask? } per row |
lawyersPerTask | integer | 10 | Default max lawyers per search task |
maxDepthPerTask | integer | 20 | Collection depth cap per task (1–100) |
minRating | number | — | Only keep lawyers rated at least this high (0–5). Leave empty for no filter |
| Lead details | |||
enableLeadDetails | boolean | true | Enrich every lawyer with profile details + firm contacts. Adds a little extra time per lawyer. Results are never filtered by lead completeness |
crawlFirmSites | boolean | true | Scan a lawyer's published firm website for emails and social profiles |
| Lawyer Profiles by URL | |||
enableProfileUrls | boolean | false | Enrich specific profile URLs |
profileUrls | string[] | — | Avvo profile URLs, one per line |
| Output & limits | |||
maxItems | integer | 10000 | Global cap on total dataset rows. Free accounts are additionally capped to a small sample (see the paid-only note) |
webhookUrl | string | — | Optional real-time POST URL — dataset is always written; webhook is additional |
webhookFormat | enum | json | json (full record) or slack (Slack message) |
| Connection | |||
proxyConfiguration | object | residential US | Apify proxy settings |
Notes
- Practice areas use directory slugs:
personal-injury,family-law,divorce,criminal-defense,dui,bankruptcy,immigration,real-estate,business,employment,estate-planning,intellectual-property,medical-malpractice,workers-compensation,social-security-disability, and more. - Locations are
City, STpairs (e.g.Chicago, IL). State-only searches also work (TX).
Output reference
Each dataset row is one lawyer. Fields present when the data is published for that lawyer:
| Field | Type | Description |
|---|---|---|
featureType | string | search (from a search task) or scrape_by_url (from profile URLs) |
type | string | Listing kind (attorney) |
name | string | Lawyer's name |
profileUrl | string | Link to the lawyer's public profile |
avvoId | string | The lawyer's identifier on Avvo |
phone | string | Phone number when published |
email | string | Email when published (typically found via the firm's own website scan) |
website | string | Firm website when published |
socials | string[] | Social profile URLs found on the firm's site |
firmName | string | Firm / office name |
practiceAreas | string[] | Listed practice areas |
specialties | string[] | Listed specialties |
city / state | string | Location from the profile or the search |
location | string | Free-form location when published |
rating | number | Star rating (0–5) |
reviewCount | integer | Number of client reviews |
avvoRating | number | Directory rating when displayed |
yearsExperience | integer | Years in practice when listed |
licensedSince | integer | Bar-admission year when listed |
imageUrl | string | Profile photo URL |
firmName / firmUrl | string | Firm identity fields |
searchPracticeArea | string | The practice area that produced this row |
searchLocation | string | The location that produced this row |
leadDetails | object | { status: "complete" | "partial" | "none", reason?: string } — outcome indicator; rows are never filtered by it |
enriched | boolean | Whether profile details were merged into this row |
scrapedAt | string | ISO timestamp |
Export formats: JSON, CSV, Excel, RSS, or via the Apify API.
The run also writes a summary record (OUTPUT) to the run's key-value storage with counts and a paywall object (detected, isPaying, pricingTier, limited, blocked) for transparency.
Webhook delivery (optional)
Every record is always saved to the Apify dataset first. If you set webhookUrl, each new record is also POSTed in real time to your service — CRMs, Slack, Zapier, Make, Google Sheets, or your own API.
| Setting | Description |
|---|---|
webhookUrl | Your service URL (http/https). Leave empty for dataset only. |
webhookFormat | json — the full record object. slack — compact Slack incoming-webhook message. |
- Delivery is best-effort: a failed webhook never stops the run or blocks dataset writes.
- Webhook payloads contain only the documented output fields — the same JSON you see in the dataset, nothing more.
Example — search with Slack alerts
{"enableSearch": true,"searchTasks": [{ "practiceArea": "family-law", "location": "Chicago, IL", "lawyersPerTask": 25 }],"webhookUrl": "https://hooks.slack.com/services/YOUR/WEBHOOK/URL","webhookFormat": "slack"}
Example — JSON webhook to a CRM
{"enableSearch": true,"searchTasks": [{ "practiceArea": "criminal-defense", "location": "Phoenix, AZ", "lawyersPerTask": 100 }],"enableLeadDetails": true,"webhookUrl": "https://your-crm.example.com/api/leads","webhookFormat": "json"}
Paid-only / free tier
- Paying Apify plans: full, uncapped output.
- Free accounts: capped to a small sample per run — 2 results by default — then the run stops gracefully with a clear message to upgrade.
- The cap is a deliberate monetization policy (also shown in the input form and the run output), configurable by the actor owner via environment variables.
- Upgrading to any paid Apify plan removes the cap — no other change needed.
FAQ
Is this Actor available for free Apify accounts? Free accounts can run it, but output is limited to a small sample (2 results by default). Upgrade to a paid Apify plan for full, unlimited data.
Why don't you filter out lawyers without contact info?
Filtering would make runtime — and therefore cost — unpredictable per 1,000 lawyers. Instead every lawyer is exported with a leadDetails status, and you filter downstream if you want.
Why is there no email on every row?
Emails are only exported when a lawyer's firm actually publishes them on its own website. Phone numbers are far more commonly published in this directory. The leadDetails object tells you what was found and what wasn't.
How current is the data?
Each run queries the directory live and stamps rows with scrapedAt.
Can I run multiple practice areas and cities in one run?
Yes — add as many rows to searchTasks as you need. Each row runs with its own cap.
Do you support state-only searches?
Yes. Use just the state (TX) in the location field, or a City, ST pair.
How do I get the results? From the run's Storage tab (JSON/CSV/Excel export), the Apify API, a webhook, or the Apify MCP server.
What do the Busy — retrying shortly… log lines mean?
The directory intermittently slows a visitor down. The Actor refreshes its connection automatically and carries on — most runs finish with a few of those lines and no missing data. If a run logs Connection preflight: challenged, the connection had to be refreshed repeatedly at startup: re-check the Proxy configuration on that run (and, if you manage your own connection, the proxy environment variables on the Actor).
What does Connection credentials were rejected mean?
The connection itself was refused before any collection happened — this is a proxy credential/entitlement problem on that run, not a data problem. Fix the proxy configuration and run again.
A task reported "No results" — is that a bug? No. It means the combination of practice area and location returned no lawyers (usually a spelling or slug mismatch). If everything looked correct and every task reports it, the connection was challenged at startup — see the preflight line above and re-run.
Is this affiliated with Avvo? No. This Actor collects publicly available directory data. Use it responsibly and in compliance with applicable laws and the directory's terms.
What does leadDetails.status mean?
complete — contact details and a firm site were found. partial — some details were found; reason explains what was missing (e.g. extra details were unavailable at collection time). none — no lead details were published for this lawyer.
LLM & MCP integration
Output is JSON Lines–friendly structured data — ideal for ChatGPT, Claude, Gemini, LangChain, LlamaIndex, and custom agents.
Apify MCP (Model Context Protocol)
Use the Apify MCP server so AI assistants can:
- Run this Actor with natural-language instructions
- Read dataset results directly in the chat
- Chain with other Actors (e.g. enrich → score → outreach)
Typical MCP tool flow:
User: "Find 30 personal injury lawyers in Texas with phone numbers and summarize each for outreach"→ MCP runs the Actor with searchTasks=[{ practiceArea: "personal-injury", location: "Austin, TX" }]→ MCP reads dataset items→ LLM summarizes and drafts emails
Example record for an LLM prompt
{"featureType": "search","name": "Jane Smith","phone": "(512) 555-0142","firmName": "Smith Law Group","practiceAreas": ["Personal Injury", "Car Accidents"],"city": "Austin","state": "TX","rating": 4.9,"reviewCount": 87,"leadDetails": { "status": "complete" },"profileUrl": "https://www.avvo.com/attorneys/..."}
API quick start
curl -X POST "https://api.apify.com/v2/acts/YOUR_ACTOR_ID/runs?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"enableSearch": true,"searchTasks": [{ "practiceArea": "personal-injury", "location": "Austin, TX", "lawyersPerTask": 50 },{ "practiceArea": "divorce", "location": "Dallas, TX", "lawyersPerTask": 50 }],"enableLeadDetails": true}'
Dataset items: GET https://api.apify.com/v2/datasets/{datasetId}/items?format=json
Quick start examples
Multi-city lead gen (default pattern)
{"enableSearch": true,"searchTasks": [{ "practiceArea": "personal-injury", "location": "Austin, TX", "lawyersPerTask": 100 },{ "practiceArea": "personal-injury", "location": "Houston, TX", "lawyersPerTask": 100 }],"enableLeadDetails": true,"crawlFirmSites": true}
Fast listing-only scan (no enrichment)
{"enableSearch": true,"searchTasks": [{ "practiceArea": "immigration", "location": "Miami, FL", "lawyersPerTask": 200 }],"enableLeadDetails": false,"maxItems": 5000}
Enrich profile URLs you already have
{"enableSearch": false,"enableProfileUrls": true,"profileUrls": ["https://www.avvo.com/attorneys/78701-tx-jane-smith-1234567.html","https://www.avvo.com/attorneys/75201-tx-john-doe-2345678.html"],"enableLeadDetails": true}
High-rated lawyers only
{"enableSearch": true,"searchTasks": [{ "practiceArea": "estate-planning", "location": "Denver, CO" }],"minRating": 4.5,"enableLeadDetails": true}
Proxy & performance
- Residential US proxy is enabled by default — no extra setup on Apify.
- Default memory: 512 MB — enough headroom for most runs.
- Results are streamed to the dataset as they are collected; long runs do not pile up data in memory.
- Run timeout: 10,000 seconds — sized for very large runs.
- Free-tier runs stop early by design (see the paid-only note).
Reliability
- Transient busy spells are retried automatically, with a fresh connection taken before each retry.
- A challenged connection is refreshed, never abandoned — one bad spell can't strand the rest of the run.
- Connection preflight runs once at the start and reports a single outcome-only line (
ready,ready after refreshing, orchallenged) so a slow run is easy to diagnose. - Failures are reported as plain sentences —
Busy,Connection trouble,Connection credentials were rejected— never as technical error dumps or stack traces. - A lawyer whose profile cannot be enriched still appears in the dataset with a
leadDetailsreason — no silent drops. - Webhook failures never stop the run.
Limitations & compliance
- Fields are present only when the lawyer or their firm has published them.
- Not affiliated with Avvo. Use responsibly and comply with applicable laws and the directory's Terms of Service.
- Respect rate limits and local regulations when collecting professional data.
Contact & custom work
Need something beyond this Actor? I build custom scrapers, data pipelines, and full-stack web applications for startups and enterprises.
- Email: dubem115@gmail.com
- GitHub: github.com/DrunkCodes
Reach out for:
- Custom Apify Actors (any website or data source)
- Legal-directory and lead-gen data projects at scale
- LLM & MCP integrations with your data stack
- Web apps, dashboards, and automation tools
Avvo Real-Time Data · by DrunkCodes