# FindLaw Scraper — Lawyer & Law Firm Leads (`b2b_leads/findlaw-real-time-data-scraper`) Actor

Search millions of US attorneys and law firms by practice area, state, and city. Get names, phones, emails, websites, addresses, ratings, and social profiles — streamed to your dataset in real time. Paid plans get full output; free plans get a sample.

- **URL**: https://apify.com/b2b\_leads/findlaw-real-time-data-scraper.md
- **Developed by:** [Emmanuel](https://apify.com/b2b_leads) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## FindLaw Real-Time Data — Attorney & Law-Firm Leads

**Turn the largest US lawyer directory into a clean, structured lead list.** Search by practice area + state + city, collect attorney and law-firm profiles with phones, emails, websites, and social links, and stream every row to your dataset in real time.

> ⚡ **Real-time streaming** • 🎯 **Phones, emails & socials** • 🏛️ **Lawyers + firms** • 🔔 **Webhooks** • 💸 **Pay per result**

***

### ⚠️ Paid only / free tier — read this first

This Actor is **paid only**. Free (Apify free-plan) accounts get a **small sample of 2 results per run** — after that the run finishes gracefully with a clear message to upgrade. There is no error, no crash, and nothing is lost: your sample is in the dataset.

| Plan | What you get |
|---|---|
| **Apify free plan** | Limited to **2 results per run**. The run exits cleanly with an upgrade message. |
| **Any paid Apify plan** (Bronze → Diamond) | ✅ **Full, unlimited output** up to your `maxItems` — no caps. |

> Free accounts: **upgrade to a paid Apify plan to get full, unlimited data.** The restriction is also shown in the input UI and logged at the start of every limited run. A `paywall` object in the run OUTPUT always shows which mode applied.

***

### What it does

FindLaw is the largest US attorney directory — millions of consumers and businesses use it monthly to find lawyers. This Actor turns that directory into structured, ready-to-use lead data:

- **Lawyer search** — pick practice areas and locations; every matched attorney becomes a dataset row.
- **Law firm search** — the same for firms, tagged separately.
- **Profile details** — fetch full structured details for specific attorney or firm URLs.
- **Lead details** *(enabled by default)* — for every result, the profile (and the practice's own website when needed) is used to collect **phone, email, website, and social profiles**. Adds a little extra time per business, but **every business goes to the dataset** — results are never filtered, so runtime per 1,000 businesses stays predictable. Lead-detail work is **pipelined in parallel pools**: profiles are read **8 at a time**, then contact hunts run **8 at a time** in their own pool, while each worker holds one exit and reuses it for every page it reads.
- **Webhooks** — push each row to your CRM, Slack, Zapier, Make, or Google Sheets the moment it is collected.
- **Real-time streaming** — rows are written to the dataset as they are found, so you can watch results arrive and export early if you want. RAM stays light during long scrapes.

#### Outcomes you get

- A **clean prospect list** of attorneys/firms in your target practice + geography with direct contact details.
- **Predictable costs** — you pay per result; the `maxItems` cap and the user spending limit are always honored.
- **Flat, stable JSON** — same field names on every row, ready for Sheets, Airtable, a database, or an LLM pipeline.

***

### Who it is for

| Persona | What they do with it |
|---|---|
| **Legal marketing agencies** | Build prospect lists of attorneys who need SEO, ads, or intake services. |
| **Legal-tech & SaaS sales teams** | Find firms by practice area and reach out with product offers. |
| **Recruiters (attorney placement)** | Track which firms are growing, which practice areas they cover, and how to reach hiring partners. |
| **Consultants & market researchers** | Map the competitive landscape of law practices by city and state. |
| **Law firms (competitive intel)** | Watch which other firms show up in your practice + city, and what they advertise. |
| **Bar associations & CLE providers** | Analyze practice-area distribution across states and cities. |
| **Lead brokers & data teams** | Enrich CRM pipelines with verified directory presence and contact URLs. |
| **AI/automation builders** | Feed flat JSON into agents, RAG pipelines, and enrichment workflows. |

***

### Use cases

**🎯 Attorney lead generation** — Export every personal-injury firm in Texas with phones and websites for an outreach campaign.

**🏢 Firm prospecting for vendors** — Sell practice-management software, answering services, or marketing to firms by practice area and firm size signals.

**📍 Market mapping** — Count attorneys per city and practice area to pick expansion markets or plan ad budgets.

**🕵️ Competitive monitoring** — Schedule weekly runs for your practice area + city and see which competitors appear.

**🤝 Recruiting pipelines** — Track firms' attorney rosters, then use profile URLs to reach candidates.

**📈 SEO/SEM prospect qualification** — Identify firms with a directory presence but no website (an easy pitch for web services).

**🧠 AI agents & RAG** — Stream flat JSON into an LLM to summarize, score, or triage law-firm leads automatically.

**🔁 CRM enrichment** — POST results through webhooks into HubSpot/Salesforce/Zapier as new records arrive.

***

### Quick start

1. Open the Actor in Apify Console.
2. In **Search tasks**, add one row per search — a practice area in a city or state.
3. Leave **Enable lead details** on to collect phones, emails, websites, and socials.
4. Press **Start** and watch rows stream into the dataset.
5. Export as **JSON / CSV / Excel**, or pull via the **Apify API**, **webhooks**, or **MCP**.

#### Example — personal injury attorneys in two cities

```json
{
    "enableSearchLawyers": true,
    "searchTasks": [
        { "practiceArea": "personal injury", "state": "New York", "city": "New York" },
        { "practiceArea": "personal injury", "state": "New York", "city": "Brooklyn" }
    ],
    "enableLeadDetails": true,
    "maxItems": 10000
}
```

#### Example — firms state-wide in two states

Leave **City** empty to cover a whole state — the task then works through the state's cities until it reaches its limit, so the rows are spread across the state instead of one market:

```json
{
    "enableSearchFirms": true,
    "firmSearchTasks": [
        { "practiceArea": "divorce", "state": "California" },
        { "practiceArea": "divorce", "state": "Florida" }
    ],
    "maxItems": 5000
}
```

#### Example — details for specific profile URLs

```json
{
    "enableScrapeByUrl": true,
    "scrapeUrls": [
        "https://lawyers.findlaw.com/profile/lawyer/…"
    ]
}
```

#### Example — several practice areas in one state

Practice areas accept natural names — common aliases are mapped automatically (`car accident` → vehicle-accident directory, `criminal defense` → criminal law, and similar).

```json
{
    "enableSearchLawyers": true,
    "searchTasks": [
        { "practiceArea": "car accident", "state": "Texas" },
        { "practiceArea": "criminal defense", "state": "Texas" },
        { "practiceArea": "divorce", "state": "Texas" }
    ]
}
```

#### Example — a different size for each market

Five cities, but you only want a handful of leads from most of them and a deep pull from two:

```json
{
    "enableSearchLawyers": true,
    "searchTasks": [
        { "practiceArea": "personal injury", "state": "Florida", "city": "Miami", "maxResults": 500 },
        { "practiceArea": "personal injury", "state": "Florida", "city": "Tampa", "maxResults": 250 },
        { "practiceArea": "personal injury", "state": "Florida", "city": "Orlando" },
        { "practiceArea": "personal injury", "state": "Florida", "city": "Jacksonville" },
        { "practiceArea": "personal injury", "state": "Florida", "city": "Hialeah" }
    ],
    "searchMaxResultsPerTask": 25,
    "maxItems": 5000
}
```

Miami exports up to 500 and Tampa up to 250 because their rows set `maxResults`. The other three have no `maxResults`, so they fall back to `searchMaxResultsPerTask` (`25`). `maxItems` still caps the whole run at 5,000.

***

### Input reference (full schema)

| Field | Type | Default | Description |
|---|---|---|---|
| `enableSearchLawyers` | boolean | `true` | Turn on attorney search by practice area and location. |
| **`searchTasks`** | object\[] | `[]` | **One row per search task** — the main input. Columns: `practiceArea` (optional, e.g. `personal injury`; aliases like `car accident` are mapped automatically), `state` (required, e.g. `Florida` or `FL`), `city` (optional — leave empty to cover the whole state, up to 60 cities per task), `maxResults` (optional — caps just this task, falls back to the default below). Duplicate rows are collapsed. |
| `searchMaxResultsPerTask` | integer | `10` | How many attorneys a task may export when its own `maxResults` is empty. Defaults to a small sample of 10 — raise it for a deeper pull per market. |
| `searchDepth` | integer | `20` | How far into each task's results to go (1–100). Usually the task's result cap should do the limiting; depth is the safety ceiling. |
| `enableSearchFirms` | boolean | `false` | Turn on law-firm search. |
| **`firmSearchTasks`** | object\[] | `[]` | **One row per firm search task** — same columns as `searchTasks`. |
| `firmSearchMaxResultsPerTask` | integer | `10` | How many firms a task may export when its own `maxResults` is empty. Defaults to a small sample of 10. |
| `firmSearchDepth` | integer | `10` | How far into each firm task's results to go (1–100). |
| `enableLawyerProfiles` | boolean | `false` | Fetch details for specific attorney profile URLs. |
| `lawyerUrls` | string\[] | `[]` | Attorney profile URLs. |
| `enableFirmProfiles` | boolean | `false` | Fetch details for specific firm profile URLs. |
| `firmUrls` | string\[] | `[]` | Firm profile URLs. |
| `enableScrapeByUrl` | boolean | `false` | Scrape any directory URL with automatic lawyer/firm detection. |
| `scrapeUrls` | string\[] | `[]` | Directory URLs to scrape. |
| `enableLeadDetails` | boolean | `true` | **Enable lead details.** Visit each profile (and, when needed, the practice's own website) to collect phone, email, website, and social profiles. Adds a little extra time per business. **Nothing is filtered** — every business goes to the dataset even if no lead details are found. |
| `maxItems` | integer | `10000` | Global cap on total dataset rows across all features — the main cost lever. Applies on top of the per-task caps. |
| `webhookUrl` | string | `""` | Optional. Each record is POSTed here as JSON as it is collected (in addition to the dataset). |
| `webhookFormat` | string | `json` | `json` = full record object; `slack` = Slack-friendly message payload. |
| `proxyConfiguration` | object | Apify residential, US | Connection settings. Apify residential proxy (US) is on by default; switch groups/country or provide your own proxy URLs if needed. |

**Validation rules:** at least one feature must actually run (a search with practice areas/states, or profile/scrape URLs). A non-empty `webhookUrl` must be a valid `http(s)` URL. Free-plan runs are capped automatically as described above.

> **Note on lead details vs filters:** there are **no lead filters** on purpose. Filtering results by lead completeness would make runtime per 1,000 businesses unpredictable and pricing impossible. With `enableLeadDetails` the Actor collects what it can for every row and returns **everything**.

***

### Output reference (field-by-field)

Every row is a flat JSON object. Rows are tagged with `featureType`: `lawyer_search`, `lawyer_profile`, `firm_search`, `firm_profile`, or `scrape_by_url`.

**Common fields (every row):**

| Field | Type | Description |
|---|---|---|
| `featureType` | string | Which feature produced the row. |
| `scrapedAt` | string | ISO 8601 UTC timestamp of collection. |
| `source` | string | The URL the row was collected from. |
| `id` | string | null | Directory record identifier when present. |
| `url` | string | null | Directory profile URL. |
| `name` | string | null | Attorney or firm name. |
| `profileImage` | string | null | Profile photo URL when listed. |
| `city` / `state` / `address` / `zipCode` | string | null | Location fields (attorneys may also have `addressLine2` on firm rows). |
| `phone` | string | null | Phone number, normalized when possible. |
| `email` | string | null | Email address when listed (lead details). |
| `website` | string | null | Practice's own website when listed (lead details). |
| `websites` | string\[] | null | All website links found. |
| `socials` | string\[] | null | Social profile URLs (Facebook, LinkedIn, X, Instagram, YouTube, …). |
| `practiceAreas` | string\[] | null | Practice areas listed on the profile. |
| `mainPracticeArea` | string | null | First / primary practice area. |
| `yearsExperience` | integer | null | Years of experience when stated. |
| `rating` / `reviewCount` | number | null | Directory rating and review count when shown. |
| `freeConsultation` | boolean | null | Whether a free consultation is offered. |
| `superLawyers` / `superLawyersCount` | boolean / integer | null | Super Lawyers recognition and selection count when shown. |
| `enriched` | boolean | null | `true` when lead details were collected for this row. |

**Lawyer-only fields:** `firmName`, `firmUrl`, `bio`, `overview`, `avPreeminent`, `leadCounsel`, `awards`, `honors`, `languages`.

**Firm-only fields:** `addressLine2`, `description`, `firmSize`, `attorneys` (list of `{name, url, title, profileImage, superLawyers}`), `attorneyCount`, `appointments`, `creditCardsAccepted`, `virtualAppointments`.

**Example — a lawyer row:**

```json
{
    "featureType": "lawyer_search",
    "scrapedAt": "2026-09-24T12:00:00.000Z",
    "source": "https://lawyers.findlaw.com/…",
    "name": "Jane Doe",
    "firmName": "Doe & Associates",
    "firmUrl": "https://lawyers.findlaw.com/…",
    "profileImage": "https://…",
    "city": "New York",
    "state": "New York",
    "address": "350 Fifth Avenue, Suite 1000",
    "zipCode": "10118",
    "phone": "(212) 555-0123",
    "email": "info@doelaw.com",
    "website": "https://doelaw.com",
    "websites": ["https://doelaw.com"],
    "socials": ["https://www.linkedin.com/in/janedoe"],
    "practiceAreas": ["Personal Injury", "Car Accidents"],
    "mainPracticeArea": "Personal Injury",
    "yearsExperience": 18,
    "rating": 4.8,
    "reviewCount": 32,
    "freeConsultation": true,
    "superLawyers": true,
    "superLawyersCount": 5,
    "enriched": true,
    "url": "https://lawyers.findlaw.com/lawyer/attorney/…"
}
```

**Example — a firm row:**

```json
{
    "featureType": "firm_profile",
    "scrapedAt": "2026-09-24T12:00:00.000Z",
    "name": "Doe & Associates",
    "city": "New York",
    "state": "New York",
    "phone": "(212) 555-0123",
    "email": "info@doelaw.com",
    "practiceAreas": ["Personal Injury"],
    "attorneyCount": 6,
    "attorneys": [
        { "name": "Jane Doe", "url": "https://lawyers.findlaw.com/…", "title": "Founding Partner", "profileImage": null, "superLawyers": true }
    ],
    "freeConsultation": true,
    "enriched": true,
    "url": "https://lawyers.findlaw.com/…"
}
```

**Run summary (OUTPUT tab → key-value store):** every run writes an `OUTPUT` record with `featuresEnabled`, `totalPushed`, `durationMs`, `errors`, `spendingLimitReached`, and the `paywall` object:

```json
{
    "paywall": {
        "detected": true,
        "isPaying": false,
        "pricingTier": "FREE",
        "limited": true,
        "blocked": false
    }
}
```

- `detected` — platform pay-status variables were present.
- `isPaying` — the account is on a paid plan (full output).
- `pricingTier` — the Apify plan tier when known.
- `limited` — free-tier cap applied (limit mode).
- `blocked` — free-tier block mode applied (owner-configurable; the default is `limit`).

***

### Webhooks

Webhooks deliver every record **in real time** as it is collected — in addition to the dataset (the dataset always gets everything; the webhook is an extra push).

#### Setup

1. Open the Actor's input in Apify Console.
2. Paste your receiver URL into **Webhook URL** (must be `http://` or `https://`).
3. Choose **Webhook format**: `JSON (full record)` or `Slack message`.
4. Start the run. Each record is POSTed right after it is saved to the dataset.

Any receiver that accepts JSON POST requests works: **Zapier, Make (Integromat), n8n, Pipedream, Slack incoming webhooks, Microsoft Teams, Google Apps Script, your own API.**

> Apify **platform webhooks** (Console → Actors → Webhooks) work too — they fire on run lifecycle events and can include the default dataset URL. The input-UI webhook here is per-record and pushes row-by-row.

#### JSON payload

`webhookFormat: json` sends the full record — exactly the same object that lands in the dataset (see Output reference):

```json
{
    "featureType": "lawyer_search",
    "scrapedAt": "2026-09-24T12:00:00.000Z",
    "name": "Jane Doe",
    "firmName": "Doe & Associates",
    "city": "New York",
    "state": "New York",
    "phone": "(212) 555-0123",
    "email": "info@doelaw.com",
    "website": "https://doelaw.com",
    "practiceAreas": ["Personal Injury"],
    "url": "https://lawyers.findlaw.com/…"
}
```

#### Slack payload

`webhookFormat: slack` sends a Slack `text` message:

```
⚖️ *Jane Doe*
*Firm:* Doe & Associates  •  *Practice:* Personal Injury
*Phone:* (212) 555-0123
350 Fifth Avenue, Suite 1000
<https://lawyers.findlaw.com/…|View listing>
```

Delivery is best-effort: a failing receiver never stops the run or blocks dataset writes (a warning is logged instead).

***

### MCP usage

Use this Actor from any MCP-compatible client (Claude Desktop, Cursor, VS Code, …) via the Apify MCP Server:

1. Get an Apify [API token](https://console.apify.com/settings/integrations).
2. Point your MCP client at `https://mcp.apify.com` (SSE) with your token, or add the Apify MCP server entry to `claude_desktop_config.json`.
3. Ask your assistant: *"Run the FindLaw Real-Time Data actor for personal injury attorneys in Dallas, Texas, max 500 results, and give me the dataset."*

The actor exposes the standard Apify tool set (`call-actor`, `get-actor-details`, `get-dataset`, …), so the agent can start runs, stream results, and read the dataset. You can also call it programmatically:

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('YOUR_USERNAME/findlaw-real-time-data').call({
    enableSearchLawyers: true,
    searchTasks: [
        { practiceArea: 'personal injury', state: 'Texas', city: 'Dallas', maxResults: 500 },
    ],
    enableLeadDetails: true,
    maxItems: 500,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0]);
```

***

### Controlling scope & cost

| Lever | Effect |
|---|---|
| **Search tasks** | One row per task — this is your scope. Add a row per practice area × city you want. |
| **Max results** on a task row | Cap for that one task. Leave it empty to use the default. |
| `searchMaxResultsPerTask` / `firmSearchMaxResultsPerTask` | The default cap for any task that does not set its own. Default `10`. |
| `maxItems` | Global ceiling on dataset rows across every feature — the final cost lever. |
| `searchDepth` / `firmSearchDepth` | Safety ceiling on how far each task goes (1–100). |
| `enableLeadDetails` | Adds a little extra time per business for contact collection. Turn off for fast, directory-only lists. |
| **User spending limit** | Apify applies your per-run max cost automatically — the Actor stops cleanly when the limit is reached, so you are never charged beyond what you approved. |

#### Three levels of control

1. **Scope — which markets you ask for.** The task list is your scope: one row per practice area and place. Delete a row and that market is skipped.
2. **Per task — how much from each.** Set **Max results** on a row to cap that task alone; leave it empty and the task uses `searchMaxResultsPerTask`. A task stops as soon as it reaches its own limit and the rest of the run continues, so no single market can consume the run.
3. **Whole run — how much in total.** `maxItems` caps the dataset across every feature and always wins over the two levels above.

The run log tells you the cap each task resolved to (`… (personal injury / Miami / Florida) — reading results (depth 1, cap 500)…`) and reports progress against it (`9 found, 5 pushed (5/500 for this task)`), so you can see exactly how each task was regulated without opening the dataset.

A capped run is also *fast*, not just small: work that cannot be exported is skipped rather than collected and discarded.

### Dataset views

The Console dataset ships with tabbed views — **Overview, Lawyers, Firms** — so you can read results without building filters. Filter on `featureType` to split lawyers from firms.

### Error handling & reliability

- Temporary data-source issues are retried automatically with fresh connection exits.
- One unreadable listing never stops the run — the Actor moves on and the summary reports what happened.
- Failure messages are always stable, user-facing sentences; raw technical errors are never written to logs or output.
- The user spending limit is honored: when reached, the run stops gracefully with everything collected so far intact.

### Free tier (Apify free plan)

Free accounts are **limited to 2 results per run**; the run then exits gracefully with an upgrade message. Paid plans have no caps. See [Paid only / free tier](#️-paid-only--free-tier--read-this-first) at the top. The owner can tune the sample size or switch to hard-blocking via the Actor's environment variables (`FREE_TIER_MODE`, `FREE_TIER_MAX_ITEMS`) without a code change.

### FAQ

**Is this Actor available on the free plan?** Only as a tiny sample — 2 results per run. To get full, unlimited data you need a paid Apify plan. Free runs stop cleanly with an upgrade message; nothing crashes.

**Why am I limited on the free plan?** Collecting complete data has real infrastructure costs (compute + connections). The sample lets you verify data quality before upgrading.

**How do I know if a run was capped?** Check the run log (a warning is printed at the start) and the `paywall` object in the run OUTPUT — `limited: true` with your pricing tier.

**Do I need API keys or login credentials?** No. The Actor works with public directory data only — no accounts, no logins, no third-party API keys.

**Are results filtered to only rows that have emails?** No — by design. **Enable lead details** collects contact information when available, and **every business goes to the dataset** regardless. This keeps runtime predictable and pricing fair.

**How fast is it?** The Actor runs everything it can **in parallel** — independent search tasks run concurrently, and lead details flow through **two pipelined pools** (profiles, then contact hunts), the same speed model as the reference Google Maps actor. Speed is not only about width, though, so these rules keep long runs moving — each one is pinned by an offline test and was chosen from live measurements:

1. **Every worker holds one exit and reuses it.** Minting a fresh residential exit per request looks parallel but costs a handshake on *every* page — measured at roughly 5 seconds per request. An exit is now swapped only when it is actually refused.
2. **A visit to a practice's own website keeps one single exit for the entire visit**, instead of re-negotiating while looking for contact details.
3. **The proxy is never given more concurrent work than it can serve.** Pushing far more parallel requests at a gateway than it has exits makes every read stall until it times out — measurably slower than simply queueing — so in-flight requests are capped at twice the exit pool.
4. **No single page can hold a batch.** Every read has a time budget: a page that cannot be read within it is reported and the row keeps the data already collected, so one dead exit costs seconds instead of minutes.
5. **The slow half no longer blocks the fast half.** Reading a row's details and hunting contacts on the practice's own site run in *separate* pools, so the slower contact hunt never queues behind the reads.

A run that is capped (a free-tier sample, or your own `maxItems`) also **stops preparing rows it will never be allowed to export**, and a state-wide task reads **one city page rather than walking the whole state** as soon as its own limit is reached.

**Can a run look stuck?** It cannot stay silent: each task logs the stage it is on — reading results, reading profiles, checking websites — as it goes, and logs when a task's limit is reached.

The result: in live testing, lead-detail runs complete at roughly **5–7 seconds per business** — a 1,000-business run finishes comfortably inside the 10,000-second timeout. Runtime does depend on exit availability at the moment you run: if a data source is temporarily refusing requests, the Actor retries on a fresh exit, keeps the row, and moves on rather than stalling.

**Will one bad URL break my run?** No. Unreadable listings are skipped with a warning; the run continues.

**Can I schedule runs?** Yes — use Apify Schedules for any cron pattern; the same input is reused.

**Can I get both lawyers and firms in one run?** Yes — enable both searches; rows are tagged by `featureType`.

**What formats can I export?** JSON, CSV, Excel, HTML, RSS — plus the Apify API, webhooks, and integrations.

**Does it respect my spending limit?** Yes. If your per-run max cost is reached, the Actor stops collecting immediately and finishes gracefully — you are never charged beyond the limit you set.

**Is any raw source content included in results or webhooks?** No. Output and webhooks contain only the structured fields documented above.

***

### Support

Open an issue on the Actor page or contact the publisher with your run ID and your (redacted) input JSON.

***

### Contact me

Need something built beyond this Actor? I take on **custom projects** — from Apify scrapers and data pipelines to full-stack web apps.

| | |
|---|---|
| **Email** | <dubem115@gmail.com> |
| **GitHub** | [github.com/DrunkCodes](https://github.com/DrunkCodes) |

Reach out with a short description of your project and timeline — happy to discuss scope and pricing.

# Actor input Schema

## `enableSearchLawyers` (type: `boolean`):

Search for attorneys by practice area and location. Enabled by default.

## `searchTasks` (type: `array`):

Add one row per search. Fill in State (e.g. "Florida" or "FL"), and optionally a City. Leave Max results empty to use the default below.

## `searchMaxResultsPerTask` (type: `integer`):

How many attorneys each task may export when its own Max results is left empty. Defaults to a small sample of 10 — raise it for a deeper pull per market. The run-wide "Max total items" still applies on top. On the Apify free plan each run exports only a small sample — see the README.

## `searchDepth` (type: `integer`):

How far into each task's results to go (1–100). Higher values collect more per task — usually you want the task's Max results to do the limiting instead.

## `enableSearchFirms` (type: `boolean`):

Search for law firms by practice area and location instead of individual attorneys.

## `firmSearchTasks` (type: `array`):

Add one row per firm search. Same columns as the lawyer search above.

## `firmSearchMaxResultsPerTask` (type: `integer`):

How many firms each task may export when its own Max results is left empty. Defaults to a small sample of 10 — raise it for a deeper pull per market.

## `firmSearchDepth` (type: `integer`):

How far into each firm task's results to go (1–100).

## `enableLawyerProfiles` (type: `boolean`):

Collect full details for specific attorney profile URLs you already have.

## `lawyerUrls` (type: `array`):

FindLaw attorney profile URLs, one per line.

## `enableFirmProfiles` (type: `boolean`):

Collect full details for specific law firm profile URLs you already have.

## `firmUrls` (type: `array`):

FindLaw firm profile URLs, one per line.

## `enableScrapeByUrl` (type: `boolean`):

Paste any directory listing URL — lawyers or firms are detected automatically.

## `scrapeUrls` (type: `array`):

Directory URLs to collect, one per line.

## `enableLeadDetails` (type: `boolean`):

Visit each attorney/firm profile for phone, email, website, and social profiles. Adds a little extra time per business, but every business still goes to the dataset. Recommended for lead generation.

## `maxItems` (type: `integer`):

Global cap on total dataset rows across all features. Set high for large runs (e.g. 10000). NOTE: On the Apify free plan the global total is capped to a small sample per run; paid plans get the full value you set here. See the README.

## `webhookUrl` (type: `string`):

Optional. Every record is always saved to the run's dataset — this webhook is an ADDITIONAL real-time push. When set, each new record is also POSTed to this URL (CRM, Slack incoming webhook, Zapier, Make, Google Sheets).

## `webhookFormat` (type: `string`):

json = full record object; slack = Slack-friendly message payload.

## `proxyConfiguration` (type: `object`):

Apify residential proxy (US) is enabled by default for reliable data collection.

## Actor input object example

```json
{
  "enableSearchLawyers": true,
  "searchTasks": [
    {
      "practiceArea": "personal injury",
      "state": "Florida",
      "city": "Miami",
      "maxResults": 10
    },
    {
      "practiceArea": "criminal defense",
      "state": "New York",
      "maxResults": 10
    }
  ],
  "searchMaxResultsPerTask": 10,
  "searchDepth": 20,
  "enableSearchFirms": false,
  "firmSearchTasks": [
    {
      "practiceArea": "personal injury",
      "state": "Texas",
      "city": "Houston",
      "maxResults": 10
    }
  ],
  "firmSearchMaxResultsPerTask": 10,
  "firmSearchDepth": 10,
  "enableLawyerProfiles": false,
  "lawyerUrls": [],
  "enableFirmProfiles": false,
  "firmUrls": [],
  "enableScrapeByUrl": false,
  "scrapeUrls": [],
  "enableLeadDetails": true,
  "maxItems": 10000,
  "webhookUrl": "",
  "webhookFormat": "json",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `allResults` (type: `string`):

Complete dataset with every field from all enabled features in this run.

## `lawyers` (type: `string`):

Attorney rows from practice-area and location search.

## `firms` (type: `string`):

Law-firm rows from firm search and firm profile URLs.

## `runSummary` (type: `string`):

Per-run metadata: features enabled, records pushed, errors, spending-limit state, and the paywall object.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "enableSearchLawyers": true,
    "searchTasks": [
        {
            "practiceArea": "personal injury",
            "state": "Florida",
            "city": "Miami",
            "maxResults": 10
        },
        {
            "practiceArea": "criminal defense",
            "state": "New York",
            "maxResults": 10
        }
    ],
    "searchMaxResultsPerTask": 10,
    "searchDepth": 20,
    "firmSearchTasks": [
        {
            "practiceArea": "personal injury",
            "state": "Texas",
            "city": "Houston",
            "maxResults": 10
        }
    ],
    "firmSearchMaxResultsPerTask": 10,
    "firmSearchDepth": 10,
    "enableLeadDetails": true,
    "maxItems": 10000,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("b2b_leads/findlaw-real-time-data-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "enableSearchLawyers": True,
    "searchTasks": [
        {
            "practiceArea": "personal injury",
            "state": "Florida",
            "city": "Miami",
            "maxResults": 10,
        },
        {
            "practiceArea": "criminal defense",
            "state": "New York",
            "maxResults": 10,
        },
    ],
    "searchMaxResultsPerTask": 10,
    "searchDepth": 20,
    "firmSearchTasks": [{
            "practiceArea": "personal injury",
            "state": "Texas",
            "city": "Houston",
            "maxResults": 10,
        }],
    "firmSearchMaxResultsPerTask": 10,
    "firmSearchDepth": 10,
    "enableLeadDetails": True,
    "maxItems": 10000,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("b2b_leads/findlaw-real-time-data-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "enableSearchLawyers": true,
  "searchTasks": [
    {
      "practiceArea": "personal injury",
      "state": "Florida",
      "city": "Miami",
      "maxResults": 10
    },
    {
      "practiceArea": "criminal defense",
      "state": "New York",
      "maxResults": 10
    }
  ],
  "searchMaxResultsPerTask": 10,
  "searchDepth": 20,
  "firmSearchTasks": [
    {
      "practiceArea": "personal injury",
      "state": "Texas",
      "city": "Houston",
      "maxResults": 10
    }
  ],
  "firmSearchMaxResultsPerTask": 10,
  "firmSearchDepth": 10,
  "enableLeadDetails": true,
  "maxItems": 10000,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call b2b_leads/findlaw-real-time-data-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,b2b_leads/findlaw-real-time-data-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pp9EtPt4K6EbB0bQF/builds/VBTlr6IFPtE0SuSLb/openapi.json
