# Thumbtack Scraper — Pro Listings, Reviews & Ratings (`oswaldocarabano/thumbtack-pros-scraper`) Actor

Export Thumbtack service professionals: ratings, review counts and full review text, jobs booked, response time, Top Pro badge, licence signals, credentials, specialties and the pro's real ZIP. Sees 3x more pros than the visible page. No login. Thumbtack has no public API; this is the closest.

- **URL**: https://apify.com/oswaldocarabano/thumbtack-pros-scraper.md
- **Developed by:** [Oswaldo Carabano](https://apify.com/oswaldocarabano) (community)
- **Categories:** Lead generation, Business, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Thumbtack Scraper — Pro Listings, Reviews & Ratings

Export public data on **Thumbtack service professionals** (US): who they are, how
well they are rated, how much work they actually win, how fast they reply, and
the **full text of their reviews**.

Thumbtack has no public API. This is the closest thing to one.

**Start for free.** You pay per professional delivered — never for a failed page,
and never twice for the same professional.

***

### What you can use it for

- **Prospecting and lead research** for anyone who sells to contractors — software,
  insurance, financing, marketing, equipment, franchising.
- **Market analysis** of a trade in a metro: how many plumbers, HVAC contractors
  or roofers compete in a ZIP, how they rate, how much they get hired.
- **Review mining**: pull every review for a professional, with rating, date, the
  job type and the pro's own reply.
- **Competitor and reputation monitoring** for home-service brands and franchises.
- **Local SEO and directory building** from a category-by-city sweep.

***

### Why this scraper sees more than the page does

Thumbtack's category pages show **seven to nine** professionals and hide the rest
behind a "See more" button. Measured across 192 city-and-category pairs:

| where the professional was found | count |
|---|---|
| on the visible page only | 243 |
| on the page **and** behind "See more" | 883 |
| **only behind "See more"** | **1,172** |
| **unique professionals** | **1,603** |

**73% of the professionals in the output are ones the visible page never shows.**

Coverage comes from Thumbtack's own sitemaps: **905,197 URLs**, of which
**406,500** are professional profiles, across **3,264 categories** and
**3,656 cities**.

And a warning about the competition: Thumbtack's `?page=2` does nothing. It
returns the same professionals as page 1, with HTTP 200, forever. Any scraper
that claims deep pagination on Thumbtack is charging you for repeats.

***

### What you get, with measured fill rates

Every percentage below was counted on **1,603 professionals** across 192
categories and 160 cities, and on **110 profile pages**. We do not quote a field
we have not measured.

#### On every row

| field | filled |
|---|---|
| `business_name`, `url`, `service_pk` | **100%** |
| `num_hires` — jobs actually booked through Thumbtack | **100%** |
| `avg_response_time_hours` — how fast they reply | **100%** |
| `business_facts` | **100%** |
| `service_area_text` | 97.3% |
| `rating` and `num_reviews` | **96.4%** |
| `highlighted_review` | 84.9% |
| `review_qualifier` | 75.4% |
| `description` | 55.1% |
| `urgency_signals` — `LICENSED`, `RESPONSIVE`, `LOW_PRICE`, 18 codes in all | 35.2% |
| `price_min` / `price_unit` — what Thumbtack publishes, when it publishes it | **27.8%** / 19.7% |
| `is_top_pro` | 19.8% |
| `is_online` | 14.5% |
| `is_licensed_signal` | 4.6% |

#### With **Also open each professional's profile page**

Measured on 110 profiles, **zero errors**:

| field | filled |
|---|---|
| `postal_code` — the professional's **real** ZIP, not the search ZIP | **90.9%** |
| `locality`, `region`, `schema_type` | 90.9% |
| `introduction` — their own pitch, in full | 88.2% |
| `years_in_business`, `num_employees` | 88.2% |
| `specialty_groups` — exactly what they do | 88.2% |
| `review_histogram` — the 1-to-5 star breakdown | 88.2% |
| `reviews` — full text, author, date, labels and the pro's reply | 88.2% |
| `credentials`, `is_background_checked` | 87.3% |
| `num_photos` | 87.3% |
| `faq` — their own Q\&A | 42.7% |
| `business_hours` | 37.3% |

Set **Reviews per professional** above 5 and it pages through the rest, five at a
time, instead of stopping at what the page shows.

***

### Questions people ask before buying

**Is there a Thumbtack API?**
No. Thumbtack publishes no public API for professional data. This actor reads the
same public pages a browser does and returns them as structured JSON or CSV.

**Does it return phone numbers or email addresses?**
No, and no scraper can: **Thumbtack does not publish them**. Contact happens
inside their platform. If a tool promises you Thumbtack phone numbers, it is
promising something that is not on the page. What this actor gives you instead is
the reputation and demand signals — ratings, review text, hires, response time,
credentials — which is what tells you which professionals are worth approaching.

**How many professionals will one city and category actually return?**
About **14**, measured as the average across the 15 example tasks, each sweeping
a whole metro area. That is not a limit of this scraper — it is how much
Thumbtack lists. Two things are worth knowing before you plan a run:

- Adding the suburbs barely helps. Austin alone returns 10 professionals;
  Austin plus five surrounding cities returns 12. **Thumbtack serves nearly the
  same regional professionals for every city slug in a metro**, so sweeping
  neighbours mostly produces duplicates, which are removed before you are charged.
- **Volume comes from breadth, not depth.** To build a large dataset, sweep many
  categories, or turn on `profilesFromSitemap` and go straight to the 406,500
  profile URLs — one request per professional, no overlap, every field filled.

There is a hard cap of 50,000 professionals per run.

**How fresh is the data?**
Every row carries `from_cache`, `fetched_at` and `data_age_hours`, so you always
know. Set `maxCacheAgeDays: 0` to force a fresh fetch of everything.

**What does it cost?**

| you pay | for |
|---|---|
| **$0.0025** | one professional: ratings, review count, hires, response time and Thumbtack's signals |
| **$0.008** | the same professional **plus the full profile**: every review, the star histogram, credentials, specialties and their real ZIP |
| **$0** | starting a run |

So 1,000 professionals cost **$2.50**, or **$8.00** with full profiles and reviews.

Three things are not billed, and they add up: **starting a run** — many scrapers
charge you before delivering anything — **error rows**, and **duplicates**, which
are removed before the meter runs. A run that finds nothing costs exactly zero.

**Does it need a Thumbtack account?**
No. No login, no session cookies, no account of any kind.

***

### Input

Start with nothing and it works. Or steer it:

```json
{
  "maxResults": 500,
  "cities": ["tx/austin", "ca/los-angeles", "ny/brooklyn"],
  "categories": ["plumbers", "electricians", "house-cleaning"],
  "scrapeProfiles": true,
  "maxReviewsPerPro": 25,
  "minRating": 4
}
```

| option | what it does |
|---|---|
| `maxResults` | Hard cap 50,000 per run. Duplicates removed **before** you are charged. |
| `cities`, `categories`, `states` | `state/city` slugs (`tx/austin`) and Thumbtack category slugs (`plumbers`). Leave empty to sweep the sitemap. |
| `scrapeProfiles` | Adds the whole profile block above. One extra request per professional. |
| `profilesFromSitemap` | Straight to the 406,500 profile URLs: one request, one professional, no overlap, every field filled. |
| `maxReviewsPerPro` | Above 5, pages through the rest of the reviews. |
| `minRating`, `minReviews`, `topProOnly`, `licensedOnly` | Filters applied before you are charged. |
| `useCache`, `maxCacheAgeDays` | Shared cache. Cached rows always declare their age. |

Popular categories: `plumbers`, `electricians`, `hvac`, `roofing`, `handyman`,
`house-cleaning`, `landscaping`, `movers`, `pest-control`, `interior-painting`,
`carpet-cleaning`, `junk-removal`, `tree-services`, `appliance-repair`,
`window-cleaning`, `home-inspection`, `deck-repair`, `lawn-mowing`.

Cities are `state/city`: `ny/new-york`, `ca/los-angeles`, `il/chicago`,
`tx/houston`, `az/phoenix`, `pa/philadelphia`, `tx/san-antonio`, `ca/san-diego`,
`tx/dallas`, `tx/austin`, `ca/san-jose`, `fl/jacksonville`, `tx/fort-worth`,
`oh/columbus`, `nc/charlotte`, and 3,641 more.

***

### Speed

- **2.8 s** per city-and-category pair.
- **1.9 s** per profile page.
- A repeat run of the same query served **100% from cache, with zero network
  requests**.

***

### Two things worth knowing before you buy

**`is_licensed_signal` is a signal, not a verification.** It means Thumbtack
displayed a licence badge. Its absence does **not** mean the professional is
unlicensed. Same for `is_background_checked`. Do not resell either as proof of
licensing.

**`badges` and `is_sponsored` came back empty** on all 1,603 rows measured. The
columns exist because the fields exist upstream, but we will not promise you data
we did not see.

***

### Good to know

- A professional whose page fails produces an **error row**, never a silent empty
  one, and you are not charged for it.
- The run resumes cleanly after a hard kill without duplicating a single row.
- Reviews are written by third parties. If one happens to contain a phone number
  or an email, it is removed and the row tells you how many times in `redactions`.
- Not affiliated with, endorsed by, or connected to Thumbtack, Inc. All
  trademarks belong to their owners. This actor collects information Thumbtack
  publishes openly to anyone with a browser.
- Removal requests: **privacy@actorstack.dev**.

# Actor input Schema

## `maxResults` (type: `integer`):

How many professionals to return. You are charged per professional delivered, never for an error row, and duplicates are removed before you are charged. Hard cap: 50,000 per run.

## `categories` (type: `array`):

Thumbtack category slugs, e.g. `plumbers`, `house-cleaning`, `electricians`, `handyman`, `movers`, `roofing`. Thumbtack publishes 3,264 categories. Leave empty to sweep whatever the sitemap offers.

## `cities` (type: `array`):

As `state/city` slugs, e.g. `tx/austin`, `ca/los-angeles`, `ny/brooklyn`. Thumbtack publishes 3,656 cities.

## `states` (type: `array`):

Two-letter codes, e.g. `tx`, `ca`. Narrows sitemap discovery without naming every city.

## `discoverFromSitemap` (type: `boolean`):

On by default, and it is how this actor gets coverage. Thumbtack publishes 489,781 city-by-category landing pages and 406,500 professional pages in its sitemaps. Paging does not work on Thumbtack: `?page=2` returns the same professionals as page 1, so any actor claiming deep pagination is charging you for repeats.

## `profilesFromSitemap` (type: `boolean`):

Reads professionals directly from the 406,500 profile URLs in the sitemap: one request equals one professional, with no overlap between cities and with every profile field filled. Slower per row, and the richest output this actor produces.

## `scrapeProfiles` (type: `boolean`):

Adds the introduction, the professional's real ZIP code, credentials, specialties, business hours, the 1-to-5 star histogram, photo count and their own Q\&A. Costs one extra request per professional and is billed as its own event.

## `maxReviewsPerPro` (type: `integer`):

The profile page carries five reviews; above that this actor pages through the rest, five at a time. Only applies when profiles are fetched.

## `minRating` (type: `integer`):

1 to 5. Professionals without a rating are excluded when you set this.

## `minReviews` (type: `integer`):

Minimum number of reviews. Professionals without a review count are excluded when you set this.

## `topProOnly` (type: `boolean`):

Thumbtack's Top Pro badge.

## `licensedOnly` (type: `boolean`):

Thumbtack surfaces a LICENSED signal on some professionals. Absence means Thumbtack did not show the signal, not that the professional is unlicensed.

## `maxConcurrency` (type: `integer`):

Leave at 4. Thumbtack's listing pages weigh about 490 KB each, so more concurrency mostly buys memory pressure.

## `useCache` (type: `boolean`):

Worth more here than on most actors: the HTML route discards 93.7% of every page downloaded, so avoiding a repeat download saves disproportionately. Cached rows always declare their age with `from_cache`, `fetched_at` and `data_age_hours`.

## `maxCacheAgeDays` (type: `integer`):

Set to 0 to force a fresh fetch of everything. Leave empty for the default of 7 days.

## `proxyConfiguration` (type: `object`):

The actor brings its own residential proxy. Set this only if you want to route through your own.

## Actor input object example

```json
{
  "maxResults": 500,
  "categories": [
    "plumbers"
  ],
  "cities": [
    "tx/austin"
  ],
  "states": [],
  "discoverFromSitemap": true,
  "profilesFromSitemap": false,
  "scrapeProfiles": false,
  "maxReviewsPerPro": 5,
  "topProOnly": false,
  "licensedOnly": false,
  "maxConcurrency": 4,
  "useCache": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

Every professional found, with ratings, hires, response time and Thumbtack's own signals.

## `reputation` (type: `string`):

Ratings, review volume, the 1-to-5 star histogram and the badges Thumbtack shows.

## `profile` (type: `string`):

What the profile page adds: introduction, the real ZIP code, credentials, specialties and hours.

## `runSummary` (type: `string`):

What was delivered, how much came from cache and how many redactions were applied. Use it to reconcile your invoice against the rows you received.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxResults": 500,
    "categories": [
        "plumbers"
    ],
    "cities": [
        "tx/austin"
    ],
    "discoverFromSitemap": true,
    "profilesFromSitemap": false,
    "scrapeProfiles": false,
    "maxReviewsPerPro": 5,
    "topProOnly": false,
    "licensedOnly": false,
    "maxConcurrency": 4,
    "useCache": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("oswaldocarabano/thumbtack-pros-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxResults": 500,
    "categories": ["plumbers"],
    "cities": ["tx/austin"],
    "discoverFromSitemap": True,
    "profilesFromSitemap": False,
    "scrapeProfiles": False,
    "maxReviewsPerPro": 5,
    "topProOnly": False,
    "licensedOnly": False,
    "maxConcurrency": 4,
    "useCache": True,
}

# Run the Actor and wait for it to finish
run = client.actor("oswaldocarabano/thumbtack-pros-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxResults": 500,
  "categories": [
    "plumbers"
  ],
  "cities": [
    "tx/austin"
  ],
  "discoverFromSitemap": true,
  "profilesFromSitemap": false,
  "scrapeProfiles": false,
  "maxReviewsPerPro": 5,
  "topProOnly": false,
  "licensedOnly": false,
  "maxConcurrency": 4,
  "useCache": true
}' |
apify call oswaldocarabano/thumbtack-pros-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,oswaldocarabano/thumbtack-pros-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Xw2Y3qOLVZQKbaYY7/builds/zgfUiBWbcIooTx6bw/openapi.json
