Thumbtack Scraper - Local Service Pros Market Data avatar

Thumbtack Scraper - Local Service Pros Market Data

Pricing

from $1.80 / 1,000 results

Go to Apify Store
Thumbtack Scraper - Local Service Pros Market Data

Thumbtack Scraper - Local Service Pros Market Data

Scrape Thumbtack service pros by category and city: ratings, review counts, hires, years in business, Top Pro status. Business-level data, no personal contact details.

Pricing

from $1.80 / 1,000 results

Rating

0.0

(0)

Developer

Brandt May

Brandt May

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

Scrape Thumbtack local service professionals by category and city, and get back clean, structured business-level market data: business name, market, star rating, review count, lifetime hire count, years in business, employee count, average response time and Thumbtack's Top Pro badge.

This is a market-research and competitive-intelligence tool. Use it to size a local service market, benchmark your own listing against the competition, track how ratings and hire counts move over time, or find which metros are crowded and which are thin.

It deliberately does not collect personal contact details. Thumbtack pros are overwhelmingly sole proprietors, which makes their phone number, email address and street address personal information under CCPA/CPRA. Other scrapers sell that data. This one does not collect it, and it actively redacts anything phone- or email-shaped that a pro typed into a free-text field. What you get is business-level market data you can defend using.


What you get

Every row is one Thumbtack pro listing.

FieldTypeDescription
businessNamestringThe business name as published on Thumbtack. Phone numbers and emails typed into the name are redacted.
categorystringThe service category you asked for, as you wrote it.
cityStatestringThe market the pro serves, e.g. Houston, TX. This is a metro, never a street address.
ratingnumber | nullAverage star rating, 1-5. null when the pro has no reviews yet.
reviewCountnumber | nullNumber of reviews behind that rating.
hireCountnumber | nullLifetime hires on Thumbtack. A strong proxy for how established the business is.
yearsInBusinessnumber | nullDerived from the founding year Thumbtack publishes.
topProbooleantrue when Thumbtack has awarded the listing its Top Pro badge.
servicesTextstring | nullServices listed for this pro in this category.
employeeCountnumber | nullEmployee count Thumbtack publishes for the business.
avgResponseTimeHoursnumber | nullAverage hours the pro takes to respond, rounded to 0.1 h, so 0 means under about 3 minutes. null when Thumbtack reports exactly 0: in our checks that value appeared only on new or low-activity pros whose response time had not been measured yet, so this Actor does not pass it off as an instant response.
profileUrlstring | nullPublic Thumbtack profile URL, with tracking parameters stripped. null in the rare case where the pro typed a phone number into their business name, because Thumbtack builds the URL slug from that name and publishing the URL would leak the number.
scrapedAtstringISO 8601 timestamp of collection.

Rows are de-duplicated across the whole run: a pro who appears on several browse pages is returned once, and you are not billed twice for them.

A SUMMARY record is also written to the key-value store with the row count, how many browse pages were requested, how many listings were skipped as already-collected duplicates, which sources returned nothing and why, and whether the run stopped on its time budget.

Input

FieldTypeDefaultDescription
categoriesstring list["House Cleaning"]Thumbtack service categories, written the way Thumbtack writes them.
locationsstring list["Houston, TX"]US markets in City, ST format. Full state names work too.
maxItemsinteger50Stop after this many pros.
maxRunSecondsinteger240Wall-clock budget. On expiry the run stops cleanly and keeps the rows collected so far.
includeNearbyAreasbooleanfalseFollow Thumbtack's own "In other nearby areas" links to neighbouring browse pages. Many more pros, many more requests.
requestDelayMsinteger1200Politeness delay between requests. Requests are always serial, never parallel.

Every category is scraped in every location, so 3 categories and 4 cities is 12 browse pages.

Example output

{
"businessName": "Brenes Pro Cleaning Services",
"category": "House Cleaning",
"cityState": "Houston, TX",
"rating": 5,
"reviewCount": 40,
"hireCount": 99,
"yearsInBusiness": 1,
"topPro": true,
"servicesText": "House Cleaning",
"employeeCount": 2,
"avgResponseTimeHours": 0.1,
"profileUrl": "https://www.thumbtack.com/tx/houston/house-cleaning/brenes-pro-cleaning-services/service/576885179750318085",
"scrapedAt": "2026-09-24T00:46:23.880Z"
}

Data source, access and limits

Where the data comes from. Public Thumbtack category-and-city browse pages at https://www.thumbtack.com/<state>/<city>/<category> - the same pages any visitor sees, and the same pages Thumbtack submits to search engines in its own sitemap. The listing data is read from the structured LocalBusiness payload Thumbtack renders into those pages. No login, no API key and no Thumbtack account is involved.

robots.txt is honoured. The Actor fetches https://www.thumbtack.com/robots.txt at the start of every run, parses the User-agent: * group, and refuses to request any path that group disallows - including /profile, /api, /graphql, /instant-results, /request/ and /bid/. If robots.txt cannot be fetched, a built-in snapshot of the same rules is used instead, so the guard fails closed. Redirects are followed one hop at a time and each hop is re-checked against robots.txt, so the Actor cannot be redirected onto a disallowed path. The browse pages this Actor reads are not disallowed.

It behaves politely. Requests are serial with a delay, never parallel bursts. The User-Agent is honest and names Mayd It LLC with a contact address; it does not impersonate a browser. HTTP 429 and Retry-After are respected. There is no CAPTCHA solving, no browser-fingerprint spoofing and no proxy rotation to evade blocks - if Thumbtack blocks this client, the Actor reports that plainly and stops.

Known limitations - please read before you buy

  • Roughly 4-10 pros per category and city. That is how many Thumbtack publishes on a public browse page. Getting more would mean using Thumbtack's internal search endpoint, which robots.txt disallows, so this Actor does not. To collect more pros, add more categories and cities.
  • includeNearbyAreas adds far less than the extra page count suggests. Neighbouring suburbs mostly list the same pros as the metro, because pros serve the whole metro. In a measured run across six Los Angeles-area pages, Thumbtack listed 47 pro cards in total but only 13 distinct businesses - and two of those six pages contributed no new pro at all. Expect a large number of requests for a modest number of extra rows. Scraping additional, genuinely separate metros is the better way to grow a dataset.
  • No prices and no background-check flag. The browse pages this Actor reads do not publish either one: the price fields in their structured listing data were blank for every pro in every city and category we tested, and there is no per-pro background-check field (Thumbtack's own pages say every account owner must pass a background check, which is a platform rule, not a per-pro data point). Rather than ship columns that are always empty, this Actor leaves them out.
  • Some categories publish less detail. On Thumbtack's Movers pages the structured listing data carries no founding year, employee count or service list, so yearsInBusiness, employeeCount and servicesText came back null for every mover we tested (Los Angeles and Miami). The other fields fill normally there, and the trade categories (plumbing, roofing, handyman, cleaning and so on) carry all three for nearly every pro.
  • No personal contact data, by design. No phone numbers, no email addresses, no street addresses. If you need those, this is not the Actor for you.
  • Category names must match Thumbtack's own wording. A category that does not exist under the name you gave returns HTTP 404; the run reports exactly which pair failed and keeps the rest.
  • US coverage. The City, ST URL structure this Actor uses is Thumbtack's US structure.

Billing. This Actor is paid per result: you are charged for the rows it returns. Runs that return no rows fail loudly with a diagnostic instead of billing you for an empty dataset.

FAQ

Why doesn't it return phone numbers or emails like other Thumbtack scrapers? Because most Thumbtack pros are sole proprietors, so their contact details are personal information about an individual under CCPA/CPRA, not neutral business data. Collecting and reselling it creates real obligations for whoever holds it. This Actor is built for market analysis, where you do not need it. That is a deliberate product decision, not a missing feature.

How many pros will I actually get? About 4-10 per category and city pair, because that is what Thumbtack publishes publicly. Ten separate cities for one category typically lands between 40 and 100 distinct pros. Turning on includeNearbyAreas sweeps the suburbs around a metro, but those pages largely repeat the same businesses - see Known limitations, where the measured numbers are - so add more metros rather than relying on nearby areas to multiply your row count.

What happens if a category or city is wrong? That one pair is reported as a failed source with the reason, and the run continues with the others. The run only fails if every pair failed - and then the error message names the likely cause and how to fix it.

What happens on a big job that cannot finish? The Actor watches the clock. When it reaches maxRunSeconds (or the platform timeout, whichever is sooner) it stops cleanly, keeps every row already collected, and logs how many sources were left. Raise maxRunSeconds or split the job across runs.

Will this get blocked? It requests only pages robots.txt permits, serially, with a delay and an honest User-Agent, so it behaves like a well-mannered crawler rather than an evasive one. If Thumbtack does block or throttle it, the run stops and tells you - it will not try to defeat the block.

Is the data live? Yes. Every run fetches the pages fresh; scrapedAt records when. Thumbtack's rankings and badges shift over time, so re-running on a schedule is a reasonable way to track a market.


Built by Mayd It LLC. Questions or a field you need added: bmay@mayd-it.com

Mayd It LLC is not affiliated with or endorsed by Thumbtack, Inc. "Thumbtack" and "Top Pro" are trademarks of their respective owner and are used here only to describe what the data is. You are responsible for using the output in line with applicable law and Thumbtack's terms.