# Trustpilot Rating Star Breakdown and Review Count Enricher (`mambalabs/review-platform-reputation-enricher`) Actor

Returns the Trustpilot TrustScore, star score, total review count and the full one to five star breakdown, read from Trustpilot own public TrustBox data endpoint. No API key. Takes a Trustpilot business unit id, or a domain when the company publishes its id on its own site.

- **URL**: https://apify.com/mambalabs/review-platform-reputation-enricher.md
- **Developed by:** [Mamba Labs](https://apify.com/mambalabs) (community)
- **Categories:** Lead generation, Automation, E-commerce
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.10 / 1,000 company checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

### 🔎 What can Trustpilot Rating Star Breakdown and Review Count Enricher do?

Give it a Trustpilot business unit id and it returns that company TrustScore, star score, total review count and the full one to five star breakdown. One flat row per company.

Give it a domain instead and it looks for the id in the company own Trustpilot widget embed. Measured on 42 domains: about one in seven publishes it, and it is concentrated in UK consumer services rather than spread evenly. When it cannot find one the row says `not_extractable`, which means nobody could look, and it is not the same claim as "this company is not on Trustpilot".

It reads Trustpilot own public TrustBox data endpoint, which Trustpilot robots.txt explicitly allows. No API key, no scraping of anything Trustpilot has asked us to leave alone.

There is a second route, and it is **off unless you turn it on**. Set `publicPageAccess` to `unblocker` and the actor reads Trustpilot public company pages instead, which resolves any domain with no id at all and adds categories, claimed status and verification. Trustpilot robots.txt disallows those pages to automated agents, so that is your call to make and not a default. It runs on your Apify account and your Unblocker units, and every row records `trustpilot_route` so your dataset says which route produced it.

| 📦 What you get | ⚙️ Features and integrations |
|---|---|
| ⭐ **TrustScore, star score and total review count**<br>📊 **The full one to five star breakdown**, not just the average<br>🧾 **31 flat fields**, `snake_case`, one row per company<br>⚪ **G2, Capterra and Glassdoor** ship as explicit skipped columns | 🆓 **No API key**, Trustpilot own public TrustBox data endpoint<br>🤖 **robots.txt read and honored** before every request<br>🚦 **Four kinds of empty**, each with its own status<br>🔓 **Optional public page route**, off by default, your call<br>⬇️ **Export** to JSON, CSV, Excel, HTML or XML |

Bought by teams qualifying on public reputation, marketplace and partnership reviewers, and agencies auditing a prospect's review posture before a pitch.

> 🚫 **A domain alone will not resolve most companies, and you should know that before you plan a list around it.** Measured 2026-08-23 across 42 domains: six published the business unit id this actor needs, about one in seven, and they were almost all UK consumer services. Trustpilot only resolves domains from hosts its robots.txt forbids us, so this is a permanent shape rather than a bug being fixed. If you hold Trustpilot ids, or you are enriching companies you can look up once by hand, this is exact and cheap. If you are feeding it a cold list of domains, most rows will come back not\_extractable unless you switch on the public page route, which is your compliance call and uses your own Unblocker units.

### 💡 Why use Trustpilot Rating Star Breakdown and Review Count Enricher?

**The star breakdown, not just the score.** A 4.2 built from a flat spread and a 4.2 built from a pile of fives and a pile of ones are different companies, and only one of them has an angry customer base worth knowing about before you call.

**Zero reviews is a real answer and it is kept as one.** A listed business nobody has reviewed returns `trustpilot_review_count: 0` with a null rating. That is a different and more useful fact than not being on Trustpilot at all, and most tools flatten both to an empty cell.

**`rating_is_meaningful` stops you comparing incomparable numbers.** A 5.0 from two reviews and a 4.2 from nine hundred are not the same measurement.

**Four kinds of empty, never one.** `not_extractable` (no id could be found), `blocked` (Trustpilot refused), `not_found` (Trustpilot answered and does not know that id) and `skipped` (nothing was asked) are separate statuses with separate meanings. Only one of them is about the company.

#### 🧭 The route this reads, and the two it refuses to

**Trustpilot, through the public TrustBox data endpoint at `widget.trustpilot.com/trustbox-data/`.** Trustpilot robots.txt names that path in an `Allow` line for the default user agent. It answers plain JSON with no key and no challenge, and it carries the TrustScore, the star score, the total and the full star breakdown.

**Two other routes exist and neither is used.**

`api.trustpilot.com`, the documented Business Units API, returns HTTP 403 with a CloudFront block on every route including the bare host root, measured from three separate network positions on 2026-08-22 and again on 2026-08-23. That is an edge block on the whole host, not a key problem, and no amount of key buying was going to change it.

`www.trustpilot.com/review/{domain}`, the public company page, is what most competing tools read. Its robots.txt says `Disallow: /` for every agent Trustpilot has not named, so **this actor does not read it unless you set `publicPageAccess` to `unblocker`**. That default is the reason a domain alone often cannot be resolved here, and it is a deliberate choice rather than a gap.

When you do switch it on, the page is fetched through Apify Unblocker, which is what clears the challenge Trustpilot puts in front of it. Measured 2026-08-23: a plain request returns an AWS WAF challenge from a direct connection, from the residential pool and from the datacenter pool, and only Unblocker got the page. What comes back carries the same TrustScore, star score, total and star breakdown as the widget, plus categories, claimed status and verification.

**G2, Capterra and Glassdoor ship as `skipped` columns.** Probed on 2026-08-22, all three refused every documented route, and G2 and Capterra terms explicitly forbid scraping their review corpora. Glassdoor rates employers, which is a different product.

Those columns exist rather than being left out, so your table does not change width the day a source is added, and so you can never mistake "we did not look" for "there is nothing there".

### 📋 What data can Trustpilot Rating Star Breakdown and Review Count Enricher extract?

Every row carries **33 flat fields**. These are the ones a buyer
actually filters and sorts on.

| Field | Type | Meaning |
|---|---|---|
| `degraded` | boolean | True when this row could not be produced normally, for example the company site was unreachable and no discovery could run. A degraded row is never charged. |
| `degradation_reason` | string | null | Why the row is degraded, in plain words. Null on a normal row. |
| `company_domain` | string | null | The company domain this row is about, normalized. Null when only a handle or a name was supplied. This is the join key across the whole Mamba Labs fleet. |
| `company_name` | string | null | The company name as supplied or derived. Improves search accuracy and is what the identity gate matches against. |
| `trustpilot_business_id` | string | null | The stable Trustpilot business unit identifier, and the key everything here is read by. Store it: a display name can change and this cannot, and having it means the next run resolves without a discovery step. Populated even on a not\_found row, so you can see which id was tried. |
| `trustpilot_url` | string | null | The public Trustpilot review page for this business. |
| `trustpilot_name` | string | null | The business name as Trustpilot displays it, which is sometimes a trading name rather than the legal entity. |
| `trustpilot_rating` | number | null | The TrustScore, 1 to 5. NULL when the business is registered and has no reviews, which is a different fact from a low rating. Read trustpilot\_review\_count before reading this number. |
| `trustpilot_review_count` | integer | null | How many reviews the business has. ZERO is a real answer and means a registered business nobody has reviewed, which is a genuinely different signal from not being on Trustpilot at all. |
| `trustpilot_stars` | number | null | The rounded star score Trustpilot displays, which is not the same number as the TrustScore and is what a visitor to the page actually sees. |
| `trustpilot_one_star` | integer | null | How many one star reviews the business has. The distribution is the part a rating hides: 4.2 from a flat spread and 4.2 from a barbell of fives and ones are very different companies to sell to. |
| `trustpilot_two_stars` | integer | null | How many two star reviews the business has. |
| `trustpilot_three_stars` | integer | null | How many three star reviews the business has. |
| `trustpilot_four_stars` | integer | null | How many four star reviews the business has. |
| `trustpilot_five_stars` | integer | null | How many five star reviews the business has. Read it against trustpilot\_one\_star: the ratio of the two ends is the fastest read on whether a rating is earned or fought over. |
| `trustpilot_country_code` | string | null | The country Trustpilot files the business under, ISO two letter. Worth checking before comparing ratings: Trustpilot review culture is not the same market to market. |
| `trustpilot_route` | string | null | Which Trustpilot surface produced this row. widget is the permitted TrustBox data endpoint and is the default. public\_page means you set publicPageAccess to unblocker and the company page was read instead. On the row so a dataset can always be traced back to the choice that produced it. |
| `trustpilot_id_source` | string | null | Where the business unit id came from: input when you supplied it, homepage\_embed when it was read out of the company own Trustpilot widget, public\_page when the company page carried it directly. Null when none was found. This is the column that tells you whether a blank row is worth retrying with an id. |
| `proxy_pool_used` | string | null | Which proxy pool served the Trustpilot request: default first, residential only if the default pool was refused, unblocker when you switched the public page route on. On the row so a run that quietly got expensive is visible rather than inferred from a bill. |
| `trustpilot_claimed_status` | string | null | not\_extractable on the default widget route, because whether a business has claimed its Trustpilot profile is published only on the public company page. ok when you switched publicPageAccess to unblocker and the page answered, in which case trustpilot\_is\_claimed carries a real true or false. A status beats a null nobody can interpret. |
| `trustpilot_categories_status` | string | null | not\_extractable on the default widget route, for the same reason as the claimed status: categories live on the public company page. ok when the public page route answered and the business is listed under at least one category, not\_found when it answered and the business is listed under none. |
| `trustpilot_is_claimed` | boolean | null | True when the company has claimed its Trustpilot profile. FALSE means checked and unclaimed, which is a real trust signal and often means the company is not managing its reputation there. Null means we could not check. |
| `trustpilot_is_verified` | boolean | null | True when Trustpilot has verified the business. FALSE means checked and unverified. Null means we could not check. |
| `trustpilot_categories` | string | null | The categories the business is listed under, comma separated. A cheap and surprisingly good proxy for what a company actually sells, which is often not what its homepage says. |
| `rating_is_meaningful` | boolean | null | True when the review count is at or above the minReviewCount threshold you set. FALSE means checked and below it, so the rating is real but thin. Null when you set no threshold or there is no rating. A 5.0 from two reviews and a 4.2 from nine hundred are not comparable numbers. |
| `trustpilot_status` | string | null | ok, not\_found, blocked, auth\_failed or skipped. NOT\_FOUND means Trustpilot was asked and has no business unit for this domain, which is true of most B2B software. BLOCKED means Trustpilot refused the request, which says nothing at all about the company. Never read the two as the same thing. |
| `source_note` | string | null | Plain words about anything that stopped a source answering, for example a refused request or a missing key. Null on a clean row. |
| `g2_status` | string | null | Always "skipped" in v1. G2 refused every documented route probed on 2026-08-22 and publishes no API we can use as documented, so this actor does not look. The column exists rather than being omitted so your table does not change width when a source is added, and so you never read "we did not look" as "there is nothing there". |
| `capterra_status` | string | null | Always "skipped" in v1, for the same reason as G2. Capterra terms explicitly forbid scraping its review corpus and it publishes no usable API. |
| `glassdoor_status` | string | null | Always "skipped" in v1. Glassdoor is EMPLOYER ratings, which is a different product from product reviews, and it belongs with a workplace or employer brand actor rather than here. Grouping it with the other three was a category error in the original plan and it is corrected rather than carried forward. |
| `coverage` | number | null | How much of what this actor can return actually came back on this row, from 0 to 1. Computed over this actor value fields only, never over the identity or status columns. Null on a degraded row, where nothing was attempted. This is a reporting field: nothing is dropped for low coverage and no event fires on it. |
| `fetch_status` | string | ok, not\_found, not\_extractable, blocked, identity\_mismatch, auth\_failed or skipped. Read this before reading any value on the row. not\_found means we looked and there is nothing there; blocked and not\_extractable mean we could not look, and they must never be read as an absence. |
| `run_date` | string | ISO 8601 timestamp of this run. Social counts move, so a row without a date is a number with no shelf life. |

> ⚠️ **How to read these values.** `trustpilot_business_id` is the field to store. Feed it back in as `trustpilotBusinessUnitId` and every later run resolves first time with no discovery step, which is both faster and the only way to get a reliable answer for the eleven companies in twelve that do not publish their id.
>
> `trustpilot_id_source` tells you whether a blank row is worth retrying. `homepage_embed` means the company published it. Null means nobody could find one, and supplying it yourself will fix that row.

### 🛠️ How to get a company Trustpilot rating from a domain

1. Put a Trustpilot business unit id in `trustpilotBusinessUnitId`. It is the 24 character hex string in `data-businessunit-id` on any page running a Trustpilot widget. This always resolves.
2. Or put a company domain in `company_domain` and let the actor look for the id on that company own site. Measured on 42 domains, it finds one for about one in seven, and far more often in UK consumer services than in DTC ecommerce or B2B software.
3. Set `minReviewCount` if you want a ready made column telling you whether a rating is thin.
4. If you would rather resolve every domain and you accept reading pages Trustpilot robots.txt disallows, set `publicPageAccess` to `unblocker`. It uses Unblocker units from your own account.
5. For a list, pass an array of objects with the same fields.

#### 🧪 Using it in Clay

Add an **Enrichment > Apify** column, pick this actor, and map `company_domain`. If you already hold Trustpilot ids, map them to `trustpilotBusinessUnitId` instead and every row resolves.

Filter on `trustpilot_status` before filtering on the rating. `not_extractable` and `blocked` rows carry no information about the company and must not be scored as bad reputations.

### 💵 How much does it cost?

Pay per event. You are charged for output, never for input.

| Event | Fires when | Price |
|---|---|---|
| `company-checked` | Once per company for which the Trustpilot business unit lookup completed and a non degraded row was produced, whether or not anything was found. A not\_found row fires this event, because looking and finding nothing is a real answer and it is the work you asked for. A degraded row, where the lookup could not run at all, fires nothing. | $0.0030 |
| `reputation-resolved` | Once per company where a Trustpilot business unit was found and its aggregate record returned. Does not fire on a company with no business unit, does not fire when Trustpilot refused the request, and does not fire when no API key was supplied, because in all three cases no reputation record was produced. | $0.0035 |
| `public-page-profile-returned` | Once per company resolved through the optional public page route, which only runs when you set publicPageAccess to unblocker. It replaces the reputation resolved charge on that row rather than adding to it, so a row is never charged twice, and a run that leaves the setting alone never fires this event at all. | $0.0050 |

> 💳 **What you are billed for.** `reputation-resolved` fires only when a business unit was actually read and returned. A row that could not find an id, a refused request and a not\_found id all charge nothing for the resolution.
>
> **If you switch `publicPageAccess` to `unblocker`, resolved rows charge `public-page-profile-returned` at $0.0050 instead of `reputation-resolved` at $0.0035, never both.** That route moves a whole page through Apify Unblocker rather than about a kilobyte of JSON, it costs us roughly ten times as much to serve, and it returns three fields the default route cannot. Leave the setting alone and nothing about your bill changes.
>
> There is no per review event, because this actor returns an aggregate rather than review rows, and an event named per review that fired on an aggregate would be dishonest.

**What the same coverage costs bought a la carte:** The Apify Store carries Trustpilot review scrapers priced per review, which is a different product: they take a Trustpilot URL and return review rows, and they get there by reading pages Trustpilot robots.txt disallows. This actor returns the aggregate and the full star breakdown from a path Trustpilot explicitly allows, which is a narrower product bought by people who care that it is. No per review event is charged here because no review rows are returned.

### ⌨️ Input

| Field | Type | Required | Meaning |
|---|---|---|---|
| `company_domain` | string | no | Bare company domain, for example monzo.com. Used to look for a Trustpilot business unit id in that company own widget embed. Supplying trustpilotBusinessUnitId instead skips this step and always resolves. |
| `company_name` | string | no | Optional. Carried through to the output row for joining. The lookup is keyed on the business unit id, so the name never changes which business is returned. |
| `publicPageAccess` | string | no | Leave at "off" (default) and this actor reads only Trustpilot public TrustBox data endpoint, which Trustpilot robots.txt explicitly allows. Set to "unblocker" and it will instead read Trustpilot public company pages through Apify Unblocker, which resolves any domain with no business unit id and also returns categories, claimed status and verification. Two things to know before you set it: Trustpilot robots.txt disallows those pages to automated agents, so this is your call and not a default, and it spends Unblocker units from your own Apify account. Every row records trustpilot\_route, so your dataset always says which route produced it. Sent as a string for Clay compatibility. |
| `trustpilotBusinessUnitId` | string | no | The 24 character Trustpilot business unit id, for example 57da77be0000ff000594bcdb. This is the fastest and most reliable input: supplied here it always resolves. Find it in the page source of any site running a Trustpilot widget, as data-businessunit-id. Leave it empty and the actor looks for it on the company own site. |
| `minReviewCount` | string | no | Sets rating\_is\_meaningful on the row. A 5.0 rating from two reviews and a 4.2 from nine hundred are not comparable numbers, and this is the column that says which one you are looking at. It never drops a row and never changes the rating returned. Sent as a string for Clay compatibility. |
| `includeCategories` | string | no | Has no effect in this version. Trustpilot publishes a business categories only on its public company page, and that page is disallowed to us by Trustpilot own robots.txt, so trustpilot\_categories is null on every row and trustpilot\_categories\_status says not\_extractable. The setting is kept so saved configurations keep working. Sent as a string for Clay compatibility. |
| `sources` | string | no | Which review sources to query. This version serves Trustpilot only. G2, Capterra and Glassdoor appear as columns and always report skipped, with the reason on the row, because all three refused every documented route and none publishes an API we can use as documented. This setting exists so a saved configuration keeps working when a source is added. Sent as a string for Clay compatibility. |
| `skipCache` | string | no | When "false" (default) a successful lookup is cached for seven days and reused, which costs you nothing on a repeated run. Set "true" to force a fresh fetch. Sent as a string for Clay compatibility. |

```json
{
  "company_domain": "monzo.com",
  "company_name": "Monzo",
  "sources": "trustpilot"
}
```

### 📤 Output

Exports to **JSON, CSV, Excel, HTML or XML**. One flat, snake\_case row per
company. No nested objects, so it drops straight into Clay, a spreadsheet or a
warehouse table without a flattening step.

```json
{
  "degraded": false,
  "degradation_reason": null,
  "company_domain": "mattressonline.co.uk",
  "company_name": "Mattress Online",
  "trustpilot_business_id": "4992d8e10000640005041a9a",
  "trustpilot_url": "https://www.trustpilot.com/review/www.mattressonline.co.uk",
  "trustpilot_name": "MattressOnline.co.uk",
  "trustpilot_rating": 4.8,
  "trustpilot_review_count": 72332,
  "trustpilot_stars": 5,
  "trustpilot_one_star": 1618,
  "trustpilot_two_stars": 768,
  "trustpilot_three_stars": 1485,
  "trustpilot_four_stars": 5850,
  "trustpilot_five_stars": 62611,
  "trustpilot_country_code": "GB",
  "trustpilot_route": "widget",
  "trustpilot_id_source": "homepage_embed",
  "proxy_pool_used": "default",
  "trustpilot_claimed_status": "not_extractable",
  "trustpilot_categories_status": "not_extractable",
  "trustpilot_is_claimed": null,
  "trustpilot_is_verified": null,
  "trustpilot_categories": null,
  "rating_is_meaningful": true,
  "trustpilot_status": "ok",
  "source_note": null,
  "g2_status": "skipped",
  "capterra_status": "skipped",
  "glassdoor_status": "skipped",
  "coverage": 1,
  "fetch_status": "ok",
  "run_date": "2026-08-23T06:39:38.824Z"
}
```

#### false versus null, and why the difference matters

`false` means we looked and the answer is no. `null` means we could not look,
or the platform withheld it. They are never interchangeable in this output. If
you filter for companies with no presence on this platform, filter on `false`,
because `null` rows are unknown rather than absent and including them will
overstate your list.

### 💡 Tips

- **Store `trustpilot_business_id` the first time you get it.** It turns a one in seven hit rate into a hundred percent one on every later run.
- **Read `trustpilot_review_count` before `trustpilot_rating`, every time.** The count is what makes the rating mean anything.
- **Read the one star count against the five star count.** The ratio of the two ends says more about a company support load than the average does.
- **Do not disqualify B2B software on `not_found`.** Most of it is not on Trustpilot, which is a consumer weighted platform.
- **Never read `not_extractable` as a bad reputation.** It means no id was published, not that anything was measured.
- **`trustpilot_is_claimed: false` is a sales signal**, not just a data point. It usually means nobody at the company is watching that page. You only get it on the public page route.
- **Filter on `trustpilot_route` before comparing rows** if you have mixed runs. A widget row and a public page row are not the same width of answer.

### ⚠️ Known limits

- **A domain alone resolves for a minority of companies, and which minority is not random.** Measured 2026-08-23 across 42 domains: six published their business unit id on their own site, about one in seven. UK consumer services embed Trustpilot widgets heavily and resolve well. DTC ecommerce and B2B software mostly do not. Trustpilot serves domain to business unit lookup only from hosts whose robots.txt forbids us, so there is no permitted way to close that gap. Supplying `trustpilotBusinessUnitId` resolves every time.
- **Claimed status, verified status and categories are not returned on the default route.** They live on the public company page. The columns are present and `trustpilot_claimed_status` and `trustpilot_categories_status` say `not_extractable` rather than leaving you a null to interpret. Switching `publicPageAccess` to `unblocker` fills all three.
- **The public page route needs Apify Unblocker and it is not free.** If Unblocker is not available on your account the actor says so in the log and falls back to the permitted route rather than pretending the company was not found.
- **Trustpilot skews consumer.** Ecommerce, financial services and consumer services are well covered; B2B software largely is not.
- **This actor reads one source.** G2, Capterra and Glassdoor are `skipped` columns and adding any of them by scraping is a decision to take on its own terms, not by accretion.
- **No review text, no review rows, no reviewer data.** This actor returns the aggregate and the star breakdown only.

### ❓ FAQ

**Why do I need a business unit id? Every other tool takes a URL.**
Because the routes that turn a domain or a URL into a business unit are on Trustpilot hosts whose robots.txt disallows us, and this actor honors that. The id is public: it is in the page source of any site running a Trustpilot widget, as `data-businessunit-id`.

**Where do I get one quickly?**
Open the company Trustpilot widget on their own site and read `data-businessunit-id` from the source. Once you have it, store it. It never changes.

**What is the difference between `not_extractable`, `not_found` and `blocked`?**
`not_extractable` means no id could be found, so nothing was asked and nothing is known. `not_found` means Trustpilot was asked and has no business unit with that id. `blocked` means Trustpilot refused the request. Only `not_found` is about the company.

**Why is the rating null when the review count is 0?**
Because there is nothing to average. A listed business with no reviews is a real and useful state.

**Can I just resolve every domain?**
Yes, by setting `publicPageAccess` to `unblocker`, which reads Trustpilot public company pages. Understand what you are switching on: Trustpilot robots.txt disallows those pages to automated agents. The actor will not do it unless you ask, it runs on your account and your Unblocker units, and `trustpilot_route` on every row records that you asked.

**Do I need a Trustpilot API key?**
No. There is no key field any more. The API this actor used to target refuses every request from every network position we can reach, key or no key, so a key field would only have made an empty column look like your fault.

**When will G2 and Capterra be added?**
Not by scraping. Both explicitly forbid it and both refused every route probed. If either opens an API, the columns are already in the row and your table will not change shape.

### 🧩 Want other GTM data?

Mamba Labs builds a fleet of GTM enrichment actors that share one flat,
Clay-ready output convention, so their rows join on `company_domain` with no
cleaning step:

| | |
|---|---|
| 🕵️ [Agent Accessibility Auditor](https://apify.com/mambalabs/agent-accessibility-auditor) | 🤖 [AI Tooling Detector](https://apify.com/mambalabs/ai-tooling-detector) |
| 📡 [B2B Buying Signals Aggregator](https://apify.com/mambalabs/b2b-buying-signals-hiring-tech-stack-intent-for-clay) | 🚀 [Prospect Engine](https://apify.com/mambalabs/b2b-prospect-engine) |
| 📝 [Publishing Frequency Tracker](https://apify.com/mambalabs/blog-publishing-frequency) | ⏩ [Sequencer Lead Push](https://apify.com/mambalabs/clay-to-instantly-smartlead-push) |
| 🔄 [Company Change-Event Feed](https://apify.com/mambalabs/company-change-event-feed) | 🧭 [Company Discovery List Builder](https://apify.com/mambalabs/company-discovery-list-builder) |
| 🏢 [Company Firmographic Enricher](https://apify.com/mambalabs/company-firmographic-enricher) | 🪪 [Company Identity Resolver](https://apify.com/mambalabs/company-identity-resolver) |
| 🌐 [Company Social Presence Mapper](https://apify.com/mambalabs/company-social-presence-mapper) | 🏷️ [Contact Classifier](https://apify.com/mambalabs/contact-classifier) |
| 📬 [Domain Deliverability Checker](https://apify.com/mambalabs/domain-deliverability-checker) | 🔗 [Domain to LinkedIn URL Resolver](https://apify.com/mambalabs/domain-to-linkedin-url-resolver) |
| ✉️ [Work Email Waterfall Finder](https://apify.com/mambalabs/email-waterfall-orchestrator) | 🎪 [Event Presence Index](https://apify.com/mambalabs/event-presence-index) |
| 💰 [Funding and Press Signal Scanner](https://apify.com/mambalabs/funding-press-signal-scanner) | 🧑‍💼 [GTM Hiring Signal Scraper](https://apify.com/mambalabs/gtm-hiring-signal-scraper) |
| 📋 [Job Posting Monitor](https://apify.com/mambalabs/gtm-job-discovery) | 🧱 [Tech Stack Detector](https://apify.com/mambalabs/gtm-tech-stack-signal-scraper) |
| 🎯 [ICP Fit Scorer](https://apify.com/mambalabs/icp-account-lead-scoring-fit-scorer-0-100-for-clay) | 🔑 [Job Board Keyword Scanner](https://apify.com/mambalabs/job-board-keyword-signal-scanner) |
| ⚖️ [Legal Entity Resolver](https://apify.com/mambalabs/legal-entity-resolver) | 💼 [LinkedIn Company Page Mapper](https://apify.com/mambalabs/linkedin-company-presence-mapper) |
| 💬 [LinkedIn Post Tracker and Comment Capture](https://apify.com/mambalabs/linkedin-post-engager-capture) | 📸 [Instagram and Facebook Brand Mapper](https://apify.com/mambalabs/meta-brand-presence-mapper) |
| 📮 [Outbound Stack Detector](https://apify.com/mambalabs/outbound-infrastructure-fingerprint) | 📄 [Page Finder and Extractor](https://apify.com/mambalabs/page-finder-extractor) |
| 👤 [People Finder and Email Verifier](https://apify.com/mambalabs/people-finder) | 📌 [Pinterest Brand Presence Mapper](https://apify.com/mambalabs/pinterest-brand-presence-mapper) |
| 🏛️ [Government Contract Award Monitor](https://apify.com/mambalabs/public-award-monitor) | 📅 [Public Company Reporting Window Finder](https://apify.com/mambalabs/public-company-reporting-window-finder) |
| 👥 [Team Page People Extractor](https://apify.com/mambalabs/team-page-people-extractor) | 🎵 [TikTok Brand Presence Mapper](https://apify.com/mambalabs/tiktok-brand-presence-mapper) |
| 🏅 [Workplace Program Detector](https://apify.com/mambalabs/workplace-program-detector) | ▶️ [YouTube Channel Stats Extractor](https://apify.com/mambalabs/youtube-channel-transcript-extractor) |

> Every actor in the suite takes a domain or a company and returns one flat row,
> so they stack in the same Clay table without reshaping anything.

> 🛠️ **Need something custom built for you or your team?** Tell us what you are
> trying to find and we will build it. [Talk to Mamba Labs](https://mambabuilt.com/contact).

### 🆘 Support

Issues, field requests and bug reports: open an issue on the actor's Issues tab.
Mamba Labs reads every one.

> ℹ️ **Sourcing and legal.** Fields come from Trustpilot's own public TrustBox data endpoint, on a path Trustpilot robots.txt explicitly allows for any user agent. robots.txt is read and honored before every request on the default route. The optional public page route reads pages Trustpilot robots.txt disallows and is off unless you switch it on, on your own account. Ratings are public and are the aggregate Trustpilot publishes, not an assessment by us. G2, Capterra and Glassdoor columns are present and explicitly skipped rather than silently empty.

Built by [Mamba Labs](https://apify.com/mambalabs).

# Actor input Schema

## `company_domain` (type: `string`):

Bare company domain, for example monzo.com. Used to look for a Trustpilot business unit id in that company own widget embed. Supplying trustpilotBusinessUnitId instead skips this step and always resolves.

## `company_name` (type: `string`):

Optional. Carried through to the output row for joining. The lookup is keyed on the business unit id, so the name never changes which business is returned.

## `publicPageAccess` (type: `string`):

Leave at "off" (default) and this actor reads only Trustpilot public TrustBox data endpoint, which Trustpilot robots.txt explicitly allows. Set to "unblocker" and it will instead read Trustpilot public company pages through Apify Unblocker, which resolves any domain with no business unit id and also returns categories, claimed status and verification. Two things to know before you set it: Trustpilot robots.txt disallows those pages to automated agents, so this is your call and not a default, and it spends Unblocker units from your own Apify account. Every row records trustpilot\_route, so your dataset always says which route produced it. Sent as a string for Clay compatibility.

## `trustpilotBusinessUnitId` (type: `string`):

The 24 character Trustpilot business unit id, for example 57da77be0000ff000594bcdb. This is the fastest and most reliable input: supplied here it always resolves. Find it in the page source of any site running a Trustpilot widget, as data-businessunit-id. Leave it empty and the actor looks for it on the company own site.

## `minReviewCount` (type: `string`):

Sets rating\_is\_meaningful on the row. A 5.0 rating from two reviews and a 4.2 from nine hundred are not comparable numbers, and this is the column that says which one you are looking at. It never drops a row and never changes the rating returned. Sent as a string for Clay compatibility.

## `includeCategories` (type: `string`):

Has no effect in this version. Trustpilot publishes a business categories only on its public company page, and that page is disallowed to us by Trustpilot own robots.txt, so trustpilot\_categories is null on every row and trustpilot\_categories\_status says not\_extractable. The setting is kept so saved configurations keep working. Sent as a string for Clay compatibility.

## `sources` (type: `string`):

Which review sources to query. This version serves Trustpilot only. G2, Capterra and Glassdoor appear as columns and always report skipped, with the reason on the row, because all three refused every documented route and none publishes an API we can use as documented. This setting exists so a saved configuration keeps working when a source is added. Sent as a string for Clay compatibility.

## `skipCache` (type: `string`):

When "false" (default) a successful lookup is cached for seven days and reused, which costs you nothing on a repeated run. Set "true" to force a fresh fetch. Sent as a string for Clay compatibility.

## `source_tag` (type: `string`):

Internal attribution tag set by Mamba Labs on published task examples. Not required, and nothing depends on it. Leave it empty.

## Actor input object example

```json
{
  "company_domain": "monzo.com",
  "company_name": "Monzo",
  "publicPageAccess": "off",
  "minReviewCount": "none",
  "includeCategories": "true",
  "sources": "trustpilot",
  "skipCache": "false"
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset of one flat row per company, with per platform status so a blocked fetch never reads as a zero.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "company_domain": "monzo.com",
    "company_name": "Monzo",
    "trustpilotBusinessUnitId": ""
};

// Run the Actor and wait for it to finish
const run = await client.actor("mambalabs/review-platform-reputation-enricher").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "company_domain": "monzo.com",
    "company_name": "Monzo",
    "trustpilotBusinessUnitId": "",
}

# Run the Actor and wait for it to finish
run = client.actor("mambalabs/review-platform-reputation-enricher").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "company_domain": "monzo.com",
  "company_name": "Monzo",
  "trustpilotBusinessUnitId": ""
}' |
apify call mambalabs/review-platform-reputation-enricher --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mambalabs/review-platform-reputation-enricher"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3lIqfCt61j4hqE1xz/builds/yz15kyf4ZlQ6HIiV0/openapi.json
