# Trustpilot Reviews Scraper (`valev-lab/trustpilot-reviews`) Actor

Extract public Trustpilot company reviews: ratings, text, replies, reviewer info, and company trust score. Filter by stars, language, and date. Deep coverage past the ~200-review wall via smart filter slicing. No login required.

- **URL**: https://apify.com/valev-lab/trustpilot-reviews.md
- **Developed by:** [Daniel Valev](https://apify.com/valev-lab) (community)
- **Categories:** Business, Marketing
- **Stats:** 2 total users, 1 monthly users, 66.7% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.25 / 1,000 review scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Trustpilot Reviews Scraper do?

Trustpilot Reviews Scraper collects **public** reviews from [Trustpilot](https://www.trustpilot.com) company profiles and saves one structured dataset item per review. Give it company domains or Trustpilot review URLs, optional star/language/date filters, and a maximum number of reviews per company.

It does **not** log in to Trustpilot, does not use mobile-app APIs or user cookies, and does not claim unlimited history. Trustpilot’s public pagination stops around page 10 (~200 reviews) per filter combination; with **deep coverage** the Actor slices filters (stars, then languages/dates) and merges unique reviews so you can go further without signing in.

### Why scrape Trustpilot reviews?

Public reviews support competitor research, reputation monitoring, and incremental sync workflows:

- schedule recurring runs and keep only reviews newer than `publishedAfter`
- compare ratings and reply behaviour across companies
- export JSON/CSV through Apify integrations, the API, or MCP

The default run scrapes **20** reviews for `pipedrive.com` so health checks stay cheap while still exercising the live browser + HTTP path.

### What data can Trustpilot Reviews Scraper extract?

| Field | Type | Meaning |
| --- | --- | --- |
| `reviewId` / `reviewUrl` | string | Public review identity and URL |
| `title` / `text` / `rating` | string / int | Review content and stars |
| `publishedDate` / `experienceDate` / `updatedDate` | ISO datetime | Trustpilot timestamps (UTC) |
| `language` / `likes` / `source` | string / int | Language, likes, source label |
| `verificationLevel` / `isVerified` | string / bool | Verification metadata |
| `authorName` / `authorId` / `country` | string | Public reviewer fields (optional anonymization) |
| `replyMessage` / `replyPublishedDate` | string | Company reply when present |
| `companyName` / `companyDomain` / `companyTrustScore` | mixed | Company profile fields |
| `scrapedAt` / `inputUrl` | string | Actor metadata |

Every field is always present; unavailable values are `null` (arrays default to `[]`).

### How to scrape Trustpilot reviews

1. Open the Actor **Input** tab.
2. Add one or more `companyUrls` (domain or full Trustpilot URL).
3. Set `maxReviewsPerCompany` (use `0` for everything reachable on the public path).
4. Optionally filter by `stars`, `languages`, `date`, `verified`, or `withReplies`.
5. Keep Apify Proxy on the **RESIDENTIAL** group — datacenter IPs are commonly blocked by Trustpilot’s WAF.
6. Run the Actor and open the **Dataset** tab (and `COVERAGE_REPORT` in the key-value store).

#### Incremental sync with `publishedAfter`

Set `publishedAfter` to an ISO date (`YYYY-MM-DD`, midnight UTC) or a full ISO date-time. Only reviews **strictly newer** than the cursor are saved and charged. With `sort=recency`, pagination stops as soon as a page crosses the cutoff.

### Pricing

Pricing is [pay-per-event](https://docs.apify.com/platform/actors/running/pay-per-event). Platform usage is included in the event price (not billed separately). Configure the same numbers in Console **Publishing → Monetization**.

| Event | Free / no discount | Description |
| --- | ---: | --- |
| **Run started** (`run-start`) | $0.005 | Once when the Actor run starts successfully |
| **Review scraped** (`review-scraped`) | $0.0005 / review ($0.50 / 1,000) | Per review saved to the dataset |

Volume discounts on higher Apify plans:

| Plan | Per-review price | Per 1,000 reviews |
| --- | ---: | ---: |
| No discount (Free) | $0.00050 | $0.50 |
| Bronze | $0.00043 | $0.43 |
| Silver | $0.00035 | $0.35 |
| Gold | $0.00025 | $0.25 |

Reviews skipped by `publishedAfter`, deduplication, or filters are **not** charged. Blocked runs that produce zero reviews do not charge review events. Set a max total charge on the run to stop cleanly with `stopReason = charge-limit-reached`.

### Input

See the **Input** tab for the full schema. Important fields:

- `companyUrls` — required list of domains or Trustpilot URLs
- `maxReviewsPerCompany` — `0` means unbounded within public/sliced reach; default/prefill is `20`
- `languages` — empty means **all languages** (the Actor sends `languages=all`; Trustpilot’s own UI often defaults to English-only)
- `deepCoverage` — slice filters to work past the ~200-review public wall
- `anonymizeReviewers` — replace `authorName` with initials; null `authorId` / `authorImage`
- `proxyConfiguration` — defaults to Apify Proxy `RESIDENTIAL`

### Output

Results are written to the default dataset. You can download JSON, CSV, Excel, and other formats, or read them through the dataset API.

```json
{
  "reviewId": "6ab516bdf6b3a0e76ca64942",
  "reviewUrl": "https://www.trustpilot.com/reviews/6ab516bdf6b3a0e76ca64942",
  "title": "Training was outstanding",
  "text": "The onboarding session was clear and useful.",
  "rating": 5,
  "publishedDate": "2026-09-24T14:25:33.000Z",
  "experienceDate": "2026-09-24T00:00:00.000Z",
  "updatedDate": null,
  "language": "en",
  "likes": 0,
  "verificationLevel": "not-verified",
  "isVerified": false,
  "source": "Organic",
  "authorName": "Alex",
  "authorId": "public-consumer-id",
  "authorImage": null,
  "authorReviewCount": 1,
  "country": "IE",
  "replyMessage": "Thank you for your feedback!",
  "replyPublishedDate": "2026-09-24T14:36:05.000Z",
  "replyUpdatedDate": null,
  "companyName": "Pipedrive",
  "companyDomain": "pipedrive.com",
  "companyTrustScore": 4.4,
  "companyStars": 4.5,
  "companyTotalReviews": 3453,
  "companyCategories": ["CRM Provider"],
  "companyUrl": "https://www.trustpilot.com/review/pipedrive.com",
  "scrapedAt": "2026-09-25T12:00:00.000Z",
  "inputUrl": "pipedrive.com"
}
```

#### Coverage report

Each run also writes `COVERAGE_REPORT` to the default key-value store (not charged): `companyDomain`, `reviewsTotalOnTrustpilot`, `reviewsReturned`, `coveragePercent`, `filterSlicesFetched`, `pagesFetched`, `stopReason`.

`stopReason` is one of: `complete`, `max-reviews-reached`, `published-after-reached`, `public-wall`, `public-wall-after-slicing`, `charge-limit-reached`, `not-found`, `blocked`, `fetch-errors`.

### Limits

- Public pagination stops at about **page 10 (~200 reviews) per filter combination**; page 11 redirects to login. This Actor never logs in.
- **Deep coverage** uses additional star/language/date slices and deduplicates by `reviewId`. Ordering is no longer a single global newest-first list once slicing is active.
- Trustpilot fronts pages with an **AWS WAF** challenge. The Actor bootstraps a real browser once per proxy session, then prefers cheap `/_next/data/...` HTTP pages. If `__NEXT_DATA__` disappears (App Router migration), the company fails loudly and raw HTML is saved as `DEBUG_{domain}_page{n}.html`.
- A run that returns **0 reviews because of blocking** ends as **failed**, not as a silent empty success.

### Integrations and API / MCP

Use Apify schedules, webhooks, and integrations to push new reviews downstream. For programmatic access, use the Apify API or an MCP client configured against Apify. Call this Actor with the same JSON input you would use in Console.

### FAQ

#### Why do I see fewer languages than on Trustpilot in a browser?

Trustpilot’s own profile UI often defaults to English. Leave `languages` empty so this Actor requests **all** languages via `languages=all`.

#### Can I get more than 200 reviews for a large company?

Yes, enable `deepCoverage` and set `maxReviewsPerCompany` above 200 (or `0`). Coverage still depends on how well star/language/date slices separate the catalogue; `COVERAGE_REPORT` records when the public wall remains after slicing.

#### Does anonymizeReviewers remove all personal data?

No. Review text can still contain personal data. Anonymization only initials the display name and drops author id/image. You remain responsible for GDPR and other applicable law.

### Responsible use

This Actor only reads **public** Trustpilot pages. Reviewer names and related fields can be personal data under the GDPR. You are responsible for complying with Trustpilot’s terms and applicable law, and for having a lawful basis before processing personal data.

# Actor input Schema

## `companyUrls` (type: `array`):

Company domains (pipedrive.com) or full Trustpilot review URLs (https://www.trustpilot.com/review/pipedrive.com or localized hosts such as de.trustpilot.com). One entry per line.

## `maxReviewsPerCompany` (type: `integer`):

Maximum number of reviews to save per company. Use 0 to collect everything reachable on the public path (subject to Trustpilot's per-filter wall and deepCoverage slicing).

## `stars` (type: `array`):

Only include reviews with these star ratings (1–5). Leave empty for all ratings.

## `languages` (type: `array`):

ISO 639-1 language codes to include (for example en, de, fr). Leave empty to request all languages. Trustpilot's own page defaults to English-only on many profiles; this Actor always sends languages=all when the list is empty.

## `sort` (type: `string`):

Trustpilot sort order for each filter window.

## `date` (type: `string`):

Optional Trustpilot server-side date preset.

## `publishedAfter` (type: `string`):

Keep only reviews strictly newer than this cursor. Accepts YYYY-MM-DD (midnight UTC) or an ISO date-time with timezone. With sort=recency, pagination stops once a page crosses the cutoff. Filtered reviews are not charged.

## `verified` (type: `boolean`):

When enabled, only verified or invited reviews are requested.

## `withReplies` (type: `boolean`):

When enabled, only reviews that have a company reply are requested.

## `includeCompanyInfo` (type: `boolean`):

Attach company metadata fields to every review. When disabled, only companyDomain and companyName remain populated.

## `deepCoverage` (type: `boolean`):

When more than 200 reviews are requested, slice by star rating (and then language/date) to get past Trustpilot's public ~200-review-per-filter wall. Never logs in. Slicing changes global ordering; for sort=recency with maxReviewsPerCompany ≤ 200 a single window is kept.

## `anonymizeReviewers` (type: `boolean`):

Replace authorName with initials and null out authorId and authorImage. Useful for GDPR-conscious workflows.

## `proxyConfiguration` (type: `object`):

Proxy settings used for outbound requests. Trustpilot's WAF blocks typical datacenter IPs — keep Apify Proxy with the RESIDENTIAL group.

## `maxConcurrency` (type: `integer`):

Reserved for future multi-company parallelism. The hybrid transport currently processes companies sequentially with low browser concurrency.

## Actor input object example

```json
{
  "companyUrls": [
    "pipedrive.com"
  ],
  "maxReviewsPerCompany": 20,
  "stars": [],
  "languages": [],
  "sort": "recency",
  "date": "",
  "verified": false,
  "withReplies": false,
  "includeCompanyInfo": true,
  "deepCoverage": true,
  "anonymizeReviewers": false,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  },
  "maxConcurrency": 3
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset containing one item per scraped review.

## `coverageReport` (type: `string`):

Per-company coverage JSON under COVERAGE\_REPORT.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "companyUrls": [
        "pipedrive.com"
    ],
    "maxReviewsPerCompany": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("valev-lab/trustpilot-reviews").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "companyUrls": ["pipedrive.com"],
    "maxReviewsPerCompany": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("valev-lab/trustpilot-reviews").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "companyUrls": [
    "pipedrive.com"
  ],
  "maxReviewsPerCompany": 20
}' |
apify call valev-lab/trustpilot-reviews --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,valev-lab/trustpilot-reviews"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/0GExAbUeMhHWKfXYV/builds/PIybpoALsZIEHaP3i/openapi.json
