# Bulk PageSpeed and Core Web Vitals (Field + Lab Data) (`madrasco/pagespeed-core-web-vitals-bulk`) Actor

PageSpeed Insights scores and real-user Core Web Vitals (Chrome UX Report p75 LCP, INP, CLS) for a list of URLs, one row per URL and device, from Google's official APIs with your own free Google API key. URL-level field data, origin fallback labelled.

- **URL**: https://apify.com/madrasco/pagespeed-core-web-vitals-bulk.md
- **Developed by:** [Jack Valmadre](https://apify.com/madrasco) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 page testeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk PageSpeed and Core Web Vitals (Field + Lab Data)

Paste a list of URLs and get one table row per page and device, with two kinds of data side by side:

- **Real-user data (field data)** from Google's Chrome UX Report: the 75th-percentile LCP, INP and CLS that real Chrome users saw over the last 28 days, each rated good / needs improvement / poor, and whether the page passes the Core Web Vitals assessment.
- **Lighthouse test data (lab data)** from the PageSpeed Insights API: the performance score, lab LCP, FCP, CLS, Total Blocking Time and Speed Index, and the failing performance audits with the largest estimated savings.

It is meant for people who check many pages at once: agencies preparing client reports, SEO and web teams watching a site after a release, and anyone who wants PageSpeed Insights results for a URL list in a spreadsheet instead of one page at a time.

The actor calls only Google's documented public APIs (PageSpeed Insights API v5 and Chrome UX Report API), with your own Google API key.

### Before you start: your own Google API key

You need a Google Cloud API key. The actor has no key of its own and never shares one between users; every call runs under your key and your Google quota. Setting one up takes a few minutes:

1. Sign in at console.cloud.google.com with a Google account.
2. Pick a project in the project list at the top of the page, or create one ("New project"; any name).
3. Go to "APIs & Services", then "Library". Search for "PageSpeed Insights API" and click "Enable". Do the same for "Chrome UX Report API".
4. Go to "APIs & Services", then "Credentials". Click "Create credentials", then "API key". Copy the key.
5. Recommended: open the new key, choose "Restrict key" under "API restrictions", tick only the two APIs above, and save. The key then can't be used for anything else. Leave "Application restrictions" at None: an IP address restriction makes Google refuse the key on Apify, whose runs come from changing addresses (every row is then `key_error`).
6. Paste the key into this actor's "Google API key" field.

If you only want field data ("Data to fetch" set to field data only), only the Chrome UX Report API needs to be enabled; for lab data only, only the PageSpeed Insights API.

**Without a key** the actor calls nothing. The run still finishes, with one `key_missing` row per URL and device and a message saying how to get a key.

**How your key is handled.** The key field is a secret input: Apify encrypts it when the input is saved, and it can only be decrypted inside the run. The actor sends it to Google in a request header, not in a URL, and removes it from any error text. It is not written to the log, the results or the run summary. You can delete or replace the key on the same Credentials page at any time.

### Input

| Field | What it does | Default |
|---|---|---|
| URLs to test | One URL per line; a pasted list separated by spaces or commas (`a.com, b.com`) is split into separate URLs, while commas inside a URL's path or query stay part of it. A bare domain such as `example.com` becomes `https://example.com/`; `#fragments` are dropped; duplicates are tested once. Anything that is not an http(s) address with a host name (for example `mailto:` links, or addresses with a user name or password in them) gets an `invalid_url` row. | two example pages |
| Google API key | Your key (see above). | none |
| Device | Mobile, desktop, or both (two rows per URL). Sets Lighthouse's device emulation and the matching Chrome UX Report form factor (phone or desktop). | Mobile |
| Data to fetch | Field and lab, field only, or lab only. | Field and lab |
| Lighthouse categories | Performance, and optionally Accessibility, Best practices and SEO scores (0-100). When lab data is on, Performance is always included whatever you pick here: the lab metrics and the top-issues list come from it. | Performance |
| Fall back to origin-level field data | If the exact page has no Chrome UX Report record, use the whole site's record and say so in `fieldDataLevel`. | On |
| Parallel requests | URLs tested at once (1-10). Lower it if your key's per-minute quota is small. | 4 |
| Top performance issues per URL | How many failing audits to list (0-20). | 5 |
| Maximum URLs | Safety cap per run (up to 5,000). | 500 |

Calling it from code or an AI agent: the URL list is also accepted as `startUrls` (strings or `{"url": ...}` objects), and a single `url` field may hold several URLs separated by spaces or commas (split the same way as the list).

### Output: one row per URL and device

The dataset's "Core Web Vitals and scores" view shows the main columns. Every row has these fields:

| Field | Meaning |
|---|---|
| `url`, `device` | The page tested and `mobile` or `desktop` |
| `status` | Overall result for the row (see below) |
| `coreWebVitals` | `passed`, `failed` or `insufficient_data`. Passed when the 75th percentiles of LCP, INP and CLS are all good; when there is not enough INP data, passed when LCP and CLS are good. This is the rule PageSpeed Insights describes for its own assessment. |
| `fieldDataLevel` | `url` (data for this exact page), `origin` (whole-site data, used because the page has none), `final_url` or `final_origin` (the address you gave has no data, but it redirects within the same site, for example `example.com` to `www.example.com`, and the data is for the page or site it redirects to), or `none` |
| `fieldLcpP75`, `fieldInpP75`, `fieldClsP75`, `fieldFcpP75`, `fieldTtfbP75` | 75th-percentile values from real Chrome users (milliseconds; CLS unitless) |
| `fieldLcpRating` ... `fieldTtfbRating` | `good`, `needs_improvement` or `poor` against Google's published thresholds |
| `fieldFirstDate`, `fieldLastDate`, `fieldKey` | The 28-day collection period, and the URL or origin the field data belongs to |
| `performanceScore` (plus `accessibilityScore`, `bestPracticesScore`, `seoScore` if chosen) | Lighthouse scores, 0-100 |
| `labLcpMs`, `labFcpMs`, `labCls`, `labTbtMs`, `labSpeedIndexMs` | Lighthouse lab metrics |
| `topIssues`, `topIssuesText` | Failing performance audits, largest estimated savings first |
| `finalUrl`, `lighthouseVersion`, `labFetchTime` | Where Lighthouse ended up after redirects, and which Lighthouse version tested it when |
| `labStatus`, `fieldStatus`, `errors` | What happened with each API; `errors` holds Google's error message when something failed, Lighthouse's warnings (for example that the page redirected or loaded too slowly), and a note when field data came from a redirect |
| `source` | Where the data came from, its licence and what this actor changed (see below) |

Row statuses:

- `ok`: everything asked for came back.
- `no_field_data`: lab data is fine, but the Chrome UX Report has no record for this page, its site, or the same-site address it redirects to. Common for small sites, and for addresses that redirect to a different site (field data from another site is never used).
- `page_error`: Lighthouse could not test the page (for example, it did not load).
- `partial`: lab data came back but field data did not (the reason is in `errors`).
- `error`: a server or network error that one retry did not fix (in a field-data-only run, a failed field-data call is reported as `error`, since nothing else was asked for).
- `quota`: your Google quota ran out.
- `key_error`: Google refused the key (not valid, the API is not enabled for its project, or the key has an IP address restriction). When the key is restricted to other IP addresses, the row error and the run message say so and tell you to set Application restrictions to None on the key: Apify runs from changing addresses, so that restriction cannot be relied on to admit a run.
- `key_missing`: no key was given, so nothing was tested.
- `invalid_url`: not an http(s) address with a host name, or one with a user name or password in it.
- `skipped`: not tested, because the run's maximum charge was reached (see Price).

When Google refuses the key or a daily quota runs out, the actor stops calling that API for the rest of the run; requests already under way still finish, so you may see a few failed calls rather than one. The run summary (key-value store record `OUTPUT`) holds the counts by status, the number of calls made to each API, which API was stopped and why, and the invalid URLs. Its `rows` count covers URL and device rows only, not `invalid_url` rows.

### Price

US$0.003 per page tested (event `page-tested`): one URL on one device that returned results from Google, that is a row with status `ok`, `partial`, or `no_field_data` when lab data came back. Testing a URL on both devices is two rows. Rows with `page_error`, `error`, `quota`, `key_error`, `key_missing`, `invalid_url` or `skipped` are free, and so is `no_field_data` in a field-data-only run (nothing came back). There is also a start charge of US$0.00005 per run (Apify's standard actor-start event). The charge is for this actor's work (running the list, combining lab and field data, ratings, one table); the data comes from Google under your own key, and Google's APIs are free within your key's quota.

If you set a maximum charge per run, the actor starts a URL only while its charge still fits, so it does not use your Google quota on results it could not return; the remaining URLs come back as free `skipped` rows, and the run summary shows `stoppedAtMaxCharge` and `chargedRows`.

### Quotas and cost on Google's side

- Chrome UX Report API: Google's published limit is 150 queries per minute per Google Cloud project, free of charge. Each run spaces its own field-data queries to at most 120 a minute. Several runs at the same time with the same key share that limit.
- PageSpeed Insights API: check your key's quota for this API in the Google Cloud Console before large runs.
- Each URL and device takes one PageSpeed Insights call for lab data, and one to four Chrome UX Report calls for field data: one for the page, a second when the page has no record and the actor falls back to the site, and up to two more for the address it redirects to when neither has data.

Run time: in our test run on Apify (30 URLs on both devices, 4 parallel requests) each URL and device took about 11 seconds of run time, mostly waiting for Google's Lighthouse. The default 4-hour run timeout covers roughly 1,300 rows at that rate; for longer lists raise the run timeout or split the list. If a run times out, rows already written are kept (and charged); the rest are not tested and not charged.

### Data sources, licence and changes

- Lab data: Google PageSpeed Insights API v5, which runs Lighthouse on Google's servers.
- Field data: Chrome UX Report (CrUX) API. CrUX data by Google, licensed under the Creative Commons Attribution 4.0 International License (https://creativecommons.org/licenses/by/4.0/); licence statement: https://developer.chrome.com/docs/crux/methodology.
- Changes made by this actor to the CrUX data: it keeps only the 75th-percentile values (INP, LCP, FCP and TTFB rounded to whole milliseconds), drops the histograms, and adds good / needs improvement / poor ratings and a Core Web Vitals pass/fail computed from Google's published thresholds (https://web.dev/articles/vitals).
- Every result row carries this attribution in its `source` field, and so does the run summary.

This actor is not affiliated with or endorsed by Google.

### Limitations

- Lab scores change from run to run, because Lighthouse loads the page again each time from Google's servers under changing network and server conditions. Field data is a 28-day summary and changes slowly.
- Field data exists only for pages and sites with enough Chrome traffic. Small sites often get `no_field_data`, and with origin fallback many pages get their whole site's numbers rather than their own: check `fieldDataLevel` before comparing pages.
- Google can only test pages it can reach from the public internet, so pages behind a login, on an intranet or on your own machine can't be tested; such pages, and addresses that don't load at all, get `page_error`.
- In a field-data-only run there is no Lighthouse test, so redirects aren't seen: give the address visitors land on (for example with `www.`), or field data may be missing.
- The actor keeps nothing outside your own run's dataset and key-value store.

Built and maintained with AI assistance by Madrasco. Questions and bug reports: the Issues tab.

# Actor input Schema

## `urls` (type: `array`):

Pages to test, one per line (https://example.com/page); a pasted list separated by spaces or commas is split into separate URLs. A bare domain becomes https://<domain>/. Also accepted as 'startUrls' or 'url'.

## `apiKey` (type: `string`):

Your own Google Cloud API key with the 'PageSpeed Insights API' and the 'Chrome UX Report API' enabled (both have a free quota; see the README for the 2-minute setup). Stored encrypted by Apify, sent to Google only in a request header, never logged or written to the results.

## `strategy` (type: `string`):

Lighthouse device emulation and the matching Chrome UX Report form factor (phone or desktop).

## `dataMode` (type: `string`):

Field data comes from the Chrome UX Report API (28-day real-user p75); lab data from a PageSpeed Insights Lighthouse run.

## `categories` (type: `array`):

Category scores to include (0-100). When lab data is on, Performance is always included: the lab metrics (LCP, FCP, CLS, Total Blocking Time, Speed Index) and the top-issues list come from it. The others add their scores.

## `originFallback` (type: `boolean`):

When the Chrome UX Report has no record for the exact URL, use the whole site's (origin's) record. If neither has data and the page redirects within the same site (e.g. example.com to www.example.com), the address it redirects to is tried too. The 'fieldDataLevel' column says which one you got: url, origin, final\_url, final\_origin or none.

## `maxConcurrency` (type: `integer`):

How many URLs are tested at once. Lower it if your key's per-minute quota is small.

## `topIssues` (type: `integer`):

Failing Lighthouse performance audits to list, largest estimated savings first.

## `maxUrls` (type: `integer`):

Safety cap on the number of URLs tested in one run.

## Actor input object example

```json
{
  "urls": [
    "https://web.dev/",
    "https://developer.chrome.com/docs/crux"
  ],
  "strategy": "mobile",
  "dataMode": "both",
  "categories": [
    "performance"
  ],
  "originFallback": true,
  "maxConcurrency": 4,
  "topIssues": 5,
  "maxUrls": 500
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://web.dev/",
        "https://developer.chrome.com/docs/crux"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("madrasco/pagespeed-core-web-vitals-bulk").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://web.dev/",
        "https://developer.chrome.com/docs/crux",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("madrasco/pagespeed-core-web-vitals-bulk").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://web.dev/",
    "https://developer.chrome.com/docs/crux"
  ]
}' |
apify call madrasco/pagespeed-core-web-vitals-bulk --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,madrasco/pagespeed-core-web-vitals-bulk"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/XR35V9PygJ6mqRD6U/builds/XE6vGD59bjtOTt3vc/openapi.json
