# Bulk PageSpeed Audit - Lighthouse & Core Web Vitals (`cuantic_data/bulk-pagespeed-audit`) Actor

Run Google Lighthouse on up to 500 URLs or a whole XML sitemap in one run. One row per page with performance, accessibility, best-practices and SEO scores, Core Web Vitals lab metrics and the top failing audits.

- **URL**: https://apify.com/cuantic\_data/bulk-pagespeed-audit.md
- **Developed by:** [Cuantic Data](https://apify.com/cuantic_data) (community)
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 page auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk PageSpeed Audit - Lighthouse & Core Web Vitals for Many URLs

Run Google Lighthouse on a whole list of pages in one run. Paste up to 500 URLs, or just give a sitemap URL, and get one dataset row per page with the four Lighthouse category scores, the Core Web Vitals lab metrics and the top failing audits. A summary record ranks every page from worst to best, so the pages that need work first are obvious.

Need a single page only? Use the sibling actor **page-speed-lighthouse-audit**.

### What it does

- Audits every URL with real Chrome and the official `lighthouse` npm package (the same engine behind PageSpeed Insights lab data).
- Accepts a list of URLs **or** an XML sitemap (a sitemap index is followed one level deep).
- Removes duplicate URLs (keeping the original order) and caps the run at 500 pages.
- Runs several audits in parallel (1 to 4), each in its own Chrome process.
- Isolates failures: a page that times out, returns 404 or cannot be resolved gets a row with an `error` message, and the run continues.
- Writes a summary to the key-value store record `OUTPUT` with counts and a worst-first ranking.

### Who it is for

- **SEO and migration audits**: capture a baseline of every page before a redesign or platform migration, then compare after.
- **Agencies**: audit a client's whole site in one run and hand over a sortable table instead of one PageSpeed report per page.
- **Whole-site Core Web Vitals checks**: find the templates and sections with the worst LCP, CLS and TBT instead of guessing from the homepage.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `urls` | array of strings | - | Pages to audit (http/https). Deduplicated, max 500. |
| `sitemapUrl` | string | - | XML sitemap or sitemap index. All `<loc>` page URLs are audited, max 500. |
| `device` | `mobile` | `desktop` | `mobile` | Lighthouse form factor. |
| `categories` | array | all four | Any of `performance`, `accessibility`, `best-practices`, `seo`. |
| `timeoutMs` | integer | `90000` | Per-URL timeout, 15000 to 300000 ms. |
| `maxConcurrency` | integer | `2` | Parallel audits, 1 to 4. |

At least one of `urls` or `sitemapUrl` is required. If both are given, `urls` is used, the sitemap is ignored and a warning is logged.

Example input:

```json
{
    "sitemapUrl": "https://www.example-shop.com/sitemap.xml",
    "device": "mobile",
    "categories": ["performance", "accessibility", "best-practices", "seo"],
    "timeoutMs": 90000,
    "maxConcurrency": 2
}
```

### Output

One dataset item per URL:

```json
{
    "url": "https://www.example-shop.com/collections/summer",
    "device": "mobile",
    "categories": ["performance", "accessibility", "best-practices", "seo"],
    "scores": {
        "performance": 41,
        "accessibility": 88,
        "bestPractices": 78,
        "seo": 92
    },
    "metrics": {
        "firstContentfulPaint": "2.9 s",
        "largestContentfulPaint": "7.4 s",
        "totalBlockingTime": "1,120 ms",
        "cumulativeLayoutShift": "0.214",
        "speedIndex": "6.8 s",
        "timeToFirstByte": "Root document took 640 ms"
    },
    "topOpportunities": [
        { "id": "render-blocking-resources", "title": "Eliminate render-blocking resources", "score": 0, "displayValue": "Est savings of 1,430 ms" },
        { "id": "layout-shifts", "title": "Avoid large layout shifts", "score": 0, "displayValue": "4 layout shifts found" },
        { "id": "unused-javascript", "title": "Reduce unused JavaScript", "score": 0.12, "displayValue": "Est savings of 412 KiB" },
        { "id": "image-alt", "title": "Image elements do not have [alt] attributes", "score": 0 },
        { "id": "modern-image-formats", "title": "Serve images in modern formats", "score": 0.5, "displayValue": "Est savings of 238 KiB" }
    ],
    "error": null
}
```

- `scores` are 0 to 100 integers. A category you did not select is `null`.
- `metrics` are Lighthouse display values (lab data). They are `null` when `performance` is not selected.
- `topOpportunities` lists up to 5 audits with a score below 0.9, worst score first (ties broken by estimated savings). The metric audits themselves (LCP, CLS, ...) are left out because they are already in `metrics`.
- `error` is `null` on success. On failure it holds the reason (for example `Lighthouse audit timed out after 90000 ms` or `ERRORED_DOCUMENT_REQUEST: ... (Status code: 404)`), and all scores and metrics are `null`.

Dataset rows are written as each audit finishes (so a run that is aborted keeps what it already did), which means their order follows completion, not input. To see the worst pages first, sort the table by `scores.performance`, or read the `OUTPUT` record:

```json
{
    "total": 137,
    "ok": 134,
    "failed": 3,
    "skipped": 0,
    "device": "mobile",
    "categories": ["performance", "accessibility", "best-practices", "seo"],
    "rankedBy": "performance",
    "worstFirst": [
        { "url": "https://www.example-shop.com/collections/summer", "score": 41 },
        { "url": "https://www.example-shop.com/blog/lookbook", "score": 47 }
    ]
}
```

The ranking uses the performance score, or the first selected category when performance is not selected.

### Combine with sitemap-url-extractor

Extract every URL from a sitemap with **sitemap-url-extractor**, filter the list (for example only `/products/` pages), then audit them all here. Or skip that step and pass the sitemap URL directly in `sitemapUrl`.

### Limitations

- **Lab data only.** Scores and metrics come from a Lighthouse run in a data center, not from real users. There is no Chrome UX Report (CrUX) field data.
- **Scores vary from run to run.** Lighthouse performance scores commonly move by several points between runs of the same page. Running more audits in parallel adds CPU contention and more variance; use `maxConcurrency: 1` when you need the most stable numbers.
- **Desktop mode is simplified.** `desktop` changes the form factor and screen size; network and CPU throttling stay at Lighthouse's default (mobile) profile, so desktop numbers are more conservative than PageSpeed Insights' desktop tab.
- **No logged-in pages.** Each audit starts a fresh browser with no cookies or credentials.
- **Heavy pages can time out.** Such pages get an error row; raise `timeoutMs` if needed.
- **Runtime grows linearly with the number of URLs.** A typical audit takes 10 to 40 seconds, so 500 URLs at concurrency 2 can take one to three hours. Set the run timeout in Apify accordingly, and give the actor about 1 GB of memory per parallel audit.
- **Max 500 URLs per run.** Extra URLs are dropped with a warning. Split larger sites across several runs.
- **Plain XML sitemaps only.** Gzipped (`.xml.gz`) sitemaps are not supported, and a sitemap index is followed one level deep (up to 50 child sitemaps).

### Pricing

Pay-per-event: $0.005 USD per successfully audited page ($5.00 per 1,000 pages). Pages that end with an error are not charged. Apify platform usage is included in this price; there is no separate platform-usage charge.

An Actor Start event costs $0.00005 USD. One start event is charged per GB of Actor memory, with a minimum of one event per run.

If the maximum charge you set for a run is reached, the actor stops starting new audits and finishes the ones in progress.

### Support

Cuantic Data - cuanticwindows@gmail.com

Terms of use: see [Terms of use](#terms-of-use) below.).

### Terms of use

Provided by Cuantic Data (cuanticwindows@gmail.com).

#### 1. What the actor does

The actor runs Google Lighthouse, through the open-source `lighthouse` npm package and a headless Chrome browser, against the URLs you provide, either directly or through a sitemap you point it to. It loads each page as a regular browser would and stores the resulting scores, metrics and failing audits in your Apify dataset.

#### 2. Your responsibility

- You are responsible for having the right to audit the sites and pages you submit. Only submit URLs of sites you own, manage, or are otherwise allowed to test.
- Each audit loads the full page with all its resources. Running many audits against one site generates real traffic on that site; choose the number of URLs and the concurrency accordingly.
- You must comply with the terms of service of the audited sites and with the Apify Terms of Service.

#### 3. About the results

- Scores, metrics and audit results come from Lighthouse and carry its known run-to-run variability. They are lab measurements from the Apify infrastructure, not measurements of real users.
- Results are provided "as is", for informational purposes, without warranty of accuracy, completeness or fitness for a particular purpose. They are not a certification of performance, accessibility or search ranking.
- Lighthouse, Chrome and PageSpeed Insights are products of Google. This actor is not affiliated with or endorsed by Google.

#### 4. Data

The actor stores only what it produces (the dataset rows and the summary record) in your own Apify storage. It does not keep copies of the audited pages or results elsewhere.

#### 5. Liability

To the extent permitted by law, Cuantic Data is not liable for any damage arising from the use of the actor or its results, including decisions made based on the scores, or any effect of the audit traffic on the audited sites.

#### 6. Changes

These terms may be updated together with the actor. The version published with the actor is the one that applies.

# Actor input Schema

## `urls` (type: `array`):

Pages to audit (http or https). Duplicates are removed, order is kept, max 500 per run. If you also fill in a sitemap URL, this list wins and the sitemap is ignored.

## `sitemapUrl` (type: `string`):

An XML sitemap (or sitemap index). Every <loc> page URL in it is audited, up to 500. Used only when the URLs list is empty.

## `device` (type: `string`):

Lighthouse form factor. Desktop switches the form factor and screen size; network/CPU throttling stays at Lighthouse's default profile.

## `categories` (type: `array`):

Lighthouse categories to run. Fewer categories make each audit a bit faster. Core Web Vitals metrics are only measured when Performance is included.

## `timeoutMs` (type: `integer`):

Maximum time for one Lighthouse audit. A page that takes longer gets an error row and the run continues.

## `maxConcurrency` (type: `integer`):

Lighthouse audits running in parallel, each with its own Chrome. Chrome is memory-hungry: plan about 1 GB of actor memory per parallel audit. Higher concurrency also makes scores noisier because audits compete for CPU.

## Actor input object example

```json
{
  "urls": [
    "https://example.com/",
    "https://www.iana.org/help/example-domains"
  ],
  "device": "mobile",
  "categories": [
    "performance",
    "accessibility",
    "best-practices",
    "seo"
  ],
  "timeoutMs": 90000,
  "maxConcurrency": 2
}
```

# Actor output Schema

## `results` (type: `string`):

One row per URL with Lighthouse category scores, Core Web Vitals lab metrics and top failing audits.

## `summary` (type: `string`):

Totals (ok/failed) and the worst-first ranking by performance score.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com/",
        "https://www.iana.org/help/example-domains"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("cuantic_data/bulk-pagespeed-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://example.com/",
        "https://www.iana.org/help/example-domains",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("cuantic_data/bulk-pagespeed-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com/",
    "https://www.iana.org/help/example-domains"
  ]
}' |
apify call cuantic_data/bulk-pagespeed-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cuantic_data/bulk-pagespeed-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/7FGkOFEa315rtvpnm/builds/MeE6MQJUk6PN3TeHf/openapi.json
