# Bulk On-Page SEO Checker (`kernfetch/onpage-seo-checker`) Actor

Audit title, meta description, canonical, noindex, hreflang, H1, Open Graph, alt text and schema.org for thousands of pages, with ready-to-filter SEO issues and duplicate titles. Works with Sitemap URL Extractor. $3 per 1,000 pages.

- **URL**: https://apify.com/kernfetch/onpage-seo-checker.md
- **Developed by:** [kernfetch](https://apify.com/kernfetch) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk On-Page SEO Checker – meta tags, headings, canonical and schema in bulk

Audit **thousands of pages** in one run and get, for each one, the **title**, **meta description**, **robots / noindex**, **canonical**, **hreflang**, **H1/H2**, **Open Graph and Twitter Card**, **images without alt**, **word count** and **schema.org types**, plus a ready-made list of **on-page SEO issues**. The run report also lists **duplicate titles and descriptions** across all pages.

Built for **SEO audits, content QA before and after a migration, agency reporting and monitoring** – at a fraction of the price of full SEO audit tools.

### Why this Actor

- 🩺 **Issues ready to filter** – `missing_title`, `title_too_long`, `missing_description`, `missing_h1`, `multiple_h1`, `noindex`, `canonical_to_other_url`, `missing_lang`, `missing_viewport`, `missing_og_image`, `images_missing_alt`, `low_word_count`, `hreflang_missing_self`, `no_structured_data`, `redirected`.
- 🔁 **Duplicate titles and descriptions** – detected across the whole run and listed in the `SUMMARY`.
- 🧩 **Works with Sitemap URL Extractor** – pass the Dataset ID of a [Sitemap URL Extractor](https://apify.com/kernfetch/sitemap-url-extractor) run and audit every page of a site, no crawling needed.
- 🎚️ **Your thresholds** – set min/max title and description length and minimum word count.
- ⚡ **Fast and cheap** – plain HTTP, no browser. **$3 per 1,000 pages**.
- 💸 **Fair billing** – pages that are blocked, disallowed by robots.txt or unreachable are **not charged**.
- ✅ **Reliable and gentle** – robots.txt respected on every URL and redirect, low load per site, protections never bypassed.

### How to use

1. Paste your page URLs in **Pages to audit** (one per line; *Bulk edit* accepts thousands), or enter a **Dataset ID** from Sitemap URL Extractor.
2. Optionally adjust the **SEO thresholds**.
3. Click **Start**. Use the **Overview**, **Meta & canonical** and **Social & structured data** views, or download JSON, CSV, Excel or via API.

#### Input example

```json
{
  "urls": ["https://www.example.com/", "https://www.example.com/blog/"],
  "titleMaxChars": 60,
  "descriptionMaxChars": 160,
  "minWords": 200,
  "maxUrls": 1000
}
```

#### Output example

```json
{
  "url": "https://www.example.com/blog/",
  "finalUrl": "https://www.example.com/blog/",
  "statusCode": 200,
  "redirectCount": 0,
  "outcome": "ok",
  "title": "Blog – Example",
  "titleLength": 14,
  "metaDescription": "News and guides from Example.",
  "metaDescriptionLength": 29,
  "metaRobots": null,
  "xRobotsTag": null,
  "indexable": true,
  "canonical": "https://www.example.com/blog/",
  "canonicalIsSelf": true,
  "lang": "en",
  "hreflang": [{ "lang": "it", "url": "https://www.example.com/it/blog/" }],
  "h1": ["Blog"],
  "h1Count": 1,
  "h2Count": 12,
  "h2Sample": ["Latest posts", "Guides"],
  "openGraph": { "title": "Blog", "description": null, "image": "https://www.example.com/og.png", "type": "website", "url": null },
  "twitterCard": { "card": "summary_large_image", "title": null, "description": null, "image": null },
  "imagesTotal": 24,
  "imagesMissingAlt": 3,
  "linksInternal": 85,
  "linksExternal": 6,
  "wordCount": 740,
  "structuredDataTypes": ["Organization", "WebSite", "BreadcrumbList"],
  "hasViewport": true,
  "issues": ["title_too_short", "description_too_short", "images_missing_alt", "hreflang_missing_self"],
  "contentType": "text/html; charset=UTF-8",
  "pageSizeKb": 96.4,
  "truncated": false,
  "responseTimeMs": 310,
  "checkedAt": "2026-09-29T10:00:00+00:00"
}
```

**Outcomes**: `ok` (page audited), `not_html` (e.g. PDF), `client_error`, `server_error`, `redirect_loop`, `too_many_redirects`, `redirect_without_location`, `invalid_redirect`. Blocked, robots-disallowed and unreachable URLs are not in the dataset: they are listed, free of charge, in the `SUMMARY` and `SKIPPED_URLS` records.

### Audit a whole site in 2 steps

1. Run [Sitemap URL Extractor](https://apify.com/kernfetch/sitemap-url-extractor) on the domain and copy the run's **Dataset ID** (Storage tab).
2. Paste it in **Dataset ID** here and start. Optionally check status codes and redirect chains first with [Bulk URL Status & Redirect Checker](https://apify.com/kernfetch/url-status-redirect-checker).

### Use with AI agents

Give your agent structured, stable on-page data for any list of URLs: titles, descriptions, headings, canonical and issues. Useful to generate SEO recommendations, rewrite titles and descriptions, or monitor changes. Works with the Apify API, Apify MCP server, Make, Zapier, n8n and LangChain.

### Pricing

Pay only for results: **$3.00 per 1,000 audited pages**. No subscription. A 500-page site costs about $1.50. Blocked, disallowed and unreachable URLs are **not charged**. Set **Maximum cost per run** in Run options to cap your spend: the Actor never audits more pages than your limit allows.

### FAQ

**Does it render JavaScript?**
No. It reads the HTML the server sends, like most search engine crawlers do on the first pass. Sites that build titles and meta tags only in the browser (pure client-side apps) may show missing values: that is itself an SEO finding.

**Why are some sites "blocked"?**
The site answered 403, 429 or an anti-bot challenge page. The Actor respects that and never tries to bypass protections: the remaining URLs of that host are skipped, listed in the `SUMMARY` and not charged.

**Does it crawl the site?**
No, it audits the URLs you give it. To get all the URLs of a site, use [Sitemap URL Extractor](https://apify.com/kernfetch/sitemap-url-extractor) first.

**How fast is it?**
Pages on different hosts are processed in parallel. On a single host the Actor waits at least 500 ms between requests (about 2 pages per second), so a 1,000-page site takes about 9 minutes.

**Does it collect personal data?**
No. It stores technical SEO fields only: no page text, no `author` meta tag, and only the `@type` of structured data blocks, never their content.

**Is it legal?**
The Actor only requests public pages you provide, follows robots.txt and stops when a site refuses access. Site owners can block it with `User-agent: kernfetch` in robots.txt. You are responsible for the URLs you submit.

### Related Actors by kernfetch

- [Sitemap URL Extractor](https://apify.com/kernfetch/sitemap-url-extractor) – every URL of a website from its sitemaps, with lastmod dates.
- [Bulk URL Status & Redirect Checker](https://apify.com/kernfetch/url-status-redirect-checker) – status codes, redirect chains and final URLs for thousands of URLs.
- [RSS Feed Finder & Reader](https://apify.com/kernfetch/rss-feed-finder) – discover the RSS, Atom and JSON feeds of any site and get the latest articles.

### Support

Found a page that doesn't behave as expected? Open an issue on the **Issues** tab with the URL: fixes are usually shipped within days.

# Actor input Schema

## `urls` (type: `array`):

One URL per line (use Bulk edit to paste thousands). URLs without a scheme get https://. Duplicates are removed.

## `datasetId` (type: `string`):

Audit the URLs stored in an Apify dataset, e.g. the output of Sitemap URL Extractor (kernfetch/sitemap-url-extractor).

## `datasetUrlField` (type: `string`):

Name of the field that contains the URL in the dataset items.

## `maxUrls` (type: `integer`):

Safety cap on the number of pages audited (and charged) in one run.

## `titleMinChars` (type: `integer`):

Titles shorter than this are flagged as title\_too\_short.

## `titleMaxChars` (type: `integer`):

Titles longer than this are flagged as title\_too\_long.

## `descriptionMinChars` (type: `integer`):

Descriptions shorter than this are flagged as description\_too\_short.

## `descriptionMaxChars` (type: `integer`):

Descriptions longer than this are flagged as description\_too\_long.

## `minWords` (type: `integer`):

Pages with fewer visible words are flagged as low\_word\_count.

## `maxConcurrency` (type: `integer`):

Maximum number of requests running at the same time across all hosts. Each host is still limited by the per-host settings below.

## `perHostConcurrency` (type: `integer`):

Kept low on purpose to be polite to each site.

## `perHostDelayMs` (type: `integer`):

Minimum pause between two requests to the same host. Default 500 ms is about 2 pages per second per host.

## `requestTimeoutSecs` (type: `integer`):

Maximum time to wait for a page. Pages that time out are listed as unreachable in the SUMMARY (not charged).

## `maxRedirects` (type: `integer`):

Redirects are followed up to this number of hops; the audited page is the final one.

## `maxPageSizeKb` (type: `integer`):

HTML beyond this size is not downloaded; the page is audited on the first part and marked truncated.

## Actor input object example

```json
{
  "urls": [
    "https://apify.com",
    "https://www.python.org/",
    "https://wordpress.org/news/"
  ],
  "datasetUrlField": "url",
  "maxUrls": 1000,
  "titleMinChars": 30,
  "titleMaxChars": 60,
  "descriptionMinChars": 70,
  "descriptionMaxChars": 160,
  "minWords": 200,
  "maxConcurrency": 10,
  "perHostConcurrency": 2,
  "perHostDelayMs": 500,
  "requestTimeoutSecs": 20,
  "maxRedirects": 5,
  "maxPageSizeKb": 3072
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `skipped` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com",
        "https://www.python.org/",
        "https://wordpress.org/news/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("kernfetch/onpage-seo-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://apify.com",
        "https://www.python.org/",
        "https://wordpress.org/news/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("kernfetch/onpage-seo-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com",
    "https://www.python.org/",
    "https://wordpress.org/news/"
  ]
}' |
apify call kernfetch/onpage-seo-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,kernfetch/onpage-seo-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pYR6GbzKVZDcdOU3h/builds/Gp7xlHrz1iB0J6MFD/openapi.json
