# Schema.org Structured Data Extractor & Rich Results Checker (`webintel/structured-data-validator`) Actor

Bulk-extract JSON-LD, Microdata and RDFa from any list of URLs and check which Google rich results each page is eligible for, with missing required/recommended properties and invalid values. HTTP-only, fast.

- **URL**: https://apify.com/webintel/structured-data-validator.md
- **Developed by:** [Deepak Ganesh](https://apify.com/webintel) (community)
- **Categories:** SEO tools, E-commerce, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Schema.org Structured Data Extractor & Rich Results Checker

![Schema.org Structured Data Extractor & Rich Results Checker](https://api.apify.com/v2/key-value-stores/21SBwDwdNtlnapIdO/records/structured-data-validator.png?v=b9eeb36b)

Paste a list of page URLs and get, for every page:

- **All structured data** in one normalized format: **JSON-LD** (including `@graph`, arrays and several scripts per page), **Microdata** (nested `itemscope`, `itemref`) and **RDFa Lite** (`vocab` / `typeof` / `property`).
- **Google rich result eligibility** per entity: Product snippet, Merchant listing, Product variants, Article, Breadcrumb, FAQ, Recipe, Event, Job posting, Video, Review snippet, Local business, Organization logo, Software app, Course info, Book actions, Site name, Profile page and more.
- **What to fix**: missing **required** properties (errors), missing **recommended** properties (warnings) and **invalid values**, such as dates that are not ISO 8601, prices like `"$25"` or `"1,299.00"`, ratings outside the rating scale, relative image URLs, bad currency codes and unknown `availability` values.
- **Broken JSON-LD detection**: invalid JSON is reported with the parser error. Blocks with comments or trailing commas are repaired and still analyzed, but flagged.

Google's Rich Results Test checks **one URL at a time**. This Actor checks hundreds of pages in a single run, using plain HTTP requests (no browser), so it is fast and cheap.

### Use cases

- **SEO audits**: find every product, recipe or article page that lost its rich result because of one missing property.
- **E-commerce QA**: check price, currency, availability and review markup across your whole catalogue after a theme or platform change.
- **Competitor research**: see which schema types and rich results competitors use.
- **Migration checks**: compare structured data before and after a CMS migration or redesign.
- **Data extraction**: use the normalized JSON (`data`) to get product prices, ratings, recipes, events or job postings from any site that publishes schema.org markup.

### Input

| Field | Default | Description |
|---|---|---|
| `urls` | – | Page URLs to check, one per line. Plain domains like `example.com` are accepted. Duplicates are removed. |
| `includeRaw` | `true` | Include the extracted, normalized JSON of each entity as `data`. Turn it off if you only need the validation report. |
| `concurrency` | `10` | Maximum number of pages fetched in parallel. |
| `requestTimeoutSecs` | `30` | Timeout per request. Failed requests are retried twice. |
| `proxyConfiguration` | off | Optional. Use it only if sites block you. |

```json
{
  "urls": [
    "https://www.bbcgoodfood.com/recipes/easy-pancakes",
    "https://www.allbirds.com/products/mens-tree-runners"
  ],
  "includeRaw": false
}
```

### Output

One dataset item per URL. The dataset has an **Overview** view (one row per page) and an **Entities** view (one row per schema.org entity). Export it as JSON, CSV or Excel, or read it through the API.

```json
{
  "inputUrl": "https://www.bbcgoodfood.com/recipes/easy-pancakes",
  "url": "https://www.bbcgoodfood.com/recipes/easy-pancakes",
  "finalUrl": "https://www.bbcgoodfood.com/recipes/easy-pancakes",
  "statusCode": 200,
  "success": true,
  "error": null,
  "formatsFound": ["json-ld"],
  "entityCount": 3,
  "entities": [
    {
      "types": ["VideoObject"],
      "format": "json-ld",
      "id": null,
      "name": "How to make perfect pancakes",
      "eligibleFor": ["Video"],
      "missingRequired": [],
      "missingRecommended": ["description", "expires", "interactionStatistic"],
      "invalidValues": [],
      "checks": [
        { "richResult": "Video", "eligible": true, "missingRequired": [], "missingRecommended": ["description", "expires", "interactionStatistic"] }
      ],
      "errors": [],
      "warnings": [
        "Missing recommended property \"description\" (Video)",
        "Missing recommended property \"expires\" (Video)",
        "Missing recommended property \"interactionStatistic\" (Video)"
      ]
    },
    {
      "types": ["Recipe"],
      "format": "json-ld",
      "id": "https://www.bbcgoodfood.com/recipes/easy-pancakes#Recipe",
      "name": "Easy pancakes",
      "eligibleFor": ["Recipe"],
      "missingRequired": [],
      "missingRecommended": ["aggregateRating", "video"],
      "invalidValues": [],
      "errors": [],
      "warnings": ["Missing recommended property \"aggregateRating\" (Recipe)", "Missing recommended property \"video\" (Recipe)"]
    },
    { "types": ["BreadcrumbList"], "format": "json-ld", "eligibleFor": ["Breadcrumb"], "errors": [], "warnings": [] }
  ],
  "parseErrors": [],
  "summary": {
    "eligibleRichResults": ["Breadcrumb", "Recipe", "Video"],
    "entityTypes": ["BreadcrumbList", "Recipe", "VideoObject"],
    "errorCount": 0,
    "warningCount": 5
  },
  "analyzedAt": "2026-10-07T16:26:48.537Z"
}
```

*(Trimmed from a real run. `data`, the extracted JSON of each entity, is left out here.)*

An entity with problems looks like this (a Microdata product):

```json
{
  "types": ["Product"],
  "format": "microdata",
  "eligibleFor": [],
  "missingRequired": [],
  "invalidValues": [
    { "property": "offers.price", "value": "1,299.00", "message": "Price must be a number using \".\" as decimal separator, without currency symbols or thousands separators", "severity": "error" },
    { "property": "aggregateRating.ratingValue", "value": "4,4", "message": "ratingValue must be a number (use \".\" as decimal separator)", "severity": "error" }
  ]
}
```

- `errors` are things that make the entity **ineligible** for a rich result: a missing required property, or an invalid value in a required property. Invalid values elsewhere are still listed as errors, but they do not block eligibility.
- `warnings` are missing recommended properties and minor problems such as relative URLs.
- `parseErrors` lists JSON-LD blocks that are not valid JSON. `recovered: true` means the block was repaired and analyzed anyway.
- A page that fails (DNS error, timeout, HTTP 4xx/5xx, bot challenge) returns `success: false` with an `error`, and **you are not charged for it**.

### Rich result rules

The rules live in a single data table in [`src/rules.ts`](src/rules.ts). They **approximate** Google's [structured data documentation](https://developers.google.com/search/docs/appearance/structured-data/search-gallery). They are not Google's official validator, and eligibility does not guarantee that Google will show a rich result.

| Rich result | schema.org types | Required (summary) |
|---|---|---|
| Article | Article, NewsArticle, BlogPosting and subtypes | none (recommended: headline, image, datePublished, dateModified, author.name/url) |
| Product snippet | Product (and subtypes) | name, and one of offers / review / aggregateRating; price in each offer; ratingValue plus ratingCount or reviewCount; review author and reviewRating |
| Merchant listing | Product with `offers` | name, image, offers.price, offers.priceCurrency |
| Product variants | ProductGroup | name, productGroupID, hasVariant (with a name or URL) |
| Offer (no standalone rich result) | Offer, AggregateOffer | price (or lowPrice / priceSpecification.price) |
| Review snippet | Review, AggregateRating (top level) | itemReviewed.name, author, reviewRating.ratingValue / ratingValue and a count |
| Breadcrumb | BreadcrumbList | itemListElement with position and name for each item |
| FAQ | FAQPage | mainEntity, each question's name, acceptedAnswer.text |
| Organization logo | Organization and subtypes | logo, url |
| Local business | LocalBusiness and 100+ subtypes | name, address |
| Event | Event and subtypes | name, startDate, location (with an address, or a URL for virtual events) |
| Recipe | Recipe | name, image |
| Job posting | JobPosting | title, description, datePosted, hiringOrganization.name, jobLocation.address or jobLocationType |
| Video | VideoObject | name, thumbnailUrl, uploadDate |
| How-to (deprecated by Google) | HowTo | name, step |
| Software app | SoftwareApplication, MobileApplication, WebApplication, VideoGame | name, offers.price, aggregateRating or review |
| Course info | Course | name, description |
| Book actions | Book | name, author, url |
| Site name | WebSite | name, url |
| Sitelinks search box (retired by Google) | WebSite with potentialAction | potentialAction.target, query-input |
| Profile page | ProfilePage | mainEntity.name |
| (no rich result) | Person | name |

Value checks run on every property anywhere inside an entity: ISO 8601 dates (`datePublished`, `startDate`, `uploadDate`, `priceValidUntil`…), ISO 8601 durations (`cookTime`, `duration`…), numeric prices, 3-letter currency codes, absolute `http(s)` URLs (`url`, `image`, `logo`, `thumbnailUrl`, `item`…), `ratingValue` within `worstRating`–`bestRating` (default 1–5), whole-number counts and positions, and known schema.org values for `availability`, `itemCondition`, `eventStatus` and `eventAttendanceMode`. `{"@id": "..."}` references are followed within the page.

### Pricing

Pay per event: you only pay for pages that were analyzed.

| Event | Price | When |
|---|---|---|
| Page | **$0.0015** ($1.50 per 1,000 pages) | Each page that loaded and was analyzed, including pages that turn out to have no structured data. |

Failed pages are **free**: invalid URLs, DNS errors, timeouts, HTTP 4xx/5xx responses and bot-challenge pages. The Actor respects your **maximum cost per run**. It analyzes as many pages as your budget allows, then stops cleanly.

### FAQ

**Is this the same as Google's Rich Results Test?** No. It follows Google's published documentation for each feature, but Google's own tools and Search Console are the final word. Use this Actor to find problems in bulk, then confirm individual pages in Google's tools. This Actor is not affiliated with or endorsed by Google.

**Does it run JavaScript?** No. It reads the HTML returned by the server. Structured data injected only by client-side JavaScript (for example by some tag managers) will not be seen. Google usually renders JavaScript, so results can differ for those sites.

**Why does a page show `HTTP 403` or `bot challenge`?** Some sites block automated requests. Try again with Apify Proxy (residential). You are not charged for blocked pages.

**Which RDFa is supported?** RDFa Lite with schema.org (`vocab`, `typeof`, `property`, `resource`, `prefix`). Non-schema.org RDFa (for example MediaWiki's `mw:` annotations) is ignored.

**Can I use it from code or AI agents?** Yes. Call it with the Apify API, the JavaScript and Python clients, or as a tool through Apify's MCP server.

### Changelog

- **0.1**: Initial release. JSON-LD, Microdata and RDFa extraction; rich result rules for 20+ Google features; value checks; repair of broken JSON-LD.

# Actor input Schema

## `urls` (type: `array`):

Pages to check, one per line (product, article, recipe, event pages...). Plain domains like "example.com" are fine. Duplicates are removed.

## `includeRaw` (type: `boolean`):

Add the extracted (normalized) JSON of every entity to the output as "data". Turn off for smaller results when you only need the validation report.

## `concurrency` (type: `integer`):

Maximum number of pages fetched in parallel.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for a page to respond before retrying (2 retries).

## `proxyConfiguration` (type: `object`):

Optional. Most sites work without a proxy. Use Apify Proxy only if pages get blocked (HTTP 403).

## Actor input object example

```json
{
  "urls": [
    "https://www.allrecipes.com/recipe/10813/best-chocolate-chip-cookies/",
    "https://apify.com"
  ],
  "includeRaw": true,
  "concurrency": 10,
  "requestTimeoutSecs": 30,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All analyzed pages (JSON): url, finalUrl, statusCode, success, error, formatsFound\[], entityCount, entities\[], parseErrors\[], summary{}.

## `overview` (type: `string`):

One row per page: formats found, entity types, eligible rich results, error and warning counts.

## `entities` (type: `string`):

One row per structured data entity with eligible rich results, missing required/recommended properties and errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.allrecipes.com/recipe/10813/best-chocolate-chip-cookies/",
        "https://apify.com"
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("webintel/structured-data-validator").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://www.allrecipes.com/recipe/10813/best-chocolate-chip-cookies/",
        "https://apify.com",
    ],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("webintel/structured-data-validator").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.allrecipes.com/recipe/10813/best-chocolate-chip-cookies/",
    "https://apify.com"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call webintel/structured-data-validator --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,webintel/structured-data-validator"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/H8xG9QGy79325ncwh/builds/NeOlW2PVeWy6FcX7A/openapi.json
