# Bulk Rich Results Audit - JSON-LD, OG & Twitter Cards (`cuantic_data/bulk-rich-results-audit`) Actor

Static audit of JSON-LD structured data, Open Graph and X/Twitter Card tags for up to 500 URLs or an XML sitemap. One row per page with detected types, parse errors and concrete metadata issues, plus a run summary.

- **URL**: https://apify.com/cuantic\_data/bulk-rich-results-audit.md
- **Developed by:** [Cuantic Data](https://apify.com/cuantic_data) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk Rich Results Audit - JSON-LD, Open Graph & Twitter Card Checker

Check the structured data and social sharing metadata of a whole list of pages in one run. Paste up to 500 URLs, or give an XML sitemap, and get one dataset row per page with its JSON-LD blocks and `@type` values, its Open Graph and X/Twitter Card tags, and a list of concrete problems found in them. A summary record totals the issues across the run by category, severity and code.

Think of it as a **bulk schema markup validator** for the static layer: a **JSON-LD audit**, an **Open Graph checker** and a **Twitter Card audit** that run together over many pages, using deterministic checks you can read and reproduce.

> **What this is not.** The findings are deterministic static checks of the HTML each server returns. They are not a Google Rich Results Test result, they do not tell you whether a page is eligible for rich results in Google Search, and they do not show how Facebook, LinkedIn, X or any other platform will actually render a link preview. For pages that matter, confirm with the official tools listed under [Verify with the official tools](#verify-with-the-official-tools).

### What it does

- Fetches each page once over plain HTTP(S) (no browser, no JavaScript) and reads the server-returned HTML.
- Parses every `<script type="application/ld+json">` block independently: a broken block is reported and the others are still analyzed.
- Lists the `@type` values of top-level entities, including `@type` arrays and `@graph` nodes.
- Reads Open Graph (`og:*`) and X/Twitter Card (`twitter:*`) meta tags, the `<title>`, the meta description and the canonical link.
- Reports issues as `{category, severity, code, message}` and counts them per page.
- Accepts a list of URLs **or** an XML sitemap (a sitemap index is followed one level deep). Only those URLs are fetched; links found on pages are never followed.
- Isolates failures: a page that times out, returns 404, is not HTML or points to a refused target gets a row with an `error` message, and the run continues.
- Writes a run summary to the key-value store record `OUTPUT`.

### Who it is for

- **Site migrations and redesigns.** Snapshot the structured data and social tags of every page before the switch, run again after, and compare the two datasets for lost `Product`, `BreadcrumbList` or `Article` markup, missing `og:image` tags or canonicals that now point somewhere else.
- **Metadata regression checks.** Run the same URL list after each release (for example on an Apify schedule) and watch the `OUTPUT` totals: a jump in `jsonld-malformed` or `og-missing-image` shows up in one number.
- **Social sharing checks.** Find pages where `og:image` is a relative path, uses `http://` on an `https://` page, is missing entirely, or where `og:title` is declared twice with different values.
- **Agencies and SEO teams.** Hand over a sortable table for a whole site instead of pasting URLs one by one into a single-page tool.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `urls` | array of strings | - | Pages to audit (http/https). Deduplicated (order kept), max 500. |
| `sitemapUrl` | string | - | XML sitemap or sitemap index. Every `<loc>` page URL is audited, max 500. |
| `sameOriginOnly` | boolean | `true` | For sitemaps: keep only child sitemaps and page URLs on the same origin as the sitemap URL. |
| `maxConcurrency` | integer | `4` | Pages fetched in parallel, 1 to 10. |
| `timeoutMs` | integer | `20000` | Timeout per request including redirects, 5000 to 60000 ms. |

At least one of `urls` or `sitemapUrl` is required. If both are given, `urls` is used, the sitemap is ignored and a warning is logged.

Example input:

```json
{
    "sitemapUrl": "https://www.example-shop.com/sitemap.xml",
    "sameOriginOnly": true,
    "maxConcurrency": 4,
    "timeoutMs": 20000
}
```

### Output

One dataset item per URL:

```json
{
    "url": "https://www.example-shop.com/products/blue-widget",
    "finalUrl": "https://www.example-shop.com/products/blue-widget",
    "statusCode": 200,
    "redirectCount": 0,
    "contentType": "text/html; charset=utf-8",
    "title": "Blue Widget - Example Shop",
    "metaDescription": "A very blue widget, shipped in 48 hours.",
    "canonical": "https://www.example-shop.com/products/blue-widget",
    "structuredData": {
        "blockCount": 2,
        "blocks": [
            { "index": 0, "valid": true, "bytes": 612, "hasContext": true, "types": ["Product"], "entityCount": 1 },
            { "index": 1, "valid": false, "bytes": 288, "hasContext": false, "types": [], "entityCount": 0 }
        ],
        "types": ["Product"],
        "parseErrors": [
            { "index": 1, "message": "Expected double-quoted property name in JSON at position 271 (line 9 column 5)" }
        ],
        "microdataItemtypes": 0
    },
    "openGraph": {
        "fields": {
            "title": "Blue Widget",
            "type": "product",
            "image": "/images/blue-widget.jpg",
            "url": "https://www.example-shop.com/products/blue-widget",
            "description": "A very blue widget.",
            "siteName": "Example Shop",
            "imageCount": 1
        },
        "issues": [
            { "category": "openGraph", "severity": "error", "code": "og-image-relative-url", "message": "og:image \"/images/blue-widget.jpg\" is not an absolute URL; sharing platforms generally do not resolve relative image paths." }
        ]
    },
    "twitterCard": {
        "fields": { "card": "summary_large_image", "title": null, "description": null, "image": null, "site": "@exampleshop", "creator": null },
        "fallbackFromOpenGraph": ["title", "description", "image"],
        "issues": []
    },
    "findings": [
        { "category": "structuredData", "severity": "error", "code": "jsonld-malformed", "message": "JSON-LD block #2 is not valid JSON: Expected double-quoted property name in JSON at position 271 (line 9 column 5)" },
        { "category": "openGraph", "severity": "error", "code": "og-image-relative-url", "message": "og:image \"/images/blue-widget.jpg\" is not an absolute URL; sharing platforms generally do not resolve relative image paths." }
    ],
    "issueCounts": {
        "total": 2, "error": 2, "warning": 0, "info": 0,
        "byCategory": { "page": 0, "structuredData": 1, "openGraph": 1, "twitterCard": 0 }
    },
    "error": null
}
```

- `findings` holds every issue on the page, errors first. `openGraph.issues` and `twitterCard.issues` repeat the ones for that category.
- `structuredData.types` lists the distinct `@type` values of top-level entities and `@graph` nodes (a `https://schema.org/` prefix is removed). Nested entities such as an `Offer` inside a `Product` are not listed separately.
- `openGraph.fields` and `twitterCard.fields` show what the page declares (first non-empty value). `twitterCard.fallbackFromOpenGraph` lists the fields X can take from Open Graph because the `twitter:*` tag is absent; those are not reported as missing.
- `error` is `null` on success. On failure it holds the reason (for example `HTTP 404 for ...`, `Request timed out after 20000 ms`, `Not an HTML page (content-type: application/pdf)` or `Refused target: ...`) and the analysis fields are `null`.

Rows are written as each page finishes, so their order follows completion, not input. The `OUTPUT` record summarizes the run:

```json
{
    "total": 137,
    "audited": 134,
    "failed": 3,
    "skipped": 0,
    "chargeLimitReached": false,
    "source": "sitemap",
    "sitemapUrl": "https://www.example-shop.com/sitemap.xml",
    "findings": { "total": 412, "error": 21, "warning": 96, "info": 295 },
    "byCategory": {
        "page": { "total": 140, "error": 0, "warning": 4, "info": 136 },
        "structuredData": { "total": 58, "error": 12, "warning": 9, "info": 37 },
        "openGraph": { "total": 97, "error": 9, "warning": 80, "info": 8 },
        "twitterCard": { "total": 117, "error": 0, "warning": 3, "info": 114 }
    },
    "byCode": [
        { "code": "canonical-missing", "category": "page", "severity": "info", "pages": 120 },
        { "code": "og-missing-type", "category": "openGraph", "severity": "warning", "pages": 77 },
        { "code": "jsonld-malformed", "category": "structuredData", "severity": "error", "pages": 12 }
    ],
    "pagesWithErrors": 19,
    "structuredData": {
        "pagesWithJsonLd": 97,
        "pagesWithoutJsonLd": 37,
        "pagesWithMalformedJsonLd": 12,
        "typesByPageCount": { "BreadcrumbList": 95, "Product": 80, "Organization": 3 }
    },
    "failedUrls": [{ "url": "https://www.example-shop.com/old-page", "error": "HTTP 404 for https://www.example-shop.com/old-page" }],
    "skippedUrls": [],
    "note": "Deterministic static checks of the HTML returned by each server (no JavaScript rendering). This is not a Google Rich Results Test result and does not predict search features or how social platforms render previews."
}
```

`byCode[].pages` counts pages, not findings. There is no overall score: the counts are the result.

### Issue codes

Severities: `error` means the markup is broken or violates the format's own rules; `warning` means it is likely to cause a problem; `info` is worth a look but can be intentional.

**Structured data (JSON-LD)**

| Code | Severity | Meaning |
|---|---|---|
| `jsonld-missing` | info | No JSON-LD block in the HTML. Mentions microdata if `itemtype` attributes exist (microdata is not analyzed). |
| `jsonld-malformed` | error | A block is not valid JSON; the parser message is included. |
| `jsonld-empty-block` | warning | A block is empty or an empty array. |
| `jsonld-missing-context` | error | A top-level item has no `@context`. |
| `jsonld-context-not-schema-org` | info | `@context` does not reference schema.org. |
| `jsonld-missing-type` | warning | An entity (top level or `@graph` node) has no `@type`. Pure `{"@id": ...}` references are ignored. |
| `jsonld-invalid-item` | error | A top-level or `@graph` value is not an object, or `@type` is not a string. |
| `jsonld-duplicate-block` | warning | Two blocks contain identical JSON. |
| `jsonld-duplicate-type` | warning | More than one `WebSite`, `WebPage` or `FAQPage` entity, which usually means markup is emitted twice. Repeated `Product`, `Offer` or other types are not flagged. |

These checks cover JSON syntax and JSON-LD structure only. They do not implement the schema.org vocabulary or Google's required and recommended properties for each rich result type.

**Open Graph**

| Code | Severity | Meaning |
|---|---|---|
| `og-missing` | warning | No `og:*` tags at all. |
| `og-missing-title` / `-type` / `-image` / `-url` | warning | One of the four properties the Open Graph protocol lists as required is absent. |
| `og-empty-title` / `-type` / `-image` / `-url` | warning | The tag exists but its `content` is empty. |
| `og-image-relative-url` | error | `og:image` is a relative or protocol-relative path instead of an absolute URL. |
| `og-image-invalid-url` | error | `og:image` is not a valid http(s) URL. |
| `og-image-insecure-url` | warning | `og:image` uses `http://` on an `https://` page. |
| `og-image-secure-url-not-https` | warning | `og:image:secure_url` is not an absolute `https://` URL. |
| `og-url-relative-url` / `og-url-invalid-url` | error | `og:url` is not an absolute, valid URL. |
| `og-url-canonical-mismatch` | info | `og:url` differs from the canonical link. |
| `og-conflicting-values` | warning | `og:title`, `og:type`, `og:url`, `og:description` or `og:site_name` is declared several times with different values. |
| `og-duplicate-tag` | info | The same property is repeated with the same value. |
| `og-name-attribute` | info | `og:*` tags use `name=` instead of `property=`. |

**X / Twitter Card**

| Code | Severity | Meaning |
|---|---|---|
| `twitter-card-missing` | info / warning | No `twitter:card`. Info when `og:title` exists (X may build a card from Open Graph), warning otherwise. |
| `twitter-card-invalid-type` | error | `twitter:card` is not `summary`, `summary_large_image`, `app` or `player`. |
| `twitter-missing-title` | warning | Neither `twitter:title` nor `og:title`. |
| `twitter-missing-description` | info | Neither `twitter:description` nor `og:description`. |
| `twitter-missing-image` | warning / info | Neither `twitter:image` nor `og:image` (warning for `summary_large_image`). |
| `twitter-image-relative-url` / `-invalid-url` / `-insecure-url` | error / error / warning | Same URL checks as for `og:image`. |
| `twitter-conflicting-values` | warning | A `twitter:*` tag is declared with different values. |
| `twitter-site-format` / `twitter-creator-format` | info | The handle does not look like `@username`. |

Title, description and image checks only run for `summary` and `summary_large_image` cards (or when no card type is declared).

**Page basics**

| Code | Severity | Meaning |
|---|---|---|
| `title-missing` / `title-empty` | warning | No `<title>` or an empty one (`<title>` inside inline SVG is ignored). |
| `title-multiple` | info | More than one `<title>`. |
| `meta-description-missing` / `meta-description-multiple` | info | No meta description, or more than one. |
| `canonical-missing` | info | No `<link rel="canonical">`. |
| `canonical-invalid-url` | error | The canonical is not a valid http(s) URL. |
| `canonical-relative-url` | info | The canonical is relative. |
| `canonical-conflicting` | warning | Several canonicals with different targets. |
| `canonical-duplicate` | info | The same canonical is repeated. |
| `response-truncated` | info | The HTML was larger than 2 MB; tags after the cut were not seen. |

### Verify with the official tools

This actor is a fast first pass over many pages. For the pages that matter, verify with the official tools:

- Google Rich Results Test: https://search.google.com/test/rich-results (what Google can read and which rich result types it detects, with JavaScript rendering).
- Schema Markup Validator: https://validator.schema.org/ (schema.org vocabulary validation of JSON-LD, microdata and RDFa).
- The Open Graph protocol: https://ogp.me/ (the reference for `og:*` properties).

Each social platform also offers its own link preview debugger; the actual preview is only ever decided by that platform.

### Safety

- Only the URLs you supply, or the `<loc>` entries of the sitemap you supply, are fetched. Links on pages are never followed.
- Targets are checked before every request and every redirect: only `http` and `https`; no embedded credentials; no `localhost`, `.local`, `.internal` or single-label hostnames; no private, loopback, link-local, carrier-grade NAT, documentation, multicast or otherwise reserved IPv4/IPv6 addresses, including hostnames that resolve to one. The address is checked again when the connection is opened.
- Limits: 5 redirects, 2 MB per page and per sitemap file (after decompression), 20 child sitemaps and 10 MB of sitemap data per run, 500 URLs, the configured timeout, one retry for network errors, timeouts, 408, 429 and 5xx.
- Requests identify as `Mozilla/5.0 (compatible; CuanticDataBot/1.0)`. No proxy is used, no cookies are sent and no login is attempted.

### Limitations

- **No JavaScript.** Only the HTML the server returns is checked. JSON-LD, Open Graph or Twitter tags injected by client-side scripts or tag managers are missed, and may be reported as missing.
- **No Google eligibility and no preview guarantee.** A page with zero findings can still be ineligible for rich results, and a social platform may still render its preview differently (cropping, caching, its own rules).
- **Structural checks only.** JSON-LD is checked for syntax, `@context` and `@type`; schema.org properties and Google's per-feature requirements are not validated. Microdata and RDFa are detected by `itemtype` count only.
- **Image URLs are checked by syntax only.** Images are not downloaded, so reachability, file type, dimensions and aspect ratio are not verified.
- **Dynamic or personalized responses.** Pages that serve different HTML by user agent, location or cookies may differ from what visitors or crawlers see. Pages behind a login, bot protection or a consent wall usually return that wall instead.
- **Max 500 URLs per run.** Extra URLs are dropped with a warning. Split larger sites across several runs.
- **Sitemaps.** Plain XML only (no `.xml.gz` files); a sitemap index is followed one level deep.
- Verify important pages with the official tools listed above.

### Pricing

Pay-per-event: $0.002 USD per page fetched and parsed as HTML ($2.00 per 1,000 pages). Failed URLs (network errors, HTTP errors, non-HTML responses, refused targets) are not charged. Apify platform usage is included in this price; there is no separate platform-usage charge.

An Actor Start event costs $0.00005 USD. One start event is charged per GB of Actor memory, with a minimum of one event per run.

If the maximum charge you set for a run is reached, the actor finishes the pages in progress, starts no new ones, and lists the URLs it did not start in `OUTPUT.skippedUrls`.

### Development

```bash
npm install
npm test          # unit tests (node:test, mocked network, no internet needed)
npm run smoke     # optional: one real request to https://example.com/ (or SMOKE_URL); prints SKIPPED without egress
```

### Support

Cuantic Data - cuanticwindows@gmail.com

Terms of use: see [Terms of use](#terms-of-use) below.

### Terms of use

Provided by Cuantic Data (cuanticwindows@gmail.com).

#### 1. What the actor does

The actor downloads the HTML of the URLs you provide, either directly or through a sitemap you point it to, with plain HTTP requests. It reads the structured data (JSON-LD), Open Graph and X/Twitter Card tags, title, meta description and canonical link in that HTML, runs deterministic static checks on them, and stores the results in your Apify dataset and key-value store. It does not follow links found on pages, execute JavaScript, log in, or download images.

#### 2. Your responsibility

- It is your responsibility to audit only URLs you have the right to inspect: sites you own, manage, or are otherwise permitted to access in this way.
- Each audited page is one request (plus redirects) to the site that serves it. Choose the number of URLs and the concurrency so that you do not put undue load on any site.
- You must comply with the terms of service of the audited sites and with the Apify Terms of Service.

#### 3. About the results

- The findings are static checks of the HTML returned at the time of the run. They are not a Google Rich Results Test result, not a determination of eligibility for any search feature, and not a prediction of how any social platform renders a link preview.
- Results are provided "as is", for informational purposes, without warranty of accuracy, completeness or fitness for a particular purpose.
- Google, schema.org, the Open Graph protocol and X/Twitter are the property of their respective owners. This actor is not affiliated with or endorsed by any of them.

#### 4. Data

The actor stores only what it produces (the dataset rows and the summary record) in your own Apify storage. It does not keep copies of the audited pages or results elsewhere.

#### 5. Liability

To the extent permitted by law, Cuantic Data is not liable for any damage arising from the use of the actor or its results, including decisions made based on the findings, or any effect of the audit requests on the audited sites.

#### 6. Changes

These terms may be updated together with the actor. The version published with the actor is the one that applies.

# Actor input Schema

## `urls` (type: `array`):

Pages to audit (http or https). Only these pages are fetched; links on them are never followed. Duplicates are removed, order is kept, max 500 per run. If you also fill in a sitemap URL, this list wins and the sitemap is ignored.

## `sitemapUrl` (type: `string`):

An XML sitemap (or sitemap index, followed one level deep). Every <loc> page URL in it is audited, up to 500. Used only when the URLs list is empty. Gzipped sitemaps are not supported.

## `sameOriginOnly` (type: `boolean`):

When reading a sitemap, keep only child sitemaps and page URLs on the same origin (scheme, host and port) as the sitemap URL. Protects against auditing unrelated sites listed in a third-party sitemap.

## `maxConcurrency` (type: `integer`):

Pages fetched in parallel. Keep it low for small sites to avoid putting load on them.

## `timeoutMs` (type: `integer`):

Maximum time for one page request, including redirects. A page that takes longer gets an error row and the run continues.

## Actor input object example

```json
{
  "urls": [
    "https://example.com/",
    "https://www.iana.org/help/example-domains"
  ],
  "sameOriginOnly": true,
  "maxConcurrency": 4,
  "timeoutMs": 20000
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com/",
        "https://www.iana.org/help/example-domains"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("cuantic_data/bulk-rich-results-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://example.com/",
        "https://www.iana.org/help/example-domains",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("cuantic_data/bulk-rich-results-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com/",
    "https://www.iana.org/help/example-domains"
  ]
}' |
apify call cuantic_data/bulk-rich-results-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cuantic_data/bulk-rich-results-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/M7WdT1XMy37scgpwY/builds/2QDsM0JySsnSae6dm/openapi.json
