# Schema Markup Extractor — JSON-LD & SEO Audit (`zenomastro/schema-markup-extractor-pro`) Actor

Schema markup extractor and structured data API for static or JavaScript-rendered pages. Extract JSON-LD, Microdata, RDFa, Schema.org types, metadata, E-E-A-T and LocalBusiness signals, plus technical SEO checks. Includes auto browser fallback, selector waits, proxy support and SSRF protection.

- **URL**: https://apify.com/zenomastro/schema-markup-extractor-pro.md
- **Developed by:** [Rosario Vitale](https://apify.com/zenomastro) (community)
- **Categories:** SEO tools, Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 processed schema pages

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Schema Markup Extractor API — JS, JSON-LD & SEO Audit

### Why use this Actor?

Schema markup extractor and structured data API for static or JavaScript-rendered pages. Extract JSON-LD, Microdata, RDFa, Schema.org types, metadata, E-E-A-T and LocalBusiness signals, plus technical SEO checks. Includes auto browser fallback, selector waits, proxy support and SSRF protection.

### Features

- **Page URLs** — Public HTTP/HTTPS pages to inspect for structured data.
- **Output mode** — Emit one consolidated page row or one row per JSON-LD schema object.
- **Include Open Graph and Twitter cards** — Extract og:\* and twitter:\* metadata alongside JSON-LD.
- **Include page metadata** — Extract title, description, canonical URL and document language.
- **Request timeout** — Maximum seconds for each page request.
- **User agent** — HTTP User-Agent sent to target pages.
- **Include Microdata** — Extract Schema.org Microdata itemscope/itemprop structures.
- **Include RDFa** — Extract RDFa typeof/property/resource evidence.
- **Include technical SEO audit** — Score title, description, canonical, H1, viewport, image alt text, robots and structured-data coverage.
- **Include author / E-E-A-T evidence** — Extract factual author, publisher, organization, byline and about/editorial-link evidence without making subjective quality claims.
- **Include LocalBusiness / NAP summaries** — Extract factual LocalBusiness-style name, address, phone, geo coordinates, opening hours, identifiers and Google Maps/Place-ID signals from JSON-LD.
- **Rendering mode** — HTTP reads server HTML. Browser executes JavaScript. Auto uses fast HTTP first and renders with Chromium when a scripted page exposes too little static content or structured data.

### Use cases

- Structured-data audits.
- Technical seo qa.
- Schema.org extraction.
- Page metadata datasets.

### Example input

```json
{
  "urls": [
    "https://www.imdb.com/title/tt0111161/"
  ],
  "emitMode": "page",
  "includeOpenGraph": true,
  "includeMeta": true,
  "requestTimeoutSecs": 20,
  "userAgent": "Mozilla/5.0 (compatible; ApifySchemaExtractor/1.0)"
}
```

### Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

### FAQ

**What is this Actor for?**\
It is designed for structured-data audits, technical SEO QA, Schema.org extraction.

**Can I run it on a schedule?**\
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

**How do I control cost and run size?**\
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

### Search keywords

schema markup extractor, schema org extractor, schema markup example, schema markup what is it, what is schema markup in seo, schema markup code, structured data extraction, structured data extraction llm, structured data extraction form, structured data extraction from pdf, structured data extractor, structured data extraction matrix, structured data extraction langchain, structured data extraction using llms

Extract structured data from public web pages for SEO audits, entity pipelines, ecommerce monitoring, search enrichment and AI/RAG workflows.

The Actor parses every `application/ld+json` block, expands JSON-LD `@graph` arrays into usable objects, reports Schema.org `@type` values, and can also collect Open Graph, Twitter card, canonical, title, description and language metadata.

### Input

```json
{"urls":["https://example.com"],"emitMode":"page","includeOpenGraph":true,"includeMeta":true,"requestTimeoutSecs":20}
```

Use `page` mode for one consolidated record per URL. Use `schema_objects` when you want each JSON-LD object as a separate dataset row.

### Reliability

The Actor validates and deduplicates URLs, isolates malformed JSON-LD blocks instead of crashing a batch, exposes parse diagnostics, follows HTTP redirects, and uses explicit request timeouts.

### Pricing

Target launch price: **$0.0015 per successfully processed page/result**, below the current leading general schema-extraction tools while maintaining room for reliable operation. Invalid inputs and request errors are diagnostic rows.

### Use cases

Technical SEO, rich-results auditing, product/entity extraction, organization metadata, job/event/product schema collection, website intelligence and structured LLM ingestion.

### Responsible use

Only process public web pages and follow applicable site terms, copyright/privacy rules, robots directives and rate limits.

### Support

For reproducible issues provide the public URL, input settings and Apify run ID. Never include private credentials.

### Extended capabilities

- Extract JSON-LD, Microdata, and RDFa plus page metadata and structured-data type inventory.
- Audit SEO/schema signals and factual author, publisher, organization, byline, and about/editorial-page evidence.
- E-E-A-T-related output reports observable signals rather than assigning subjective quality scores.

# Actor input Schema

## `urls` (type: `array`):

Public HTTP/HTTPS pages to inspect for structured data.

## `emitMode` (type: `string`):

Emit one consolidated page row or one row per JSON-LD schema object.

## `includeOpenGraph` (type: `boolean`):

Extract og:\* and twitter:\* metadata alongside JSON-LD.

## `includeMeta` (type: `boolean`):

Extract title, description, canonical URL and document language.

## `requestTimeoutSecs` (type: `integer`):

Maximum seconds for each page request.

## `userAgent` (type: `string`):

HTTP User-Agent sent to target pages.

## `includeMicrodata` (type: `boolean`):

Extract Schema.org Microdata itemscope/itemprop structures.

## `includeRdfa` (type: `boolean`):

Extract RDFa typeof/property/resource evidence.

## `includeSeoAudit` (type: `boolean`):

Score title, description, canonical, H1, viewport, image alt text, robots and structured-data coverage.

## `includeEeatSignals` (type: `boolean`):

Extract factual author, publisher, organization, byline and about/editorial-link evidence without making subjective quality claims.

## `includeLocalBusinessSignals` (type: `boolean`):

Extract factual LocalBusiness-style name, address, phone, geo coordinates, opening hours, identifiers and Google Maps/Place-ID signals from JSON-LD.

## `renderMode` (type: `string`):

HTTP reads server HTML. Browser executes JavaScript. Auto uses fast HTTP first and renders with Chromium when a scripted page exposes too little static content or structured data.

## `minStaticTextChars` (type: `integer`):

In auto mode, browser rendering can be used for thin scripted pages with no structured-data evidence.

## `browserWaitUntil` (type: `string`):

Readiness strategy for JavaScript-rendered pages.

## `browserDelayMs` (type: `integer`):

Extra wait after browser navigation before structured data is extracted.

## `browserWaitForSelector` (type: `string`):

Optional CSS selector that must become visible before extraction.

## `proxyConfiguration` (type: `object`):

Optional Apify/custom proxy for browser-rendered pages. Private/local/reserved destinations remain blocked.

## Actor input object example

```json
{
  "urls": [
    "https://www.imdb.com/title/tt0111161/"
  ],
  "emitMode": "page",
  "includeOpenGraph": true,
  "includeMeta": true,
  "requestTimeoutSecs": 20,
  "userAgent": "Mozilla/5.0 (compatible; ApifySchemaExtractor/1.0)",
  "includeMicrodata": true,
  "includeRdfa": true,
  "includeSeoAudit": true,
  "includeEeatSignals": true,
  "includeLocalBusinessSignals": true,
  "renderMode": "auto",
  "minStaticTextChars": 200,
  "browserWaitUntil": "domcontentloaded",
  "browserDelayMs": 400,
  "browserWaitForSelector": "",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("zenomastro/schema-markup-extractor-pro").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("zenomastro/schema-markup-extractor-pro").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call zenomastro/schema-markup-extractor-pro --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zenomastro/schema-markup-extractor-pro"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/GGdmz9hPD86c3WrYZ/builds/obP3p14CQpi4F8qT7/openapi.json
