Schema Markup Extractor — JSON-LD & SEO Audit avatar

Schema Markup Extractor — JSON-LD & SEO Audit

Pricing

from $1.50 / 1,000 processed schema pages

Go to Apify Store
Schema Markup Extractor — JSON-LD & SEO Audit

Schema Markup Extractor — JSON-LD & SEO Audit

Schema markup extractor and structured data API for static or JavaScript-rendered pages. Extract JSON-LD, Microdata, RDFa, Schema.org types, metadata, E-E-A-T and LocalBusiness signals, plus technical SEO checks. Includes auto browser fallback, selector waits, proxy support and SSRF protection.

Pricing

from $1.50 / 1,000 processed schema pages

Rating

0.0

(0)

Developer

Rosario Vitale

Rosario Vitale

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

18 hours ago

Last modified

Share

Schema Markup Extractor API — JS, JSON-LD & SEO Audit

Why use this Actor?

Schema markup extractor and structured data API for static or JavaScript-rendered pages. Extract JSON-LD, Microdata, RDFa, Schema.org types, metadata, E-E-A-T and LocalBusiness signals, plus technical SEO checks. Includes auto browser fallback, selector waits, proxy support and SSRF protection.

Features

  • Page URLs — Public HTTP/HTTPS pages to inspect for structured data.
  • Output mode — Emit one consolidated page row or one row per JSON-LD schema object.
  • Include Open Graph and Twitter cards — Extract og:* and twitter:* metadata alongside JSON-LD.
  • Include page metadata — Extract title, description, canonical URL and document language.
  • Request timeout — Maximum seconds for each page request.
  • User agent — HTTP User-Agent sent to target pages.
  • Include Microdata — Extract Schema.org Microdata itemscope/itemprop structures.
  • Include RDFa — Extract RDFa typeof/property/resource evidence.
  • Include technical SEO audit — Score title, description, canonical, H1, viewport, image alt text, robots and structured-data coverage.
  • Include author / E-E-A-T evidence — Extract factual author, publisher, organization, byline and about/editorial-link evidence without making subjective quality claims.
  • Include LocalBusiness / NAP summaries — Extract factual LocalBusiness-style name, address, phone, geo coordinates, opening hours, identifiers and Google Maps/Place-ID signals from JSON-LD.
  • Rendering mode — HTTP reads server HTML. Browser executes JavaScript. Auto uses fast HTTP first and renders with Chromium when a scripted page exposes too little static content or structured data.

Use cases

  • Structured-data audits.
  • Technical seo qa.
  • Schema.org extraction.
  • Page metadata datasets.

Example input

{
"urls": [
"https://www.imdb.com/title/tt0111161/"
],
"emitMode": "page",
"includeOpenGraph": true,
"includeMeta": true,
"requestTimeoutSecs": 20,
"userAgent": "Mozilla/5.0 (compatible; ApifySchemaExtractor/1.0)"
}

Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

FAQ

What is this Actor for?
It is designed for structured-data audits, technical SEO QA, Schema.org extraction.

Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

Search keywords

schema markup extractor, schema org extractor, schema markup example, schema markup what is it, what is schema markup in seo, schema markup code, structured data extraction, structured data extraction llm, structured data extraction form, structured data extraction from pdf, structured data extractor, structured data extraction matrix, structured data extraction langchain, structured data extraction using llms

Extract structured data from public web pages for SEO audits, entity pipelines, ecommerce monitoring, search enrichment and AI/RAG workflows.

The Actor parses every application/ld+json block, expands JSON-LD @graph arrays into usable objects, reports Schema.org @type values, and can also collect Open Graph, Twitter card, canonical, title, description and language metadata.

Input

{"urls":["https://example.com"],"emitMode":"page","includeOpenGraph":true,"includeMeta":true,"requestTimeoutSecs":20}

Use page mode for one consolidated record per URL. Use schema_objects when you want each JSON-LD object as a separate dataset row.

Reliability

The Actor validates and deduplicates URLs, isolates malformed JSON-LD blocks instead of crashing a batch, exposes parse diagnostics, follows HTTP redirects, and uses explicit request timeouts.

Pricing

Target launch price: $0.0015 per successfully processed page/result, below the current leading general schema-extraction tools while maintaining room for reliable operation. Invalid inputs and request errors are diagnostic rows.

Use cases

Technical SEO, rich-results auditing, product/entity extraction, organization metadata, job/event/product schema collection, website intelligence and structured LLM ingestion.

Responsible use

Only process public web pages and follow applicable site terms, copyright/privacy rules, robots directives and rate limits.

Support

For reproducible issues provide the public URL, input settings and Apify run ID. Never include private credentials.

Extended capabilities

  • Extract JSON-LD, Microdata, and RDFa plus page metadata and structured-data type inventory.
  • Audit SEO/schema signals and factual author, publisher, organization, byline, and about/editorial-page evidence.
  • E-E-A-T-related output reports observable signals rather than assigning subjective quality scores.