Schema Markup Extractor — JSON-LD & SEO Audit
Pricing
from $1.50 / 1,000 processed schema pages
Schema Markup Extractor — JSON-LD & SEO Audit
Schema markup extractor and structured data API for static or JavaScript-rendered pages. Extract JSON-LD, Microdata, RDFa, Schema.org types, metadata, E-E-A-T and LocalBusiness signals, plus technical SEO checks. Includes auto browser fallback, selector waits, proxy support and SSRF protection.
Pricing
from $1.50 / 1,000 processed schema pages
Rating
0.0
(0)
Developer
Rosario Vitale
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
18 hours ago
Last modified
Categories
Share
Schema Markup Extractor API — JS, JSON-LD & SEO Audit
Why use this Actor?
Schema markup extractor and structured data API for static or JavaScript-rendered pages. Extract JSON-LD, Microdata, RDFa, Schema.org types, metadata, E-E-A-T and LocalBusiness signals, plus technical SEO checks. Includes auto browser fallback, selector waits, proxy support and SSRF protection.
Features
- Page URLs — Public HTTP/HTTPS pages to inspect for structured data.
- Output mode — Emit one consolidated page row or one row per JSON-LD schema object.
- Include Open Graph and Twitter cards — Extract og:* and twitter:* metadata alongside JSON-LD.
- Include page metadata — Extract title, description, canonical URL and document language.
- Request timeout — Maximum seconds for each page request.
- User agent — HTTP User-Agent sent to target pages.
- Include Microdata — Extract Schema.org Microdata itemscope/itemprop structures.
- Include RDFa — Extract RDFa typeof/property/resource evidence.
- Include technical SEO audit — Score title, description, canonical, H1, viewport, image alt text, robots and structured-data coverage.
- Include author / E-E-A-T evidence — Extract factual author, publisher, organization, byline and about/editorial-link evidence without making subjective quality claims.
- Include LocalBusiness / NAP summaries — Extract factual LocalBusiness-style name, address, phone, geo coordinates, opening hours, identifiers and Google Maps/Place-ID signals from JSON-LD.
- Rendering mode — HTTP reads server HTML. Browser executes JavaScript. Auto uses fast HTTP first and renders with Chromium when a scripted page exposes too little static content or structured data.
Use cases
- Structured-data audits.
- Technical seo qa.
- Schema.org extraction.
- Page metadata datasets.
Example input
{"urls": ["https://www.imdb.com/title/tt0111161/"],"emitMode": "page","includeOpenGraph": true,"includeMeta": true,"requestTimeoutSecs": 20,"userAgent": "Mozilla/5.0 (compatible; ApifySchemaExtractor/1.0)"}
Pricing & cost control
Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.
FAQ
What is this Actor for?
It is designed for structured-data audits, technical SEO QA, Schema.org extraction.
Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.
How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.
Search keywords
schema markup extractor, schema org extractor, schema markup example, schema markup what is it, what is schema markup in seo, schema markup code, structured data extraction, structured data extraction llm, structured data extraction form, structured data extraction from pdf, structured data extractor, structured data extraction matrix, structured data extraction langchain, structured data extraction using llms
Extract structured data from public web pages for SEO audits, entity pipelines, ecommerce monitoring, search enrichment and AI/RAG workflows.
The Actor parses every application/ld+json block, expands JSON-LD @graph arrays into usable objects, reports Schema.org @type values, and can also collect Open Graph, Twitter card, canonical, title, description and language metadata.
Input
{"urls":["https://example.com"],"emitMode":"page","includeOpenGraph":true,"includeMeta":true,"requestTimeoutSecs":20}
Use page mode for one consolidated record per URL. Use schema_objects when you want each JSON-LD object as a separate dataset row.
Reliability
The Actor validates and deduplicates URLs, isolates malformed JSON-LD blocks instead of crashing a batch, exposes parse diagnostics, follows HTTP redirects, and uses explicit request timeouts.
Pricing
Target launch price: $0.0015 per successfully processed page/result, below the current leading general schema-extraction tools while maintaining room for reliable operation. Invalid inputs and request errors are diagnostic rows.
Use cases
Technical SEO, rich-results auditing, product/entity extraction, organization metadata, job/event/product schema collection, website intelligence and structured LLM ingestion.
Responsible use
Only process public web pages and follow applicable site terms, copyright/privacy rules, robots directives and rate limits.
Support
For reproducible issues provide the public URL, input settings and Apify run ID. Never include private credentials.
Extended capabilities
- Extract JSON-LD, Microdata, and RDFa plus page metadata and structured-data type inventory.
- Audit SEO/schema signals and factual author, publisher, organization, byline, and about/editorial-page evidence.
- E-E-A-T-related output reports observable signals rather than assigning subjective quality scores.