Structured Data & JSON-LD Extractor (Schema.org, Open Graph) avatar

Structured Data & JSON-LD Extractor (Schema.org, Open Graph)

Pricing

from $1.20 / 1,000 results

Go to Apify Store
Structured Data & JSON-LD Extractor (Schema.org, Open Graph)

Structured Data & JSON-LD Extractor (Schema.org, Open Graph)

Structured Data & JSON-LD Extractor reads every Schema.org JSON-LD block and Open Graph tag on a page and returns one clean row per URL — product, price, rating, article, job posting, event, recipe, FAQ and breadcrumb data, ready for RAG pipelines and SEO rich-result audits.

Pricing

from $1.20 / 1,000 results

Rating

0.0

(0)

Developer

Murat Uzun

Murat Uzun

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What is Structured Data & JSON-LD Extractor?

Structured Data & JSON-LD Extractor is an Apify Actor that reads every <script type="application/ld+json"> block and the Open Graph / Twitter Card meta tags on a page and turns them into one clean JSON row per URL. Point it at a product page, a news article, a job posting, an event, a recipe or an FAQ page and it returns the price, rating, headline, salary, breadcrumbs or Q&A pairs the site already publishes for Google's rich results — no scraper to write, no CSS selectors to maintain. It runs on plain HTTP fetches (no browser, no proxy needed), so it is fast and cheap even for hundreds of URLs, and works from the Apify Console, API, schedules or the Apify MCP server for AI agents.

What data does Structured Data & JSON-LD Extractor extract?

Structured Data & JSON-LD Extractor extracts up to 24 fields per URL, grouped by entity type — only the ones the page actually publishes are filled in, the rest are null:

FieldTypeDescription
url, finalUrl, statusCodestring, numberURL requested, URL after redirects, HTTP status
title, metaDescription, canonical, langstring<title>, meta description, canonical link and page language
schemaTypesarrayEvery distinct Schema.org @type found on the page, e.g. ["Product", "BreadcrumbList", "FAQPage"]
productobject{name, sku, brand, description, image, price, priceCurrency, availability, url}
aggregateRating, reviewsobject, array{ratingValue, reviewCount, ratingCount} and up to 20 {author, rating, datePublished, body}
organizationobject{name, url, logo, sameAs} — the site's brand identity
articleobject{headline, author, datePublished, dateModified, publisher}
breadcrumbsarrayBreadcrumb trail names, in order
jobPostingobject{title, hiringOrganization, location, datePosted, validThrough, employmentType, salary}
event, recipe, videoObjectobject{name, startDate, endDate, location}, {name, totalTime, calories, ingredientsCount}, {name, duration, uploadDate}
faqarray{question, answer} pairs from FAQPage / QAPage nodes
openGraphobjectog:title, og:description, og:image, og:type, og:site_name, twitter:card and related tags
jsonLd, jsonLdCountarray, numberEvery raw parsed JSON-LD document (optional) and how many parsed successfully
error, scrapedAtstringWhy a row has no data, and when it was fetched

How to use Structured Data & JSON-LD Extractor

  1. Paste the page URLs you want into URLs — product pages, articles, job ads, anything.
  2. Leave Include raw JSON-LD and Include Open Graph on for the full picture, or turn them off to keep the dataset small.
  3. Click Start. Each row lands in the dataset with the entity fields already mapped — export as JSON, CSV, Excel or HTML, or pull it via the API.

Example input

{
"urls": ["https://www.apify.com", "https://www.allbirds.com/products/mens-tree-dashers"],
"includeRawJsonLd": true,
"includeOpenGraph": true,
"followRedirects": true,
"maxConcurrency": 5
}

Example output

{
"url": "https://www.allbirds.com/products/mens-tree-dashers",
"statusCode": 200,
"title": "Men's Tree Dasher 2 - Running Shoes",
"schemaTypes": ["ProductGroup", "Brand", "Offer", "AggregateRating", "Product"],
"product": {
"name": "Men's Tree Dasher 2",
"brand": "Allbirds",
"price": 140,
"priceCurrency": "USD",
"availability": "OutOfStock",
"url": "https://www.allbirds.com/products/mens-tree-dashers"
},
"aggregateRating": { "ratingValue": 4.5, "reviewCount": 3449, "ratingCount": null },
"error": null,
"scrapedAt": "2026-09-12T17:40:00.000Z"
}

You can download the dataset in various formats such as JSON, HTML, CSV or Excel.

Input parameters

ParameterTypeDefaultDescription
urlsarray["https://www.apify.com"]Page URLs to extract structured data from, one row each
includeRawJsonLdbooleantrueAlso return every raw parsed JSON-LD document in jsonLd (capped at 400 KB/row)
includeOpenGraphbooleantrueAlso return Open Graph / Twitter Card meta tags in openGraph
followRedirectsbooleantrueFollow HTTP redirects; turn off to get an error row on a 3xx instead
maxConcurrencyinteger5URLs fetched in parallel (1-20)

Pricing

Structured Data & JSON-LD Extractor uses pay-per-event pricing: $0.002 per URL result, i.e. $2 per 1,000 pages, plus a negligible actor-start fee, platform usage included. Every page is a single HTTP GET with no browser and no proxy, so even runs of thousands of URLs stay cheap. Set Maximum cost per run and the Actor trims the list to what the budget covers instead of overspending.

Structured Data & JSON-LD Extractor vs. manually inspecting "View Source"

Checking a handful of pages for rich-result eligibility by hand — opening dev tools, finding the ld+json script, pretty-printing it — does not scale past a handful of URLs and misses pages where the JSON is malformed (trailing commas, HTML comments) or split across several <script> tags. This Actor parses leniently, flattens @graph containers, maps ten entity types to clean fields and returns a flat dataset you can filter, join or feed into a pipeline — for one product page or fifty thousand.

Using Structured Data & JSON-LD Extractor with AI agents and MCP

Structured Data & JSON-LD Extractor is pay-per-event with limited permissions — the two requirements for an Actor to be callable through the Apify MCP server at mcp.apify.com. An agent passes a list of urls and gets back one structured row per page — ready to drop into a RAG index, a price-monitoring pipeline, or an SEO audit step — without writing or maintaining a page-specific scraper. The same run works from n8n, Make, Zapier and LangChain through Apify's integrations.

FAQ

What if a page has no JSON-LD or Open Graph at all? The row still comes back with title, metaDescription, canonical and lang filled from plain HTML, schemaTypes: [], all entity fields null, and error: "No JSON-LD or Open Graph data on the page".

Does this work on JavaScript-rendered sites? Only for what is present in the initial HTML response. Structured data injected client-side after the page loads (rare — most sites render JSON-LD server-side for SEO) is invisible to a plain HTTP fetch. If a page is behind an anti-bot wall (Cloudflare, DataDome, PerimeterX), the row reports error: "Blocked by anti-bot (...)" instead of guessing.

Why is product.price sometimes null on a page that clearly has a price? Some sites price only their variants (hasVariant) with no headline price on the parent ProductGroup; the Actor falls back to the first priced variant it finds, but a page with no priced offer anywhere returns null rather than a wrong number.

Is this legal to run? Yes. JSON-LD and Open Graph tags are structured metadata that site owners deliberately publish for search engines and social previews to read — this Actor reads exactly what a browser or crawler already sees.

Can I export to CSV or Excel? Yes, from the Output tab or the API, with ready-made Overview and All fields views.

Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:

Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.

Website & domain intelligence

Content for AI, LLMs and RAG

Search, video and social

Leads, jobs and company data

Developer, app and research data

Support and feedback

Found a schema type worth mapping, or a page whose JSON-LD is parsed incorrectly? Open an issue on the Issues tab.