Schema.org Structured Data Extractor & Rich Results Checker avatar

Schema.org Structured Data Extractor & Rich Results Checker

Pricing

from $1.50 / 1,000 pages

Go to Apify Store
Schema.org Structured Data Extractor & Rich Results Checker

Schema.org Structured Data Extractor & Rich Results Checker

Bulk-extract JSON-LD, Microdata and RDFa from any list of URLs and check which Google rich results each page is eligible for, with missing required/recommended properties and invalid values. HTTP-only, fast.

Pricing

from $1.50 / 1,000 pages

Rating

0.0

(0)

Developer

Deepak Ganesh

Deepak Ganesh

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Schema.org Structured Data Extractor & Rich Results Checker

Paste a list of page URLs and get, for every page:

  • All structured data in one normalized format: JSON-LD (including @graph, arrays and several scripts per page), Microdata (nested itemscope, itemref) and RDFa Lite (vocab / typeof / property).
  • Google rich result eligibility per entity: Product snippet, Merchant listing, Product variants, Article, Breadcrumb, FAQ, Recipe, Event, Job posting, Video, Review snippet, Local business, Organization logo, Software app, Course info, Book actions, Site name, Profile page and more.
  • What to fix: missing required properties (errors), missing recommended properties (warnings) and invalid values, such as dates that are not ISO 8601, prices like "$25" or "1,299.00", ratings outside the rating scale, relative image URLs, bad currency codes and unknown availability values.
  • Broken JSON-LD detection: invalid JSON is reported with the parser error. Blocks with comments or trailing commas are repaired and still analyzed, but flagged.

Google's Rich Results Test checks one URL at a time. This Actor checks hundreds of pages in a single run, using plain HTTP requests (no browser), so it is fast and cheap.

Use cases

  • SEO audits: find every product, recipe or article page that lost its rich result because of one missing property.
  • E-commerce QA: check price, currency, availability and review markup across your whole catalogue after a theme or platform change.
  • Competitor research: see which schema types and rich results competitors use.
  • Migration checks: compare structured data before and after a CMS migration or redesign.
  • Data extraction: use the normalized JSON (data) to get product prices, ratings, recipes, events or job postings from any site that publishes schema.org markup.

Input

FieldDefaultDescription
urls–Page URLs to check, one per line. Plain domains like example.com are accepted. Duplicates are removed.
includeRawtrueInclude the extracted, normalized JSON of each entity as data. Turn it off if you only need the validation report.
concurrency10Maximum number of pages fetched in parallel.
requestTimeoutSecs30Timeout per request. Failed requests are retried twice.
proxyConfigurationoffOptional. Use it only if sites block you.
{
"urls": [
"https://www.bbcgoodfood.com/recipes/easy-pancakes",
"https://www.allbirds.com/products/mens-tree-runners"
],
"includeRaw": false
}

Output

One dataset item per URL. The dataset has an Overview view (one row per page) and an Entities view (one row per schema.org entity). Export it as JSON, CSV or Excel, or read it through the API.

{
"inputUrl": "https://www.bbcgoodfood.com/recipes/easy-pancakes",
"url": "https://www.bbcgoodfood.com/recipes/easy-pancakes",
"finalUrl": "https://www.bbcgoodfood.com/recipes/easy-pancakes",
"statusCode": 200,
"success": true,
"error": null,
"formatsFound": ["json-ld"],
"entityCount": 3,
"entities": [
{
"types": ["VideoObject"],
"format": "json-ld",
"id": null,
"name": "How to make perfect pancakes",
"eligibleFor": ["Video"],
"missingRequired": [],
"missingRecommended": ["description", "expires", "interactionStatistic"],
"invalidValues": [],
"checks": [
{ "richResult": "Video", "eligible": true, "missingRequired": [], "missingRecommended": ["description", "expires", "interactionStatistic"] }
],
"errors": [],
"warnings": [
"Missing recommended property \"description\" (Video)",
"Missing recommended property \"expires\" (Video)",
"Missing recommended property \"interactionStatistic\" (Video)"
]
},
{
"types": ["Recipe"],
"format": "json-ld",
"id": "https://www.bbcgoodfood.com/recipes/easy-pancakes#Recipe",
"name": "Easy pancakes",
"eligibleFor": ["Recipe"],
"missingRequired": [],
"missingRecommended": ["aggregateRating", "video"],
"invalidValues": [],
"errors": [],
"warnings": ["Missing recommended property \"aggregateRating\" (Recipe)", "Missing recommended property \"video\" (Recipe)"]
},
{ "types": ["BreadcrumbList"], "format": "json-ld", "eligibleFor": ["Breadcrumb"], "errors": [], "warnings": [] }
],
"parseErrors": [],
"summary": {
"eligibleRichResults": ["Breadcrumb", "Recipe", "Video"],
"entityTypes": ["BreadcrumbList", "Recipe", "VideoObject"],
"errorCount": 0,
"warningCount": 5
},
"analyzedAt": "2026-10-07T16:26:48.537Z"
}

(Trimmed from a real run. data, the extracted JSON of each entity, is left out here.)

An entity with problems looks like this (a Microdata product):

{
"types": ["Product"],
"format": "microdata",
"eligibleFor": [],
"missingRequired": [],
"invalidValues": [
{ "property": "offers.price", "value": "1,299.00", "message": "Price must be a number using \".\" as decimal separator, without currency symbols or thousands separators", "severity": "error" },
{ "property": "aggregateRating.ratingValue", "value": "4,4", "message": "ratingValue must be a number (use \".\" as decimal separator)", "severity": "error" }
]
}
  • errors are things that make the entity ineligible for a rich result: a missing required property, or an invalid value in a required property. Invalid values elsewhere are still listed as errors, but they do not block eligibility.
  • warnings are missing recommended properties and minor problems such as relative URLs.
  • parseErrors lists JSON-LD blocks that are not valid JSON. recovered: true means the block was repaired and analyzed anyway.
  • A page that fails (DNS error, timeout, HTTP 4xx/5xx, bot challenge) returns success: false with an error, and you are not charged for it.

Rich result rules

The rules live in a single data table in src/rules.ts. They approximate Google's structured data documentation. They are not Google's official validator, and eligibility does not guarantee that Google will show a rich result.

Rich resultschema.org typesRequired (summary)
ArticleArticle, NewsArticle, BlogPosting and subtypesnone (recommended: headline, image, datePublished, dateModified, author.name/url)
Product snippetProduct (and subtypes)name, and one of offers / review / aggregateRating; price in each offer; ratingValue plus ratingCount or reviewCount; review author and reviewRating
Merchant listingProduct with offersname, image, offers.price, offers.priceCurrency
Product variantsProductGroupname, productGroupID, hasVariant (with a name or URL)
Offer (no standalone rich result)Offer, AggregateOfferprice (or lowPrice / priceSpecification.price)
Review snippetReview, AggregateRating (top level)itemReviewed.name, author, reviewRating.ratingValue / ratingValue and a count
BreadcrumbBreadcrumbListitemListElement with position and name for each item
FAQFAQPagemainEntity, each question's name, acceptedAnswer.text
Organization logoOrganization and subtypeslogo, url
Local businessLocalBusiness and 100+ subtypesname, address
EventEvent and subtypesname, startDate, location (with an address, or a URL for virtual events)
RecipeRecipename, image
Job postingJobPostingtitle, description, datePosted, hiringOrganization.name, jobLocation.address or jobLocationType
VideoVideoObjectname, thumbnailUrl, uploadDate
How-to (deprecated by Google)HowToname, step
Software appSoftwareApplication, MobileApplication, WebApplication, VideoGamename, offers.price, aggregateRating or review
Course infoCoursename, description
Book actionsBookname, author, url
Site nameWebSitename, url
Sitelinks search box (retired by Google)WebSite with potentialActionpotentialAction.target, query-input
Profile pageProfilePagemainEntity.name
(no rich result)Personname

Value checks run on every property anywhere inside an entity: ISO 8601 dates (datePublished, startDate, uploadDate, priceValidUntil…), ISO 8601 durations (cookTime, duration…), numeric prices, 3-letter currency codes, absolute http(s) URLs (url, image, logo, thumbnailUrl, item…), ratingValue within worstRating–bestRating (default 1–5), whole-number counts and positions, and known schema.org values for availability, itemCondition, eventStatus and eventAttendanceMode. {"@id": "..."} references are followed within the page.

Pricing

Pay per event: you only pay for pages that were analyzed.

EventPriceWhen
Page$0.0015 ($1.50 per 1,000 pages)Each page that loaded and was analyzed, including pages that turn out to have no structured data.

Failed pages are free: invalid URLs, DNS errors, timeouts, HTTP 4xx/5xx responses and bot-challenge pages. The Actor respects your maximum cost per run. It analyzes as many pages as your budget allows, then stops cleanly.

FAQ

Is this the same as Google's Rich Results Test? No. It follows Google's published documentation for each feature, but Google's own tools and Search Console are the final word. Use this Actor to find problems in bulk, then confirm individual pages in Google's tools. This Actor is not affiliated with or endorsed by Google.

Does it run JavaScript? No. It reads the HTML returned by the server. Structured data injected only by client-side JavaScript (for example by some tag managers) will not be seen. Google usually renders JavaScript, so results can differ for those sites.

Why does a page show HTTP 403 or bot challenge? Some sites block automated requests. Try again with Apify Proxy (residential). You are not charged for blocked pages.

Which RDFa is supported? RDFa Lite with schema.org (vocab, typeof, property, resource, prefix). Non-schema.org RDFa (for example MediaWiki's mw: annotations) is ignored.

Can I use it from code or AI agents? Yes. Call it with the Apify API, the JavaScript and Python clients, or as a tool through Apify's MCP server.

Changelog

  • 0.1: Initial release. JSON-LD, Microdata and RDFa extraction; rich result rules for 20+ Google features; value checks; repair of broken JSON-LD.