Meta Tags & Open Graph Checker + JSON-LD Extractor (Bulk URLs)
Pricing
from $2.00 / 1,000 successful lookups
Meta Tags & Open Graph Checker + JSON-LD Extractor (Bulk URLs)
A schema.org extractor for a list of URLs without an API key: pulls every JSON-LD object a page publishes, plus meta title, description, canonical and Open Graph tags, from the same HTML, no rendering. Where a structured data testing tool checks one page, this does your whole list, to Google Sheets.
Pricing
from $2.00 / 1,000 successful lookups
Rating
0.0
(0)
Developer
Adrian Voss
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
Two audits in one fetch, for a whole list of URLs at once. First, a full meta tags
and Open Graph checker: title, meta description, canonical, robots, hreflang
alternates, Open Graph, Twitter Card, favicon, H1s, and a flagged list of common SEO
issues (missing description, an oversized title, a missing og:image, a canonical
mismatch). Second, the original JSON-LD extractor: every
schema.org Product, Article, JobPosting, Organization,
BreadcrumbList, FAQPage and more object already embedded in the page. No
third-party API, no rendering — one HTML fetch per URL, everything above read straight
out of it and handed back as clean, structured rows.
Features
- Meta tag & Open Graph audit. Title, description, canonical, robots, hreflang,
Open Graph (
og:title/description/image/type/url/site_name) and Twitter Card (twitter:card/title/description/image) tags, plus favicon and H1 count — from the same fetch, no extra requests. - Built-in issue detection. Every row carries an
issueslist:missing-description,title-too-long(over 60 characters),missing-og-image,canonical-mismatch. - Full JSON-LD object extraction. Every JSON-LD block on the page is parsed and
returned, including objects nested inside a top-level
@graph. - Type summary. A deduplicated list of every
@typefound on the page (Product,Article,JobPosting, etc.), so you can see what's there without scanning the raw objects. - Resilient parsing. A malformed JSON-LD block is skipped on its own — one bad block doesn't fail the whole page.
- Pick your columns. Every field above is its own output column, on by default — turn off what you don't need in the Console's "Which columns do you want?" input.
- Bulk-friendly. Feed in a list of URLs with adjustable
maxConcurrency.
How to use JSON-LD Structured Data Lookup — Schema.org Extractor API
- In the Apify Console. Open the actor page and click Start — the
itemsfield is already pre-filled with a working example. Results land in the run's dataset as soon as each item is found. - Via the API. Call it directly with a POST request — no Console needed once you have an API token:
curl "https://api.apify.com/v2/acts/accountable_eel~jsonld-structured-data-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \-X POST \-H "Content-Type: application/json" \-d '{"items":["https://www.nytimes.com"]}'
- On a schedule. Save this actor as an Apify Task with the input you want, then add a Schedule (hourly, daily, weekly) so it runs on its own — no server of your own required.
Input
{"items": ["https://www.nytimes.com"]}
One URL per line. A scheme is added automatically if missing (defaults to https://). Accepted formats: https://example.com/product/123, example.com/blog/post.
{"items": ["https://example.com/product/123"],"columns": ["metaTitle", "metaDescription", "og", "issues", "objectCount", "types"],"maxConcurrency": 5,"proxyConfiguration": { "useApifyProxy": true }}
items is the only required field — a list of page URLs (a scheme is added
automatically if missing, defaulting to https://). One dataset row is returned per
URL; rows with "found": false are never charged. columns picks which fields come
back (all are included by default — leave it out to get everything). maxConcurrency
(default 5, up to 20) controls how many requests run in parallel. proxyConfiguration
routes requests through Apify Proxy.
Output
| query | found | status | url | metaTitle | titleLength | metaDescription | descriptionLength | canonical | robotsMeta | hreflang | og | favicon | h1Count | firstH1 | issues | objectCount | types | objects | truncated | scrapedAt | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| https://www.nytimes.com | true | OK | <title length (characters)> | <meta description length (characters)> | <json-ld @types found> | <json-ld objects (raw, up to 25)> | <more than 25 json-ld objects were found> | 2026-09-23T06:10:13.898Z |
{"query": "https://example.com/product/123","found": true,"status": "OK","url": "https://example.com/product/123","metaTitle": "Wireless Headphones | Example Store","titleLength": 32,"metaDescription": "Premium wireless headphones with 30-hour battery life.","descriptionLength": 56,"canonical": "https://example.com/product/123","robotsMeta": "index, follow","hreflang": [{ "hreflang": "en", "href": "https://example.com/product/123" }],"og": {"title": "Wireless Headphones","description": "Premium wireless headphones with 30-hour battery life.","image": "https://example.com/images/headphones-og.jpg","type": "product","url": "https://example.com/product/123","siteName": "Example Store"},"twitter": { "card": "summary_large_image", "title": "Wireless Headphones", "description": "Premium wireless headphones with 30-hour battery life.", "image": "https://example.com/images/headphones-og.jpg" },"favicon": "https://example.com/favicon.ico","h1Count": 1,"firstH1": "Wireless Headphones","issues": [],"objectCount": 2,"types": ["Product", "BreadcrumbList"],"objects": [{"@context": "https://schema.org","@type": "Product","name": "Wireless Headphones","sku": "WH-1000","offers": { "@type": "Offer", "price": "199.00", "priceCurrency": "USD" }},{"@context": "https://schema.org","@type": "BreadcrumbList","itemListElement": [ "..." ]}],"truncated": false,"scrapedAt": "2026-08-20T12:00:00.000Z"}
Every field is null (or []/0 for a list/count) when the page doesn't publish it —
a missing og:image is exactly why issues would include "missing-og-image", not a
sign of a broken run. objects holds up to 25 raw JSON-LD objects found on the page;
truncated: true means more than 25 were present and the rest were cut off. A page
with none of the above — no title, no meta tags, no Open Graph, no JSON-LD — gets a
dataset row with "found": false, so you always get one row per input URL, but you're
never charged for it.
Use cases
- Meta tag & Open Graph audits — check title length, missing descriptions, and broken or missing social-preview tags across a whole site in one pass.
- Pre-launch QA — catch a canonical pointing at the wrong page, a missing
og:image, or duplicate H1s before a redesign goes live. - SEO audits — confirm your product, article, or job pages actually emit the structured data you think they do, across a whole site.
- Competitive SEO research — see exactly what schema types, meta tags, and Open Graph fields a competitor publishes for rich results and social previews.
- Content pipeline enrichment — pull structured
Product/Article/JobPostingdata, plus meta tags, out of pages without writing a custom parser per site. - Schema & metadata coverage monitoring — track structured-data and meta-tag presence across a URL list over time to catch regressions after a site redesign.
Pricing
$4 per 1,000 URLs, plus a $0.00005 start fee. Misses (found:false) are never charged.
Use it from Clay, n8n, Make, or an AI agent
This actor runs synchronously over plain HTTP — call it directly from a script, a workflow tool, or an AI agent, no Apify Console needed once you have an API token.
curl "https://api.apify.com/v2/acts/accountable_eel~jsonld-structured-data-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \-X POST \-H "Content-Type: application/json" \-d '{"items":["https://www.nytimes.com"]}'
n8n. Add an HTTP Request node: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~jsonld-structured-data-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body Content Type JSON, JSON Body {"items":["https://www.nytimes.com"]} (swap in an expression from an earlier node for a real value).
Clay. Add an "HTTP API" column: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~jsonld-structured-data-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body {"items":["{{URL}}"]}, mapping the row's URL into the items array.
MCP. In Claude, Cursor, or any MCP client with the Apify MCP server, ask for "JSON-LD Checker: Schema.org Structured Data for a URL List" — the agent will find and run this actor.
FAQ
What counts as "not found"?
A page with none of the following returns "found": false and is never charged: no
<title>, no meta description, no canonical, no Open Graph tags, no Twitter Card tags,
no <h1>, and no <script type="application/ld+json"> block (or every block on the
page fails to parse as JSON). Real content on a page with, say, a title and meta
description but zero JSON-LD is still a paid "found": true row — this actor checks
both, not just JSON-LD.
What's in the issues list?
Up to four codes, whichever apply: missing-description (no meta description),
title-too-long (the <title> tag is over 60 characters), missing-og-image (no
og:image), and canonical-mismatch (the page's <link rel="canonical"> resolves to
a different URL than the one you checked, ignoring protocol/trailing-slash/hash
differences). An empty list means none of those four were found.
Why are og and twitter sometimes null instead of an object with empty fields?
null means the page has no Open Graph (or Twitter Card) tags at all. If it has even
one, you get an object with every field, null for the ones the page doesn't set — so
you can tell "no social tags" from "some social tags."
Does this render JavaScript to find dynamically-injected tags or JSON-LD? No — this reads the raw HTML response as delivered by the server. If a site injects meta tags or JSON-LD client-side after page load via JavaScript, they won't be captured.
Not every page type carries rich JSON-LD — is that expected?
Yes. JSON-LD coverage depends entirely on what the site chooses to publish: product,
article, recipe, and job-posting pages tend to carry the richest markup, while general
informational or homepage-type URLs often have little or none — objectCount: 0 on
those pages is normal, and the row is still billed if the page has meta tags. A
found: false result on a page you expected to have data is worth spot-checking in
your browser's "View Page Source" first.
What happens if one JSON-LD block on the page is broken? It's skipped individually — a malformed block doesn't prevent extraction of the other valid blocks on the same page, or of the meta tags.
Are objects inside a top-level @graph included?
Yes — @graph arrays are flattened automatically, so every object inside one appears
in objects and contributes to types.
Can I check thousands of URLs at once?
Yes — maxConcurrency controls how many requests run in parallel (default 5, up to
20).