Meta Tags & Open Graph Checker + JSON-LD Extractor (Bulk URLs) avatar

Meta Tags & Open Graph Checker + JSON-LD Extractor (Bulk URLs)

Pricing

from $2.00 / 1,000 successful lookups

Go to Apify Store
Meta Tags & Open Graph Checker + JSON-LD Extractor (Bulk URLs)

Meta Tags & Open Graph Checker + JSON-LD Extractor (Bulk URLs)

A schema.org extractor for a list of URLs without an API key: pulls every JSON-LD object a page publishes, plus meta title, description, canonical and Open Graph tags, from the same HTML, no rendering. Where a structured data testing tool checks one page, this does your whole list, to Google Sheets.

Pricing

from $2.00 / 1,000 successful lookups

Rating

0.0

(0)

Developer

Adrian Voss

Adrian Voss

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

1

Monthly active users

5 days ago

Last modified

Share

Two audits in one fetch, for a whole list of URLs at once. First, a full meta tags and Open Graph checker: title, meta description, canonical, robots, hreflang alternates, Open Graph, Twitter Card, favicon, H1s, and a flagged list of common SEO issues (missing description, an oversized title, a missing og:image, a canonical mismatch). Second, the original JSON-LD extractor: every schema.org Product, Article, JobPosting, Organization, BreadcrumbList, FAQPage and more object already embedded in the page. No third-party API, no rendering — one HTML fetch per URL, everything above read straight out of it and handed back as clean, structured rows.

Features

  • Meta tag & Open Graph audit. Title, description, canonical, robots, hreflang, Open Graph (og:title/description/image/type/url/site_name) and Twitter Card (twitter:card/title/description/image) tags, plus favicon and H1 count — from the same fetch, no extra requests.
  • Built-in issue detection. Every row carries an issues list: missing-description, title-too-long (over 60 characters), missing-og-image, canonical-mismatch.
  • Full JSON-LD object extraction. Every JSON-LD block on the page is parsed and returned, including objects nested inside a top-level @graph.
  • Type summary. A deduplicated list of every @type found on the page (Product, Article, JobPosting, etc.), so you can see what's there without scanning the raw objects.
  • Resilient parsing. A malformed JSON-LD block is skipped on its own — one bad block doesn't fail the whole page.
  • Pick your columns. Every field above is its own output column, on by default — turn off what you don't need in the Console's "Which columns do you want?" input.
  • Bulk-friendly. Feed in a list of URLs with adjustable maxConcurrency.

How to use JSON-LD Structured Data Lookup — Schema.org Extractor API

  1. In the Apify Console. Open the actor page and click Start — the items field is already pre-filled with a working example. Results land in the run's dataset as soon as each item is found.
  2. Via the API. Call it directly with a POST request — no Console needed once you have an API token:
    curl "https://api.apify.com/v2/acts/accountable_eel~jsonld-structured-data-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
    -X POST \
    -H "Content-Type: application/json" \
    -d '{"items":["https://www.nytimes.com"]}'
  3. On a schedule. Save this actor as an Apify Task with the input you want, then add a Schedule (hourly, daily, weekly) so it runs on its own — no server of your own required.

Input

{
"items": [
"https://www.nytimes.com"
]
}

One URL per line. A scheme is added automatically if missing (defaults to https://). Accepted formats: https://example.com/product/123, example.com/blog/post.

{
"items": ["https://example.com/product/123"],
"columns": ["metaTitle", "metaDescription", "og", "issues", "objectCount", "types"],
"maxConcurrency": 5,
"proxyConfiguration": { "useApifyProxy": true }
}

items is the only required field — a list of page URLs (a scheme is added automatically if missing, defaulting to https://). One dataset row is returned per URL; rows with "found": false are never charged. columns picks which fields come back (all are included by default — leave it out to get everything). maxConcurrency (default 5, up to 20) controls how many requests run in parallel. proxyConfiguration routes requests through Apify Proxy.

Output

queryfoundstatusurlmetaTitletitleLengthmetaDescriptiondescriptionLengthcanonicalrobotsMetahreflangogtwitterfaviconh1CountfirstH1issuesobjectCounttypesobjectstruncatedscrapedAt
https://www.nytimes.comtrueOK <title length (characters)><meta description length (characters)>

<json-ld @types found><json-ld objects (raw, up to 25)><more than 25 json-ld objects were found>2026-09-23T06:10:13.898Z
{
"query": "https://example.com/product/123",
"found": true,
"status": "OK",
"url": "https://example.com/product/123",
"metaTitle": "Wireless Headphones | Example Store",
"titleLength": 32,
"metaDescription": "Premium wireless headphones with 30-hour battery life.",
"descriptionLength": 56,
"canonical": "https://example.com/product/123",
"robotsMeta": "index, follow",
"hreflang": [{ "hreflang": "en", "href": "https://example.com/product/123" }],
"og": {
"title": "Wireless Headphones",
"description": "Premium wireless headphones with 30-hour battery life.",
"image": "https://example.com/images/headphones-og.jpg",
"type": "product",
"url": "https://example.com/product/123",
"siteName": "Example Store"
},
"twitter": { "card": "summary_large_image", "title": "Wireless Headphones", "description": "Premium wireless headphones with 30-hour battery life.", "image": "https://example.com/images/headphones-og.jpg" },
"favicon": "https://example.com/favicon.ico",
"h1Count": 1,
"firstH1": "Wireless Headphones",
"issues": [],
"objectCount": 2,
"types": ["Product", "BreadcrumbList"],
"objects": [
{
"@context": "https://schema.org",
"@type": "Product",
"name": "Wireless Headphones",
"sku": "WH-1000",
"offers": { "@type": "Offer", "price": "199.00", "priceCurrency": "USD" }
},
{
"@context": "https://schema.org",
"@type": "BreadcrumbList",
"itemListElement": [ "..." ]
}
],
"truncated": false,
"scrapedAt": "2026-08-20T12:00:00.000Z"
}

Every field is null (or []/0 for a list/count) when the page doesn't publish it — a missing og:image is exactly why issues would include "missing-og-image", not a sign of a broken run. objects holds up to 25 raw JSON-LD objects found on the page; truncated: true means more than 25 were present and the rest were cut off. A page with none of the above — no title, no meta tags, no Open Graph, no JSON-LD — gets a dataset row with "found": false, so you always get one row per input URL, but you're never charged for it.

Use cases

  • Meta tag & Open Graph audits — check title length, missing descriptions, and broken or missing social-preview tags across a whole site in one pass.
  • Pre-launch QA — catch a canonical pointing at the wrong page, a missing og:image, or duplicate H1s before a redesign goes live.
  • SEO audits — confirm your product, article, or job pages actually emit the structured data you think they do, across a whole site.
  • Competitive SEO research — see exactly what schema types, meta tags, and Open Graph fields a competitor publishes for rich results and social previews.
  • Content pipeline enrichment — pull structured Product/Article/JobPosting data, plus meta tags, out of pages without writing a custom parser per site.
  • Schema & metadata coverage monitoring — track structured-data and meta-tag presence across a URL list over time to catch regressions after a site redesign.

Pricing

$4 per 1,000 URLs, plus a $0.00005 start fee. Misses (found:false) are never charged.

Use it from Clay, n8n, Make, or an AI agent

This actor runs synchronously over plain HTTP — call it directly from a script, a workflow tool, or an AI agent, no Apify Console needed once you have an API token.

curl "https://api.apify.com/v2/acts/accountable_eel~jsonld-structured-data-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
-X POST \
-H "Content-Type: application/json" \
-d '{"items":["https://www.nytimes.com"]}'

n8n. Add an HTTP Request node: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~jsonld-structured-data-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body Content Type JSON, JSON Body {"items":["https://www.nytimes.com"]} (swap in an expression from an earlier node for a real value).

Clay. Add an "HTTP API" column: Method POST, URL https://api.apify.com/v2/acts/accountable_eel~jsonld-structured-data-lookup/run-sync-get-dataset-items?token=<YOUR_TOKEN>, Body {"items":["{{URL}}"]}, mapping the row's URL into the items array.

MCP. In Claude, Cursor, or any MCP client with the Apify MCP server, ask for "JSON-LD Checker: Schema.org Structured Data for a URL List" — the agent will find and run this actor.

FAQ

What counts as "not found"? A page with none of the following returns "found": false and is never charged: no <title>, no meta description, no canonical, no Open Graph tags, no Twitter Card tags, no <h1>, and no <script type="application/ld+json"> block (or every block on the page fails to parse as JSON). Real content on a page with, say, a title and meta description but zero JSON-LD is still a paid "found": true row — this actor checks both, not just JSON-LD.

What's in the issues list? Up to four codes, whichever apply: missing-description (no meta description), title-too-long (the <title> tag is over 60 characters), missing-og-image (no og:image), and canonical-mismatch (the page's <link rel="canonical"> resolves to a different URL than the one you checked, ignoring protocol/trailing-slash/hash differences). An empty list means none of those four were found.

Why are og and twitter sometimes null instead of an object with empty fields? null means the page has no Open Graph (or Twitter Card) tags at all. If it has even one, you get an object with every field, null for the ones the page doesn't set — so you can tell "no social tags" from "some social tags."

Does this render JavaScript to find dynamically-injected tags or JSON-LD? No — this reads the raw HTML response as delivered by the server. If a site injects meta tags or JSON-LD client-side after page load via JavaScript, they won't be captured.

Not every page type carries rich JSON-LD — is that expected? Yes. JSON-LD coverage depends entirely on what the site chooses to publish: product, article, recipe, and job-posting pages tend to carry the richest markup, while general informational or homepage-type URLs often have little or none — objectCount: 0 on those pages is normal, and the row is still billed if the page has meta tags. A found: false result on a page you expected to have data is worth spot-checking in your browser's "View Page Source" first.

What happens if one JSON-LD block on the page is broken? It's skipped individually — a malformed block doesn't prevent extraction of the other valid blocks on the same page, or of the meta tags.

Are objects inside a top-level @graph included? Yes — @graph arrays are flattened automatically, so every object inside one appears in objects and contributes to types.

Can I check thousands of URLs at once? Yes — maxConcurrency controls how many requests run in parallel (default 5, up to 20).