SEO Metadata Extractor — OG, JSON-LD, Twitter avatar

SEO Metadata Extractor — OG, JSON-LD, Twitter

Pricing

from $0.50 / 1,000 url metadata extracteds

Go to Apify Store
SEO Metadata Extractor — OG, JSON-LD, Twitter

SEO Metadata Extractor — OG, JSON-LD, Twitter

Bulk extract Open Graph, Twitter Cards, JSON-LD, microdata, RDFa and basic meta tags from a URL list. Pure structured JSON — not an SEO audit score. Chain from sitemap / URL status. Default 256 MB. No browser, no AI keys.

Pricing

from $0.50 / 1,000 url metadata extracteds

Rating

0.0

(0)

Developer

新世紀書僮

新世紀書僮

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

13 hours ago

Last modified

Share

Bulk-extract Open Graph, Twitter Cards, JSON-LD, microdata and basic meta tags into clean structured JSON. Paste URLs or chain from a Sitemap / URL Status dataset. Pure extraction — not an SEO audit score (0–100). Default memory: 256 MB. No browser, no AI keys.

What you get

  • 🏷️ Open Graph — og:title, og:description, og:image, og:type, …
  • 🐦 Twitter Cards — twitter:card, twitter:title, twitter:image, …
  • 📦 JSON-LD — Schema.org and other <script type="application/ld+json"> blocks
  • 🧩 Microdata (optional RDFa) — HTML embedded structured data
  • 📄 Basic meta — <title>, description, keywords, robots, canonical, hreflang
  • 📊 SUMMARY report — counts by status and by format found
  • 🔌 Chain-friendly — urls accepts {"url":…} objects; or pass another Actor's datasetId / DOC_TO_MARKDOWN_INPUT KV record
  • 💾 Light — HTTP + extruct (BSD-3-Clause); 256 MB default

Measured results

Local smoke measured 2026-09-30 Asia/Taipei. Cloud benches and PPE lock pending (draft price only).

TestResult
Local smoke: example.com + w3.org + HTTP 404 + invalid DNS4/4 processed in 0.2 s, peak 90 MB; ok 2 / error 2; charged 2 free 2; formats meta×2, openGraph×1, twitter×1
Local smoke: schema.org + python.org2/2 ok in 0.1 s, peak 90 MB; jsonLd×2, openGraph×1, meta×2
Cloud smokepending private push

Use cases

  • RAG / knowledge-base enrichment — attach OG title/description/image to crawled URLs
  • Schema inventory — list JSON-LD types present across a sitemap
  • Social preview QA — check OG / Twitter fields without an audit scorecard
  • Post-status hygiene — run after URL Status Checker on ok URLs only

How to use

  1. Add URLs in URLs to extract, and/or a Source dataset ID / key-value store from another run.
  2. Optional: toggle formats, set Max URLs, concurrency and timeout.
  3. Click Start. Results appear in the Dataset; SUMMARY and OUTPUT are in the Key-value store.

Input example

{
"urls": [
{ "url": "https://example.com/" },
{ "url": "https://www.w3.org/" }
],
"maxUrls": 100,
"extractOpenGraph": true,
"extractTwitter": true,
"extractJsonLd": true,
"extractMicrodata": true,
"extractBasicMeta": true
}

Chain from a Sitemap Actor dataset:

{
"datasetId": "<sitemap-run-default-dataset-id>",
"maxUrls": 200
}

Output example (one dataset item)

{
"url": "https://example.com/",
"finalUrl": "https://example.com/",
"httpStatus": 200,
"ok": true,
"status": "ok",
"title": "Example Domain",
"description": null,
"canonical": null,
"robots": null,
"openGraph": null,
"twitter": null,
"jsonLd": null,
"microdata": null,
"meta": { "title": "Example Domain" },
"extracted": ["meta"],
"errorClass": null,
"error": null,
"durationMs": 120,
"host": "example.com"
}

Key-value store records

KeyContent
SUMMARYCounts: totalProcessed, charged, free, byStatus, byFormat, byErrorClass, duration, peak memory
OUTPUTSame summary (kept for consistency with sibling Actors)

Pricing

Pay per event (DRAFT stub — not locked):

EventPrice
URL metadata extracted (primary)$0.0005 per URL (draft) (= $0.50 per 1,000)
Actor start (Apify synthetic)$0.00005 per GB (platform default)

Worked example (draft): 1,000 successful URLs ≈ $0.50 + one start event.

By default only ok / partial rows are charged. Fetch/parse failures (status: error) are free unless chargeFailedUrls is enabled. Invalid inputs that never become a row are not charged. Details: docs/PRICING.md.

Chaining: Sitemap → metadata

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("ingenious_quip_bxq/sitemap-url-discovery").call(run_input={
"startUrls": [{"url": "https://www.example.com"}],
"maxUrls": 50,
})
run2 = client.actor("ingenious_quip_bxq/seo-metadata-extract").call(run_input={
"datasetId": run["defaultDatasetId"],
"maxUrls": 50,
})
for item in client.dataset(run2["defaultDatasetId"]).iterate_items():
print(item["url"], item.get("title"), item.get("extracted"))

Known limits

  • Does not render JavaScript — SPA-only meta injected client-side will be missing.
  • Does not compute an SEO score, ranking grade, or “issues” checklist (by design; sell pure JSON).
  • Does not validate Schema.org against Google’s rich-result rules.
  • Some hosts rate-limit or block data-center IPs — use the proxy input if needed.
  • HTML is truncated at maxHtmlBytes (default 2 MB); metadata is almost always in the head.
  • RDFa is off by default (enable extractRdfa if you need it).

FAQ

Why not an SEO audit score? Competitors already sell 0–100 audits. This Actor returns the raw structured fields so you can score, filter or store them yourself.

Empty openGraph / jsonLd? Many pages have no OG or JSON-LD. That is still a successful extract (status: ok, extracted may only contain meta).

Failures? DNS / timeout / SSL / HTTP 4xx–5xx rows are saved with status: error and are free by default.

License & source code

This Actor is open source under the GNU Affero General Public License v3.0 (AGPL-3.0) — see LICENSE. Public GitHub mirror is planned after private validation (not published yet).

Libraries: extruct (BSD-3-Clause).

See CHANGELOG.md for version history.