SEO Metadata Extractor — OG, JSON-LD, Twitter
Pricing
from $0.50 / 1,000 url metadata extracteds
SEO Metadata Extractor — OG, JSON-LD, Twitter
Bulk extract Open Graph, Twitter Cards, JSON-LD, microdata, RDFa and basic meta tags from a URL list. Pure structured JSON — not an SEO audit score. Chain from sitemap / URL status. Default 256 MB. No browser, no AI keys.
Pricing
from $0.50 / 1,000 url metadata extracteds
Rating
0.0
(0)
Developer
新世紀書僮
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
13 hours ago
Last modified
Categories
Share
Bulk-extract Open Graph, Twitter Cards, JSON-LD, microdata and basic meta tags into clean structured JSON. Paste URLs or chain from a Sitemap / URL Status dataset. Pure extraction — not an SEO audit score (0–100). Default memory: 256 MB. No browser, no AI keys.
What you get
- 🏷️ Open Graph —
og:title,og:description,og:image,og:type, … - 🐦 Twitter Cards —
twitter:card,twitter:title,twitter:image, … - 📦 JSON-LD — Schema.org and other
<script type="application/ld+json">blocks - 🧩 Microdata (optional RDFa) — HTML embedded structured data
- 📄 Basic meta —
<title>, description, keywords, robots, canonical, hreflang - 📊 SUMMARY report — counts by status and by format found
- 🔌 Chain-friendly —
urlsaccepts{"url":…}objects; or pass another Actor'sdatasetId/DOC_TO_MARKDOWN_INPUTKV record - 💾 Light — HTTP +
extruct(BSD-3-Clause); 256 MB default
Measured results
Local smoke measured 2026-09-30 Asia/Taipei. Cloud benches and PPE lock pending (draft price only).
| Test | Result |
|---|---|
| Local smoke: example.com + w3.org + HTTP 404 + invalid DNS | 4/4 processed in 0.2 s, peak 90 MB; ok 2 / error 2; charged 2 free 2; formats meta×2, openGraph×1, twitter×1 |
| Local smoke: schema.org + python.org | 2/2 ok in 0.1 s, peak 90 MB; jsonLd×2, openGraph×1, meta×2 |
| Cloud smoke | pending private push |
Use cases
- RAG / knowledge-base enrichment — attach OG title/description/image to crawled URLs
- Schema inventory — list JSON-LD types present across a sitemap
- Social preview QA — check OG / Twitter fields without an audit scorecard
- Post-status hygiene — run after URL Status Checker on
okURLs only
How to use
- Add URLs in URLs to extract, and/or a Source dataset ID / key-value store from another run.
- Optional: toggle formats, set Max URLs, concurrency and timeout.
- Click Start. Results appear in the Dataset;
SUMMARYandOUTPUTare in the Key-value store.
Input example
{"urls": [{ "url": "https://example.com/" },{ "url": "https://www.w3.org/" }],"maxUrls": 100,"extractOpenGraph": true,"extractTwitter": true,"extractJsonLd": true,"extractMicrodata": true,"extractBasicMeta": true}
Chain from a Sitemap Actor dataset:
{"datasetId": "<sitemap-run-default-dataset-id>","maxUrls": 200}
Output example (one dataset item)
{"url": "https://example.com/","finalUrl": "https://example.com/","httpStatus": 200,"ok": true,"status": "ok","title": "Example Domain","description": null,"canonical": null,"robots": null,"openGraph": null,"twitter": null,"jsonLd": null,"microdata": null,"meta": { "title": "Example Domain" },"extracted": ["meta"],"errorClass": null,"error": null,"durationMs": 120,"host": "example.com"}
Key-value store records
| Key | Content |
|---|---|
SUMMARY | Counts: totalProcessed, charged, free, byStatus, byFormat, byErrorClass, duration, peak memory |
OUTPUT | Same summary (kept for consistency with sibling Actors) |
Pricing
Pay per event (DRAFT stub — not locked):
| Event | Price |
|---|---|
| URL metadata extracted (primary) | $0.0005 per URL (draft) (= $0.50 per 1,000) |
| Actor start (Apify synthetic) | $0.00005 per GB (platform default) |
Worked example (draft): 1,000 successful URLs ≈ $0.50 + one start event.
By default only ok / partial rows are charged. Fetch/parse failures (status: error) are free unless chargeFailedUrls is enabled. Invalid inputs that never become a row are not charged. Details: docs/PRICING.md.
Chaining: Sitemap → metadata
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("ingenious_quip_bxq/sitemap-url-discovery").call(run_input={"startUrls": [{"url": "https://www.example.com"}],"maxUrls": 50,})run2 = client.actor("ingenious_quip_bxq/seo-metadata-extract").call(run_input={"datasetId": run["defaultDatasetId"],"maxUrls": 50,})for item in client.dataset(run2["defaultDatasetId"]).iterate_items():print(item["url"], item.get("title"), item.get("extracted"))
Known limits
- Does not render JavaScript — SPA-only meta injected client-side will be missing.
- Does not compute an SEO score, ranking grade, or “issues” checklist (by design; sell pure JSON).
- Does not validate Schema.org against Google’s rich-result rules.
- Some hosts rate-limit or block data-center IPs — use the proxy input if needed.
- HTML is truncated at
maxHtmlBytes(default 2 MB); metadata is almost always in the head. - RDFa is off by default (enable
extractRdfaif you need it).
FAQ
Why not an SEO audit score? Competitors already sell 0–100 audits. This Actor returns the raw structured fields so you can score, filter or store them yourself.
Empty openGraph / jsonLd? Many pages have no OG or JSON-LD. That is still a successful extract (status: ok, extracted may only contain meta).
Failures? DNS / timeout / SSL / HTTP 4xx–5xx rows are saved with status: error and are free by default.
License & source code
This Actor is open source under the GNU Affero General Public License v3.0 (AGPL-3.0) — see LICENSE. Public GitHub mirror is planned after private validation (not published yet).
Libraries: extruct (BSD-3-Clause).
See CHANGELOG.md for version history.