Page Metadata Extractor - SEO & RAG JSON Feed avatar

Page Metadata Extractor - SEO & RAG JSON Feed

Pricing

from $0.50 / 1,000 extracted results

Go to Apify Store
Page Metadata Extractor - SEO & RAG JSON Feed

Page Metadata Extractor - SEO & RAG JSON Feed

Extract structured JSON metadata (title, meta description, canonical URL, Open Graph tags, H1, word count) from any site. Built-in change detection surfaces only what changed; automatic retries and session rotation need no upkeep. Pay only per successfully extracted result.

Pricing

from $0.50 / 1,000 extracted results

Rating

0.0

(0)

Developer

Stefano Seggio

Stefano Seggio

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

11 days ago

Last modified

Categories

Share

Page Metadata Extractor — SEO & RAG JSON Feed

Every SEO audit and RAG pipeline needs the same five fields. Most crawlers make you pay for the whole page to get them.

You don't need the full HTML of a page — you need its title, meta description, canonical URL, Open Graph tags, H1, and word count, structured and consistent across thousands of URLs. This Actor gives you exactly that: a clean metadata feed built for SEO audits and LLM/RAG ingestion pipelines, with built-in retries and session rotation so a flaky page doesn't kill your crawl.


Why this outperforms a standard scraper

  • Pay per result, never per wasted run. A crawl that fails to extract usable metadata from a page doesn't charge you for it — you pay $0.0005 per successfully extracted result, full stop.
  • Change detection built in. Enable onlyChanged and a page is still crawled and its links still followed, but it only lands in your dataset — and only gets billed — when its metadata actually differs from last time.
  • Resilience as a default, not an upgrade. Session rotation and automatic retries are on out of the box, so one broken page or rate-limited request doesn't take down a 500-URL crawl.

See it before you trust it

{
"url": "https://example.com/blog/post",
"title": "Example Post Title",
"metaDescription": "A clean, structured summary of the page.",
"canonicalUrl": "https://example.com/blog/post",
"ogTags": { "og:title": "Example Post Title", "og:type": "article" },
"h1": "Example Post Title",
"wordCount": 1420
}

Every field here is one your SEO or RAG pipeline can ingest directly — no HTML parsing on your end.

Zero-risk trial

Unchanged pages (with onlyChanged on) cost $0.00. Run it once against a real URL before you commit to anything:

curl -X POST "https://api.apify.com/v2/acts/U9fUBHDngX6IyjzzF/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \
-H "Content-Type: application/json" \
-d '{"startUrls": [{"url": "https://example.com"}]}'
import requests
response = requests.post(
"https://api.apify.com/v2/acts/U9fUBHDngX6IyjzzF/run-sync-get-dataset-items",
params={"token": "<YOUR_API_TOKEN>"},
json={"startUrls": [{"url": "https://example.com"}]},
)
records = response.json()
print(f"{len(records)} records returned")
const response = await fetch(
"https://api.apify.com/v2/acts/U9fUBHDngX6IyjzzF/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>",
{
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ startUrls: [{ url: "https://example.com" }] }),
}
);
const records = await response.json();
console.log(`${records.length} records returned`);

Pricing

EventWhat it meansPrice
Extracted resultOne page's metadata successfully extracted and delivered (or changed, in onlyChanged mode).$0.0005

Actor-start fee: $0.00005/GB-memory (one-time per run, not per page).

What you get on every record

  • Title, meta description, canonical URL
  • Full Open Graph tag set
  • H1 and total word count
  • Optional pagination following via a CSS next-page selector

Input parameters

FieldTypeDescriptionDefault
startUrlsarrayURLs to start crawling from. Required.—
maxRequestsPerCrawlintegerHard limit on pages fetched per run, across start URLs and discovered links.100
paginationSelectorstringCSS selector for a next-page link (e.g. a[rel=next]) to follow pagination automatically.none
onlyChangedbooleanWhen enabled, a page is still crawled but only delivered — and billed — if its metadata changed.false

Source & reliability

Works against any publicly reachable site — no proprietary source to depend on. Runs on Apify's proxy infrastructure with automatic retries and session rotation, so transient failures on individual pages don't abort the whole crawl.

Output & Pricing

Real JSON output schema

Every dataset item follows this shape (as shown in the Actor's own example):

{
"url": "https://example.com/blog/post",
"title": "Example Post Title",
"metaDescription": "A clean, structured summary of the page.",
"canonicalUrl": "https://example.com/blog/post",
"ogTags": { "og:title": "Example Post Title", "og:type": "article" },
"h1": "Example Post Title",
"wordCount": 1420
}

No HTML parsing required on your end — every field is ready to load directly into your SEO or RAG ingestion pipeline.

Use cases

  • SEO audits at scale: crawl a full site (with optional pagination following via a CSS next-page selector) to inventory titles, meta descriptions, canonical URLs and H1s across thousands of pages in one consistent JSON feed.
  • Change-aware monitoring for RAG/LLM pipelines: enable onlyChanged to re-crawl a known set of pages on a schedule and have only pages whose metadata actually changed land in (and be billed to) your dataset — keeping a RAG index or content-monitoring feed fresh without re-ingesting unchanged pages.

PAY_PER_EVENT pricing, explained

This Actor bills per event, not per run duration:

EventCharged whenPrice
Actor StartActor starts running; one event per GB of memory allocated (minimum one event), charged once per run$0.00005 / event
Extracted result (primary event)One page's metadata is successfully extracted and pushed to the dataset (or, in onlyChanged mode, only when it differs from the prior run)$0.0005 / event

A page that fails to yield usable metadata is not charged as a result — you only pay for data you actually receive. In onlyChanged mode, unchanged pages cost $0.00 per the Actor's own documentation.

Native Alerting & Monitoring, Ready to Wire Up

You don't need to poll this Actor's dataset or write custom code to find out when new results land — it runs on the Apify platform, so its runs and dataset connect directly to Apify's own automation tools:

  • Apify Webhooks — add a webhook on this Actor (or on a saved Task/Schedule) for the ACTOR.RUN.SUCCEEDED event, and it fires automatically every time a run finishes, so downstream systems can react the moment new data is ready. Configure it under the Actor's Integrations tab in the Apify Console.
  • apify/slack-alert — chain this Actor's webhook to apify/slack-alert to post a Slack message whenever a run succeeds or fails, so your team sees new results (or a break in monitoring) without checking the Console.
  • apify/send-email — use the same trigger with apify/send-email to email a run summary or dataset link to stakeholders who don't use Slack.
  • Google Sheets export — export this Actor's dataset to Google Sheets directly from the Console (Storage → Export) or via Apify's Google Sheets integration, giving non-technical teams a live, spreadsheet-native view with no code required.

These are standard Apify platform capabilities available to any Actor's runs and dataset — not custom code built into this Actor — but they work with it out of the box, with no development required.

Deprecation notice (2026-09-22)

This Actor is deprecated and will not receive further feature development. It remains live and billable for existing users; no new capabilities or source coverage will be added. See the developer's other Actors at apify.com/stefano_seggio for actively maintained procurement, compliance and regulatory-monitoring feeds.