Page Metadata Extractor - SEO & RAG JSON Feed
Pricing
from $0.50 / 1,000 extracted results
Page Metadata Extractor - SEO & RAG JSON Feed
Extract structured JSON metadata (title, meta description, canonical URL, Open Graph tags, H1, word count) from any site. Built-in change detection surfaces only what changed; automatic retries and session rotation need no upkeep. Pay only per successfully extracted result.
Pricing
from $0.50 / 1,000 extracted results
Rating
0.0
(0)
Developer
Stefano Seggio
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
11 days ago
Last modified
Categories
Share
Page Metadata Extractor — SEO & RAG JSON Feed
Every SEO audit and RAG pipeline needs the same five fields. Most crawlers make you pay for the whole page to get them.
You don't need the full HTML of a page — you need its title, meta description, canonical URL, Open Graph tags, H1, and word count, structured and consistent across thousands of URLs. This Actor gives you exactly that: a clean metadata feed built for SEO audits and LLM/RAG ingestion pipelines, with built-in retries and session rotation so a flaky page doesn't kill your crawl.
Why this outperforms a standard scraper
- Pay per result, never per wasted run. A crawl that fails to extract usable metadata from a page doesn't charge you for it — you pay $0.0005 per successfully extracted result, full stop.
- Change detection built in. Enable
onlyChangedand a page is still crawled and its links still followed, but it only lands in your dataset — and only gets billed — when its metadata actually differs from last time. - Resilience as a default, not an upgrade. Session rotation and automatic retries are on out of the box, so one broken page or rate-limited request doesn't take down a 500-URL crawl.
See it before you trust it
{"url": "https://example.com/blog/post","title": "Example Post Title","metaDescription": "A clean, structured summary of the page.","canonicalUrl": "https://example.com/blog/post","ogTags": { "og:title": "Example Post Title", "og:type": "article" },"h1": "Example Post Title","wordCount": 1420}
Every field here is one your SEO or RAG pipeline can ingest directly — no HTML parsing on your end.
Zero-risk trial
Unchanged pages (with onlyChanged on) cost $0.00. Run it once against a real URL before you commit to anything:
curl -X POST "https://api.apify.com/v2/acts/U9fUBHDngX6IyjzzF/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \-H "Content-Type: application/json" \-d '{"startUrls": [{"url": "https://example.com"}]}'
import requestsresponse = requests.post("https://api.apify.com/v2/acts/U9fUBHDngX6IyjzzF/run-sync-get-dataset-items",params={"token": "<YOUR_API_TOKEN>"},json={"startUrls": [{"url": "https://example.com"}]},)records = response.json()print(f"{len(records)} records returned")
const response = await fetch("https://api.apify.com/v2/acts/U9fUBHDngX6IyjzzF/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>",{method: "POST",headers: { "Content-Type": "application/json" },body: JSON.stringify({ startUrls: [{ url: "https://example.com" }] }),});const records = await response.json();console.log(`${records.length} records returned`);
Pricing
| Event | What it means | Price |
|---|---|---|
| Extracted result | One page's metadata successfully extracted and delivered (or changed, in onlyChanged mode). | $0.0005 |
Actor-start fee: $0.00005/GB-memory (one-time per run, not per page).
What you get on every record
- Title, meta description, canonical URL
- Full Open Graph tag set
- H1 and total word count
- Optional pagination following via a CSS next-page selector
Input parameters
| Field | Type | Description | Default |
|---|---|---|---|
startUrls | array | URLs to start crawling from. Required. | — |
maxRequestsPerCrawl | integer | Hard limit on pages fetched per run, across start URLs and discovered links. | 100 |
paginationSelector | string | CSS selector for a next-page link (e.g. a[rel=next]) to follow pagination automatically. | none |
onlyChanged | boolean | When enabled, a page is still crawled but only delivered — and billed — if its metadata changed. | false |
Source & reliability
Works against any publicly reachable site — no proprietary source to depend on. Runs on Apify's proxy infrastructure with automatic retries and session rotation, so transient failures on individual pages don't abort the whole crawl.
Output & Pricing
Real JSON output schema
Every dataset item follows this shape (as shown in the Actor's own example):
{"url": "https://example.com/blog/post","title": "Example Post Title","metaDescription": "A clean, structured summary of the page.","canonicalUrl": "https://example.com/blog/post","ogTags": { "og:title": "Example Post Title", "og:type": "article" },"h1": "Example Post Title","wordCount": 1420}
No HTML parsing required on your end — every field is ready to load directly into your SEO or RAG ingestion pipeline.
Use cases
- SEO audits at scale: crawl a full site (with optional pagination following via a CSS next-page selector) to inventory titles, meta descriptions, canonical URLs and H1s across thousands of pages in one consistent JSON feed.
- Change-aware monitoring for RAG/LLM pipelines: enable
onlyChangedto re-crawl a known set of pages on a schedule and have only pages whose metadata actually changed land in (and be billed to) your dataset — keeping a RAG index or content-monitoring feed fresh without re-ingesting unchanged pages.
PAY_PER_EVENT pricing, explained
This Actor bills per event, not per run duration:
| Event | Charged when | Price |
|---|---|---|
| Actor Start | Actor starts running; one event per GB of memory allocated (minimum one event), charged once per run | $0.00005 / event |
| Extracted result (primary event) | One page's metadata is successfully extracted and pushed to the dataset (or, in onlyChanged mode, only when it differs from the prior run) | $0.0005 / event |
A page that fails to yield usable metadata is not charged as a result — you only pay for data you actually receive. In onlyChanged mode, unchanged pages cost $0.00 per the Actor's own documentation.
Native Alerting & Monitoring, Ready to Wire Up
You don't need to poll this Actor's dataset or write custom code to find out when new results land — it runs on the Apify platform, so its runs and dataset connect directly to Apify's own automation tools:
- Apify Webhooks — add a webhook on this Actor (or on a saved Task/Schedule) for the
ACTOR.RUN.SUCCEEDEDevent, and it fires automatically every time a run finishes, so downstream systems can react the moment new data is ready. Configure it under the Actor's Integrations tab in the Apify Console. - apify/slack-alert — chain this Actor's webhook to
apify/slack-alertto post a Slack message whenever a run succeeds or fails, so your team sees new results (or a break in monitoring) without checking the Console. - apify/send-email — use the same trigger with
apify/send-emailto email a run summary or dataset link to stakeholders who don't use Slack. - Google Sheets export — export this Actor's dataset to Google Sheets directly from the Console (Storage → Export) or via Apify's Google Sheets integration, giving non-technical teams a live, spreadsheet-native view with no code required.
These are standard Apify platform capabilities available to any Actor's runs and dataset — not custom code built into this Actor — but they work with it out of the box, with no development required.
Deprecation notice (2026-09-22)
This Actor is deprecated and will not receive further feature development. It remains live and billable for existing users; no new capabilities or source coverage will be added. See the developer's other Actors at apify.com/stefano_seggio for actively maintained procurement, compliance and regulatory-monitoring feeds.