# SEO Metadata Extractor — OG, JSON-LD, Twitter (`ingenious_quip_bxq/seo-metadata-extract`) Actor

Bulk extract Open Graph, Twitter Cards, JSON-LD, microdata, RDFa and basic meta tags from a URL list. Pure structured JSON — not an SEO audit score. Chain from sitemap / URL status. Default 256 MB. No browser, no AI keys.

- **URL**: https://apify.com/ingenious\_quip\_bxq/seo-metadata-extract.md
- **Developed by:** [新世紀書僮](https://apify.com/ingenious_quip_bxq) (community)
- **Categories:** SEO tools, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 url metadata extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## SEO Metadata Extractor — OG, JSON-LD, Twitter

**Bulk-extract Open Graph, Twitter Cards, JSON-LD, microdata and basic meta tags into clean structured JSON.**
Paste URLs or chain from a Sitemap / URL Status dataset. Pure extraction — **not** an SEO audit score (0–100). Default memory: 256 MB. No browser, no AI keys.

### What you get

- 🏷️ **Open Graph** — `og:title`, `og:description`, `og:image`, `og:type`, …
- 🐦 **Twitter Cards** — `twitter:card`, `twitter:title`, `twitter:image`, …
- 📦 **JSON-LD** — Schema.org and other `<script type="application/ld+json">` blocks
- 🧩 **Microdata** (optional RDFa) — HTML embedded structured data
- 📄 **Basic meta** — `<title>`, description, keywords, robots, canonical, hreflang
- 📊 **SUMMARY report** — counts by status and by format found
- 🔌 **Chain-friendly** — `urls` accepts `{"url":…}` objects; or pass another Actor's `datasetId` / `DOC_TO_MARKDOWN_INPUT` KV record
- 💾 **Light** — HTTP + `extruct` (BSD-3-Clause); 256 MB default

### Measured results

Local smoke measured **2026-09-30 Asia/Taipei**. Cloud benches and PPE lock **pending** (draft price only).

| Test | Result |
|---|---|
| Local smoke: example.com + w3.org + HTTP 404 + invalid DNS | **4/4 processed in 0.2 s**, peak **90 MB**; ok 2 / error 2; charged 2 free 2; formats meta×2, openGraph×1, twitter×1 |
| Local smoke: schema.org + python.org | **2/2 ok in 0.1 s**, peak **90 MB**; jsonLd×2, openGraph×1, meta×2 |
| Cloud smoke | *pending private push* |

### Use cases

- **RAG / knowledge-base enrichment** — attach OG title/description/image to crawled URLs
- **Schema inventory** — list JSON-LD types present across a sitemap
- **Social preview QA** — check OG / Twitter fields without an audit scorecard
- **Post-status hygiene** — run after URL Status Checker on `ok` URLs only

### How to use

1. Add URLs in **URLs to extract**, and/or a **Source dataset ID** / **key-value store** from another run.
2. Optional: toggle formats, set **Max URLs**, concurrency and timeout.
3. Click **Start**. Results appear in the **Dataset**; `SUMMARY` and `OUTPUT` are in the **Key-value store**.

#### Input example

```json
{
  "urls": [
    { "url": "https://example.com/" },
    { "url": "https://www.w3.org/" }
  ],
  "maxUrls": 100,
  "extractOpenGraph": true,
  "extractTwitter": true,
  "extractJsonLd": true,
  "extractMicrodata": true,
  "extractBasicMeta": true
}
```

Chain from a Sitemap Actor dataset:

```json
{
  "datasetId": "<sitemap-run-default-dataset-id>",
  "maxUrls": 200
}
```

#### Output example (one dataset item)

```json
{
  "url": "https://example.com/",
  "finalUrl": "https://example.com/",
  "httpStatus": 200,
  "ok": true,
  "status": "ok",
  "title": "Example Domain",
  "description": null,
  "canonical": null,
  "robots": null,
  "openGraph": null,
  "twitter": null,
  "jsonLd": null,
  "microdata": null,
  "meta": { "title": "Example Domain" },
  "extracted": ["meta"],
  "errorClass": null,
  "error": null,
  "durationMs": 120,
  "host": "example.com"
}
```

#### Key-value store records

| Key | Content |
|---|---|
| `SUMMARY` | Counts: `totalProcessed`, `charged`, `free`, `byStatus`, `byFormat`, `byErrorClass`, duration, peak memory |
| `OUTPUT` | Same summary (kept for consistency with sibling Actors) |

### Pricing

Pay per event (**DRAFT stub — not locked**):

| Event | Price |
|---|---|
| URL metadata extracted (primary) | **$0.0005** per URL (draft) (= $0.50 per 1,000) |
| Actor start (Apify synthetic) | $0.00005 per GB (platform default) |

**Worked example (draft):** 1,000 successful URLs ≈ **$0.50** + one start event.

By default only `ok` / `partial` rows are charged. Fetch/parse failures (`status: error`) are free unless `chargeFailedUrls` is enabled. Invalid inputs that never become a row are not charged. Details: `docs/PRICING.md`.

### Chaining: Sitemap → metadata

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")

run = client.actor("ingenious_quip_bxq/sitemap-url-discovery").call(run_input={
    "startUrls": [{"url": "https://www.example.com"}],
    "maxUrls": 50,
})

run2 = client.actor("ingenious_quip_bxq/seo-metadata-extract").call(run_input={
    "datasetId": run["defaultDatasetId"],
    "maxUrls": 50,
})
for item in client.dataset(run2["defaultDatasetId"]).iterate_items():
    print(item["url"], item.get("title"), item.get("extracted"))
```

### Known limits

- Does **not** render JavaScript — SPA-only meta injected client-side will be missing.
- Does **not** compute an SEO score, ranking grade, or “issues” checklist (by design; sell pure JSON).
- Does **not** validate Schema.org against Google’s rich-result rules.
- Some hosts rate-limit or block data-center IPs — use the proxy input if needed.
- HTML is truncated at `maxHtmlBytes` (default 2 MB); metadata is almost always in the head.
- RDFa is **off by default** (enable `extractRdfa` if you need it).

### FAQ

**Why not an SEO audit score?** Competitors already sell 0–100 audits. This Actor returns the raw structured fields so you can score, filter or store them yourself.

**Empty `openGraph` / `jsonLd`?** Many pages have no OG or JSON-LD. That is still a successful extract (`status: ok`, `extracted` may only contain `meta`).

**Failures?** DNS / timeout / SSL / HTTP 4xx–5xx rows are saved with `status: error` and are free by default.

### License & source code

This Actor is open source under the **GNU Affero General Public License v3.0 (AGPL-3.0)** — see `LICENSE`. Public GitHub mirror is planned after private validation (not published yet).

Libraries: [extruct](https://github.com/scrapinghub/extruct) (BSD-3-Clause).

See `CHANGELOG.md` for version history.

# Changelog

This Actor's version history is a separate document: https://apify.com/ingenious\_quip\_bxq/seo-metadata-extract/changelog.md

# Actor input Schema

## `urls` (type: `array`):

One or more URLs. Accepts `{"url": "..."}` objects (same shape as Sitemap Actor dataset rows and DOC\_TO\_MARKDOWN\_INPUT). You can also paste plain URL strings or use a remote text list (`requestsFromUrl`).

## `datasetId` (type: `string`):

Optional. Apify dataset ID from another Actor run (e.g. Sitemap URL Extractor or URL Status Checker). Reads the `url` field from each item. Combined with `urls` if both are set.

## `keyValueStoreId` (type: `string`):

Optional. Apify key-value store ID. Used with Record key below to load a JSON list such as DOC\_TO\_MARKDOWN\_INPUT (`{"urls":[{"url":...}]}`) or a plain URL array.

## `keyValueRecordKey` (type: `string`):

Key inside the source key-value store (e.g. `DOC_TO_MARKDOWN_INPUT`). Ignored unless a store ID is set.

## `maxUrls` (type: `integer`):

Stop after extracting this many unique URLs (0 = no limit). Also caps cost.

## `maxHtmlBytes` (type: `integer`):

Truncate the response body after this many bytes before parsing (guards memory). Metadata is usually in the first few hundred KB.

## `extractOpenGraph` (type: `boolean`):

Parse `og:*` properties (via extruct).

## `extractTwitter` (type: `boolean`):

Parse `twitter:*` meta tags.

## `extractJsonLd` (type: `boolean`):

Parse `<script type="application/ld+json">` blocks (Schema.org and others).

## `extractMicrodata` (type: `boolean`):

Parse HTML microdata (`itemscope` / `itemprop`).

## `extractRdfa` (type: `boolean`):

Parse RDFa attributes. Turn off if you only need OG / Twitter / JSON-LD (slightly faster).

## `extractBasicMeta` (type: `boolean`):

Parse `<title>`, meta description / keywords / robots, and `<link rel="canonical">`.

## `chargeFailedUrls` (type: `boolean`):

When false (default), only ok/partial rows are billed. When true, every attempted URL that produced a dataset row is billed once (including network/HTTP errors).

## `maxConcurrency` (type: `integer`):

Parallel fetches overall.

## `maxConcurrencyPerHost` (type: `integer`):

Parallel fetches to any single host.

## `minDelayMsPerHost` (type: `integer`):

Minimum milliseconds between starting requests to the same host. 0 = no extra delay.

## `requestTimeoutSecs` (type: `integer`):

Timeout for each HTTP request (connect + read).

## `maxRetries` (type: `integer`):

Retries for timeouts, network blips, HTTP 429 and selected 5xx (exponential backoff). DNS and certificate failures are not retried.

## `userAgent` (type: `string`):

Optional custom User-Agent header.

## `proxyConfiguration` (type: `object`):

Optional. Use a proxy if a site blocks data-center IPs. Not needed for most public sites.

## Actor input object example

```json
{
  "urls": [
    {
      "url": "https://example.com/"
    },
    {
      "url": "https://www.w3.org/"
    }
  ],
  "keyValueRecordKey": "DOC_TO_MARKDOWN_INPUT",
  "maxUrls": 50,
  "maxHtmlBytes": 2097152,
  "extractOpenGraph": true,
  "extractTwitter": true,
  "extractJsonLd": true,
  "extractMicrodata": true,
  "extractRdfa": false,
  "extractBasicMeta": true,
  "chargeFailedUrls": false,
  "maxConcurrency": 6,
  "maxConcurrencyPerHost": 2,
  "minDelayMsPerHost": 0,
  "requestTimeoutSecs": 25,
  "maxRetries": 2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Default dataset items, one per URL: url, finalUrl, httpStatus, status, title, description, canonical, openGraph, twitter, jsonLd, microdata, meta, extracted, errorClass, durationMs, host. Views: overview, structured.

## `summary` (type: `string`):

Key-value store record SUMMARY (JSON): totals by status, format counts, durationSecs, peakMemoryMb, charged.

## `output` (type: `string`):

Key-value store record OUTPUT (JSON): same run summary as SUMMARY. Kept for consistency with sibling Actors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        {
            "url": "https://example.com/"
        },
        {
            "url": "https://www.w3.org/"
        }
    ],
    "maxUrls": 50,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("ingenious_quip_bxq/seo-metadata-extract").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        { "url": "https://example.com/" },
        { "url": "https://www.w3.org/" },
    ],
    "maxUrls": 50,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("ingenious_quip_bxq/seo-metadata-extract").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    {
      "url": "https://example.com/"
    },
    {
      "url": "https://www.w3.org/"
    }
  ],
  "maxUrls": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call ingenious_quip_bxq/seo-metadata-extract --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ingenious_quip_bxq/seo-metadata-extract"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/7eZuqGeDGfdXotB9C/builds/zyUCeBZo0jmz25Gm4/openapi.json
