# Website Structured Data (JSON-LD) Extractor (`parseforge/website-structured-data-extractor`) Actor

- **URL**: https://apify.com/parseforge/website-structured-data-extractor.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.62 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

![ParseForge Banner](https://github.com/ParseForge/apify-assets/blob/main/banner.jpg?raw=true)

## 🧩 Website Structured Data (JSON-LD) Extractor

> 🚀 **Turn any list of web pages into clean schema.org structured data in seconds.**

Give this Actor a list of URLs and it returns the schema.org JSON-LD embedded in each page: Product, Article, JobPosting, Event, FAQPage, Organization, BreadcrumbList, and every other type sites publish for search engines. One row per structured-data object, with the full JSON preserved.

Most modern sites ship rich JSON-LD for SEO. This Actor reads it directly, so you get the site's own clean, structured facts without writing per-site selectors.

| For | Use it to |
|---|---|
| SEO & content teams | Audit structured data across pages, catch missing or malformed markup |
| Data & RevOps teams | Pull product, job, event, or article facts from pages you already track |
| Developers | Normalize schema.org data from any set of URLs into one dataset |

### 📋 What it does

- Fetches each URL you provide (US residential proxy by default, so arbitrary sites do not block a datacenter IP).
- Extracts every `<script type="application/ld+json">` block, flattening `@graph` and arrays so each schema.org object is its own row.
- Preserves the complete JSON-LD object alongside its type and name.

> 💡 **Why it matters:** you supply the URLs, so there is no anti-bot guesswork. The Actor reads the structured data the site already publishes.

### 📊 Output

| Field | Description |
|---|---|
| 🔗 `sourceUrl` | The page the record came from |
| 🏷️ `type` | schema.org `@type` (e.g. Product, Article, JobPosting) |
| 📝 `name` | Name / headline / title of the object, if present |
| 📦 `data` | The full JSON-LD object |
| 🕓 `scrapedAt` | When this row was collected |
| ⚠️ `error` | Null on success; a message when a page could not be read or had no JSON-LD |

Sample record:

```json
{
  "sourceUrl": "https://www.bbc.com/news",
  "type": "WebPage",
  "name": "BBC News - Breaking news, video and the latest top stories",
  "data": { "@context": "https://schema.org", "@type": "WebPage", "name": "BBC News …" },
  "scrapedAt": "2026-08-23T14:00:00.000Z",
  "error": null
}
```

### 🚀 How to use

1. [Create a free account w/ $5 credit](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Paste your URLs into `startUrls`, set `maxItems`.
3. Run it and download the dataset as JSON, CSV, Excel, or XML.

### ❓ FAQ

**Which structured-data formats are supported?** schema.org JSON-LD (the format almost every site uses for SEO).

**What if a page has no JSON-LD?** You get one row for that URL with an `error` note, so nothing fails silently.

**Do I need proxies or keys?** No. A US residential proxy is used by default; you can change it in the input.

**Can it read product / job / event / article / FAQ pages?** Yes. It returns whatever schema.org types the page publishes (Product, JobPosting, Event, Article, FAQPage, and more).

### 🔗 Recommended Actors

- [DOAJ Journals & Subject Classification Scraper](https://apify.com/parseforge/doaj-subject-classification-scraper)
- [PubMed Article Metadata Scraper](https://apify.com/parseforge/pubmed-article-metadata-scraper)

> 💡 **Pro Tip:** browse the complete [ParseForge collection](https://apify.com/parseforge) for more data-extraction Actors.

***

*This Actor extracts publicly available structured data from the URLs you provide, for analysis and research. Use it in line with each site's terms.*

# Actor input Schema

## `startUrls` (type: `array`):

Web pages to extract schema.org JSON-LD structured data from. Paste product, article, job, event, or any pages.

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `proxyConfiguration` (type: `object`):

Proxy used to fetch pages. US residential by default so arbitrary sites do not block a datacenter IP.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://blog.apify.com"
    },
    {
      "url": "https://www.bbc.com/news"
    }
  ],
  "maxItems": 10,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all extracted records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://blog.apify.com"
        },
        {
            "url": "https://www.bbc.com/news"
        }
    ],
    "maxItems": 10,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/website-structured-data-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        { "url": "https://blog.apify.com" },
        { "url": "https://www.bbc.com/news" },
    ],
    "maxItems": 10,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/website-structured-data-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://blog.apify.com"
    },
    {
      "url": "https://www.bbc.com/news"
    }
  ],
  "maxItems": 10,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call parseforge/website-structured-data-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/website-structured-data-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zBIZEafw3P3kcKTU0/builds/fo0bz9RDSp95nD7fO/openapi.json
