# Recipe, Event & Job Data Extractor (schema.org JSON-LD) (`swiftkit/structured-data`) Actor

Extract clean recipes (ingredients, steps, times, nutrition), events (dates, venue, tickets), job postings (salary, location, remote), products, articles, businesses, FAQs and more from any website that publishes schema.org data for Google. One row per item, normalized fields. $1 per 1,000 items.

- **URL**: https://apify.com/swiftkit/structured-data.md
- **Developed by:** [SwiftKit](https://apify.com/swiftkit) (community)
- **Categories:** Developer tools, SEO tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Recipe, Event & Job Data Extractor (schema.org JSON-LD)

Millions of websites describe their content in **schema.org structured data** so Google can show
rich results: recipe cards, event listings, Google for Jobs, product prices. This tool reads that
data from any page and turns it into **clean, normalized rows**, one per item:

- **Recipes:** name, author, ingredients list, step-by-step instructions, prep / cook / total
  minutes, yield, category, cuisine, keywords, calories and nutrition, rating
- **Events:** name, start and end, status (scheduled, cancelled, moved online), venue, address,
  coordinates, online URL, organizer, performers, ticket price, currency, availability, ticket link
- **Job postings:** title, company, date posted, valid through, employment type, remote flag,
  locations, salary range with currency and unit, description
- **Products:** name, brand, SKU, GTIN, price range, currency, availability, rating
- **Articles:** headline, authors, publisher, published and modified dates, section, keywords
- **Businesses:** name, address, coordinates, phone, opening hours, price range, cuisine, rating
- **FAQs** (question and answer pairs), **how-to guides**, **courses**, **apps & software**

Paste the pages, or start from an index page (a recipe category, an events calendar, a careers
page) and let it follow links on the same site.

### Who it's for

- **Food apps and meal planners:** ingredients, steps and times from recipe sites.
- **Event aggregators and city guides:** dates, venues and tickets from venue and organizer sites.
- **Job boards and recruiters:** postings from company career pages that publish Google for Jobs data.
- **SEO teams:** see exactly what structured data a page exposes (turn on raw JSON-LD).

### Input

| Option | Default | What it does |
|---|---|---|
| Pages | – | Any URLs |
| Only these types | all | Recipes, events, jobs, products, articles… |
| Follow links on the same site | off | Crawl from index pages |
| Only follow links containing | – | e.g. `/recipe/` |
| Max pages / Max items | 50 / 1,000 | Limits |
| Include raw JSON-LD | off | The original object too |
| Respect robots.txt | on | Skip disallowed pages |

### Output

A real recipe (loveandlemons.com):

```json
{
  "status": "ok",
  "itemType": "Recipe",
  "name": "BEST Hummus",
  "author": "Jeanine Donofrio",
  "prepMinutes": 5,
  "totalMinutes": 5,
  "ingredients": ["…9 items…"],
  "instructions": ["…2 steps…"],
  "rating": 4.96
}
```

A real job posting (a Lever career page):

```json
{
  "itemType": "JobPosting",
  "title": "Client Partner - Emerging & Scaled, Independent Agency (UK)",
  "company": "Spotify",
  "datePosted": "2026-09-17",
  "employmentType": ["Permanent"],
  "remote": false,
  "locations": [{ "city": "London" }]
}
```

Generic "Article" wrappers and site-wide organization boilerplate are skipped automatically when a
page has a specific item (a recipe, a job…), so you don't pay for noise.

| status | Meaning |
|---|---|
| `ok` | Item extracted. |
| `no_structured_data` | The page has no supported structured data. Not charged. |
| `blocked_by_robots_txt` / `blocked_by_bot_protection` | The site doesn't want bots. Not charged. |
| `unreachable` | The page didn't load (some big publishers answer bots with 402/403). Not charged. |

### Use it from your AI agent (MCP)

Claude, Cursor and other MCP clients can call this tool directly through Apify's MCP server. Add it as
an MCP server / connector:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=swiftkit/structured-data",
      "headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }
    }
  }
}
```

Then just ask, for example: *"Get the ingredients and total time from this recipe: https://…"*. Each call is billed like a normal run.

### Pricing

You pay **per extracted item**. Every other status is free. See the Pricing tab.

### Limits, honestly

- Only data the site publishes as JSON-LD in its HTML. Pages that add it with JavaScript aren't covered.
- Some large publishers block automated visitors entirely (for example with HTTP 402 or 403).
  The tool respects that and moves on.
- Recipe instructions and descriptions are the site's own text: use them in line with the site's
  terms and copyright (ingredients and facts are generally fine; republishing whole recipes may not be).
- People are never extracted as items on their own; author and organizer names appear only as part
  of the item they belong to.

### More tools from SwiftKit

- [Product Page Scraper](https://apify.com/swiftkit/product-pages): prices and stock from any shop, with price tracking
- [Job Search API](https://apify.com/swiftkit/job-search): 10,000+ company career sites in one search
- [Website to Markdown for AI](https://apify.com/swiftkit/web-to-markdown): clean page content for LLMs

### Questions?

Open an issue on the Issues tab.

# Actor input Schema

## `urls` (type: `array`):

Recipe, event, job, product, article or business pages from any website.

## `types` (type: `array`):

Empty = all supported types.

## `followLinks` (type: `boolean`):

Also visit links on the pages (same website), e.g. to go from a recipe index or events calendar to the detail pages.

## `linkPatterns` (type: `array`):

Text or regular expressions, e.g. /recipe/ or /events/. Recommended with "Follow links".

## `maxPages` (type: `integer`):

Pages to visit at most.

## `maxItems` (type: `integer`):

Stop after this many extracted items.

## `includeRaw` (type: `boolean`):

Add the original schema.org object to each row.

## `respectRobotsTxt` (type: `boolean`):

Skip pages the site asks bots not to visit.

## `maxConcurrency` (type: `integer`):

Keep it low to be gentle with websites.

## Actor input object example

```json
{
  "urls": [
    "https://www.loveandlemons.com/hummus-recipe/",
    "https://www.mozilla.org/en-US/firefox/"
  ],
  "followLinks": false,
  "maxPages": 50,
  "maxItems": 1000,
  "includeRaw": false,
  "respectRobotsTxt": true,
  "maxConcurrency": 4
}
```

# Actor output Schema

## `results` (type: `string`):

One row per recipe, event, job, product or other item.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.loveandlemons.com/hummus-recipe/",
        "https://www.mozilla.org/en-US/firefox/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("swiftkit/structured-data").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://www.loveandlemons.com/hummus-recipe/",
        "https://www.mozilla.org/en-US/firefox/",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("swiftkit/structured-data").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.loveandlemons.com/hummus-recipe/",
    "https://www.mozilla.org/en-US/firefox/"
  ]
}' |
apify call swiftkit/structured-data --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,swiftkit/structured-data"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CFl83nu7ttlqI5I1Y/builds/VEi2nzUqnghG999nk/openapi.json
