# Restaurant Menu Scraper (`axiomworks/restaurant-menus-scraper`) Actor

Extract menu items with section, name, description, price and currency from restaurant website URLs. Reads schema.org JSON-LD, PDF menus and HTML pages, following menu links from homepages (up to 12 pages and 6 PDFs per site). Results vary by how each site publishes its menu.

- **URL**: https://apify.com/axiomworks/restaurant-menus-scraper.md
- **Developed by:** [Axiom Works](https://apify.com/axiomworks) (community)
- **Categories:** E-commerce, Lead generation
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Restaurant Menu Scraper

### What does Restaurant Menu Scraper do?

Restaurant Menu Scraper turns restaurant websites into structured menu data: one dataset row per menu item, with the section, item name, description, price and currency. You give it a restaurant homepage, a menu page or a link to a menu PDF, and it works out where the menu is and how it is published. It works on any restaurant site you supply, not on a fixed list of platforms.

For each URL the Actor reads the page, then follows same-site links that look like menu pages (text or URL containing "menu", "lunch", "dinner", "drinks", "carte", "speisekarte" and similar) up to two levels deep, and reads menu PDFs linked from those pages. Per URL it opens at most 12 pages and 6 PDFs, and repeated items are removed. Items come from three sources: schema.org JSON-LD menu data, text on HTML pages, and text inside PDF files. The `source` field on every item says which one produced it, so you can judge how much to trust it.

If no menu item is found for any of your URLs, the run fails with an error message that lists what was tried, instead of finishing with an empty dataset.

People use it for competitor price monitoring, local market research, building menu databases, and onboarding restaurants to delivery or point-of-sale products. No login and no API key are needed.

### What data can you get?

Each dataset item is one menu item. Optional fields are left out when the source has no value for them.

| Field | Description | Example |
|---|---|---|
| `restaurantUrl` | The URL you supplied that the item was scraped from | `https://www.taquerialosperez.com/` |
| `restaurantName` | Restaurant name from JSON-LD, `og:site_name` or the page title (generic segments like "Menu" removed), or the domain as a fallback | `Taqueria 86` |
| `menuUrl` | The linked menu page or PDF the item was read from. Only present when it differs from `restaurantUrl` | `https://www.taquerialosperez.com/en/menu` |
| `section` | Menu section or category heading (max 100 characters). Absent if none was detected | `Alambres` |
| `itemName` | Name of the dish or drink (max 200 characters) | `Alambre De Pastor` |
| `description` | Item description (max 500 characters). Absent if there is none. Mostly available from JSON-LD menus | `Griddled meat with peppers, onions, and melted cheese` |
| `price` | Price as a number, without currency symbol. Can be `null` if the source lists no price | `14.99` |
| `currency` | ISO code (`USD`, `EUR`, `GBP`). Taken from the symbol next to the price or from JSON-LD. If the page shows bare numbers, it is inferred from other prices on the site, the page language or the domain. Null if none of these works | `USD` |
| `sourceUrl` | The exact page or PDF the item was read from | `https://www.taqueria86.com/drinks` |
| `id` | Stable 16-character identifier for the item | `3f9a1c0b7d2e4a58` |
| `source` | Strategy that produced the item: `json-ld`, `pdf` or `html` | `json-ld` |
| `scrapedAt` | ISO 8601 UTC timestamp of the scrape | `2026-09-29T18:48:33.221252+00:00` |

The dataset has three views in the Console: "Menu items" (the main columns), "Full data" (every field) and "Items with descriptions".

### How to use Restaurant Menu Scraper

1. Open the Actor in Apify Console and go to the Input tab.
2. Paste one or more restaurant URLs into "Restaurant URLs". Use a homepage, a menu page or a direct link to a menu PDF. You can add up to 20 URLs per run.
3. Set "Max items" to cap the total number of items extracted in the run. Lower it for a quick test.
4. Leave the proxy off unless a site returns HTTP 403 to datacenter traffic. In that case enable Apify Proxy and pick the residential group.
5. Click Start. Small runs finish in under a minute because the Actor fetches plain HTML and PDFs and does not start a browser.
6. Open the Output tab, choose a view, and export the dataset as JSON, CSV, Excel, XML or HTML. You can also fetch it through the API.

### Input

| Field | Type | Default / prefill | Description |
|---|---|---|---|
| `startUrls` | array of strings | prefilled with two example taqueria menu pages | Restaurant website URLs (homepage or menu page) or direct menu PDF links. Same-site menu pages and menu PDFs linked from them are followed (max 12 pages and 6 PDFs per URL). Maximum 20 URLs per run; extra URLs are ignored. |
| `maxItems` | integer, 1 to 500 | default 200, prefill 10 | Maximum menu items to extract in total across all URLs. Charges apply per item, so lower this to cap cost. |
| `proxyConfiguration` | object | Apify Proxy off | Optional. Some restaurant sites block datacenter IPs with HTTP 403. Enable Apify Proxy (residential works best) for those. |

`startUrls` must contain at least one URL. If it is empty or contains an invalid URL, the run fails immediately with a clear message. URLs without a scheme, such as `example.com/menu`, get `https://` added automatically.

Complete input example:

```json
{
  "startUrls": [
    "https://www.taquerialosperez.com/",
    "https://www.taqueria86.com/menu"
  ],
  "maxItems": 100,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

### Output

This item comes from a real local run on a restaurant menu page that publishes JSON-LD (first item of the dataset):

```json
{
    "restaurantUrl": "https://www.taquerialosperez.com/en/menu",
    "restaurantName": "Taquería Los Pérez",
    "itemName": "alambre mixto",
    "price": 16.99,
    "currency": "USD",
    "source": "json-ld",
    "sourceUrl": "https://www.taquerialosperez.com/en/menu",
    "scrapedAt": "2026-09-30T05:48:16.035682+00:00",
    "section": "Alambres",
    "id": "172246e809a89e97"
}
```

A menu page with no structured data falls back to text extraction. Here is a real item from that path:

```json
{
  "restaurantUrl": "https://www.taqueria86.com/menu",
  "restaurantName": "Taqueria 86",
  "itemName": "GUACAMOLE & CHIPS",
  "price": 16.95,
  "currency": "USD",
  "source": "html",
  "sourceUrl": "https://www.taqueria86.com/menu",
  "scrapedAt": "2026-09-30T05:48:17.038983+00:00",
  "section": "APPETIZERS",
  "description": "Freshly made Guacamole with onions, cilantro and tomato.",
  "id": "d34c281b03af8009"
}
```

`restaurantName` comes from the page's structured data, `og:site_name` or the page title with generic segments such as "Menu" removed. `description` is left out when it only repeats the item name, and fee lines such as a bag charge are skipped. HTML and PDF items often carry no description, and section names can be off on unusual layouts.

### How much does it cost?

Pricing is per result, which means per menu item written to the dataset. A typical restaurant has a few dozen to about a hundred items, so a single restaurant is a very small run. Apify's free plan includes monthly platform credit that covers small runs and testing. Exact per-item prices for each plan tier are shown on the Pricing tab of the Actor page. Use `maxItems` to set an upper bound on what a run can produce.

### Use with the API

Python, using the `apify-client` package:

```python
import os

from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("axiomworks/restaurant-menus-scraper").call(run_input={
    "startUrls": ["https://www.taqueria86.com/menu"],
    "maxItems": 50,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["itemName"], item["price"])
```

JavaScript, using the `apify-client` package:

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('axiomworks/restaurant-menus-scraper').call({
    startUrls: ['https://www.taqueria86.com/menu'],
    maxItems: 50,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

cURL, running the Actor synchronously and returning the dataset items:

```bash
curl -X POST "https://api.apify.com/v2/acts/axiomworks~restaurant-menus-scraper/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"startUrls": ["https://www.taqueria86.com/menu"], "maxItems": 50}'
```

You can connect the dataset to other tools with Apify integrations: Zapier, Make, n8n, Google Sheets, and webhooks that fire when a run finishes. Schedule the Actor to re-scrape menus daily or weekly and compare prices over time.

### Use with AI agents (MCP)

You can call Restaurant Menu Scraper from any MCP-capable client through Apify's MCP server at mcp.apify.com. The server exposes the tools `search-actors` and `call-actor`, so an agent can find this Actor and run it. To load only this Actor as a tool, use this configuration in Claude Desktop, Cursor, VS Code or another MCP client:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=axiomworks/restaurant-menus-scraper"
    }
  }
}
```

Example prompts you could type to an agent:

- "Get the menu items and prices from https://www.taqueria86.com/menu and put them in a table."
- "Compare taco and burrito prices on https://www.taqueria86.com/menu and https://www.taquerialosperez.com/en/menu."
- "Scrape the menu PDF at <pdf url> and list every item under 15 dollars."

The Actor has typed input and output schemas, so the agent knows the exact input fields (`startUrls`, `maxItems`, `proxyConfiguration`) and the shape of every returned item. Remember to set a small `maxItems` for exploratory agent calls.

### FAQ

**Does it work on JavaScript-heavy menus?**
No. The Actor reads plain HTML and PDFs and does not run a browser. Menus that load only after the page renders (many ordering widgets, for example order pages hosted by Toast, which also answer datacenter traffic with HTTP 403) and menus that are just images produce no items. If that is the case for every URL you gave, the run fails with an error message.

**How accurate are the results?**
We test the Actor on a benchmark of 20 real restaurant websites in the US, UK, Italy and Germany, covering menus on the homepage, separate menu pages, PDF menus, JSON-LD, and pages built with WordPress, Squarespace, Wix and Square. For each site we counted by hand roughly how many priced items the menu shows. In the last run, 18 of 20 sites returned at least 80% of those items with section, name, price and currency. The two misses were a site whose wine list is a multi-column table (about 45% found) and a site that shows prices without any currency symbol. Counts are approximate (about ±10%), and returned items were spot-checked, not all verified. Expect some wrong section names and a few non-dish lines (for example a service-charge note) on unusual layouts. Check the `source` field to see which method produced each row.

**Why did I get no results for a URL?**
Common reasons are a menu rendered by JavaScript, a menu embedded as an image or a scanned PDF, a page that returns HTTP 403 or 404, or a menu that lists no prices. Try the direct menu page or PDF link instead of the homepage, or enable Apify Proxy if the site blocks datacenter IPs. If no URL in the run yields an item, the run fails and the error message names each URL and what happened.

**Do I need a proxy?**
Usually not. Enable Apify Proxy only for sites that answer 403 to datacenter traffic. Proxy traffic may add platform costs.

**How fast is it, and what are the limits?**
A run handles up to 20 URLs and up to 500 items per run. For each URL the Actor opens at most 12 pages and 6 PDFs, waits between requests to the same host, and reads at most the first 15 pages of a PDF. Only same-site links are followed for pages; PDFs linked from those pages may be hosted elsewhere.

**Does one bad URL stop the run?**
No. Errors are logged per URL and the run continues with the next one. The run only fails when no URL produced any item.

**Can it scrape delivery platforms such as Uber Eats or DoorDash?**
No. Those sites are not supported. The Actor is designed for restaurants' own websites and menu PDFs.

**How do I keep data fresh?**
The Actor reads the live page on every run. Use Apify schedules to re-run it at the frequency you need.

### Is it legal to scrape restaurant websites?

This Actor only reads publicly available pages and files that anyone can open in a browser, and it does not log in or bypass access controls. Menu and price data is generally public, but you remain responsible for how you use it. Respect each site's terms of service, copyright in menu text and photos, and applicable privacy and data protection laws such as GDPR and CCPA. Do not use the data in ways that would infringe someone's rights, and seek legal advice if you are unsure about your specific use case.

### Feedback

Found a site where extraction fails or returns wrong items? Open an issue from the Issues tab of this Actor and include the URL and what you expected. Reports with a concrete example help the most.

# Actor input Schema

## `startUrls` (type: `array`):

Restaurant website URLs as plain strings, e.g. 'https://www.taqueria86.com/' (homepage or menu page) or a direct menu PDF link. Same-site menu pages and menu PDFs linked from them are followed (max 12 pages, 6 PDFs per URL). Max 20 URLs per run.

## `maxItems` (type: `integer`):

Maximum menu items to output in total across all URLs, 1-500 (default 200). Lower it to cap cost and run time.

## `proxyConfiguration` (type: `object`):

Optional. Some restaurant sites block datacenter IPs (HTTP 403); enable Apify Proxy (residential group works best) for those.

## Actor input object example

```json
{
  "startUrls": [
    "https://www.taquerialosperez.com/en/menu",
    "https://www.taqueria86.com/menu"
  ],
  "maxItems": 10,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

All results in the dataset

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.taquerialosperez.com/en/menu",
        "https://www.taqueria86.com/menu"
    ],
    "maxItems": 10,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("axiomworks/restaurant-menus-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        "https://www.taquerialosperez.com/en/menu",
        "https://www.taqueria86.com/menu",
    ],
    "maxItems": 10,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("axiomworks/restaurant-menus-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.taquerialosperez.com/en/menu",
    "https://www.taqueria86.com/menu"
  ],
  "maxItems": 10,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call axiomworks/restaurant-menus-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,axiomworks/restaurant-menus-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9AiY6XS61vVanUmsM/builds/yWkMDWyEUWu5Wh3PB/openapi.json
