# Gluecrawl Universal Scraper (`gluecrawl/gluecrawl-apify-actor`) Actor

Turn any public website into structured data with your Gluecrawl API key.

- **URL**: https://apify.com/gluecrawl/gluecrawl-apify-actor.md
- **Developed by:** [Valentino Arbelaiz](https://apify.com/gluecrawl) (community)
- **Categories:** AI, Agents, Open source
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Gluecrawl Universal Scraper

Turn a public webpage into structured data without selectors, without paying for an idle Apify container. This Actor submits to Gluecrawl first, then retrieves completed results in a separate short run.

### Before you start

You need a Gluecrawl account with API access and available credits. Paste your personal Gluecrawl API key into the secret **Gluecrawl API key** input. Apify encrypts secret inputs; the Actor never adds your key to the dataset or logs.

Gluecrawl—not Apify—runs the extraction and charges the applicable Gluecrawl credits. This Actor does not support login-protected targets, cookies, custom proxies, schedules, or multi-URL batches in v1.

### Inputs

- **Mode** — use **start** to submit a new extraction, then **fetch** to retrieve its result. Both runs finish promptly; neither polls or waits for Gluecrawl.
- **Target URL** — required in start mode: one public webpage or listing page.
- **Gluecrawl API key** — your personal Gluecrawl key, stored as an Apify secret.
- **Goal** — required in start mode; describe the data you want, for example: `Extract each travel book's title and price.`
- **Fields** — in start mode, alternatively provide reproducible field definitions such as `[{"name":"title","description":"Book title"}]`.
- **Maximum pages** — start mode only; defaults to 2; choose 1–100 within your Gluecrawl plan limit.
- **Gluecrawl job ID** — fetch mode only; copy this from the prior start run's receipt.

Use **Goal** or **Fields**, not both.

### How to run it

1. Run in **start** mode. It writes a small receipt such as `{"status":"submitted","jobId":"..."}` to the default dataset and exits.
2. After Gluecrawl has finished, run the Actor in **fetch** mode with the same API key and that `jobId`.
3. Fetch writes every extracted item to its default dataset. If Gluecrawl is still processing, fetch succeeds quickly with a `pending` receipt; retry it later.

### Output

Completed fetches write each extracted item unchanged to the default Apify dataset. Start and pending fetches write only their status receipt. Mapping errors, scraping errors, expired/invalid keys, insufficient credits, and zero-item completions fail clearly instead of producing a misleading empty dataset.

### Example

Start with `https://books.toscrape.com/catalogue/category/books/travel_2/index.html` and the goal `Extract each travel book's title and price.` Then fetch with the returned job ID. The completed fetch dataset contains one row per book.

### API use

Run this Actor through the Apify Console, API, CLI, or an Apify task. Read the default dataset of the completed fetch run as you would any other Actor output.

# Actor input Schema

## `mode` (type: `string`):

Start submits a new job and exits. Fetch retrieves one existing job without waiting.

## `url` (type: `string`):

The public page to extract data from.

## `gluecrawlApiKey` (type: `string`):

Create this from your Gluecrawl account. It is encrypted by Apify and never written to the dataset.

## `goal` (type: `string`):

Use this or Fields, not both.

## `fields` (type: `array`):

Alternative to Goal: \[{"name": "title", "description": "Book title"}].

## `maxPages` (type: `integer`):

Maximum pages Gluecrawl may scrape.

## `jobId` (type: `string`):

Required only in fetch mode. Copy it from the start run's dataset receipt.

## Actor input object example

```json
{
  "mode": "start",
  "url": "https://books.toscrape.com/catalogue/category/books/travel_2/index.html",
  "goal": "Extract each travel book's title and price.",
  "maxPages": 2
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "start",
    "url": "https://books.toscrape.com/catalogue/category/books/travel_2/index.html",
    "goal": "Extract each travel book's title and price.",
    "maxPages": 2
};

// Run the Actor and wait for it to finish
const run = await client.actor("gluecrawl/gluecrawl-apify-actor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "start",
    "url": "https://books.toscrape.com/catalogue/category/books/travel_2/index.html",
    "goal": "Extract each travel book's title and price.",
    "maxPages": 2,
}

# Run the Actor and wait for it to finish
run = client.actor("gluecrawl/gluecrawl-apify-actor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "start",
  "url": "https://books.toscrape.com/catalogue/category/books/travel_2/index.html",
  "goal": "Extract each travel book'\''s title and price.",
  "maxPages": 2
}' |
apify call gluecrawl/gluecrawl-apify-actor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gluecrawl/gluecrawl-apify-actor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/d4Dom9Coa9SUI01ax/builds/jNDa5nRFBJOsSg6sR/openapi.json
