# Web Scraping Data Export API - Grepsr Reports, Runs and Files (`nabeelbaghoor/web-scraping-data-export-api`) Actor

Export the data of your own Grepsr web scraping projects: every scraped record of a report's latest run or any history, run history with item counts, output file links, run parameters and schedules, plus your projects. Read only: nothing is run or stopped. Bring your own key.

- **URL**: https://apify.com/nabeelbaghoor/web-scraping-data-export-api.md
- **Developed by:** [Nabeel Hassan](https://apify.com/nabeelbaghoor) (community)
- **Categories:** Developer tools, E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$5.00 / 1,000 web scraping data record returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Web Scraping Data Export API - Grepsr Reports, Runs and Files

Export what your own Grepsr web scraping projects have already collected: every scraped record of a report's latest run or of any past history, the run history with item counts, output file links, run parameters and schedules, as clean rows ready for a warehouse, a spreadsheet or the next step of a pipeline.

### What it collects

- **Scraped records**: the newest finished history of a report, or any history by id, one row per scraped record with every field under `data`, keyed by the report's own column names. The actor reads the history's JSON output file, or its CSV file when there is no JSON one.
- **Run history**: the histories of a report with status, item count, request count, upload and download bandwidth, start and end time and the CSV link; or its runs with run parameters and every output file link.
- **Run search**: the runs of a report whose run parameters match the values you give, such as one file name.
- **Output files**: the XLSX, XML, CSV and JSON download links of a history, with their expiry time.
- **Run status**: the status of the run behind a history.
- **Parameters and schedules**: the run parameters a report takes, its schedules with start, end and next update times, and the parameters set on a schedule.
- **Projects**: every project in the account, with status, type and schedule type.
- Read only, pay per result, bring your own key.

### Input

| Field | What it does |
| --- | --- |
| What to read | Projects (default, needs nothing else), latest scraped records of a report, scraped records of a history, histories, runs, search runs, files of a history, status of a history, parameters of a report, schedules, or parameters of a schedule. |
| Ids | One number per line: report ids, history ids or schedule ids, depending on the service. |
| Histories or runs per report | Histories and runs: how many to return per report. |
| Search parameters | Search runs: a JSON object of run parameter name to value. |
| Report id for schedule parameters | Parameters of a schedule: the report the schedules belong to. |
| Project type | Projects: WEBAPP (default) or any. |
| Maximum results | Row cap for the run. |
| Requests per minute | Pacing for calls to the provider, which allows 100 per minute. |
| API key | Your own API key, as a secret input. |

### Example output

```json
{
  "service": "latestRecords",
  "serviceLabel": "Latest scraped records of a report",
  "requested": "104522",
  "found": true,
  "reportId": "104522",
  "historyId": "8841207",
  "status": "SUCCESS",
  "startedAt": "2026-09-28T06:00:12",
  "endedAt": "2026-09-28T06:41:55",
  "fileFormat": "JSON",
  "rowIndex": 0,
  "data": {
    "Product Name": "Trail Running Shoe 1042",
    "Price": "89.99",
    "Currency": "EUR",
    "Availability": "In stock",
    "Product URL": "https://shop.example.com/products/trail-running-shoe-1042"
  },
  "retrievedAt": "2026-09-29T09:14:52.118Z",
  "note": null
}
```

Values are illustrative; the columns under `data` are whatever your report extracts. A history row looks like this:

```json
{
  "service": "histories",
  "requested": "104522",
  "found": true,
  "reportId": "104522",
  "historyId": "8841207",
  "status": "SUCCESS",
  "itemCount": 2488,
  "startedAt": "2026-09-28T06:00:12",
  "endedAt": "2026-09-28T06:41:55",
  "fileFormat": "CSV",
  "fileUrl": "https://files.example.com/oJElX0e",
  "details": { "requestCount": 126, "bandwidthUploadBytes": 25022, "bandwidthDownloadBytes": 1984907 }
}
```

### FAQ

#### What is this web scraping data export API used for?

Getting data out of managed scraping projects that already run on a schedule and into the systems that use it. A pricing team pulls the latest competitor price and product scrape every morning into BigQuery, Snowflake or Google Sheets. An analyst lists last month's histories of a report to see item counts and which runs failed. An operations team checks request counts and bandwidth per run. An integration finds the run for one file name with a parameter search.

#### Which data source does this actor read?

The Grepsr API at api.grepsr.com, through the read routes of its public documentation at api-docs.grepsr.com: `POST /v1/project/list`, `POST /v1/report/history/list`, `POST /v1/report/run/list`, `POST /v1/report/run/search`, `POST /v1/history/files`, `POST /v1/history/run/status`, `GET /v1/report/parameters/list`, `POST /v1/report/schedule/list` and `GET /v1/report/schedule/parameters`. The provider documents most reads as POST requests with a JSON body; none of them writes anything. Scraped records come from the JSON or CSV output file a history links to.

#### Is this a scraping API that takes a URL?

No. The provider's documented API does not take a URL and return a page. It exports the output of the crawlers (reports) set up in your own account, which is what this actor reads.

#### Do I need an API key?

Yes. This actor is bring-your-own-key and never ships one. The key is shown under Personal Details on the Profile page of your Grepsr account. Paste it into the input, or set it once as the `DATA_API_KEY` environment secret. It is sent in the `X-Api-Key` header to the API host only, and never to the file download links. A missing or refused key ends the run cleanly with a message saying which it was.

#### How do I find report ids and history ids?

The provider has no route that lists the reports of a project, so a report id is the number in the report's address in the web app. With a report id, the histories or runs service lists the history ids. The latest scraped records service needs only the report id, which is usually all you need.

#### Can this actor run a report or change anything?

No. It never runs or stops a report, never stops a schedule and never calls the instant crawler. It reads what your reports have already collected.

#### What happens when there is nothing for an id?

It becomes its own row with `found: false` and a note: a report with no finished history yet, a history id the account cannot see, an expired file link, or a history delivered only as XLSX or XML. Those rows are never charged.

#### How is it priced?

Pay per result: one flat price per record returned, whether a scraped record, a project, a history, a run, a file link, a status, a parameter set or a schedule. Rows for ids that produced nothing are free. Your Grepsr subscription applies separately.

### Pricing

| Event | Price |
| --- | --- |
| Web scraping data record returned | $0.005 per record |

### Keyword map

web scraping data export, web scraping API, scraped data API, managed web scraping export, crawler output to database, scheduled scraping export, ecommerce product data export, competitor pricing data feed, scraping run history, Grepsr API, Grepsr reports, Grepsr histories, web data extraction export.

# Actor input Schema

## `service` (type: `string`):

Projects lists every project in the account and needs nothing else, so it is the default. Latest scraped records reads the newest finished history of each report id and returns every scraped record in its JSON or CSV file. Scraped records of a history does the same for one history id. Histories and runs list the times a report ran, which is where history ids come from. The other services read output file links, run status, run parameters and schedules.

## `identifiers` (type: `array`):

One per line, as numbers. Report ids for latest scraped records, histories, runs, search runs, report parameters and schedules; history ids for scraped records of a history, files and status; schedule ids (or scheduler job ids) for parameters of a schedule. A report id is the number in the report's address in the web app. Projects reads none.

## `runCount` (type: `integer`):

Histories and runs only: how many the provider returns for each report, sent as its size field. The provider has no paging, so this is the whole list.

## `searchParams` (type: `object`):

Search runs only: the run parameters to match, as a JSON object of parameter name to value, such as {"file\_name": "2023\_12\_15\_export.xlsx"}. Parameter names are the ones parameters of a report lists.

## `reportId` (type: `string`):

Parameters of a schedule only: the report the schedule ids belong to.

## `projectType` (type: `string`):

Projects only: WEBAPP, the type the provider documents and the default, or any to send no type.

## `maxResults` (type: `integer`):

Stop after this many rows. One output file can hold tens of thousands of scraped records.

## `requestsPerMinute` (type: `integer`):

Pacing ceiling for calls to the provider, which allows 100 per minute per user across all routes. Rate limited answers are retried after a pause.

## `apiKey` (type: `string`):

Your own API key, shown under Personal Details on the Profile page of your Grepsr account. This actor is bring-your-own-key and never ships one. Leave blank to use the DATA\_API\_KEY environment secret. The key is sent only to the API host, never to file links, and is never written to a row or logged.

## `baseUrl` (type: `string`):

Override the host the actor calls. Only useful for testing against a different environment.

## Actor input object example

```json
{
  "service": "projects",
  "runCount": 20,
  "projectType": "WEBAPP",
  "maxResults": 1000,
  "requestsPerMinute": 60
}
```

# Actor output Schema

## `records` (type: `string`):

One row per scraped record, project, history, run, file, status, parameter set or schedule.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("nabeelbaghoor/web-scraping-data-export-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("nabeelbaghoor/web-scraping-data-export-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call nabeelbaghoor/web-scraping-data-export-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nabeelbaghoor/web-scraping-data-export-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/lefxg8aC7myTT5BJS/builds/xYxaVwUCBwUlXAdd9/openapi.json
