# OpenDataSoft Scraper (`aurenic/opendatasoft-scraper`) Actor

Extract datasets and records from any of 3,000+ OpenDataSoft public portals — public.opendatasoft.com, data.paris.fr, data.economie.gouv.fr, BODACC, BOAMP, and thousands more. ODSQL where/select/group\_by/order\_by, facets, auto-pagination. No API key, no browser.

- **URL**: https://apify.com/aurenic/opendatasoft-scraper.md
- **Developed by:** [Aurenic](https://apify.com/aurenic) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.30 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## OpenDataSoft Scraper

Extract datasets and records from any of 3,000+ OpenDataSoft public portals — public.opendatasoft.com, data.paris.fr, data.economie.gouv.fr, BODACC, BOAMP, and thousands more. ODSQL where/select/group\_by/order\_by, facets, auto-pagination. No API key, no browser.

### What does OpenDataSoft Scraper do?

OpenDataSoft (now Huwise) powers the open data portals of **3,000+ organizations** — from the City of Paris and the French Ministry of Economy to the BODACC legal notices registry, RTE, and thousands of local governments worldwide. Every portal runs the same Explore API v2.1.

This actor is the generic reader:

- **Catalog search** — search a portal's dataset catalog by keyword, list all datasets, get metadata (title, description, theme, publisher, license, record count).
- **Dataset metadata** — full schema for any dataset: field names, types, labels, descriptions.
- **Records** — pull actual data rows with ODSQL `where`, `select`, `group_by`, `order_by`, and `refine` (facet filters). Auto-paginates up to 10,000 rows per dataset.
- **Facets** — list facet definitions and their distinct values.

**No API key required.** Anonymous access on `public.opendatasoft.com` allows up to **10 million calls per day**. Other portals have their own quotas.

### Output fields

#### Dataset catalog / metadata

| Field | Description |
|---|---|
| datasetId | Dataset identifier |
| title / description | Dataset title and description |
| theme / keyword | Classification tags |
| publisher / license | Publisher and license |
| recordsCount | Number of rows |
| modified | Last modification date |
| fields | Schema (metadata mode): `{ name, label, type, description }` |
| url | Direct portal link |

#### Record

Every row is emitted **as-is** from the dataset, with three internal fields added:

| Field | Description |
|---|---|
| recordType | `record` |
| portal | Portal hostname |
| datasetId | Source dataset |
| …all dataset columns | Whatever columns the dataset defines |

#### Facets

| Field | Description |
|---|---|
| facets | Array of facet names available on the dataset |

### Who is it for?

- **Civic tech builders** pulling permits, transit, environmental, and municipal data at scale
- **Data journalists** aggregating datasets across multiple French and international portals
- **Market researchers** combining business registries, legal notices, and tender data
- **Compliance teams** monitoring BODACC insolvency and BOAMP public procurement feeds
- **Real estate analysts** sourcing building permits, zoning, and property data
- **Data scientists** building pipelines across thousands of open data sources

### Pricing

**$0.30 per 1,000 results.** No subscription.

| Results | Cost |
|---|---|
| 100 | $0.03 |
| 1,000 | $0.30 |
| 10,000 | $3.00 |

### How to use it

1. Pick a **Mode**.
2. Set the **Portal** hostname (e.g. `public.opendatasoft.com`).
3. For records: enter **Dataset ID** and optionally **where**, **select**, **group\_by**, **order\_by**, **refine**.
4. Set **Max Items** (default 10000).
5. Click **Start**.

### Output example

```json
{
  "recordType": "record",
  "portal": "data.paris.fr",
  "datasetId": "les-arbres",
  "id": "12345",
  "arrondissement": "PARIS 5E ARRDT",
  "genre": "Alignement",
  "espece": "Platanus x hispanica",
  "hauteur": 15.5,
  "circonference": 210,
  "annee_plantation": 1950,
  "adresse": "BOULEVARD SAINT-GERMAIN",
  "geo_point_2d": { "lat": 48.8501, "lon": 2.3448 },
  "scrapedAt": "2026-09-26T12:00:00.000Z"
}
```

### Technical details

- **Source: OpenDataSoft Explore API v2.1** — `https://{portal}/api/explore/v2.1`.
- **No API key required for public datasets.** Anonymous access is allowed on all portals.
- **Rate limits** — 10M calls/day on `public.opendatasoft.com`. Other portals have their own quotas communicated via `X-RateLimit-*` headers. The actor backs off on 429.
- **ODSQL query language** — same syntax across all endpoints. Supports `where`, `select`, `group_by`, `order_by`, `refine`, and the `search()` full-text function.
- **Pagination** — `limit` per request caps at 100. Offset pagination caps at 10,000 records per dataset.
- **3,000+ portals** — any OpenDataSoft instance works: `public.opendatasoft.com`, `data.paris.fr`, `data.economie.gouv.fr`, `bodacc-datadila.opendatasoft.com`, `boamp-datadila.opendatasoft.com`, `opendata.paris.fr`, and thousands more.
- **No browser, no proxy** — pure REST JSON.

### Known limits

- **Records per request cap at 100.** Offset pagination caps at 10,000 rows per dataset. For deeper extraction, use the dataset's export endpoints or filter by date to split queries.
- **Anonymous quotas vary by portal.** `public.opendatasoft.com` allows 10M calls/day. Smaller portals may have lower limits. The actor respects `X-RateLimit-*` headers.
- **Some datasets are restricted.** Public datasets work anonymously; authenticated-only datasets return 403.
- **Field names are portal- and dataset-specific.** Each dataset has its own schema. The `datasetId` and `portal` fields are attached to every record for downstream filtering.
- **No PDF or file attachments.** The actor returns structured record data. Dataset file exports (CSV, GeoJSON, Shapefile) are separate endpoints.

### FAQ

**Do I need an API key?** No. Anonymous access works for all public datasets.

**Do I need a proxy?** No. Datacenter IPs work.

**How do I find a dataset ID?** Browse the portal, open a dataset, click **API** — the `dataset_id` is in the endpoint URL.

**What's the difference between catalog and records mode?** Catalog returns dataset metadata (title, description, record count). Records returns the actual rows.

**What ODSQL functions are available?** `search()`, `count()`, `sum()`, `avg()`, `min()`, `max()`, `year()`, `month()`, `day()`, `distance()`, and standard comparison operators.

**How do I export data?** After a run, go to Storage → Export as JSON, CSV, Excel.

### Support

Open an issue on the Actor's page for bugs or feature requests.

# Actor input Schema

## `mode` (type: `string`):

What to fetch.

## `portal` (type: `string`):

OpenDataSoft portal hostname (e.g. public.opendatasoft.com, opendata.paris.fr, data.economie.gouv.fr, bodacc-datadila.opendatasoft.com). 3,000+ portals supported.

## `datasetId` (type: `string`):

Dataset identifier on the portal (e.g. world-administrative-boundaries on public.opendatasoft.com, les-arbres on opendata.paris.fr, annonces-commerciales on bodacc-datadila.opendatasoft.com).

## `searchQuery` (type: `string`):

Full-text search in catalog mode. Leave empty to list all datasets.

## `where` (type: `string`):

ODSQL filter clause. Examples: "borough='BROOKLYN'", "magnitude > 5", "search('earthquake')".

## `select` (type: `string`):

Comma-separated column list or expressions. Examples: 'unique\_key, created\_date', 'borough, count(\*)'.

## `groupBy` (type: `string`):

Group-by clause. Required when using aggregate functions in select. Example: 'borough'.

## `orderBy` (type: `string`):

Sort clause. Examples: 'created\_date DESC', 'magnitude DESC'.

## `refine` (type: `string`):

Facet refine clause. Example: 'theme:environment' or 'borough:BROOKLYN'.

## `limitPerPage` (type: `integer`):

Records per request. OpenDataSoft caps at 100 per request.

## `maxRecordsPerDataset` (type: `integer`):

Hard cap on records per dataset. Offset pagination caps at 10,000 by OpenDataSoft.

## `maxItems` (type: `integer`):

Hard cap on records per run.

## `requestDelayMs` (type: `integer`):

Delay between requests. OpenDataSoft has no hard rate limit but is polite.

## Actor input object example

```json
{
  "mode": "records",
  "portal": "public.opendatasoft.com",
  "datasetId": "world-administrative-boundaries",
  "searchQuery": "",
  "where": "",
  "select": "",
  "groupBy": "",
  "orderBy": "",
  "refine": "",
  "limitPerPage": 100,
  "maxRecordsPerDataset": 10000,
  "maxItems": 10000,
  "requestDelayMs": 250
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("aurenic/opendatasoft-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("aurenic/opendatasoft-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call aurenic/opendatasoft-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,aurenic/opendatasoft-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kk6Je4frhjXLd8ogu/builds/MTDvz630ZiRqhUUXp/openapi.json
