# Japan data.go.jp Scraper (`aurenic/japan-datagojp-scraper`) Actor

Extract datasets and CSV records from Japan's national open-data catalog (data.go.jp / data.e-gov.go.jp, CKAN v3). 18,000+ datasets across all ministries and municipalities. Shift\_JIS aware. No API key, no browser, no proxy.

- **URL**: https://apify.com/aurenic/japan-datagojp-scraper.md
- **Developed by:** [Aurenic](https://apify.com/aurenic) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Japan data.go.jp Scraper

Extract datasets and CSV records from Japan's national open-data catalog (data.go.jp / data.e-gov.go.jp, CKAN v3). 18,000+ datasets across all ministries and municipalities. Shift\_JIS aware. No API key, no browser, no proxy.

### What does Japan data.go.jp Scraper do?

Scrape Japan's central government open-data catalog — the データカタログサイト operated by the Digital Agency (デジタル庁) — in five modes:

- **Search datasets** — full CKAN `package_search` with filters for keyword, publisher, category, tag, and resource format. Auto-paginates up to the full catalog.
- **Dataset detail** — full metadata for one dataset by ID, including every resource with its URL, format, size, and per-resource license.
- **Extract CSV records** — download a dataset's CSV/TSV resources and return each row as structured JSON. **Shift\_JIS aware** — many Japanese CSVs are Shift\_JIS even when the HTTP header claims UTF-8.
- **List organizations** — every ministry, agency, and municipality publishing data, with dataset counts.
- **List groups** — all catalog categories with dataset counts.

The API is **public, keyless, and unauthenticated**. Data is published under 公共データ利用規約 第1.0版 (PDL1.0), which permits commercial reuse with attribution.

### Output fields

#### Dataset

| Field | Description |
|---|---|
| datasetId | CKAN dataset slug |
| title / name | Dataset title and slug |
| notes | Full description |
| publisher / organizationId | Publishing organization |
| groups | Categories |
| tags | Tags |
| licenseId / licenseTitle / licenseUrl | Dataset-level license |
| isOpen | Open-license flag |
| metadataCreated / metadataModified | Timestamps |
| resourceCount | Number of resources |
| resources | Array of `{ name, format, url, mimetype, size, licenseId, lastModified }` |
| datasetUrl | Deep link to data.go.jp |

#### Organization

| Field | Description |
|---|---|
| id / name / title | Organization identity |
| description | Description |
| packageCount | Number of datasets published |
| url | Deep link |

#### Record (CSV row)

| Field | Description |
|---|---|
| datasetId / resourceName / resourceUrl | Provenance |
| licenseId | Inherited from the resource |
| rowIndex | Row position |
| …all CSV columns | Keyed by header |

#### Diagnostic

Emitted when a dataset has no CSV/TSV resources, or when a run fails.

### Who is it for?

- **Data journalists and researchers** analyzing Japanese government statistics
- **Civic tech builders** integrating Japanese open data into products
- **Market researchers** extracting population, economic, and industry datasets
- **AI/ML teams** sourcing Japanese-language structured data for training
- **Policy analysts** monitoring datasets published by specific ministries
- **Data scientists** combining e-Stat, ministry, and municipal datasets

### Pricing

**$1.50 per 1,000 results.** No subscription.

| Results | Cost |
|---|---|
| 100 | $0.15 |
| 1,000 | $1.50 |
| 10,000 | $15.00 |

### How to use it

1. Pick a **Mode**.
2. For search: enter a **Search Query** (Japanese works best) and/or set **Organization**, **Group**, **Tags**, or **Format** filters.
3. For show/records: enter **Dataset ID** (from a search run) or a direct **Resource URL**.
4. For records: pick an **Encoding** (default auto).
5. Click **Start**.

### Output example

```json
{
  "recordType": "dataset",
  "datasetId": "mhlw_20260624_2024",
  "title": "食中毒統計調査＿令和６年食中毒統計調査＿年次＿2024年",
  "notes": "食中毒統計調査は、食品衛生法に基づき...",
  "publisher": "厚生労働省",
  "organizationId": "org_1600",
  "groups": [{ "name": "health", "title": "健康・医療" }],
  "tags": [{ "name": "食中毒", "displayName": "食中毒" }],
  "licenseId": "",
  "isOpen": false,
  "metadataCreated": "2026-06-24T02:15:33.000Z",
  "metadataModified": "2026-06-24T02:15:33.000Z",
  "resourceCount": 3,
  "resources": [
    {
      "name": "第1表 食中毒事件一覧",
      "format": "CSV",
      "url": "https://www.e-stat.go.jp/.../file-download?...",
      "mimetype": "text/csv",
      "size": 24576,
      "licenseId": "cc-by",
      "lastModified": "2026-06-24T02:15:33.000Z"
    }
  ],
  "datasetUrl": "https://data.go.jp/data/dataset/mhlw_20260624_2024",
  "scrapedAt": "2026-09-26T12:00:00.000Z"
}
```

### Technical details

- **Source: CKAN Action API v3** at `https://www.data.go.jp/data/api/3/action`. Public, keyless, unauthenticated.
- **~18,366 datasets** across all Japanese central-government ministries and many municipalities.
- **Endpoints used:** `package_search`, `package_show`, `organization_list`, `group_list`.
- **No published rate limit.** The actor spaces requests ≥1.2s apart and backs off exponentially on 429/5xx.
- **Shift\_JIS handling** — CSV bodies are decoded via `iconv-lite`. Auto mode sniffs UTF-8 BOM, tries UTF-8, and falls back to Shift\_JIS on replacement characters.
- **CSV parser** — minimal RFC-4180 parser handles quoted fields, embedded commas, and multi-line values.
- **No browser, no proxy** — pure REST API.

### Known limits

- **Not a full-catalog dump.** `package_search` requires at least one filter per CKAN convention — pass a query, organization, group, tags, or format.
- **Records mode only reads CSV/TSV resources.** Datasets with only PDF, XLSX, or other formats emit a diagnostic. Add `format: "CSV"` in search mode to find CSV datasets.
- **Shift\_JIS detection is heuristic.** Auto mode may mis-detect malformed encodings. Force `shift_jis` or `utf-8` if output has garbled characters.
- **License varies per dataset.** The default is PDL1.0 (commercial use with attribution), but some datasets carry `cc-by-nc`, `cc-by-nc-nd`, or `other-closed`. Check `licenseId` and `resources[].licenseId` before commercial use.
- **Some resources are hosted off-site.** CSV files often live on `e-stat.go.jp` or ministry servers. Those hosts may have their own rate limits.

### FAQ

**Do I need an API key?** No. The CKAN API is public and unauthenticated.

**Do I need a proxy?** No. Datacenter IPs work.

**Why do I get garbled Japanese characters?** Force `encoding: "shift_jis"` in records mode. Many Japanese open-data CSVs are Shift\_JIS despite claiming UTF-8 in headers.

**How do I find a dataset ID?** Run search mode first — the `datasetId` field is in every result.

**What license applies?** Default is PDL1.0 (公共データ利用規約 第1.0版), commercial use permitted with attribution 「出典：データカタログサイト（data.go.jp）」. Per-dataset licenses vary — check `licenseId`.

**How do I export data?** After a run, go to Storage → Export as JSON, CSV, Excel.

### Support

Open an issue on the Actor's page for bugs or feature requests.

# Actor input Schema

## `mode` (type: `string`):

What to fetch.

## `query` (type: `string`):

Free-text search in Japanese works best (e.g. 人口, 医療機関, 気象).

## `organization` (type: `string`):

Restrict to a publisher (e.g. org\_1600 for 厚生労働省, org\_1100 for 総務省). Use organizations mode to list all.

## `group` (type: `string`):

Restrict to a category (use groups mode to list all).

## `tags` (type: `string`):

Restrict to a specific tag.

## `format` (type: `string`):

Restrict to datasets containing a resource in this format (CSV, XLSX, PDF, JSON).

## `datasetId` (type: `string`):

Dataset name/slug (e.g. mhlw\_20260624\_2024). Used in show and records modes. Take it from a search run.

## `resourceUrl` (type: `string`):

Fetch a specific CSV resource directly. Used in records mode as an alternative to datasetId.

## `encoding` (type: `string`):

Character encoding of CSV bodies. Japanese open-data CSVs are frequently Shift\_JIS even when the HTTP header claims UTF-8.

## `maxItems` (type: `integer`):

Hard cap on records per run.

## `requestDelayMs` (type: `integer`):

Delay between requests. No published rate limit; default 1200ms matches the CKAN harvesting convention.

## Actor input object example

```json
{
  "mode": "search",
  "query": "人口",
  "organization": "",
  "group": "",
  "tags": "",
  "format": "",
  "datasetId": "",
  "resourceUrl": "",
  "encoding": "auto",
  "maxItems": 1000,
  "requestDelayMs": 1200
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("aurenic/japan-datagojp-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("aurenic/japan-datagojp-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call aurenic/japan-datagojp-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,aurenic/japan-datagojp-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yEXsyqB32f1zIXY5m/builds/skl0nHh2AaBXkR8xS/openapi.json
