# Japan data.go.jp Open Data Catalog (CKAN API) (`jpopendata/japan-datagojp-catalog`) Actor

Unofficial client for Japan's government open-data catalog data.go.jp (CKAN API v3, no key). Search datasets across all ministries and municipalities, or download a dataset's CSV records as structured JSON (Shift\_JIS aware). Per-dataset license shown via licenseId; default PDL1.0.

- **URL**: https://apify.com/jpopendata/japan-datagojp-catalog.md
- **Developed by:** [JP Open Data](https://apify.com/jpopendata) (community)
- **Categories:**
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 record scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Japan data.go.jp Open Data Catalog (CKAN API)

**Search Japan's government open-data catalog (data.go.jp) across every ministry and municipality by keyword, publisher, category or format — then actually download a dataset's CSV rows as clean, structured JSON (Shift\_JIS handled automatically).**

This Actor is a thin, polite client for the **CKAN API v3** (`data.e-gov.go.jp/data/api/3/action/`) behind the データカタログサイト (data.go.jp), the nationwide open-data catalog operated by the **Digital Agency (デジタル庁)** — the legacy `data.go.jp` host now redirects to the e-Gov data portal that serves it. It is **not a scraper** — it calls the CKAN JSON API, needs **no API key**, and returns each result as a flat record with a legal envelope (`source` / `sourceUrl` / `license` / `retrievedAt`) on every item.

Unlike a single-city catalog, this covers **all central-government ministries and local governments in one search** (~18,000 datasets). The catalogue as a whole is a directory, not a product. The value here is `get_records`: point it at a dataset and it fetches the actual CSV distribution and returns its rows — so you get the *data*, not just a link to it.

> **Unofficial tool.** Independently built and maintained. **Not affiliated with, endorsed by, or connected to the Japan Digital Agency** or any publishing ministry/municipality. Catalogue content is provided by default under the **公共データ利用規約 第1.0版 (PDL1.0)**, which permits commercial reuse with attribution; the required 出典 statement **「出典：データカタログサイト（data.go.jp）」** is embedded in the `license` of every record.

***

### Quick start — verified input

Copy, paste, run. This exact input is verified on the platform (SUCCEEDED, items > 0):

```json
{
  "mode": "search_datasets",
  "query": "人口"
}
```

Running with **no input at all** also works (same defaults). To download a dataset's rows, take a `datasetId` from the search output and run `{"mode": "get_records", "datasetId": "<that id>"}` (only CSV/TSV resources are read — add `"format": "CSV"` to the search to find them).

### No API key required

The data.go.jp CKAN API is public and unauthenticated. Just pick a mode and run. The Actor keeps a single connection, spaces requests ≥ 1.2 s apart, backs off exponentially on HTTP 429/5xx, and **fails visibly on a persistent block — it never attempts rate-limit evasion.**

### License, per-dataset diversity & source-display obligation (出典記載義務)

The 利用規約 states (原文): 「本サイトで掲載・発信している情報……の著作権は、特記されていない限りデジタル庁に帰属し、権利表記の記載がない限り「**公共データ利用規約（第1.0版）**」（PDL1.0）が適用されています。」「本コンテンツを利用する際は出典を記載してください。」 — i.e. the default license is **PDL1.0** (CC BY 4.0-compatible), commercial use permitted, with attribution.

- The credit **「出典：データカタログサイト（data.go.jp）」** is embedded verbatim in the `license` field of every record. If you display or republish the data, keep that credit visible.
- **Per-dataset licenses vary and are your responsibility.** The terms explicitly carve out content that is 「特記されている」 / carries its own 権利表記. This catalogue's CKAN license vocabulary includes not just open licenses (`cc-by`, `cc-by-4.0`, `cc-by-sa`, `cc-zero`, `pdl-1.0`, `gjstu`) but also **non-commercial** ones (`cc-by-nc`, `cc-by-nc-sa`, `cc-by-nc-nd`) and **non-open** ones (`other-closed`). At data.go.jp the real per-dataset signal usually lives on each **resource**, so the Actor surfaces:

  - `licenseId` — the dataset-level CKAN license id (often empty here).
  - `resourceLicenseIds` — the distinct license ids across the dataset's resources.
  - `resources[].licenseId` — each individual resource's own license id.
  - In `get_records`, every output row inherits its source resource's `licenseId`.

  **You must confirm and comply with each dataset's own license and content — including whether it is non-commercial or non-open, and whether it contains personal data — before use.**

### Two modes

#### 1. `search_datasets` — find datasets (CKAN `package_search`)

Give a `query` (free text) and/or an `organization`, `tags`, `group`, or `format` filter, and get back matching datasets: `datasetId`, title, publisher (ministry/municipality), update date, license ids, a catalogue deep link, and the full `resources[]` list (each resource's URL, format, size and license). At least one filter is required (a targeted search, not a full-catalogue dump).

#### 2. `get_records` — download a dataset's CSV rows

Give a `datasetId` (from a `search_datasets` run) and the Actor resolves that dataset's CSV/TSV resources, downloads them, and returns **one record per CSV row** — the columns keyed by header, plus provenance (`datasetId`, `resourceName`, `resourceUrl`, `licenseId`, `rowIndex`). Or pass a `resourceUrl` directly. Japanese open-data CSVs are frequently **Shift\_JIS** (confirmed live: an e-stat resource served `charset=UTF-8` in its header yet was Shift\_JIS on the wire); `encoding: "auto"` (default) sniffs UTF-8 (incl. BOM) then falls back to Shift\_JIS, or force `utf-8` / `shift_jis`.

### Sample output

**search\_datasets** (dataset item):

```json
{
  "datasetId": "mhlw_20260624_2024",
  "title": "食中毒統計調査＿令和６年食中毒統計調査＿年次＿2024年",
  "organization": "厚生労働省",
  "organizationId": "org_1600",
  "publisher": "厚生労働省",
  "licenseId": null,
  "resourceLicenseIds": ["cc-by"],
  "isOpen": false,
  "lastModified": "2026-06-24T...",
  "datasetUrl": "https://data.e-gov.go.jp/data/dataset/mhlw_20260624_2024",
  "resources": [
    { "name": "第1表 ...", "format": "CSV", "url": "https://www.e-stat.go.jp/.../file-download?...", "licenseId": "cc-by" }
  ],
  "source": "データカタログサイト（data.go.jp）/ Japan Data Catalog site (data.go.jp, e-Govデータポータル, CKAN)",
  "sourceUrl": "https://www.data.go.jp/",
  "license": "出典：データカタログサイト（data.go.jp）… PDL1.0 … licenseId … Unofficial …",
  "retrievedAt": "2026-08-26T12:00:00.000Z"
}
```

**get\_records** (dataset item — one CSV row):

```json
{
  "datasetId": "mhlw_20260624_2024",
  "datasetTitle": "食中毒統計調査 ...",
  "resourceName": "第1表 ...",
  "resourceUrl": "https://www.e-stat.go.jp/.../file-download?...",
  "licenseId": "cc-by",
  "rowIndex": 1,
  "fields": { "都道府県名等": "全国", "...": "..." },
  "source": "データカタログサイト（data.go.jp）...",
  "sourceUrl": "https://www.data.go.jp/",
  "license": "出典：データカタログサイト（data.go.jp）… PDL1.0 … Unofficial …",
  "retrievedAt": "2026-08-26T12:00:00.000Z"
}
```

### Input reference

| Field | Mode | Description |
| --- | --- | --- |
| `mode` | both | `search_datasets` (default) or `get_records`. Case-insensitive; if you pass only a `datasetId`/`resourceUrl` the Actor switches to `get_records` for you. |
| `query` | search | Free-text query (CKAN `q`), Japanese works best, e.g. `人口`, `医療機関`. |
| `organization` | search | CKAN organization name, e.g. `org_1600` (厚生労働省). |
| `tags` | search | Restrict to a CKAN tag. |
| `group` | search | Category/group name (see the `groups` field of search results). |
| `format` | search | Resource format filter, e.g. `CSV`, `XLSX`, `PDF`, `JSON`. |
| `datasetId` | get\_records | Dataset name/slug to read (from the search output's `datasetId`). |
| `resourceUrl` | get\_records | Fetch a specific CSV resource URL directly. |
| `encoding` | get\_records | `auto` (default), `utf-8`, or `shift_jis` (`sjis`, `cp932` … accepted). |
| `maxItems` | both | Max records to output, 1–100000 (default 1000). |
| `maxApiRequests` | both | Hard per-run upstream request cap, 1–25 (default 10). |
| `proxyConfiguration` | both | Optional Apify proxy; default is a direct connection. |

At least one of `query` / `organization` / `tags` / `group` / `format` is required in `search_datasets` mode; `datasetId` or `resourceUrl` is required in `get_records` mode. Values are validated **before** the first request; an invalid value fails the run immediately with a message that lists the valid values.

#### Common input mistakes

| Mistake | Correct |
|---------|---------|
| `{"mode": "search_datasets"}` with no filter | add `"query": "人口"` (or `organization` / `tags` / `group` / `format`) |
| `{"mode": "search_datasets", "datasetId": "mhlw_20260624_2024"}` | `"mode": "get_records"` — datasetId is a get\_records field |
| `{"mode": "get_records", "query": "人口"}` | get\_records needs `datasetId` or `resourceUrl`, not a query |
| `"datasetId": "https://data.go.jp/data/dataset/mhlw_20260624_2024"` | just the slug: `"mhlw_20260624_2024"` |
| `"organization": "厚生労働省"` | the CKAN name: `"org_1600"` (find it in the `organizationId` field of a search) |
| `"encoding": "euc-jp"` | `"auto"`, `"utf-8"` or `"shift_jis"` |

#### Empty results?

A run that finds nothing completes with 0 items and a warning in the log (not a failure). Typical causes: an English `query` (the catalogue is Japanese — try `人口` instead of `population`), a `tags`/`group`/`organization` value that is not an exact CKAN name, or a `get_records` dataset whose resources are all XLSX/PDF (only CSV/TSV are read — search with `"format": "CSV"` or pass a `resourceUrl` to a CSV). Broaden the query or drop a filter and retry.

### Notes & limits

- `get_records` reads **CSV/TSV** resources. Non-CSV distributions (XLSX, PDF, HTML, etc.) are skipped — pass a `resourceUrl` to a CSV, or read the other format yourself. Catalogue-wide, PDF/HTML/XLS dominate and CSV is a minority, so use the `format: "CSV"` search filter to find machine-readable datasets.
- Resource files are hosted on the publishers' own servers (ministry / e-Stat / municipal domains). A resource that responds with a redirect surfaces as an error (the Actor does not follow redirects to keep the request budget honest) — re-run with the final URL.
- `search_datasets` pages at up to 1000 datasets per request under `maxItems`.

### Disclaimer

1. **Unofficial** — not affiliated with, endorsed by, or connected to the Japan Digital Agency or any publishing ministry / municipality. This tool uses the public data.go.jp CKAN API.
2. **Public data, your responsibility** — content is provided by default under PDL1.0 (公共データ利用規約 第1.0版). Keep the 出典 credit visible. **Each dataset's own license (`licenseId` / `resourceLicenseIds`) — which may be non-commercial or non-open — and its content, including any personal data it may contain, are the user's responsibility to verify and comply with before use.** Compliance with the source's terms, and applicable law (incl. GDPR / Japan's APPI where relevant), is your responsibility.
3. **Polite by design** — one connection, ≥ 1.2 s spacing, exponential backoff, a hard request budget, and a visible failure on a persistent block. No rate-limit evasion.
4. **No warranty** — provided "as is"; verify anything important against the source catalogue.

### Search terms

japan government data catalog, data.go.jp api, japan open data search english, japan open data csv, CKAN japan, japanese government datasets, e-gov data portal, japan ministry open data, digital agency japan data, japan municipal open data, 政府 オープンデータ, data.go.jp csv download

# Actor input Schema

## `mode` (type: `string`):

`search_datasets` (default) finds datasets on data.go.jp (CKAN package\_search) and returns their metadata incl. resource URLs and per-resource license ids — set at least one of query / organization / tags / group / format. `get_records` downloads a dataset's CSV resource(s) and returns the rows as structured JSON — set datasetId or resourceUrl. Example: "search\_datasets".

## `query` (type: `string`):

search\_datasets mode: free-text query over dataset titles/descriptions (CKAN `q`), Japanese works best, e.g. "人口" (population), "医療機関", "気象". At least one of query / organization / tags / group / format is required in search\_datasets mode.

## `organization` (type: `string`):

search\_datasets mode (optional): restrict to a publisher by its CKAN organization name, e.g. "org\_1600" (厚生労働省 / MHLW) or "org\_1100" (総務省 / MIC). Find names in a search\_datasets run (the `organizationId` field).

## `tags` (type: `string`):

search\_datasets mode (optional): restrict to datasets carrying this CKAN tag (exact tag text).

## `group` (type: `string`):

search\_datasets mode (optional): restrict to a category by its CKAN group name (find group names in a search\_datasets run's `groups` field).

## `format` (type: `string`):

search\_datasets mode (optional): restrict to datasets that have a resource in this format, e.g. "CSV", "XLSX", "PDF", "JSON" (case-insensitive). Useful to find datasets you can then pull with get\_records.

## `datasetId` (type: `string`):

get\_records mode (REQUIRED there unless resourceUrl is given): the CKAN dataset name/slug to read, e.g. "mhlw\_20260624\_2024" — take it from a search\_datasets run (the `datasetId` field). The Actor resolves the dataset's CSV/TSV resources and downloads their rows.

## `resourceUrl` (type: `string`):

get\_records mode (alternative to datasetId): fetch this specific CSV resource URL directly (an absolute http(s) URL from a dataset's `resources[].url`).

## `encoding` (type: `string`):

get\_records mode: character encoding of the CSV body. `auto` (default) sniffs UTF-8 (incl. BOM) then falls back to Shift\_JIS — many Japanese open-data CSVs are Shift\_JIS even when the HTTP header claims UTF-8. Spellings such as "utf8", "Shift-JIS", "sjis", "cp932" are accepted.

## `maxItems` (type: `integer`):

Maximum number of records (datasets or CSV rows) to output, 1-100000. PPE charges per record. Example: 1000.

## `maxApiRequests` (type: `integer`):

Hard safety cap on upstream requests per run, 1-25 (search pages 1000 datasets per request; get\_records may fetch several resources). Politeness (1 connection, >= 1.2 s spacing, exponential backoff on 429/5xx) is enforced in code. Example: 10.

## `proxyConfiguration` (type: `object`):

Apify proxy settings. Default is NO proxy (direct connection) — a public API rarely needs one. The Actor backs off exponentially on 429/5xx and fails visibly on a persistent block; it never attempts rate-limit evasion.

## Actor input object example

```json
{
  "mode": "search_datasets",
  "query": "人口",
  "organization": "org_1600",
  "format": "CSV",
  "datasetId": "mhlw_20260624_2024",
  "encoding": "auto",
  "maxItems": 1000,
  "maxApiRequests": 10,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `records` (type: `string`):

Dataset metadata records (search\_datasets) or CSV-row records (get\_records), with source attribution (source, sourceUrl, license, retrievedAt) on every item. Per-dataset license is surfaced via licenseId and resourceLicenseIds.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "search_datasets",
    "query": "人口",
    "encoding": "auto",
    "maxItems": 1000,
    "maxApiRequests": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("jpopendata/japan-datagojp-catalog").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "search_datasets",
    "query": "人口",
    "encoding": "auto",
    "maxItems": 1000,
    "maxApiRequests": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("jpopendata/japan-datagojp-catalog").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "search_datasets",
  "query": "人口",
  "encoding": "auto",
  "maxItems": 1000,
  "maxApiRequests": 10
}' |
apify call jpopendata/japan-datagojp-catalog --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jpopendata/japan-datagojp-catalog"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/6XR6FhSYnVfS2Md8h/builds/2JNjGA2bTQPuopQ1r/openapi.json
