Japan data.go.jp Scraper avatar

Japan data.go.jp Scraper

Pricing

from $1.50 / 1,000 results

Go to Apify Store
Japan data.go.jp Scraper

Japan data.go.jp Scraper

Extract datasets and CSV records from Japan's national open-data catalog (data.go.jp / data.e-gov.go.jp, CKAN v3). 18,000+ datasets across all ministries and municipalities. Shift_JIS aware. No API key, no browser, no proxy.

Pricing

from $1.50 / 1,000 results

Rating

0.0

(0)

Developer

Aurenic

Aurenic

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Extract datasets and CSV records from Japan's national open-data catalog (data.go.jp / data.e-gov.go.jp, CKAN v3). 18,000+ datasets across all ministries and municipalities. Shift_JIS aware. No API key, no browser, no proxy.

What does Japan data.go.jp Scraper do?

Scrape Japan's central government open-data catalog — the データカタログサイト operated by the Digital Agency (デジタル庁) — in five modes:

  • Search datasets — full CKAN package_search with filters for keyword, publisher, category, tag, and resource format. Auto-paginates up to the full catalog.
  • Dataset detail — full metadata for one dataset by ID, including every resource with its URL, format, size, and per-resource license.
  • Extract CSV records — download a dataset's CSV/TSV resources and return each row as structured JSON. Shift_JIS aware — many Japanese CSVs are Shift_JIS even when the HTTP header claims UTF-8.
  • List organizations — every ministry, agency, and municipality publishing data, with dataset counts.
  • List groups — all catalog categories with dataset counts.

The API is public, keyless, and unauthenticated. Data is published under 公共データ利用規約 第1.0版 (PDL1.0), which permits commercial reuse with attribution.

Output fields

Dataset

FieldDescription
datasetIdCKAN dataset slug
title / nameDataset title and slug
notesFull description
publisher / organizationIdPublishing organization
groupsCategories
tagsTags
licenseId / licenseTitle / licenseUrlDataset-level license
isOpenOpen-license flag
metadataCreated / metadataModifiedTimestamps
resourceCountNumber of resources
resourcesArray of { name, format, url, mimetype, size, licenseId, lastModified }
datasetUrlDeep link to data.go.jp

Organization

FieldDescription
id / name / titleOrganization identity
descriptionDescription
packageCountNumber of datasets published
urlDeep link

Record (CSV row)

FieldDescription
datasetId / resourceName / resourceUrlProvenance
licenseIdInherited from the resource
rowIndexRow position
…all CSV columnsKeyed by header

Diagnostic

Emitted when a dataset has no CSV/TSV resources, or when a run fails.

Who is it for?

  • Data journalists and researchers analyzing Japanese government statistics
  • Civic tech builders integrating Japanese open data into products
  • Market researchers extracting population, economic, and industry datasets
  • AI/ML teams sourcing Japanese-language structured data for training
  • Policy analysts monitoring datasets published by specific ministries
  • Data scientists combining e-Stat, ministry, and municipal datasets

Pricing

$1.50 per 1,000 results. No subscription.

ResultsCost
100$0.15
1,000$1.50
10,000$15.00

How to use it

  1. Pick a Mode.
  2. For search: enter a Search Query (Japanese works best) and/or set Organization, Group, Tags, or Format filters.
  3. For show/records: enter Dataset ID (from a search run) or a direct Resource URL.
  4. For records: pick an Encoding (default auto).
  5. Click Start.

Output example

{
"recordType": "dataset",
"datasetId": "mhlw_20260624_2024",
"title": "食中毒統計調査_令和6年食中毒統計調査_年次_2024年",
"notes": "食中毒統計調査は、食品衛生法に基づき...",
"publisher": "厚生労働省",
"organizationId": "org_1600",
"groups": [{ "name": "health", "title": "健康・医療" }],
"tags": [{ "name": "食中毒", "displayName": "食中毒" }],
"licenseId": "",
"isOpen": false,
"metadataCreated": "2026-06-24T02:15:33.000Z",
"metadataModified": "2026-06-24T02:15:33.000Z",
"resourceCount": 3,
"resources": [
{
"name": "第1表 食中毒事件一覧",
"format": "CSV",
"url": "https://www.e-stat.go.jp/.../file-download?...",
"mimetype": "text/csv",
"size": 24576,
"licenseId": "cc-by",
"lastModified": "2026-06-24T02:15:33.000Z"
}
],
"datasetUrl": "https://data.go.jp/data/dataset/mhlw_20260624_2024",
"scrapedAt": "2026-09-26T12:00:00.000Z"
}

Technical details

  • Source: CKAN Action API v3 at https://www.data.go.jp/data/api/3/action. Public, keyless, unauthenticated.
  • ~18,366 datasets across all Japanese central-government ministries and many municipalities.
  • Endpoints used: package_search, package_show, organization_list, group_list.
  • No published rate limit. The actor spaces requests ≥1.2s apart and backs off exponentially on 429/5xx.
  • Shift_JIS handling — CSV bodies are decoded via iconv-lite. Auto mode sniffs UTF-8 BOM, tries UTF-8, and falls back to Shift_JIS on replacement characters.
  • CSV parser — minimal RFC-4180 parser handles quoted fields, embedded commas, and multi-line values.
  • No browser, no proxy — pure REST API.

Known limits

  • Not a full-catalog dump. package_search requires at least one filter per CKAN convention — pass a query, organization, group, tags, or format.
  • Records mode only reads CSV/TSV resources. Datasets with only PDF, XLSX, or other formats emit a diagnostic. Add format: "CSV" in search mode to find CSV datasets.
  • Shift_JIS detection is heuristic. Auto mode may mis-detect malformed encodings. Force shift_jis or utf-8 if output has garbled characters.
  • License varies per dataset. The default is PDL1.0 (commercial use with attribution), but some datasets carry cc-by-nc, cc-by-nc-nd, or other-closed. Check licenseId and resources[].licenseId before commercial use.
  • Some resources are hosted off-site. CSV files often live on e-stat.go.jp or ministry servers. Those hosts may have their own rate limits.

FAQ

Do I need an API key? No. The CKAN API is public and unauthenticated.

Do I need a proxy? No. Datacenter IPs work.

Why do I get garbled Japanese characters? Force encoding: "shift_jis" in records mode. Many Japanese open-data CSVs are Shift_JIS despite claiming UTF-8 in headers.

How do I find a dataset ID? Run search mode first — the datasetId field is in every result.

What license applies? Default is PDL1.0 (公共データ利用規約 第1.0版), commercial use permitted with attribution 「出典:データカタログサイト(data.go.jp)」. Per-dataset licenses vary — check licenseId.

How do I export data? After a run, go to Storage → Export as JSON, CSV, Excel.

Support

Open an issue on the Actor's page for bugs or feature requests.