Japan data.go.jp Scraper
Pricing
from $1.50 / 1,000 results
Japan data.go.jp Scraper
Extract datasets and CSV records from Japan's national open-data catalog (data.go.jp / data.e-gov.go.jp, CKAN v3). 18,000+ datasets across all ministries and municipalities. Shift_JIS aware. No API key, no browser, no proxy.
Pricing
from $1.50 / 1,000 results
Rating
0.0
(0)
Developer
Aurenic
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Extract datasets and CSV records from Japan's national open-data catalog (data.go.jp / data.e-gov.go.jp, CKAN v3). 18,000+ datasets across all ministries and municipalities. Shift_JIS aware. No API key, no browser, no proxy.
What does Japan data.go.jp Scraper do?
Scrape Japan's central government open-data catalog — the データカタログサイト operated by the Digital Agency (デジタル庁) — in five modes:
- Search datasets — full CKAN
package_searchwith filters for keyword, publisher, category, tag, and resource format. Auto-paginates up to the full catalog. - Dataset detail — full metadata for one dataset by ID, including every resource with its URL, format, size, and per-resource license.
- Extract CSV records — download a dataset's CSV/TSV resources and return each row as structured JSON. Shift_JIS aware — many Japanese CSVs are Shift_JIS even when the HTTP header claims UTF-8.
- List organizations — every ministry, agency, and municipality publishing data, with dataset counts.
- List groups — all catalog categories with dataset counts.
The API is public, keyless, and unauthenticated. Data is published under 公共データ利用規約 第1.0版 (PDL1.0), which permits commercial reuse with attribution.
Output fields
Dataset
| Field | Description |
|---|---|
| datasetId | CKAN dataset slug |
| title / name | Dataset title and slug |
| notes | Full description |
| publisher / organizationId | Publishing organization |
| groups | Categories |
| tags | Tags |
| licenseId / licenseTitle / licenseUrl | Dataset-level license |
| isOpen | Open-license flag |
| metadataCreated / metadataModified | Timestamps |
| resourceCount | Number of resources |
| resources | Array of { name, format, url, mimetype, size, licenseId, lastModified } |
| datasetUrl | Deep link to data.go.jp |
Organization
| Field | Description |
|---|---|
| id / name / title | Organization identity |
| description | Description |
| packageCount | Number of datasets published |
| url | Deep link |
Record (CSV row)
| Field | Description |
|---|---|
| datasetId / resourceName / resourceUrl | Provenance |
| licenseId | Inherited from the resource |
| rowIndex | Row position |
| …all CSV columns | Keyed by header |
Diagnostic
Emitted when a dataset has no CSV/TSV resources, or when a run fails.
Who is it for?
- Data journalists and researchers analyzing Japanese government statistics
- Civic tech builders integrating Japanese open data into products
- Market researchers extracting population, economic, and industry datasets
- AI/ML teams sourcing Japanese-language structured data for training
- Policy analysts monitoring datasets published by specific ministries
- Data scientists combining e-Stat, ministry, and municipal datasets
Pricing
$1.50 per 1,000 results. No subscription.
| Results | Cost |
|---|---|
| 100 | $0.15 |
| 1,000 | $1.50 |
| 10,000 | $15.00 |
How to use it
- Pick a Mode.
- For search: enter a Search Query (Japanese works best) and/or set Organization, Group, Tags, or Format filters.
- For show/records: enter Dataset ID (from a search run) or a direct Resource URL.
- For records: pick an Encoding (default auto).
- Click Start.
Output example
{"recordType": "dataset","datasetId": "mhlw_20260624_2024","title": "食中毒統計調査_令和6年食中毒統計調査_年次_2024年","notes": "食中毒統計調査は、食品衛生法に基づき...","publisher": "厚生労働省","organizationId": "org_1600","groups": [{ "name": "health", "title": "健康・医療" }],"tags": [{ "name": "食中毒", "displayName": "食中毒" }],"licenseId": "","isOpen": false,"metadataCreated": "2026-06-24T02:15:33.000Z","metadataModified": "2026-06-24T02:15:33.000Z","resourceCount": 3,"resources": [{"name": "第1表 食中毒事件一覧","format": "CSV","url": "https://www.e-stat.go.jp/.../file-download?...","mimetype": "text/csv","size": 24576,"licenseId": "cc-by","lastModified": "2026-06-24T02:15:33.000Z"}],"datasetUrl": "https://data.go.jp/data/dataset/mhlw_20260624_2024","scrapedAt": "2026-09-26T12:00:00.000Z"}
Technical details
- Source: CKAN Action API v3 at
https://www.data.go.jp/data/api/3/action. Public, keyless, unauthenticated. - ~18,366 datasets across all Japanese central-government ministries and many municipalities.
- Endpoints used:
package_search,package_show,organization_list,group_list. - No published rate limit. The actor spaces requests ≥1.2s apart and backs off exponentially on 429/5xx.
- Shift_JIS handling — CSV bodies are decoded via
iconv-lite. Auto mode sniffs UTF-8 BOM, tries UTF-8, and falls back to Shift_JIS on replacement characters. - CSV parser — minimal RFC-4180 parser handles quoted fields, embedded commas, and multi-line values.
- No browser, no proxy — pure REST API.
Known limits
- Not a full-catalog dump.
package_searchrequires at least one filter per CKAN convention — pass a query, organization, group, tags, or format. - Records mode only reads CSV/TSV resources. Datasets with only PDF, XLSX, or other formats emit a diagnostic. Add
format: "CSV"in search mode to find CSV datasets. - Shift_JIS detection is heuristic. Auto mode may mis-detect malformed encodings. Force
shift_jisorutf-8if output has garbled characters. - License varies per dataset. The default is PDL1.0 (commercial use with attribution), but some datasets carry
cc-by-nc,cc-by-nc-nd, orother-closed. ChecklicenseIdandresources[].licenseIdbefore commercial use. - Some resources are hosted off-site. CSV files often live on
e-stat.go.jpor ministry servers. Those hosts may have their own rate limits.
FAQ
Do I need an API key? No. The CKAN API is public and unauthenticated.
Do I need a proxy? No. Datacenter IPs work.
Why do I get garbled Japanese characters? Force encoding: "shift_jis" in records mode. Many Japanese open-data CSVs are Shift_JIS despite claiming UTF-8 in headers.
How do I find a dataset ID? Run search mode first — the datasetId field is in every result.
What license applies? Default is PDL1.0 (公共データ利用規約 第1.0版), commercial use permitted with attribution 「出典:データカタログサイト(data.go.jp)」. Per-dataset licenses vary — check licenseId.
How do I export data? After a run, go to Storage → Export as JSON, CSV, Excel.
Support
Open an issue on the Actor's page for bugs or feature requests.