Data.gov Dataset Search API: US Government Open Data Catalog
Pricing
from $1.40 / 1,000 catalog row returneds
Data.gov Dataset Search API: US Government Open Data Catalog
Search the US federal open-data catalog by keyword, agency, government level, tag, theme or file format and get one row per dataset: title, publisher, description, licence, issue and harvest dates, popularity and every resource download URL. Dictionary modes list the agencies and the tags.
Pricing
from $1.40 / 1,000 catalog row returneds
Rating
0.0
(0)
Developer
Samat Makatov
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Search the US federal open-data catalog by keyword, agency, government level, tag, theme or file format and get one row per dataset: title, publisher, description, licence, issue and harvest dates, popularity and every resource download URL. Dictionary modes list the agencies and the tags.
The catalog at catalog.data.gov held 570,111 datasets on 29.09.2026 — federal departments plus state, county, city, university and tribal publishers. It stores the metadata: what a dataset is, who publishes it, when it was last refreshed and where the publisher's CSV, JSON, XML or GeoJSON files live. This actor turns a search over that catalog into rows you can filter, schedule and hand to a script, so the download link is a field instead of two clicks. No API key, no login, no proxy, no browser.
Use cases
- Find the data behind a question: search
medicare spendingorbroadbandwith Only datasets with downloadable files on and get the datasets plus the direct file links, instead of landing pages you have to open one by one. - Inventory one agency: put
nasa,census,epaorhhsinto Publishing organization and export everything that agency publishes, with tags, themes, licence and dates — the base table for a coverage audit or an internal catalog. - Work the local half of the catalog: city, county and state portals harvest into the same index. Level of government =
City Government+State Governmentturns the federal catalog into a municipal-data search. - Watch what is new: sort by
last harvested, set Harvested in the last N hours to 168 and switch on Only datasets not seen before — a weekly digest that pays only for entries that appeared since the last run. - Feed an agent or an MCP tool: a stable, keyless index of US government data. Use
fieldsto keep rows small, anddatasetmode when your agent already holds a catalog URL. - Build a correct filter first:
organizationsmode writes the 121 publishers with their slugs and levels,keywordsmode writes the tags with the number of datasets behind each one — reference tables, one request each.
Input
Every field is optional: the actor runs on its prefilled example with all defaults.
| Field | Type | Default | Allowed values / notes |
|---|---|---|---|
mode | string | search | search, dataset, organizations, keywords — see Modes |
query | string | (prefill air quality) | Full-text search over title, description and tags. Empty + sortBy: popularity = the most-used datasets of the whole catalog. search mode; in keywords mode it narrows the tag list |
slugs | string[] | (prefill ["air-quality","electric-vehicle-population-data"]) | dataset mode only: catalog slugs or whole page URLs, one request and one row each |
organizationSlug | string | (none) | One publisher slug, e.g. nasa — see Organizations |
organizationTypes | string[] | (none) | Levels of government, OR-combined — see Levels of government |
keywords | string[] | (none) | Catalog tags; a dataset must carry all of them. Matched without regard to case |
themes | string[] | (none) | Publisher themes, OR-combined, case-insensitive — see Themes |
publisher | string | (none) | Exact source portal, e.g. data.cityofnewyork.us |
spatialFilter | string | any | any, geospatial, non-geospatial |
onlyWithDownloads | boolean | false | Keep only datasets the catalog marks as having a downloadable file |
formats | string[] | (none) | File formats, at least one must be present — see File formats. Applied here, after the fetch |
sinceHours | integer | (none) | 1–8760. Keep datasets re-harvested inside this window (UTC). Applied here, after the fetch |
onlyNew | boolean | false | Write only datasets this input has not produced before — see Only new |
sortBy | string | relevance | relevance, popularity, lastHarvested |
includeResources | boolean | true | Add resources, the full file list. Costs no extra request |
includeRawDcat | boolean | false | Add dcat, the publisher's own metadata record |
maxDescriptionChars | integer | 1200 | 0–20000. 0 leaves description out of the row |
maxItems | integer | 25 | 1–5000. The real cost control of a run |
fields | string[] | (all) | Keep only these fields, in this order. rowType is always kept |
Reference
Modes
| Mode | Requests | One row is | Filters used |
|---|---|---|---|
search | 1 per 1000 rows | a dataset | all of them |
dataset | 1 per slug | a dataset, or a found: false marker | none — the slugs decide |
organizations | 1 | a publishing organization | none |
keywords | 1 | a tag with its dataset count | none; query narrows the list |
Filters the actor cannot do: there is no state or place filter (the catalog has a geography parameter, but it returned a national FEMA dataset and an Indiana model for California on 29.09.2026, so it is not exposed), and distance sorting needs a geometry this actor does not take.
Levels of government
Federal Government, State Government, City Government, County Government, University, Tribal, Non-Profit. Several values are OR-combined. Two of the 121 organizations carry no level at all; their rows keep organizationType: null.
Organizations
organizationSlug takes the catalog's own publisher slug, not a name. Run the actor once in organizations mode for all 121 with their dataset counts. Slugs seen on 29.09.2026:
- Federal:
census(296,099 datasets),noaa(87,848),doi(54,416),nasa,hhs,epa,va,doj,usda,energy,dot,dol,ed,treasury,hud,dhs,commerce,state,nist,ssa,americorps - States:
california,washington,maryland,new-york,connecticut,oregon,iowa,district-of-columbia - Counties and cities:
cook-county-il,lake-county-il,montgomery-county-md,nyc-ny,austin-tx,chicago-il,seattle-wa,san-francisco-ca,los-angeles-ca,louisville-ky,honolulu-hi,tempe-az - Other:
opentopography
A slug that does not exist matches nothing: the run succeeds with 0 rows and says so in its status message. It is never widened to something broader.
Themes
Themes are the publisher's own free text, matched case-insensitively (Education and education are the same filter). Seen live on 29.09.2026: Environment, Education, Health, Agriculture, Energy, Public Safety, Natural Resources & Environment, Management/Operations, geospatial, City Government. Many datasets carry no theme at all, so a theme filter is a narrow one.
Tags
Tags (keywords) are stricter than a text search — a dataset must carry every tag you list — and they are the publishers' own vocabulary, so check them in keywords mode before you rely on one. Be warned that the biggest tags of this catalog are geographic boilerplate, not topics: county or equivalent entity (277,724 datasets), united states (165,279), state fips code (155,354), county fips code (154,799), u.s. (152,622). Useful topical tags are much smaller: climate had 1,994 datasets.
File formats
The catalog has no format facet, so this filter runs on the rows that came back. A file's format is read from the format name the publisher declared, from its media type and from the extension of its download link, and normalized to one lowercase code, so text/csv, CSV and a .csv link all become csv. Codes seen in the catalog: csv, json, geojson, xml, rdf, xlsx, xls, zip, gz, pdf, html, txt, tsv, kml, kmz, netcdf, gpkg, gdb, shp, api, wms, tiff, jpeg, png, doc, docx, bin.
Media types that describe a request rather than a file (application/http, placeholder/value) and script endpoints (.cgi, .aspx) produce no code. A dataset whose publisher registered no file at all gets formats: [], resourceCount: 0 and hasDownloads: false — that is the source, not a gap in the row.
Only new
onlyNew remembers the catalog slugs a run delivered in a named key-value store and skips them next time. The memory is keyed by the filters of the input, so a daily climate schedule and a weekly nasa schedule keep separate memories and never swallow each other's rows. The first run writes everything it finds. Only rows the run really wrote are remembered, so a dataset cut off by maxItems comes back next time. onlyNew applies to search mode.
Examples
Most used datasets of the whole catalog — no search words, ranked by the catalog's own visit counter.
{ "mode": "search", "sortBy": "popularity", "maxItems": 20 }
Air-quality datasets that really end in a CSV file
{ "mode": "search", "query": "air quality", "onlyWithDownloads": true, "formats": ["csv"], "sortBy": "popularity", "maxItems": 15 }
Everything one federal agency publishes
{ "mode": "search", "organizationSlug": "nasa", "sortBy": "popularity", "maxItems": 20 }
City and state budget data inside the federal catalog
{ "mode": "search", "query": "budget", "organizationTypes": ["City Government", "State Government"], "sortBy": "popularity", "maxItems": 20 }
A weekly digest of newly harvested climate records — put this on a schedule.
{ "mode": "search", "query": "climate", "sortBy": "lastHarvested", "sinceHours": 168, "onlyNew": true, "maxItems": 50 }
Datasets carrying both tags, not just the words
{ "mode": "search", "keywords": ["finance", "budget"], "sortBy": "popularity", "maxItems": 15 }
Catalog URLs you already hold → full records
{ "mode": "dataset", "slugs": ["air-quality", "https://catalog.data.gov/dataset/electric-vehicle-population-data"], "maxItems": 10 }
The publisher dictionary
{ "mode": "organizations", "maxItems": 40 }
Output
One real row, from run oxsnimBahZGQh6QDu on 29.09.2026 (the air-quality + CSV example above):
{"rowType": "dataset","found": true,"slug": "air-quality","title": "Air Quality and Health Impacts","description": "Dataset contains information on New York City air quality surveillance data. Air pollution is one of the most important environmental threats to urban populations …","descriptionTruncated": false,"organizationName": "City of New York","organizationSlug": "nyc-ny","organizationType": "City Government","publisher": "data.cityofnewyork.us","identifier": "https://data.cityofnewyork.us/api/views/c3uy-2p5r","themes": ["Environment"],"keywords": ["2018od4a-video", "air quality", "climate", "dohmh", "health", "surveillance"],"license": null,"accessLevel": "public","issued": "2020-12-09","modified": "2026-06-18","landingPage": "https://data.cityofnewyork.us/d/c3uy-2p5r","lastHarvestedAt": "2026-09-10T18:31:43.436Z","hasDownloads": true,"hasSpatial": false,"latitude": null,"longitude": null,"popularity": 172,"resourceCount": 3,"formats": ["csv", "json", "xml"],"primaryDownloadUrl": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/query.json?accessType=DOWNLOAD","parentIdentifier": null,"harvestRecordUrl": "https://catalog.data.gov/harvest_record/92d8908a-e4ca-4043-bc66-0c2b02a1e623","url": "https://catalog.data.gov/dataset/air-quality","fetchedAt": "2026-09-29T21:42:07.309Z","resources": [{ "title": null, "format": "json", "mediaType": "application/json", "url": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/query.json?accessType=DOWNLOAD", "describedBy": "https://data.cityofnewyork.us/api/views/c3uy-2p5r/columns.json" },{ "title": null, "format": "xml", "mediaType": "application/xml", "url": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/query.xml?accessType=DOWNLOAD", "describedBy": "https://data.cityofnewyork.us/api/views/c3uy-2p5r/columns.xml" },{ "title": null, "format": "csv", "mediaType": "text/csv", "url": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/export.csv?accessType=DOWNLOAD", "describedBy": null }]}
Dataset rows
| Field | Type | Meaning |
|---|---|---|
rowType | string | dataset — always present, also when fields is set |
found | boolean | false on the marker row of a slug the catalog does not hold |
slug | string | Catalog slug, the last part of the catalog page URL |
title | string | Dataset title as the publisher wrote it |
description | string | Plain text, markup removed, cut at maxDescriptionChars |
descriptionTruncated | boolean | Something was cut off |
organizationName | string | Publishing organization, e.g. City of New York |
organizationSlug | string | Its catalog slug — the value for the organization filter |
organizationType | string | Level of government; null for two organizations |
publisher | string | The source portal the record was harvested from |
identifier | string | The publisher's own identifier for the dataset |
themes | string[] | Publisher themes; often empty |
keywords | string[] | Catalog tags |
license | string | Licence URL when the publisher gave one — about half do |
accessLevel | string | public, restricted public or non-public metadata |
issued | string | First publication, day precision, as the publisher wrote it |
modified | string | Last change at the publisher, day precision |
landingPage | string | The dataset's page on the publisher's own portal |
lastHarvestedAt | string | When the catalog last re-read the record, ISO 8601 UTC |
hasDownloads | boolean | The catalog marks at least one downloadable file |
hasSpatial | boolean | The record carries a geometry |
latitude / longitude | number | Centroid of that geometry, when there is one |
popularity | number | The catalog's own visit counter for the dataset page |
resourceCount | number | Number of files and endpoints in the record |
formats | string[] | Lowercase format codes, sorted and de-duplicated |
primaryDownloadUrl | string | Link of the first file — the one-field answer to "where is the data" |
parentIdentifier | string | Set on datasets that belong to a collection |
harvestRecordUrl | string | The catalog's harvest record for this dataset |
url | string | The catalog page of the dataset |
fetchedAt | string | Time of the request, ISO 8601 UTC |
resources | object[] | With includeResources: title, format, mediaType, url, describedBy per file |
dcat | object | With includeRawDcat: the publisher's own metadata record |
message | string | On a found: false row: what to check |
Organization and tag rows
| Field | Type | Meaning |
|---|---|---|
rowType | string | organization or keyword |
organizationSlug | string | The value the organization filter takes |
organizationName | string | Full name of the organization |
organizationType | string | Level of government |
organizationId | string | The catalog's internal id |
datasetCount | number | Datasets of this organization, or datasets behind this tag |
sourceCount | number | Source portals feeding this organization |
aliases | string[] | Alternative names the catalog knows |
keyword | string | The tag itself, on keyword rows |
url | string | A catalog search for this organization or tag |
fetchedAt | string | Time of the request, ISO 8601 UTC |
Dataset views: Datasets (title, organization, level, themes, tags, dates, formats, files, popularity, catalog page), Files and licence (title, organization, source portal, licence, formats, first file, publisher page), Organizations and tags (the two dictionary modes).
Dates come from two different places: issued and modified are the publisher's own, at day precision and exactly as written; lastHarvestedAt and fetchedAt are full UTC timestamps.
Use it from code / agents
curl -X POST "https://api.apify.com/v2/acts/yadroo~data-gov-datasets/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"query":"air quality","onlyWithDownloads":true,"formats":["csv"],"maxItems":15}'
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('yadroo/data-gov-datasets').call({ organizationSlug: 'nasa', sortBy: 'popularity', maxItems: 20 });const { items } = await client.dataset(run.defaultDatasetId).listItems();
from apify_client import ApifyClientclient = ApifyClient(os.environ["APIFY_TOKEN"])run = client.actor("yadroo/data-gov-datasets").call(run_input={"mode": "dataset", "slugs": ["air-quality"]})items = client.dataset(run["defaultDatasetId"]).list_items().items
Small rows for an LLM — five fields instead of thirty:
{ "query": "medicare spending", "maxItems": 10, "maxDescriptionChars": 300, "includeResources": false,"fields": ["title", "organizationName", "formats", "primaryDownloadUrl", "url"] }
MCP: add https://mcp.apify.com to Claude / Cursor / any MCP client and call the yadroo/data-gov-datasets tool with the same JSON input. When your agent already holds a catalog URL, dataset mode is the right call — it costs one request and returns the full record.
Pricing
Pay per event: $0.001 per run start + $0.002 per row at the FREE tier. The start event is charged on every run, including runs that end with no matching dataset. Apify plans above FREE pay less per event (−10 % on Bronze, −20 % on Silver, −30 % on Gold and above).
| Run | Rows | What you pay (FREE tier) |
|---|---|---|
The default search (maxItems: 25) | 25 | $0.001 + 25 × $0.002 = $0.051 |
| Three catalog slugs looked up | 3 | $0.001 + 3 × $0.002 = $0.007 |
| The 40 biggest publishers | 40 | $0.001 + 40 × $0.002 = $0.081 |
| A 500-row agency export | 500 | $0.001 + 500 × $0.002 = $1.001 |
maxItems decides the bill, not the filters. Two things are worth knowing: a formats list or a sinceHours window throws rows away after the request, so such a run can end below maxItems while still having paid the start event; and the dictionary modes charge per organization or tag written, one request in total.
Limits & FAQ
How many datasets matched my filters? The catalog's search API returns no total, so the actor cannot tell you — and does not invent a number. keywords mode gives the dataset count per tag, which is the closest thing to a count before a run.
Does it filter by state or city? Not by geography. organizationSlug and publisher get you the datasets of a named state or city portal (washington, nyc-ny, data.cityofnewyork.us); the catalog's own place parameter did not filter correctly on 29.09.2026 and is deliberately not exposed.
Are the download links checked? No. The catalog stores what the publisher last published, so a downloadURL can be stale or moved. Verifying every link would mean one request per file; the actor reports what the catalog holds.
Why is a run with many rows slow? The catalog's robots.txt asks for a 10-second crawl delay, and the actor honours it between consecutive requests. Rows come in pages of up to 1000, so every run up to 1000 rows makes a single request and waits not at all; a 5000-row run makes five requests with four pauses. dataset mode makes one request per slug, so ten slugs take about a minute and a half.
Why does my format filter return fewer rows than maxItems? Because the catalog has no format facet: the filter runs on the rows that came back. The actor fetches larger pages when you use it, and follows up to six extra pages, but a rare format over a narrow search can still come up short. The status message says how many rows were skipped and why.
Contact details. The source record carries the name and mailbox of the publishing office. The actor removes them from every row, including from the raw dcat record.
Empty results. A filter combination that matches nothing is a successful run with 0 rows and a status message saying so — never an error, and never silently widened. A dataset slug that does not exist becomes one row with found: false.
Source stability. This catalog was rebuilt on a new stack: its old CKAN-style action endpoints answer 404 as of 29.09.2026, and the current API is versioned 0.1.0, so parameters can still move. The actor fails with a clear message if the search stops returning a result list, and unit tests run against saved real responses so a change in shape breaks a test instead of a paying run.
Terms. The catalog's robots.txt disallows nothing and asks only for the crawl delay above; no term forbidding automated access was found on the site on 29.09.2026. The catalog metadata is the work of the US federal government. The files themselves belong to the publishing organizations — check license and the publisher's page before you redistribute data.
Made by Yadroo. Sibling actors: world-bank-indicators, federal-register-documents, ecfr-regulations, openfda-records, nasa-eonet-events.