Data.gov Dataset Search API: US Government Open Data Catalog avatar

Data.gov Dataset Search API: US Government Open Data Catalog

Pricing

from $1.40 / 1,000 catalog row returneds

Go to Apify Store
Data.gov Dataset Search API: US Government Open Data Catalog

Data.gov Dataset Search API: US Government Open Data Catalog

Search the US federal open-data catalog by keyword, agency, government level, tag, theme or file format and get one row per dataset: title, publisher, description, licence, issue and harvest dates, popularity and every resource download URL. Dictionary modes list the agencies and the tags.

Pricing

from $1.40 / 1,000 catalog row returneds

Rating

0.0

(0)

Developer

Samat Makatov

Samat Makatov

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Search the US federal open-data catalog by keyword, agency, government level, tag, theme or file format and get one row per dataset: title, publisher, description, licence, issue and harvest dates, popularity and every resource download URL. Dictionary modes list the agencies and the tags.

The catalog at catalog.data.gov held 570,111 datasets on 29.09.2026 — federal departments plus state, county, city, university and tribal publishers. It stores the metadata: what a dataset is, who publishes it, when it was last refreshed and where the publisher's CSV, JSON, XML or GeoJSON files live. This actor turns a search over that catalog into rows you can filter, schedule and hand to a script, so the download link is a field instead of two clicks. No API key, no login, no proxy, no browser.

Use cases

  • Find the data behind a question: search medicare spending or broadband with Only datasets with downloadable files on and get the datasets plus the direct file links, instead of landing pages you have to open one by one.
  • Inventory one agency: put nasa, census, epa or hhs into Publishing organization and export everything that agency publishes, with tags, themes, licence and dates — the base table for a coverage audit or an internal catalog.
  • Work the local half of the catalog: city, county and state portals harvest into the same index. Level of government = City Government + State Government turns the federal catalog into a municipal-data search.
  • Watch what is new: sort by last harvested, set Harvested in the last N hours to 168 and switch on Only datasets not seen before — a weekly digest that pays only for entries that appeared since the last run.
  • Feed an agent or an MCP tool: a stable, keyless index of US government data. Use fields to keep rows small, and dataset mode when your agent already holds a catalog URL.
  • Build a correct filter first: organizations mode writes the 121 publishers with their slugs and levels, keywords mode writes the tags with the number of datasets behind each one — reference tables, one request each.

Input

Every field is optional: the actor runs on its prefilled example with all defaults.

FieldTypeDefaultAllowed values / notes
modestringsearchsearch, dataset, organizations, keywords — see Modes
querystring(prefill air quality)Full-text search over title, description and tags. Empty + sortBy: popularity = the most-used datasets of the whole catalog. search mode; in keywords mode it narrows the tag list
slugsstring[](prefill ["air-quality","electric-vehicle-population-data"])dataset mode only: catalog slugs or whole page URLs, one request and one row each
organizationSlugstring(none)One publisher slug, e.g. nasa — see Organizations
organizationTypesstring[](none)Levels of government, OR-combined — see Levels of government
keywordsstring[](none)Catalog tags; a dataset must carry all of them. Matched without regard to case
themesstring[](none)Publisher themes, OR-combined, case-insensitive — see Themes
publisherstring(none)Exact source portal, e.g. data.cityofnewyork.us
spatialFilterstringanyany, geospatial, non-geospatial
onlyWithDownloadsbooleanfalseKeep only datasets the catalog marks as having a downloadable file
formatsstring[](none)File formats, at least one must be present — see File formats. Applied here, after the fetch
sinceHoursinteger(none)1–8760. Keep datasets re-harvested inside this window (UTC). Applied here, after the fetch
onlyNewbooleanfalseWrite only datasets this input has not produced before — see Only new
sortBystringrelevancerelevance, popularity, lastHarvested
includeResourcesbooleantrueAdd resources, the full file list. Costs no extra request
includeRawDcatbooleanfalseAdd dcat, the publisher's own metadata record
maxDescriptionCharsinteger12000–20000. 0 leaves description out of the row
maxItemsinteger251–5000. The real cost control of a run
fieldsstring[](all)Keep only these fields, in this order. rowType is always kept

Reference

Modes

ModeRequestsOne row isFilters used
search1 per 1000 rowsa datasetall of them
dataset1 per sluga dataset, or a found: false markernone — the slugs decide
organizations1a publishing organizationnone
keywords1a tag with its dataset countnone; query narrows the list

Filters the actor cannot do: there is no state or place filter (the catalog has a geography parameter, but it returned a national FEMA dataset and an Indiana model for California on 29.09.2026, so it is not exposed), and distance sorting needs a geometry this actor does not take.

Levels of government

Federal Government, State Government, City Government, County Government, University, Tribal, Non-Profit. Several values are OR-combined. Two of the 121 organizations carry no level at all; their rows keep organizationType: null.

Organizations

organizationSlug takes the catalog's own publisher slug, not a name. Run the actor once in organizations mode for all 121 with their dataset counts. Slugs seen on 29.09.2026:

  • Federal: census (296,099 datasets), noaa (87,848), doi (54,416), nasa, hhs, epa, va, doj, usda, energy, dot, dol, ed, treasury, hud, dhs, commerce, state, nist, ssa, americorps
  • States: california, washington, maryland, new-york, connecticut, oregon, iowa, district-of-columbia
  • Counties and cities: cook-county-il, lake-county-il, montgomery-county-md, nyc-ny, austin-tx, chicago-il, seattle-wa, san-francisco-ca, los-angeles-ca, louisville-ky, honolulu-hi, tempe-az
  • Other: opentopography

A slug that does not exist matches nothing: the run succeeds with 0 rows and says so in its status message. It is never widened to something broader.

Themes

Themes are the publisher's own free text, matched case-insensitively (Education and education are the same filter). Seen live on 29.09.2026: Environment, Education, Health, Agriculture, Energy, Public Safety, Natural Resources & Environment, Management/Operations, geospatial, City Government. Many datasets carry no theme at all, so a theme filter is a narrow one.

Tags

Tags (keywords) are stricter than a text search — a dataset must carry every tag you list — and they are the publishers' own vocabulary, so check them in keywords mode before you rely on one. Be warned that the biggest tags of this catalog are geographic boilerplate, not topics: county or equivalent entity (277,724 datasets), united states (165,279), state fips code (155,354), county fips code (154,799), u.s. (152,622). Useful topical tags are much smaller: climate had 1,994 datasets.

File formats

The catalog has no format facet, so this filter runs on the rows that came back. A file's format is read from the format name the publisher declared, from its media type and from the extension of its download link, and normalized to one lowercase code, so text/csv, CSV and a .csv link all become csv. Codes seen in the catalog: csv, json, geojson, xml, rdf, xlsx, xls, zip, gz, pdf, html, txt, tsv, kml, kmz, netcdf, gpkg, gdb, shp, api, wms, tiff, jpeg, png, doc, docx, bin.

Media types that describe a request rather than a file (application/http, placeholder/value) and script endpoints (.cgi, .aspx) produce no code. A dataset whose publisher registered no file at all gets formats: [], resourceCount: 0 and hasDownloads: false — that is the source, not a gap in the row.

Only new

onlyNew remembers the catalog slugs a run delivered in a named key-value store and skips them next time. The memory is keyed by the filters of the input, so a daily climate schedule and a weekly nasa schedule keep separate memories and never swallow each other's rows. The first run writes everything it finds. Only rows the run really wrote are remembered, so a dataset cut off by maxItems comes back next time. onlyNew applies to search mode.

Examples

Most used datasets of the whole catalog — no search words, ranked by the catalog's own visit counter.

{ "mode": "search", "sortBy": "popularity", "maxItems": 20 }

Air-quality datasets that really end in a CSV file

{ "mode": "search", "query": "air quality", "onlyWithDownloads": true, "formats": ["csv"], "sortBy": "popularity", "maxItems": 15 }

Everything one federal agency publishes

{ "mode": "search", "organizationSlug": "nasa", "sortBy": "popularity", "maxItems": 20 }

City and state budget data inside the federal catalog

{ "mode": "search", "query": "budget", "organizationTypes": ["City Government", "State Government"], "sortBy": "popularity", "maxItems": 20 }

A weekly digest of newly harvested climate records — put this on a schedule.

{ "mode": "search", "query": "climate", "sortBy": "lastHarvested", "sinceHours": 168, "onlyNew": true, "maxItems": 50 }

Datasets carrying both tags, not just the words

{ "mode": "search", "keywords": ["finance", "budget"], "sortBy": "popularity", "maxItems": 15 }

Catalog URLs you already hold → full records

{ "mode": "dataset", "slugs": ["air-quality", "https://catalog.data.gov/dataset/electric-vehicle-population-data"], "maxItems": 10 }

The publisher dictionary

{ "mode": "organizations", "maxItems": 40 }

Output

One real row, from run oxsnimBahZGQh6QDu on 29.09.2026 (the air-quality + CSV example above):

{
"rowType": "dataset",
"found": true,
"slug": "air-quality",
"title": "Air Quality and Health Impacts",
"description": "Dataset contains information on New York City air quality surveillance data. Air pollution is one of the most important environmental threats to urban populations …",
"descriptionTruncated": false,
"organizationName": "City of New York",
"organizationSlug": "nyc-ny",
"organizationType": "City Government",
"publisher": "data.cityofnewyork.us",
"identifier": "https://data.cityofnewyork.us/api/views/c3uy-2p5r",
"themes": ["Environment"],
"keywords": ["2018od4a-video", "air quality", "climate", "dohmh", "health", "surveillance"],
"license": null,
"accessLevel": "public",
"issued": "2020-12-09",
"modified": "2026-06-18",
"landingPage": "https://data.cityofnewyork.us/d/c3uy-2p5r",
"lastHarvestedAt": "2026-09-10T18:31:43.436Z",
"hasDownloads": true,
"hasSpatial": false,
"latitude": null,
"longitude": null,
"popularity": 172,
"resourceCount": 3,
"formats": ["csv", "json", "xml"],
"primaryDownloadUrl": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/query.json?accessType=DOWNLOAD",
"parentIdentifier": null,
"harvestRecordUrl": "https://catalog.data.gov/harvest_record/92d8908a-e4ca-4043-bc66-0c2b02a1e623",
"url": "https://catalog.data.gov/dataset/air-quality",
"fetchedAt": "2026-09-29T21:42:07.309Z",
"resources": [
{ "title": null, "format": "json", "mediaType": "application/json", "url": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/query.json?accessType=DOWNLOAD", "describedBy": "https://data.cityofnewyork.us/api/views/c3uy-2p5r/columns.json" },
{ "title": null, "format": "xml", "mediaType": "application/xml", "url": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/query.xml?accessType=DOWNLOAD", "describedBy": "https://data.cityofnewyork.us/api/views/c3uy-2p5r/columns.xml" },
{ "title": null, "format": "csv", "mediaType": "text/csv", "url": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/export.csv?accessType=DOWNLOAD", "describedBy": null }
]
}

Dataset rows

FieldTypeMeaning
rowTypestringdataset — always present, also when fields is set
foundbooleanfalse on the marker row of a slug the catalog does not hold
slugstringCatalog slug, the last part of the catalog page URL
titlestringDataset title as the publisher wrote it
descriptionstringPlain text, markup removed, cut at maxDescriptionChars
descriptionTruncatedbooleanSomething was cut off
organizationNamestringPublishing organization, e.g. City of New York
organizationSlugstringIts catalog slug — the value for the organization filter
organizationTypestringLevel of government; null for two organizations
publisherstringThe source portal the record was harvested from
identifierstringThe publisher's own identifier for the dataset
themesstring[]Publisher themes; often empty
keywordsstring[]Catalog tags
licensestringLicence URL when the publisher gave one — about half do
accessLevelstringpublic, restricted public or non-public metadata
issuedstringFirst publication, day precision, as the publisher wrote it
modifiedstringLast change at the publisher, day precision
landingPagestringThe dataset's page on the publisher's own portal
lastHarvestedAtstringWhen the catalog last re-read the record, ISO 8601 UTC
hasDownloadsbooleanThe catalog marks at least one downloadable file
hasSpatialbooleanThe record carries a geometry
latitude / longitudenumberCentroid of that geometry, when there is one
popularitynumberThe catalog's own visit counter for the dataset page
resourceCountnumberNumber of files and endpoints in the record
formatsstring[]Lowercase format codes, sorted and de-duplicated
primaryDownloadUrlstringLink of the first file — the one-field answer to "where is the data"
parentIdentifierstringSet on datasets that belong to a collection
harvestRecordUrlstringThe catalog's harvest record for this dataset
urlstringThe catalog page of the dataset
fetchedAtstringTime of the request, ISO 8601 UTC
resourcesobject[]With includeResources: title, format, mediaType, url, describedBy per file
dcatobjectWith includeRawDcat: the publisher's own metadata record
messagestringOn a found: false row: what to check

Organization and tag rows

FieldTypeMeaning
rowTypestringorganization or keyword
organizationSlugstringThe value the organization filter takes
organizationNamestringFull name of the organization
organizationTypestringLevel of government
organizationIdstringThe catalog's internal id
datasetCountnumberDatasets of this organization, or datasets behind this tag
sourceCountnumberSource portals feeding this organization
aliasesstring[]Alternative names the catalog knows
keywordstringThe tag itself, on keyword rows
urlstringA catalog search for this organization or tag
fetchedAtstringTime of the request, ISO 8601 UTC

Dataset views: Datasets (title, organization, level, themes, tags, dates, formats, files, popularity, catalog page), Files and licence (title, organization, source portal, licence, formats, first file, publisher page), Organizations and tags (the two dictionary modes).

Dates come from two different places: issued and modified are the publisher's own, at day precision and exactly as written; lastHarvestedAt and fetchedAt are full UTC timestamps.

Use it from code / agents

curl -X POST "https://api.apify.com/v2/acts/yadroo~data-gov-datasets/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"query":"air quality","onlyWithDownloads":true,"formats":["csv"],"maxItems":15}'
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('yadroo/data-gov-datasets').call({ organizationSlug: 'nasa', sortBy: 'popularity', maxItems: 20 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("yadroo/data-gov-datasets").call(run_input={"mode": "dataset", "slugs": ["air-quality"]})
items = client.dataset(run["defaultDatasetId"]).list_items().items

Small rows for an LLM — five fields instead of thirty:

{ "query": "medicare spending", "maxItems": 10, "maxDescriptionChars": 300, "includeResources": false,
"fields": ["title", "organizationName", "formats", "primaryDownloadUrl", "url"] }

MCP: add https://mcp.apify.com to Claude / Cursor / any MCP client and call the yadroo/data-gov-datasets tool with the same JSON input. When your agent already holds a catalog URL, dataset mode is the right call — it costs one request and returns the full record.

Pricing

Pay per event: $0.001 per run start + $0.002 per row at the FREE tier. The start event is charged on every run, including runs that end with no matching dataset. Apify plans above FREE pay less per event (−10 % on Bronze, −20 % on Silver, −30 % on Gold and above).

RunRowsWhat you pay (FREE tier)
The default search (maxItems: 25)25$0.001 + 25 × $0.002 = $0.051
Three catalog slugs looked up3$0.001 + 3 × $0.002 = $0.007
The 40 biggest publishers40$0.001 + 40 × $0.002 = $0.081
A 500-row agency export500$0.001 + 500 × $0.002 = $1.001

maxItems decides the bill, not the filters. Two things are worth knowing: a formats list or a sinceHours window throws rows away after the request, so such a run can end below maxItems while still having paid the start event; and the dictionary modes charge per organization or tag written, one request in total.

Limits & FAQ

How many datasets matched my filters? The catalog's search API returns no total, so the actor cannot tell you — and does not invent a number. keywords mode gives the dataset count per tag, which is the closest thing to a count before a run.

Does it filter by state or city? Not by geography. organizationSlug and publisher get you the datasets of a named state or city portal (washington, nyc-ny, data.cityofnewyork.us); the catalog's own place parameter did not filter correctly on 29.09.2026 and is deliberately not exposed.

Are the download links checked? No. The catalog stores what the publisher last published, so a downloadURL can be stale or moved. Verifying every link would mean one request per file; the actor reports what the catalog holds.

Why is a run with many rows slow? The catalog's robots.txt asks for a 10-second crawl delay, and the actor honours it between consecutive requests. Rows come in pages of up to 1000, so every run up to 1000 rows makes a single request and waits not at all; a 5000-row run makes five requests with four pauses. dataset mode makes one request per slug, so ten slugs take about a minute and a half.

Why does my format filter return fewer rows than maxItems? Because the catalog has no format facet: the filter runs on the rows that came back. The actor fetches larger pages when you use it, and follows up to six extra pages, but a rare format over a narrow search can still come up short. The status message says how many rows were skipped and why.

Contact details. The source record carries the name and mailbox of the publishing office. The actor removes them from every row, including from the raw dcat record.

Empty results. A filter combination that matches nothing is a successful run with 0 rows and a status message saying so — never an error, and never silently widened. A dataset slug that does not exist becomes one row with found: false.

Source stability. This catalog was rebuilt on a new stack: its old CKAN-style action endpoints answer 404 as of 29.09.2026, and the current API is versioned 0.1.0, so parameters can still move. The actor fails with a clear message if the search stops returning a result list, and unit tests run against saved real responses so a change in shape breaks a test instead of a paying run.

Terms. The catalog's robots.txt disallows nothing and asks only for the crawl delay above; no term forbidding automated access was found on the site on 29.09.2026. The catalog metadata is the work of the US federal government. The files themselves belong to the publishing organizations — check license and the publisher's page before you redistribute data.


Made by Yadroo. Sibling actors: world-bank-indicators, federal-register-documents, ecfr-regulations, openfda-records, nasa-eonet-events.