# Data.gov Dataset Search API: US Government Open Data Catalog (`yadroo/data-gov-datasets`) Actor

Search the US federal open-data catalog by keyword, agency, government level, tag, theme or file format and get one row per dataset: title, publisher, description, licence, issue and harvest dates, popularity and every resource download URL. Dictionary modes list the agencies and the tags.

- **URL**: https://apify.com/yadroo/data-gov-datasets.md
- **Developed by:** [Samat Makatov](https://apify.com/yadroo) (community)
- **Categories:** Developer tools, Business, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.40 / 1,000 catalog row returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Data.gov Dataset Search API: US Government Open Data Catalog

Search the US federal open-data catalog by keyword, agency, government level, tag, theme or file format and get one row per dataset: title, publisher, description, licence, issue and harvest dates, popularity and every resource download URL. Dictionary modes list the agencies and the tags.

The catalog at `catalog.data.gov` held **570,111 datasets** on 29.09.2026 — federal departments plus state, county, city, university and tribal publishers. It stores the *metadata*: what a dataset is, who publishes it, when it was last refreshed and where the publisher's CSV, JSON, XML or GeoJSON files live. This actor turns a search over that catalog into rows you can filter, schedule and hand to a script, so the download link is a field instead of two clicks. No API key, no login, no proxy, no browser.

### Use cases

- **Find the data behind a question**: search `medicare spending` or `broadband` with *Only datasets with downloadable files* on and get the datasets plus the direct file links, instead of landing pages you have to open one by one.
- **Inventory one agency**: put `nasa`, `census`, `epa` or `hhs` into *Publishing organization* and export everything that agency publishes, with tags, themes, licence and dates — the base table for a coverage audit or an internal catalog.
- **Work the local half of the catalog**: city, county and state portals harvest into the same index. *Level of government* = `City Government` + `State Government` turns the federal catalog into a municipal-data search.
- **Watch what is new**: sort by `last harvested`, set *Harvested in the last N hours* to 168 and switch on *Only datasets not seen before* — a weekly digest that pays only for entries that appeared since the last run.
- **Feed an agent or an MCP tool**: a stable, keyless index of US government data. Use `fields` to keep rows small, and `dataset` mode when your agent already holds a catalog URL.
- **Build a correct filter first**: `organizations` mode writes the 121 publishers with their slugs and levels, `keywords` mode writes the tags with the number of datasets behind each one — reference tables, one request each.

### Input

Every field is optional: the actor runs on its prefilled example with all defaults.

| Field | Type | Default | Allowed values / notes |
|---|---|---|---|
| `mode` | string | `search` | `search`, `dataset`, `organizations`, `keywords` — see [Modes](#modes) |
| `query` | string | *(prefill `air quality`)* | Full-text search over title, description and tags. Empty + `sortBy: popularity` = the most-used datasets of the whole catalog. `search` mode; in `keywords` mode it narrows the tag list |
| `slugs` | string\[] | *(prefill `["air-quality","electric-vehicle-population-data"]`)* | `dataset` mode only: catalog slugs or whole page URLs, one request and one row each |
| `organizationSlug` | string | *(none)* | One publisher slug, e.g. `nasa` — see [Organizations](#organizations) |
| `organizationTypes` | string\[] | *(none)* | Levels of government, OR-combined — see [Levels of government](#levels-of-government) |
| `keywords` | string\[] | *(none)* | Catalog tags; a dataset must carry **all** of them. Matched without regard to case |
| `themes` | string\[] | *(none)* | Publisher themes, OR-combined, case-insensitive — see [Themes](#themes) |
| `publisher` | string | *(none)* | Exact source portal, e.g. `data.cityofnewyork.us` |
| `spatialFilter` | string | `any` | `any`, `geospatial`, `non-geospatial` |
| `onlyWithDownloads` | boolean | `false` | Keep only datasets the catalog marks as having a downloadable file |
| `formats` | string\[] | *(none)* | File formats, at least one must be present — see [File formats](#file-formats). Applied **here**, after the fetch |
| `sinceHours` | integer | *(none)* | 1–8760. Keep datasets re-harvested inside this window (UTC). Applied here, after the fetch |
| `onlyNew` | boolean | `false` | Write only datasets this input has not produced before — see [Only new](#only-new) |
| `sortBy` | string | `relevance` | `relevance`, `popularity`, `lastHarvested` |
| `includeResources` | boolean | `true` | Add `resources`, the full file list. Costs no extra request |
| `includeRawDcat` | boolean | `false` | Add `dcat`, the publisher's own metadata record |
| `maxDescriptionChars` | integer | `1200` | 0–20000. `0` leaves `description` out of the row |
| `maxItems` | integer | `25` | 1–5000. The real cost control of a run |
| `fields` | string\[] | *(all)* | Keep only these fields, in this order. `rowType` is always kept |

### Reference

#### Modes

| Mode | Requests | One row is | Filters used |
|---|---|---|---|
| `search` | 1 per 1000 rows | a dataset | all of them |
| `dataset` | 1 per slug | a dataset, or a `found: false` marker | none — the slugs decide |
| `organizations` | 1 | a publishing organization | none |
| `keywords` | 1 | a tag with its dataset count | none; `query` narrows the list |

Filters the actor cannot do: there is no state or place filter (the catalog has a geography parameter, but it returned a national FEMA dataset and an Indiana model for `California` on 29.09.2026, so it is not exposed), and `distance` sorting needs a geometry this actor does not take.

#### Levels of government

`Federal Government`, `State Government`, `City Government`, `County Government`, `University`, `Tribal`, `Non-Profit`. Several values are OR-combined. Two of the 121 organizations carry no level at all; their rows keep `organizationType: null`.

#### Organizations

`organizationSlug` takes the catalog's own publisher slug, not a name. Run the actor once in `organizations` mode for all 121 with their dataset counts. Slugs seen on 29.09.2026:

- Federal: `census` (296,099 datasets), `noaa` (87,848), `doi` (54,416), `nasa`, `hhs`, `epa`, `va`, `doj`, `usda`, `energy`, `dot`, `dol`, `ed`, `treasury`, `hud`, `dhs`, `commerce`, `state`, `nist`, `ssa`, `americorps`
- States: `california`, `washington`, `maryland`, `new-york`, `connecticut`, `oregon`, `iowa`, `district-of-columbia`
- Counties and cities: `cook-county-il`, `lake-county-il`, `montgomery-county-md`, `nyc-ny`, `austin-tx`, `chicago-il`, `seattle-wa`, `san-francisco-ca`, `los-angeles-ca`, `louisville-ky`, `honolulu-hi`, `tempe-az`
- Other: `opentopography`

A slug that does not exist matches nothing: the run succeeds with 0 rows and says so in its status message. It is never widened to something broader.

#### Themes

Themes are the publisher's own free text, matched case-insensitively (`Education` and `education` are the same filter). Seen live on 29.09.2026: `Environment`, `Education`, `Health`, `Agriculture`, `Energy`, `Public Safety`, `Natural Resources & Environment`, `Management/Operations`, `geospatial`, `City Government`. Many datasets carry no theme at all, so a theme filter is a narrow one.

#### Tags

Tags (`keywords`) are stricter than a text search — a dataset must carry **every** tag you list — and they are the publishers' own vocabulary, so check them in `keywords` mode before you rely on one. Be warned that the biggest tags of this catalog are geographic boilerplate, not topics: `county or equivalent entity` (277,724 datasets), `united states` (165,279), `state fips code` (155,354), `county fips code` (154,799), `u.s.` (152,622). Useful topical tags are much smaller: `climate` had 1,994 datasets.

#### File formats

The catalog has no format facet, so this filter runs on the rows that came back. A file's format is read from the format name the publisher declared, from its media type and from the extension of its download link, and normalized to one lowercase code, so `text/csv`, `CSV` and a `.csv` link all become `csv`. Codes seen in the catalog: `csv`, `json`, `geojson`, `xml`, `rdf`, `xlsx`, `xls`, `zip`, `gz`, `pdf`, `html`, `txt`, `tsv`, `kml`, `kmz`, `netcdf`, `gpkg`, `gdb`, `shp`, `api`, `wms`, `tiff`, `jpeg`, `png`, `doc`, `docx`, `bin`.

Media types that describe a request rather than a file (`application/http`, `placeholder/value`) and script endpoints (`.cgi`, `.aspx`) produce no code. A dataset whose publisher registered no file at all gets `formats: []`, `resourceCount: 0` and `hasDownloads: false` — that is the source, not a gap in the row.

#### Only new

`onlyNew` remembers the catalog slugs a run delivered in a named key-value store and skips them next time. The memory is keyed by the filters of the input, so a daily `climate` schedule and a weekly `nasa` schedule keep separate memories and never swallow each other's rows. The first run writes everything it finds. Only rows the run really wrote are remembered, so a dataset cut off by `maxItems` comes back next time. `onlyNew` applies to `search` mode.

### Examples

**Most used datasets of the whole catalog** — no search words, ranked by the catalog's own visit counter.

```json
{ "mode": "search", "sortBy": "popularity", "maxItems": 20 }
```

**Air-quality datasets that really end in a CSV file**

```json
{ "mode": "search", "query": "air quality", "onlyWithDownloads": true, "formats": ["csv"], "sortBy": "popularity", "maxItems": 15 }
```

**Everything one federal agency publishes**

```json
{ "mode": "search", "organizationSlug": "nasa", "sortBy": "popularity", "maxItems": 20 }
```

**City and state budget data inside the federal catalog**

```json
{ "mode": "search", "query": "budget", "organizationTypes": ["City Government", "State Government"], "sortBy": "popularity", "maxItems": 20 }
```

**A weekly digest of newly harvested climate records** — put this on a schedule.

```json
{ "mode": "search", "query": "climate", "sortBy": "lastHarvested", "sinceHours": 168, "onlyNew": true, "maxItems": 50 }
```

**Datasets carrying both tags, not just the words**

```json
{ "mode": "search", "keywords": ["finance", "budget"], "sortBy": "popularity", "maxItems": 15 }
```

**Catalog URLs you already hold → full records**

```json
{ "mode": "dataset", "slugs": ["air-quality", "https://catalog.data.gov/dataset/electric-vehicle-population-data"], "maxItems": 10 }
```

**The publisher dictionary**

```json
{ "mode": "organizations", "maxItems": 40 }
```

### Output

One real row, from run `oxsnimBahZGQh6QDu` on 29.09.2026 (the air-quality + CSV example above):

```json
{
  "rowType": "dataset",
  "found": true,
  "slug": "air-quality",
  "title": "Air Quality and Health Impacts",
  "description": "Dataset contains information on New York City air quality surveillance data. Air pollution is one of the most important environmental threats to urban populations …",
  "descriptionTruncated": false,
  "organizationName": "City of New York",
  "organizationSlug": "nyc-ny",
  "organizationType": "City Government",
  "publisher": "data.cityofnewyork.us",
  "identifier": "https://data.cityofnewyork.us/api/views/c3uy-2p5r",
  "themes": ["Environment"],
  "keywords": ["2018od4a-video", "air quality", "climate", "dohmh", "health", "surveillance"],
  "license": null,
  "accessLevel": "public",
  "issued": "2020-12-09",
  "modified": "2026-06-18",
  "landingPage": "https://data.cityofnewyork.us/d/c3uy-2p5r",
  "lastHarvestedAt": "2026-09-10T18:31:43.436Z",
  "hasDownloads": true,
  "hasSpatial": false,
  "latitude": null,
  "longitude": null,
  "popularity": 172,
  "resourceCount": 3,
  "formats": ["csv", "json", "xml"],
  "primaryDownloadUrl": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/query.json?accessType=DOWNLOAD",
  "parentIdentifier": null,
  "harvestRecordUrl": "https://catalog.data.gov/harvest_record/92d8908a-e4ca-4043-bc66-0c2b02a1e623",
  "url": "https://catalog.data.gov/dataset/air-quality",
  "fetchedAt": "2026-09-29T21:42:07.309Z",
  "resources": [
    { "title": null, "format": "json", "mediaType": "application/json", "url": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/query.json?accessType=DOWNLOAD", "describedBy": "https://data.cityofnewyork.us/api/views/c3uy-2p5r/columns.json" },
    { "title": null, "format": "xml", "mediaType": "application/xml", "url": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/query.xml?accessType=DOWNLOAD", "describedBy": "https://data.cityofnewyork.us/api/views/c3uy-2p5r/columns.xml" },
    { "title": null, "format": "csv", "mediaType": "text/csv", "url": "https://data.cityofnewyork.us/api/v3/views/c3uy-2p5r/export.csv?accessType=DOWNLOAD", "describedBy": null }
  ]
}
```

#### Dataset rows

| Field | Type | Meaning |
|---|---|---|
| `rowType` | string | `dataset` — always present, also when `fields` is set |
| `found` | boolean | `false` on the marker row of a slug the catalog does not hold |
| `slug` | string | Catalog slug, the last part of the catalog page URL |
| `title` | string | Dataset title as the publisher wrote it |
| `description` | string | Plain text, markup removed, cut at `maxDescriptionChars` |
| `descriptionTruncated` | boolean | Something was cut off |
| `organizationName` | string | Publishing organization, e.g. `City of New York` |
| `organizationSlug` | string | Its catalog slug — the value for the organization filter |
| `organizationType` | string | Level of government; `null` for two organizations |
| `publisher` | string | The source portal the record was harvested from |
| `identifier` | string | The publisher's own identifier for the dataset |
| `themes` | string\[] | Publisher themes; often empty |
| `keywords` | string\[] | Catalog tags |
| `license` | string | Licence URL when the publisher gave one — about half do |
| `accessLevel` | string | `public`, `restricted public` or `non-public` metadata |
| `issued` | string | First publication, day precision, as the publisher wrote it |
| `modified` | string | Last change at the publisher, day precision |
| `landingPage` | string | The dataset's page on the publisher's own portal |
| `lastHarvestedAt` | string | When the catalog last re-read the record, ISO 8601 **UTC** |
| `hasDownloads` | boolean | The catalog marks at least one downloadable file |
| `hasSpatial` | boolean | The record carries a geometry |
| `latitude` / `longitude` | number | Centroid of that geometry, when there is one |
| `popularity` | number | The catalog's own visit counter for the dataset page |
| `resourceCount` | number | Number of files and endpoints in the record |
| `formats` | string\[] | Lowercase format codes, sorted and de-duplicated |
| `primaryDownloadUrl` | string | Link of the first file — the one-field answer to "where is the data" |
| `parentIdentifier` | string | Set on datasets that belong to a collection |
| `harvestRecordUrl` | string | The catalog's harvest record for this dataset |
| `url` | string | The catalog page of the dataset |
| `fetchedAt` | string | Time of the request, ISO 8601 UTC |
| `resources` | object\[] | With `includeResources`: `title`, `format`, `mediaType`, `url`, `describedBy` per file |
| `dcat` | object | With `includeRawDcat`: the publisher's own metadata record |
| `message` | string | On a `found: false` row: what to check |

#### Organization and tag rows

| Field | Type | Meaning |
|---|---|---|
| `rowType` | string | `organization` or `keyword` |
| `organizationSlug` | string | The value the organization filter takes |
| `organizationName` | string | Full name of the organization |
| `organizationType` | string | Level of government |
| `organizationId` | string | The catalog's internal id |
| `datasetCount` | number | Datasets of this organization, or datasets behind this tag |
| `sourceCount` | number | Source portals feeding this organization |
| `aliases` | string\[] | Alternative names the catalog knows |
| `keyword` | string | The tag itself, on `keyword` rows |
| `url` | string | A catalog search for this organization or tag |
| `fetchedAt` | string | Time of the request, ISO 8601 UTC |

Dataset views: **Datasets** (title, organization, level, themes, tags, dates, formats, files, popularity, catalog page), **Files and licence** (title, organization, source portal, licence, formats, first file, publisher page), **Organizations and tags** (the two dictionary modes).

Dates come from two different places: `issued` and `modified` are the publisher's own, at day precision and exactly as written; `lastHarvestedAt` and `fetchedAt` are full UTC timestamps.

### Use it from code / agents

```bash
curl -X POST "https://api.apify.com/v2/acts/yadroo~data-gov-datasets/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"query":"air quality","onlyWithDownloads":true,"formats":["csv"],"maxItems":15}'
```

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('yadroo/data-gov-datasets').call({ organizationSlug: 'nasa', sortBy: 'popularity', maxItems: 20 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
```

```python
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("yadroo/data-gov-datasets").call(run_input={"mode": "dataset", "slugs": ["air-quality"]})
items = client.dataset(run["defaultDatasetId"]).list_items().items
```

Small rows for an LLM — five fields instead of thirty:

```json
{ "query": "medicare spending", "maxItems": 10, "maxDescriptionChars": 300, "includeResources": false,
  "fields": ["title", "organizationName", "formats", "primaryDownloadUrl", "url"] }
```

MCP: add `https://mcp.apify.com` to Claude / Cursor / any MCP client and call the `yadroo/data-gov-datasets` tool with the same JSON input. When your agent already holds a catalog URL, `dataset` mode is the right call — it costs one request and returns the full record.

### Pricing

Pay per event: **$0.001 per run start + $0.002 per row** at the FREE tier. The start event is charged on every run, including runs that end with no matching dataset. Apify plans above FREE pay less per event (−10 % on Bronze, −20 % on Silver, −30 % on Gold and above).

| Run | Rows | What you pay (FREE tier) |
|---|---|---|
| The default search (`maxItems: 25`) | 25 | $0.001 + 25 × $0.002 = **$0.051** |
| Three catalog slugs looked up | 3 | $0.001 + 3 × $0.002 = **$0.007** |
| The 40 biggest publishers | 40 | $0.001 + 40 × $0.002 = **$0.081** |
| A 500-row agency export | 500 | $0.001 + 500 × $0.002 = **$1.001** |

`maxItems` decides the bill, not the filters. Two things are worth knowing: a `formats` list or a `sinceHours` window throws rows away after the request, so such a run can end below `maxItems` while still having paid the start event; and the dictionary modes charge per organization or tag written, one request in total.

### Limits & FAQ

**How many datasets matched my filters?** The catalog's search API returns no total, so the actor cannot tell you — and does not invent a number. `keywords` mode gives the dataset count per tag, which is the closest thing to a count before a run.

**Does it filter by state or city?** Not by geography. `organizationSlug` and `publisher` get you the datasets of a named state or city portal (`washington`, `nyc-ny`, `data.cityofnewyork.us`); the catalog's own place parameter did not filter correctly on 29.09.2026 and is deliberately not exposed.

**Are the download links checked?** No. The catalog stores what the publisher last published, so a `downloadURL` can be stale or moved. Verifying every link would mean one request per file; the actor reports what the catalog holds.

**Why is a run with many rows slow?** The catalog's robots.txt asks for a 10-second crawl delay, and the actor honours it between consecutive requests. Rows come in pages of up to 1000, so every run up to 1000 rows makes a single request and waits not at all; a 5000-row run makes five requests with four pauses. `dataset` mode makes one request per slug, so ten slugs take about a minute and a half.

**Why does my format filter return fewer rows than `maxItems`?** Because the catalog has no format facet: the filter runs on the rows that came back. The actor fetches larger pages when you use it, and follows up to six extra pages, but a rare format over a narrow search can still come up short. The status message says how many rows were skipped and why.

**Contact details.** The source record carries the name and mailbox of the publishing office. The actor removes them from every row, including from the raw `dcat` record.

**Empty results.** A filter combination that matches nothing is a successful run with 0 rows and a status message saying so — never an error, and never silently widened. A `dataset` slug that does not exist becomes one row with `found: false`.

**Source stability.** This catalog was rebuilt on a new stack: its old CKAN-style action endpoints answer 404 as of 29.09.2026, and the current API is versioned 0.1.0, so parameters can still move. The actor fails with a clear message if the search stops returning a result list, and unit tests run against saved real responses so a change in shape breaks a test instead of a paying run.

**Terms.** The catalog's robots.txt disallows nothing and asks only for the crawl delay above; no term forbidding automated access was found on the site on 29.09.2026. The catalog metadata is the work of the US federal government. The files themselves belong to the publishing organizations — check `license` and the publisher's page before you redistribute data.

***

Made by **Yadroo**. Sibling actors: [world-bank-indicators](https://apify.com/yadroo/world-bank-indicators), [federal-register-documents](https://apify.com/yadroo/federal-register-documents), [ecfr-regulations](https://apify.com/yadroo/ecfr-regulations), [openfda-records](https://apify.com/yadroo/openfda-records), [nasa-eonet-events](https://apify.com/yadroo/nasa-eonet-events).

# Actor input Schema

## `mode` (type: `string`):

`search` runs the catalog search and writes a row per dataset. `dataset` looks up the exact catalog slugs listed below, one row each, and marks a slug the catalog does not hold with `found: false` instead of failing the run. `organizations` writes the publishing organizations of the catalog (121 on 29.09.2026) with their slug and government level - that slug is what *Publishing organization* takes. `keywords` writes the tag list with the number of datasets behind each tag, so you can see what a filter will return before you pay for rows. The two dictionary modes ignore the filters and cost one request.

## `query` (type: `string`):

Full-text search over dataset title, description and tags, e.g. `air quality`, `electric vehicle`, `medicare spending`, `broadband`. Empty = no text condition: with *Sort rows by* set to `popularity` that returns the most-used datasets of the whole catalog, which is the cheap way to explore before you narrow down. Used in `search` mode only.

## `slugs` (type: `array`):

Used in `dataset` mode only: the slugs to look up, e.g. `air-quality`, `motor-vehicle-collisions-crashes`. The slug is the last part of a catalog dataset page URL (`https://catalog.data.gov/dataset/air-quality` -> `air-quality`); a whole page URL is accepted and the slug is read out of it. One request per slug. In `search` mode this field is ignored.

## `organizationSlug` (type: `string`):

Keep datasets of one organization, by its catalog slug: `census`, `noaa`, `nasa`, `epa`, `hhs`, `usda`, `energy`, `dot`, `dol`, `doi`, `ed`, `treasury`, `doj`, `hud` for federal departments, `california`, `washington` for states, `nyc-ny` for a city. Run the actor once in `organizations` mode for the complete list - a slug that is not in it matches nothing and yields an empty run, not an error.

## `organizationTypes` (type: `array`):

Keep datasets published by organizations of these levels; several values = OR. The catalog is far from federal-only - city and state portals harvest into it too, so this is the filter that separates "what does Washington publish" from "what does my state publish". Empty = every level.

## `keywords` (type: `array`):

Keep datasets tagged with every one of these catalog tags, e.g. \["finance", "budget"] returns datasets that carry both. Tags are lowercase free text the publisher chose, so check the spelling in `keywords` mode first: the catalog's most common tags are geographic boilerplate (`county or equivalent entity`, `state fips code`) rather than topics.

## `themes` (type: `array`):

Keep datasets filed under these themes; several values = OR. Themes come from the publisher's own metadata and are matched without regard to case, so `Education` and `education` are the same value. Values seen live on 29.09.2026: `Environment`, `Education`, `Health`, `Agriculture`, `Energy`, `Management/Operations`, `geospatial`. A theme is a coarser bucket than a tag and many datasets carry none.

## `publisher` (type: `string`):

Keep datasets harvested from one source portal, matched exactly, e.g. `data.cityofnewyork.us`. This is finer than *Publishing organization*: a single organization often feeds the catalog from several portals, and the field on every row tells you which portal a dataset really came from.

## `spatialFilter` (type: `string`):

Large parts of the catalog are map layers and boundary files: a popularity search inside one mapping agency comes back as shapefiles almost entirely. Pick `non-geospatial` when you want tables and statistics, `geospatial` when you want the map layers, and expect a narrow combination of both this and an organization filter to return nothing.

## `onlyWithDownloads` (type: `boolean`):

Keep only datasets the catalog marks as having at least one downloadable file. Part of the catalog describes datasets that are behind a request form, an interactive map or an API you have to sign up for; switch this on when the run has to end in files you can actually fetch.

## `formats` (type: `array`):

Keep datasets that offer at least one file in these formats, e.g. \["csv"], \["json", "geojson"], \["xlsx"]. The format of a file is read from its declared media type and from the file extension of its download link, so `text/csv` and a `.csv` link both count as `csv`. The catalog search itself has no format filter, so this one runs here, on the rows that came back: a run with a narrow format list can end with fewer rows than *Max rows* while still costing the requests it made. Empty = no format condition.

## `sinceHours` (type: `integer`):

Keep datasets whose catalog record was last refreshed from the publisher inside this window, in UTC. Pair it with *Sort rows by* = `last harvested` to watch what the catalog picked up overnight. Careful: this is the moment the catalog re-read the publisher's metadata, not the moment the data changed - a re-harvest with unchanged content updates it too.

## `onlyNew` (type: `boolean`):

Remember the catalog identifiers this input has already produced in the actor's key-value store and write only datasets that were not there on the previous run. The first run writes everything it finds, later runs write what appeared since. Built for a schedule: a daily run over a search you care about, paying only for the new entries.

## `sortBy` (type: `string`):

`relevance` ranks by how well a dataset matches the search words and is the sensible default when you typed some. Without search words use `popularity`, which orders by the catalog's own page-visit counter and puts the datasets people actually use on top. `lastHarvested` orders by the moment the catalog last refreshed the record, newest first - the ordering the monitoring filter above is built for.

## `includeResources` (type: `boolean`):

Add `resources`, the full list of files and endpoints of a dataset with title, format, media type, download link and, where the publisher gives one, a link to the column description. The list arrives in the same request, so it costs nothing extra - switch it off only to keep rows small when you already have `formats`, `resourceCount` and `primaryDownloadUrl`.

## `includeRawDcat` (type: `boolean`):

Add `dcat`, the publisher's metadata record exactly as the catalog stores it (DCAT-US vocabulary: `@type`, `accessLevel`, `identifier`, `issued`, `modified`, `distribution`, `license`, `theme` and whatever else that publisher filled in). For pipelines that want the original vocabulary instead of our flat row. Contact names and mailbox addresses of the publishing office are removed from this record, like everywhere else in the output.

## `maxDescriptionChars` (type: `integer`):

Cut `description` after this many characters at a word boundary and set `descriptionTruncated` when something was cut. Catalog descriptions range from one line to several pages of HTML; the field arrives as plain text with markup removed. `0` leaves the description out of the row altogether.

## `maxItems` (type: `integer`):

Stop after this many rows. The catalog held 570,118 datasets on 29.09.2026 and a broad search matches tens of thousands of them, so this cap - not the filters - is what decides what a run costs. Rows are fetched in pages of up to 1000, and the run stops as soon as the cap is reached.

## `fields` (type: `array`):

Keep only these fields, in this order, e.g. \["title", "organizationName", "formats", "primaryDownloadUrl", "url"]. Empty = every field the mode produces.

## Actor input object example

```json
{
  "mode": "search",
  "query": "air quality",
  "slugs": [
    "air-quality",
    "electric-vehicle-population-data"
  ],
  "spatialFilter": "any",
  "onlyWithDownloads": false,
  "onlyNew": false,
  "sortBy": "relevance",
  "includeResources": true,
  "includeRawDcat": false,
  "maxDescriptionChars": 1200,
  "maxItems": 25
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "air quality",
    "slugs": [
        "air-quality",
        "electric-vehicle-population-data"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("yadroo/data-gov-datasets").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "air quality",
    "slugs": [
        "air-quality",
        "electric-vehicle-population-data",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("yadroo/data-gov-datasets").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "air quality",
  "slugs": [
    "air-quality",
    "electric-vehicle-population-data"
  ]
}' |
apify call yadroo/data-gov-datasets --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,yadroo/data-gov-datasets"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/V853qnLtZX2TpLckv/builds/fkIEZmSn8bsOdIRj4/openapi.json
