Socrata Open Data Scraper | SoQL Filters, No API Key avatar

Socrata Open Data Scraper | SoQL Filters, No API Key

Pricing

from $0.60 / 1,000 records

Go to Apify Store
Socrata Open Data Scraper | SoQL Filters, No API Key

Socrata Open Data Scraper | SoQL Filters, No API Key

Search and pull datasets from any Socrata government portal (CDC, HHS, CMS, NY, NYC, Texas, 100s more). SoQL filtered queries, structured JSON, no API key, no anti bot. Works in Claude, ChatGPT and any MCP agent for instant public data access.

Pricing

from $0.60 / 1,000 records

Rating

0.0

(0)

Developer

The Mine Works

The Mine Works

Maintained by Community

Actor stats

0

Bookmarked

5

Total users

1

Monthly active users

2 hours ago

Last modified

Share

2,000 NYC inspection rows in 8 seconds

From The Mine Works, makers of Threads Scraper and B2B Leads Finder, with over 140,000 runs across 170+ public actors. This actor ranks #2 for "socrata" in Apify Store search.

Hundreds of US government bodies publish open data on Socrata: the CDC, New York City, New York State, Washington, Colorado, Chicago and many counties and universities. They all share one URL pattern and one SQL-like query language, SoQL. This actor gives you both halves of the job: discovery mode finds datasets on a topic across the Socrata catalog, and data mode pulls rows from any dataset, filtered, trimmed and sorted on the portal before they reach you. No API key.

Why choose this actor?

  • Thousands of filtered rows in seconds. In a recorded run on 2 October 2026, NYC's restaurant inspection dataset filtered to Manhattan inspections scoring over 20, newest first, returned 2,000 rows in 8 seconds: every row in Manhattan, every score 21 or more, covering 338 restaurants.
  • Filters run on the portal, not on your machine. where, select and order are sent to the portal as SoQL $where, $select and $order, so you download and pay for only the rows and columns you asked for.
  • No key, nothing wasted. No Socrata app token is needed. A bad dataset ID, a SoQL error or a filter that matches nothing returns no rows and charges nothing but the start fee; the portal's own error message is written to the log.

Run it on Apify

Part of The Mine Works Science, health and government data family: CourtListener Scraper, data.gov.in Scraper, Academic Research MCP, OpenAlex Scraper, FDA 510(k) Clearances Scraper, Crossref Scraper.

Try it in one minute

Paste this into the JSON tab of the input page and press Start. It returns 50 rows from a CDC COVID-19 deaths dataset in a couple of seconds.

{
"domain": "data.cdc.gov",
"datasetId": "9bhg-hcku",
"where": "state='New York' AND covid_19_deaths > 100",
"select": "state,sex,age_group,covid_19_deaths",
"order": "covid_19_deaths DESC",
"maxResults": 50
}

There are two ways to give the input. For data mode, give domain (the portal's host name, such as data.cityofnewyork.us; a full URL is trimmed to the host) and datasetId (the four-by-four ID from the dataset's address, such as 43nn-pn8j). For discovery mode, leave those two empty and give searchQuery (keywords such as restaurant inspections). If you fill in all three, data mode wins and the search is ignored.

Apify's free plan includes $5 of credit every month, which covers about 4,900 records at this actor's Free plan price ($0.001 a record plus the $0.005 start fee, in runs of 1,000).

Copy to your AI assistant

Paste this block into ChatGPT, Claude, Cursor or any assistant that can write code, and it can run the actor for you.

themineworks/socrata-open-data on Apify. Two modes: discovery finds datasets across Socrata open data portals by keyword (one row per dataset with domain, dataset_id, name, description, updated_at, permalink); data mode pulls rows from one dataset with SoQL filtering (rows unchanged plus _domain, _dataset_id, _scraped_at). Call ApifyClient("TOKEN").actor("themineworks/socrata-open-data").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items. Data mode: domain (e.g. "data.cityofnewyork.us") and datasetId (four-by-four, e.g. "43nn-pn8j"), optional where (SoQL $where), select (comma-separated columns), order (e.g. "inspection_date DESC"). Discovery mode: searchQuery with domain and datasetId empty. Optional: maxResults (1 to 100000, default 1000), appToken (Socrata app token, only raises rate limits). Rows with _type "summary" or "info" are run reports and are never billed. Full spec: GET https://api.apify.com/v2/acts/themineworks~socrata-open-data/builds/default (Bearer TOKEN), which returns inputSchema and readme. Token: https://console.apify.com/account/integrations?fpr=ymnoit&utm_source=apify-readme&utm_medium=referral

Key features

  • Two modes, one actor. Discovery mode returns 10 fields per dataset (portal domain, dataset ID, name, description, type, last update, permalink and link). Data mode returns the dataset's own columns, unchanged, plus 3 provenance fields.
  • Full SoQL pass-through. Anything the portal accepts in $where works here: comparisons, AND and OR, like, between, date literals and geo functions such as within_circle.
  • Up to 100,000 rows per run, fetched 1,000 at a time, so even a large pull stays light on memory. Discovery pages 100 datasets at a time.
  • Any Socrata portal. Data mode works against any Socrata domain you name, whether or not discovery found it.
  • Retries built in. Rate-limit (429) and server (5xx) answers are retried up to 4 times with back-off capped at 30 seconds. An optional appToken raises your rate limit for very large pulls.
  • Provenance on every row. _domain, _dataset_id and _scraped_at tell you exactly where and when each data row came from, which matters once you mix several datasets.

How to use it

Discovery: find datasets on a topic

{
"searchQuery": "restaurant inspections",
"maxResults": 100
}

On 2 October this returned 56 datasets (all the catalog had for that search) in 5 seconds, from NYC, New York State, Santa Clara County, Montgomery County, Cambridge, Colorado and others. Take domain and dataset_id from the row you want and use them in data mode.

Data: pull a whole dataset

{
"domain": "data.cityofnewyork.us",
"datasetId": "43nn-pn8j",
"maxResults": 5000
}

Rows come back exactly as the portal publishes them, column names and all.

Filter and sort on the portal

{
"domain": "data.cityofnewyork.us",
"datasetId": "43nn-pn8j",
"where": "boro='Manhattan' AND score > 20",
"order": "inspection_date DESC",
"maxResults": 2000
}

This is the exact input of our 2 October test: 2,000 rows in 8 seconds, all from Manhattan, all scoring over 20, newest inspections first.

Only the columns you need

{
"domain": "data.cdc.gov",
"datasetId": "9bhg-hcku",
"select": "state,sex,age_group,covid_19_deaths",
"where": "state='New York'",
"maxResults": 500
}

A dataset with 15 or 60 columns comes back with just the four you name, which keeps exports small and spreadsheets readable.

Daily feed of new records

{
"domain": "data.cityofnewyork.us",
"datasetId": "43nn-pn8j",
"where": "inspection_date > '2026-09-25T00:00:00'",
"order": "inspection_date DESC",
"maxResults": 5000
}

Save this as a task and schedule it daily (for example 0 6 * * *). There is no built-in "only new rows" mode, so move the date in where forward between runs (through the API or a Make, Zapier or n8n step), or deduplicate on the dataset's own ID column after loading.

Building a dataset inventory for research

{
"searchQuery": "building permits",
"maxResults": 500
}

Discovery is the quickest way to see which cities and counties publish a kind of data before you plan a project. Group the results by domain and sort by updated_at to find the portals that keep theirs current.

Input parameters

ParameterTypeDefaultWhat it does
searchQuerystringnone (form prefill: covid deaths)Discovery mode: keywords to search the Socrata catalog for. Used only when domain or datasetId is empty.
domainstringnone (form prefill: data.cdc.gov)Data mode: the portal host, such as data.cityofnewyork.us, data.ny.gov or data.cdc.gov. https:// and any path are removed.
datasetIdstringnone (form prefill: 9bhg-hcku)Data mode: the dataset's four-by-four ID, from discovery or from the dataset's web address.
wherestringnoneSoQL filter sent as $where, such as state='New York' AND covid_19_deaths > 100. Text values go in single quotes.
selectstringnoneComma-separated columns sent as $select. Leave empty for every column.
orderstringnoneSort sent as $order, such as inspection_date DESC.
appTokenstring (secret)noneOptional Socrata app token, sent as X-App-Token. It only raises your rate limit.
maxResultsinteger (1 to 100,000)1000 (form prefill: 50)Most records to return: rows in data mode, datasets in discovery mode.

"Form prefill" values fill the Console form for you; an API call that leaves a field out gets the default shown, or nothing. Note that the form's prefill puts values in both searchQuery and domain plus datasetId, so a run started straight from the form uses data mode; clear domain or datasetId to search.

There is no proxy setting: the portals are public APIs and are called directly. The default run timeout is 300 seconds; for pulls close to 100,000 rows, give the run more time in Run options.

What data do you get?

The shape depends on the mode.

Discovery mode, one row per dataset: _type (always dataset), domain (the portal), dataset_id (the four-by-four ID), name, description (the publisher's text, sometimes with HTML), dataset_type, updated_at (last update, ISO timestamp), permalink and link (addresses of the dataset page), and scraped_at. The catalog no longer reports a row count: rows_count is defined in the code but came back empty for all 56 datasets in our 2 October test, so it is left out of the rows.

Data mode, one row per dataset row: every column the dataset publishes, under the dataset's own column names, plus _domain, _dataset_id and _scraped_at. Socrata's API returns numbers as text (for example "score": "23"), so convert them before doing maths. Columns that are empty in a row are left out of that row. Geographic columns come back as GeoJSON objects.

Each run ends with a _type: "summary" row (mode, records, charged_for, scraped_at) and, when records were delivered, a _type: "info" row with a scheduling tip. Neither is billed. Skip rows whose _type is summary or info when you load the data.

Stable fields for automations

In data mode, the dataset decides the columns, so only these 3 fields are guaranteed on every row: _domain, _dataset_id and _scraped_at.

In discovery mode, these 10 fields were present in every one of the 56 dataset rows in our 2 October test:

FieldWhat it holds
_typeAlways dataset in discovery rows
domainPortal host, such as data.cityofnewyork.us
dataset_idFour-by-four ID to use as datasetId in data mode
nameDataset title
descriptionPublisher's description
dataset_typeSocrata asset type (dataset with the current search)
updated_atLast update, ISO timestamp
permalinkStable link to the dataset page
linkPortal link to the dataset page
scraped_atISO timestamp when the row was captured

We will not rename these fields. New fields may be added over time; existing ones keep their names.

Output examples

Real rows from our 2 October 2026 runs, with long text trimmed.

A data-mode row (NYC restaurant inspections, Manhattan, score over 20):

{
"camis": "50141510",
"dba": "EL TEPEYAC",
"boro": "Manhattan",
"building": "1505",
"street": "LEXINGTON AVENUE",
"zipcode": "10029",
"cuisine_description": "Mexican",
"inspection_date": "2026-09-30T00:00:00.000",
"action": "Violations were cited in the following area(s).",
"violation_code": "04N",
"violation_description": "Filth flies or food/refuse/sewage associated with (FRSA) flies or other nuisance pests in establishment’s food and/or non-food areas…",
"critical_flag": "Critical",
"score": "23",
"inspection_type": "Cycle Inspection / Initial Inspection",
"latitude": "40.786669299016",
"longitude": "-73.950324719525",
"location": { "type": "Point", "coordinates": [-73.950324719525, 40.786669299016] },
"_domain": "data.cityofnewyork.us",
"_dataset_id": "43nn-pn8j",
"_scraped_at": "2026-10-02T08:16:35.029Z"
}

A data-mode row with select (CDC, four columns):

{
"state": "New York",
"sex": "All Sexes",
"age_group": "All Ages",
"covid_19_deaths": "42273",
"_domain": "data.cdc.gov",
"_dataset_id": "9bhg-hcku",
"_scraped_at": "2026-10-02T08:16:47.738Z"
}

A discovery row (search restaurant inspections):

{
"_type": "dataset",
"domain": "data.cityofnewyork.us",
"dataset_id": "4dx7-axux",
"name": "Open Restaurants Inspections (Historical)",
"description": "<b>This is historical dataset</b> program, is no longer accepting applications…",
"dataset_type": "dataset",
"updated_at": "2026-07-22T15:32:55.000Z",
"permalink": "https://data.cityofnewyork.us/d/4dx7-axux",
"link": "https://data.cityofnewyork.us/Transportation/Open-Restaurants-Inspections-Historical-/4dx7-axux",
"scraped_at": "2026-10-02T08:16:44.911Z"
}

The summary row from a platform run on 30 September (CDC dataset, 5 rows, 4.3 seconds; never charged itself):

{
"_type": "summary",
"mode": "data",
"records": 5,
"charged_for": 5,
"scraped_at": "2026-09-30T11:00:27.488Z"
}

Pricing

Pay per event: you pay for each record delivered to your dataset (a data row or a discovered dataset), plus a small start fee per run.

EventFreeBronzeSilverGold and above
record-scraped, per record$0.001$0.0009$0.00075$0.0006
record-scraped, per 1,000 records$1.00$0.90$0.75$0.60
apify-actor-start, per run$0.005 per GB of run memory, minimum one eventsamesamesame

The start fee, exactly. Apify's apify-actor-start event is charged once when a run starts, at $0.005 for each GB of memory the run uses, with a minimum of one event. This actor runs on 512 MB by default, so a default run pays one event: $0.005.

Worked examples on the Gold tier: 1,000 rows cost $0.60 plus $0.005. 10,000 rows cost $6.00 plus $0.005. A discovery search returning 56 datasets costs about $0.034 plus $0.005. On the Free tier, 1,000 rows cost $1.00 plus $0.005.

Never charged: a filter that matches nothing, a wrong domain or dataset ID, a SoQL error, retried requests, and the summary and info rows. Such a run pays only the start fee.

There is no scheduled price change for this actor; these prices have applied since 22 September 2026. The Pricing tab on this page always shows the rate for your own plan; if it and this table ever differ, the Pricing tab is right.

FAQ

What is Socrata? Socrata is the open data platform (now part of Tyler Technologies) behind hundreds of government data portals, such as data.cdc.gov, data.cityofnewyork.us, data.ny.gov and data.wa.gov. Every portal exposes the same public API, which is what this actor uses.

Do I need a Socrata API key or app token? No. Keyless access works for normal runs. A free app token from the portal only raises rate limits; put it in appToken if you pull very large volumes often.

How do I find a dataset ID? Run discovery mode, or open the dataset on its portal: the ID is the four-by-four code at the end of the address, such as 43nn-pn8j.

What is SoQL? Socrata's SQL-like query language. where filters rows, select picks columns and order sorts, and the portal runs all three before sending anything. Text values go in single quotes (boro='Manhattan'); column names are the API field names shown on the dataset's API page, which are often lower case with underscores.

How many rows can I get? Up to 100,000 per run, fetched 1,000 at a time. For a bigger dataset, split it with where (for example by year or borough) across several runs.

How fresh is the data? Every run reads the portal live. How current a dataset is depends on its publisher: the NYC inspection rows in our 2 October test included inspections from 30 September, while some datasets are updated yearly or are frozen. updated_at in discovery mode tells you which.

Which portals does discovery cover? Discovery searches Socrata's US catalog service, which indexes datasets across the portals that list themselves in it. Data mode works on any Socrata domain you name, including ones that discovery does not show.

Why are numbers in quotes? Socrata's JSON API returns numbers as text. Convert them in your spreadsheet or code; sorting with order on the portal is still numeric for number columns.

What happens if my dataset ID or SoQL is wrong? The portal's error is written to the run log (for example 404 dataset.missing for an ID that no longer exists), the run ends with no rows, and nothing but the start fee is charged.

Can I run discovery and data mode together? No. When domain and datasetId are both set, the run is in data mode and searchQuery is ignored. Run discovery first, then data mode.

How do I export the data? From the run's Storage tab as JSON, CSV, Excel, XML or HTML, or through the Apify API. Remove rows whose _type is summary or info if you want data only.

Can I use it from Claude, ChatGPT or another AI assistant?

  • Connector URL: https://mcp.apify.com/?tools=themineworks/socrata-open-data.
  • Claude: Settings > Connectors > Add custom connector, paste the URL, sign in with Apify.
  • ChatGPT: developer mode, add an MCP connector with the URL, sign in with Apify.
  • Cursor or VS Code: add it as an HTTP MCP server with that URL.
  • Claude Code: claude mcp add -t http socrata-open-data "https://mcp.apify.com/?tools=themineworks/socrata-open-data".

Is it legal to use this data? The actor reads public open data APIs that government bodies publish for reuse. Each dataset carries its publisher's own terms, and some datasets (inspections, permits, complaints) name businesses or people, so check the dataset's terms and data protection laws such as GDPR and CCPA before you republish or contact anyone. This is general information, not legal advice. This actor is independent and not affiliated with Tyler Technologies or Socrata.

Integrations

  • Google Sheets: export a run straight to a sheet, or use Apify's Google Sheets integration to refresh a sheet on a schedule.
  • Make, Zapier and n8n: use the Apify app or node to start a run with a fresh where date and push new rows to Slack, email or a database.
  • Webhooks: have Apify call your URL when a run succeeds, then fetch the dataset.
  • API and client libraries: start runs and read datasets from Python, JavaScript or any HTTP client. See the "Copy to your AI assistant" block above for the exact call.
  • MCP clients: Claude, ChatGPT, Cursor, VS Code and other MCP clients can call the actor through https://mcp.apify.com.

More from The Mine Works

Science, health and government data

Social media and video

Leads and business directories

Marketing, SEO and reviews

LinkedIn

Real estate

Jobs and hiring

E-commerce and marketplaces

Company and business data

Food and local services

Developer and AI tools

More tools

Support

Found a bug or need a field we do not return yet? Open an issue on the Issues tab of this page and we will reply there. To ask for a new data source, email dmineworks@gmail.com. A guide for this actor lives at themineworks.com, and a longer tutorial at Socrata API: pull CDC, NYC and other open data portals.

Socrata Open Data Scraper finds government datasets across Socrata portals and pulls their rows with SoQL filters, billed per record delivered.