Socrata Open Data API Scraper avatar

Socrata Open Data API Scraper

Pricing

Pay per event

Go to Apify Store
Socrata Open Data API Scraper

Socrata Open Data API Scraper

Pull raw records from any Socrata Open Data (SODA) portal -- NYC, Chicago, and hundreds more US city, county, and state portals. Give us portal domains and dataset IDs; we page through results and retry on throttling, returning typed rows for 311 complaints, crime data, inspections, or any dataset.

Pricing

Pay per event

Rating

0.0

(0)

Developer

DevilScrapes

DevilScrapes

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 hours ago

Last modified

Categories

Share


🎯 What this scrapes

Several hundred US (and some EU) city, county, and state governments publish their open data through the Socrata Open Data API (SODA) — a standard /resource/<dataset-id>.json REST interface with a SQL-like query language (SoQL). Every portal exposes a different dataset: NYC's 311 complaints, Chicago's crime reports, Seattle's permits, LA's budget lines — same protocol, completely different columns.

This Actor is a generic, cross-portal harvester, not a single-vertical scraper. You supply a list of {portal_domain, dataset_id} pairs (plus optional SoQL filters), and we return every matching row from every target in one run — whatever shape that dataset happens to be.

🔥 Features — what we handle for you

  • 🔁 We retry with exponential backoff on 429 and 5xx responses — up to 5 attempts per page, honouring Retry-After when a portal sends it.
  • 📄 We page every target for you$limit/$offset walked automatically until the dataset is exhausted or your cap is hit.
  • 🧩 We fan out across portals in one run — mix a 311 feed and a crime feed and get valid, disjoint-column rows for both, no schema conflicts.
  • 🧱 We isolate failures per target — an unknown or private dataset is logged and skipped; your other targets still complete.
  • 🧊 We keep the dataset clean — Pydantic-validated rows, ISO-8601 timestamps, a stable record_id when the dataset provides one.
  • 💰 You pay only for results that land. No data → no charge beyond the small warm-up fee.

💡 Use cases

  • Civic analytics — pull a city's 311 or crime feed on a schedule and track trends over time.
  • Journalism / research — cross-reference multiple portals' datasets without writing a client for each one.
  • Proptech / real-estate diligence — pull permit, inspection, or code violation datasets for a market you're evaluating.
  • Academic / policy research — bulk-export any Socrata dataset into a typed, analysis-ready format.
  • Compliance monitoring — watch an inspections or recalls dataset for new entries.

⚙️ How to use it

  1. Click Try for free at the top of the page.
  2. Add one or more targets — each is {portal_domain, dataset_id} plus optional soql_where / soql_select / soql_order / label. The prefilled example already scrapes NYC 311 (NYPD-filtered) and Chicago crime so you can see multi-portal fan-out on the first run.
  3. Click Start. Output streams into the run's dataset as each target finishes paging.
  4. Export from Storage → Dataset as JSON, CSV, or Excel — or fetch via the API.

📥 Input

FieldTypeRequiredDefaultNotes
targetsarrayyes2-item NYC 311 + Chicago crime prefillList of {portal_domain, dataset_id, soql_where?, soql_select?, soql_order?, label?}
page_sizeintegerno1000Internal $limit per page request (1–2000)
max_items_per_targetintegerno5000Row cap per target; 0 = unlimited
proxy_configurationobjectno{"useApifyProxy": true}Apify Proxy spec

Example input

{
"targets": [
{
"portal_domain": "data.cityofnewyork.us",
"dataset_id": "erm2-nwe9",
"soql_where": "agency='NYPD'",
"label": "NYC 311"
},
{
"portal_domain": "data.cityofchicago.org",
"dataset_id": "ijzp-q8t2",
"label": "Chicago crime"
}
],
"page_size": 1000,
"max_items_per_target": 20,
"proxy_configuration": { "useApifyProxy": true }
}

📤 Output

Every row is one dataset item. data carries the raw Socrata row exactly as returned — untyped by design, since the target dataset's columns are unknown at build time and change from portal to portal.

FieldTypeNotes
portal_domainstringSocrata hostname the row came from
dataset_idstringSocrata 4x4 dataset id
target_labelstring | nullEchoed from the target's label input field
record_idstring | nullThe row's :id, when the dataset provides one
scraped_atstringISO-8601 UTC fetch timestamp
dataobjectThe raw Socrata row, untouched

Example output

{
"portal_domain": "data.cityofnewyork.us",
"dataset_id": "erm2-nwe9",
"target_label": "NYC 311",
"record_id": "46882077",
"scraped_at": "2026-08-31T09:15:03.000Z",
"data": {
"unique_key": "46882077",
"agency": "NYPD",
"complaint_type": "Noise - Vehicle",
"status": "Closed"
}
}

💰 Pricing

Pay-Per-Event — you pay only when these events fire. Pricing is not yet live on this listing; the figures below are the provisional target, matching our existing Socrata Actor (us-building-permit-leads-scraper), and will be confirmed in the Apify Console before this Actor is published.

EventUSDWhat it is
actor-start$0.20One-off warm-up charge per run
result$0.002Per dataset row emitted

Example: 1 000 results at the rates above ≈ $2.20. No subscription, no minimum, no card to start — Apify gives every new account free credit.

🚧 Limitations

Only the /resource/<dataset-id>.json SODA2 endpoint is supported — metadata/catalog APIs and CKAN portals (a different protocol some cities use instead of Socrata) are out of scope. record_id is null whenever the source dataset has no :id column, which is common and expected, not a bug. This Actor does not normalize columns across datasets — data is the raw row for whatever dataset you point it at.

❓ FAQ

Do I need a Socrata account or API key?

No. Public Socrata datasets return data with no $app_token — this Actor runs entirely keyless.

What happens if I give it a bad dataset_id?

That target is logged and skipped; the run still finishes successfully and your other targets complete normally.

What if my soql_where filter matches nothing?

The run finishes successfully with zero rows for that target — an empty result is not an error.

Can I scrape any Socrata portal, or only NYC/Chicago?

Any Socrata portal — several hundred US city, county, and state governments run one. The prefill uses NYC and Chicago only to demonstrate multi-portal fan-out in a single run.

Why is data a raw untyped object instead of fixed columns?

Because the dataset is unknown at build time — a 311 feed and a crime feed share zero columns. Forcing them into one fixed schema would mean dropping real data; data keeps every column the source portal returned.

💬 Your feedback

Spotted a bug, hit a weird edge case, or need a new field? Open an issue on the Actor's Issues tab on Apify Console — we ship fixes weekly and we read every report.