data.gov.in Scraper | Any Dataset ID, Up to 100k Rows
DeprecatedPricing
from $1.00 / 1,000 records
data.gov.in Scraper | Any Dataset ID, Up to 100k Rows
DeprecatedScrape any India open government dataset from data.gov.in via the official OGD API: foreign trade, mandi prices, census, agriculture. Filter and paginate to clean JSON. As of Oct 2026 data.gov.in is refusing Apify servers; nothing is charged when it refuses.
Pricing
from $1.00 / 1,000 records
Rating
0.0
(0)
Developer
The Mine Works
Maintained by CommunityActor stats
0
Bookmarked
17
Total users
6
Monthly active users
2 hours ago
Last modified
Categories
Share
From The Mine Works, makers of Threads Scraper and B2B Leads Finder, with over 140,000 runs across 170+ public actors.
data.gov.in is the Government of India's open data portal: thousands of datasets from ministries, states and public bodies, from daily mandi (wholesale market) prices to foreign trade, census tables, health and energy figures. Every dataset has a resource ID and is served by the official OGD (Open Government Data) API. This actor takes one resource ID, pages through that API for you, has the API apply your field filters, and returns the dataset as flat rows with the resource ID and a timestamp added to each.
Current status, 2 Oct 2026. data.gov.in has been refusing requests from Apify's servers. On 30 Sep 2026 our runs could not connect to api.data.gov.in at all, and on 2 Oct a request through Apify's datacenter proxy was refused as well. The shared sample key was rate limited on every run we recorded from 1 Sep 2026. While this lasts, a run ends with a summary row that says exactly what happened, and no record is charged. Your own free key solves the rate limit, not a refused connection. We could not test a run with a personal key, because we have none and the API refuses our servers.
Why choose this actor?
- Nothing charged when data.gov.in says no. A refused connection, a refused or rate limited key, an error page or an empty result ends the run with a
statusand a plainnotein the summary row, and $0 for records. Our 2 Oct 2026 test against the live API ended in 1.7 seconds withstatus: "api_unreachable"and nothing charged. - Up to 100,000 rows per run, 100 per request. The actor follows the API's paging until it reaches your
maxResultsor the dataset's own total, and the summary row reports that total, so you know how much of the dataset you pulled. - Filters run on the government's side.
field=valuefilters become the API's ownfilters[field]parameters, so a state or commodity filter narrows what is sent and you pay only for rows you keep. No login, no browser, no cookies.
Part of The Mine Works Science, health and government data family: CourtListener Scraper, Socrata Open Data Scraper, Academic Research MCP, OpenAlex Scraper, FDA 510(k) Clearances Scraper, Crossref Scraper.
Try it in one minute
In the input page's JSON view, paste this, put your own data.gov.in key in apiKey, and press Start. It asks for 10 onion price rows from Maharashtra's mandis.
{"resourceId": "9ef84268-d588-465a-a308-a864a43d0070","apiKey": "YOUR_FREE_DATA_GOV_IN_KEY","filters": ["state=Maharashtra", "commodity=Onion"],"maxResults": 10}
The actor takes one dataset per run. Give resourceId as the 36-character resource ID from the dataset's API tab on data.gov.in, or paste the dataset's API URL (such as https://api.data.gov.in/resource/9ef84268-d588-465a-a308-a864a43d0070?format=json) and the ID inside it is used. Text with no resource ID in it stops the run at once, before any request or charge.
Apify's free plan includes $5 of credit every month, which covers about 2,497 results at this actor's price (one run at the default 512 MB).
Copy to your AI assistant
themineworks/india-data-gov-scraper on Apify. It pulls one dataset from India's data.gov.in open data API by resource ID and returns its records as flat JSON rows with _resource_id and _scraped_at added, plus a summary row whose status and note say whether the API answered. Call ApifyClient("TOKEN").actor("themineworks/india-data-gov-scraper").call(run_input={...}), then client.dataset(run["defaultDatasetId"]).list_items().items. Required: resourceId (the 36-character data.gov.in resource ID, or a URL that contains it). Optional: apiKey (your free data.gov.in key; blank uses the shared sample key, which data.gov.in rate limits), filters (list of "field=value" strings, default none), maxResults (default 1000, up to 100000), monitorMode (default false; true delivers and charges only records not delivered by an earlier run with the same input). Full spec: GET https://api.apify.com/v2/acts/themineworks~india-data-gov-scraper/builds/default (Bearer TOKEN), which returns inputSchema and readme. Token: https://console.apify.com/account/integrations?fpr=ymnoit&utm_source=apify-readme&utm_medium=referral
Key features
- Any dataset with an API resource ID on data.gov.in, one per run: mandi prices, foreign trade, census, agriculture, health, energy, transport and more.
- Every column the dataset publishes, unchanged, plus 2 audit fields on every record:
_resource_idand_scraped_at. - 100 records per request, up to 100,000 per run, paged automatically and stopped at your
maxResultsor the dataset's total. - AND filters as
field=value, applied by the API before records are sent. Entries without that form are ignored and listed in the summary row. - A summary row on every run with one of 14
statusvalues (ok,partial,no_records,sample_key_rate_limited,key_refused,api_unreachableand others) and anotein plain words whenever something went wrong. - Fails fast on a missing ID: an empty
resourceId, or text with no resource ID in it, stops the run before any request, with a message showing what an ID looks like.
How to use it
Basic: one dataset
{"resourceId": "9ef84268-d588-465a-a308-a864a43d0070","apiKey": "YOUR_FREE_DATA_GOV_IN_KEY","maxResults": 1000}
This is the daily mandi price dataset (current daily price of various commodities from various markets). Records come back with that dataset's own columns, such as state, district, market, commodity, variety, arrival_date and the min_price, max_price and modal_price values.
One state and one commodity
{"resourceId": "9ef84268-d588-465a-a308-a864a43d0070","apiKey": "YOUR_FREE_DATA_GOV_IN_KEY","filters": ["state=Maharashtra", "commodity=Onion"],"maxResults": 500}
Field names must match the dataset's column names exactly. Run once without filters and a small maxResults to see them.
Size a dataset before a big pull
{"resourceId": "9ef84268-d588-465a-a308-a864a43d0070","apiKey": "YOUR_FREE_DATA_GOV_IN_KEY","maxResults": 1}
One record plus the summary row, whose total_available and dataset_title tell you how big the dataset is. On the Free plan that costs $0.002 plus the $0.005 start fee.
Daily mandi price feed
Save the filtered input above as a task and add a schedule in Apify Console (Schedules, then Create new), for example every morning with the cron 0 9 * * *. Each run pulls that day's rows; the arrival_date column tells you which day a price belongs to. A run that data.gov.in refuses costs only the start fee, and its summary row says why, so a broken day is easy to spot in an integration.
A large backfill with your own key
{"resourceId": "9ef84268-d588-465a-a308-a864a43d0070","apiKey": "YOUR_FREE_DATA_GOV_IN_KEY","maxResults": 100000}
100,000 records take 1,000 requests. The default run timeout is 300 seconds; for a pull this size, raise the timeout in the run options. If a run reaches its time limit first, it ends cleanly: the records it saved stay in your dataset, are the only ones charged, and the summary row says status: "partial" or "timed_out".
Input parameters
| Parameter | Type | Default | What it does |
|---|---|---|---|
resourceId | string, required | none (prefill 9ef84268-d588-465a-a308-a864a43d0070) | The data.gov.in resource ID, or a URL that contains it. Text with no resource ID stops the run before any request. |
apiKey | string, secret | blank (shared sample key) | Your free data.gov.in API key. Blank uses data.gov.in's shared sample key, which data.gov.in rate limits. |
filters | array of strings | none (prefill state=Maharashtra, commodity=Onion) | field=value filters, combined with AND and applied by data.gov.in. |
maxResults | integer, 1 to 100,000 | 1000 (prefill 50) | Most records to return in this run. |
monitorMode | boolean | false | Deliver only records not delivered by an earlier run with the same input (a record with changed values counts as new). Skipped records are not charged. Made for daily feeds. |
What data do you get?
Records: one row per dataset record, with every column the dataset publishes, unchanged. The columns depend on the dataset you ask for: the mandi price dataset has state, district, market, commodity, variety, grade, arrival_date, min_price, max_price and modal_price; a trade or census dataset has its own. Two fields are added to every record:
_resource_id: the resource ID the row came from, so rows from several datasets can share one table._scraped_at: when the row was saved, as an ISO timestamp.
The summary row: every run ends with one _type: "summary" row. It always has records (rows saved), charged_for (rows charged), status and scraped_at. With a valid ID it has resource_id, and when the API answered it adds dataset_title and total_available (the dataset's own record count). When something went wrong it has a note that says what happened and what to do. Filters that were not in field=value form are listed in ignored_filters.
What status means:
status | What happened | Charged |
|---|---|---|
ok | The run finished normally. | Records saved |
no_records | The API answered but sent no records for this ID and these filters. | Nothing |
no_new_records | Monitor mode: the API answered, and every record was already delivered by an earlier run with this input. | Nothing |
partial | Records were saved, then the API stopped answering or the time limit came. | Records saved |
sample_key_rate_limited | data.gov.in refused the shared sample key (HTTP 429). | Nothing |
rate_limited | data.gov.in rate limited your own key (HTTP 429) before the first record. | Nothing |
key_refused | data.gov.in refused the API key (HTTP 401 or 403, or a key error message). | Nothing |
resource_not_found | data.gov.in has no dataset with this ID. | Nothing |
api_unreachable | No answer at all: connection refused, timed out or DNS failed. | Nothing |
api_error, unexpected_response | Another HTTP error, or a page that is not the usual JSON. | Nothing |
timed_out | The run reached its time limit before the first record. | Nothing |
invalid_input | resourceId was empty or had no resource ID in it. The run stops before any request. | Nothing |
error | Something unexpected stopped the run; the note says what. | Records saved before it |
When at least one record was saved, the run also adds a _type: "info" row with a scheduling tip. Neither row is charged. Filter on _type to keep only records.
Stable fields for automations
| Field | Where | What it is |
|---|---|---|
_resource_id | every record | The resource ID the record came from |
_scraped_at | every record | ISO timestamp of the save |
_type | summary row | Always "summary" |
records | summary row | Records saved in this run |
charged_for | summary row | Records charged in this run |
status | summary row | One of the values in the table above |
scraped_at | summary row | ISO timestamp of the summary |
note | summary row, whenever status is not ok | What happened, in plain words |
These names will not change. The dataset's own columns are data.gov.in's and can change when the publishing body changes the dataset.
Output examples
A record from the daily mandi price dataset (a real row, captured on 15 Jul 2026, when the API still answered Apify):
{"state": "Keralam","district": "Idukki","market": "Kattappana Market","commodity": "Water Melon","variety": "Other","grade": "Medium","arrival_date": "15/07/2026","min_price": 2000,"max_price": 2600,"modal_price": 2300,"_resource_id": "9ef84268-d588-465a-a308-a864a43d0070","_scraped_at": "2026-07-15T04:14:45.668Z"}
The summary row when data.gov.in refuses the connection (a real row from our test of this version on 2 Oct 2026, through Apify's datacenter proxy):
{"_type": "summary","resource_id": "9ef84268-d588-465a-a308-a864a43d0070","records": 0,"charged_for": 0,"status": "api_unreachable","note": "Could not reach api.data.gov.in (the proxy could not connect to it, ERR_GOT_REQUEST_ERROR). Nothing was charged. Since 30 Sep 2026 the data.gov.in API has often refused connections from Apify's servers; run it again later.","scraped_at": "2026-10-02T11:33:16.981Z"}
The summary row when resourceId has no ID in it (a real row from the same day; the run stopped before any request):
{"_type": "summary","records": 0,"charged_for": 0,"status": "invalid_input","note": "\"mandi prices\" is not a data.gov.in resource ID. A resource ID is a 36-character code such as 9ef84268-d588-465a-a308-a864a43d0070; copy it from the dataset's API tab on data.gov.in. Nothing was requested or charged.","scraped_at": "2026-10-02T11:30:46.094Z"}
Pricing
A small fee when a run starts, then a price per record saved; higher Apify plans pay less per record.
| Event | Free plan | Bronze (Starter) | Silver (Scale) | Gold and above (Business) |
|---|---|---|---|---|
record-scraped, per record | $0.002 | $0.0016 | $0.00125 | $0.001 |
| Per 1,000 records | $2.00 | $1.60 | $1.25 | $1.00 |
apify-actor-start, per run | $0.005 per GB of run memory, minimum one | same | same | same |
- Start fee: once per run, Apify's start event bills $0.005 per GB of memory with a minimum of one event, so $0.005 at the default 512 MB. It has applied since 14 Sep 2026, and it is the only charge on a run that data.gov.in refuses.
- Never charged: the
summaryandinforows, refused or rate limited keys, failed requests, error pages, empty results and inputs with no resource ID. Compute costs are already in these prices. - Spending cap: a maximum cost per run set in Apify's run options is respected; once it is used up, no further records are saved or charged.
- No price change is scheduled.
Worked examples: sizing a dataset with maxResults: 1 costs $0.007 on the Free plan. 10,000 records on Gold cost $10.005.
Run it on a schedule
Switch on monitorMode and the actor remembers which records it has delivered for this input. Each later run delivers only records you have not received before, and you pay only for those. data.gov.in records have no ID of their own, so a record is recognised by all of its values: a record whose values changed (a revised price) comes back as a new row.
- Fill in the input, tick Monitor mode, and save it as a task.
- In Apify Console open Schedules, click Add schedule, and pick Daily (or a cron such as
0 9 * * *). - Add the saved task to the schedule and click Save.
{"resourceId": "9ef84268-d588-465a-a308-a864a43d0070","apiKey": "YOUR_FREE_DATA_GOV_IN_KEY","filters": ["state=Maharashtra", "commodity=Onion"],"maxResults": 1000,"monitorMode": true}
The first run delivers every record it reads. Each later run with the same input reads the dataset again and delivers only records it has not delivered before, up to maxResults of them. Skipped records are never charged; the start fee applies to every run as usual, and a run with nothing new says status: "no_new_records" in its summary row. The history keeps the last 50,000 records delivered. For a large dataset that grows over time, add a filter (for example on a date or state column) so each run reads a small slice. The last row of each run (_type: "info") gives new_this_run and skipped_already_seen. The history belongs to the input: change resourceId or filters and a new history starts; change maxResults or apiKey and it carries on.
FAQ
What is data.gov.in? The Open Government Data Platform India, run by the Government of India. Ministries, states and public bodies publish datasets there, and most of them can be read through its OGD API by resource ID. This actor reads that API; it is not affiliated with or endorsed by the Government of India.
Do I need an API key?
In practice, yes. Without one the actor uses data.gov.in's shared public sample key, and data.gov.in refused that key with HTTP 429 (rate limit exceeded) on every run we recorded from 1 Sep 2026. A personal key is free: sign in at data.gov.in, open My Account and generate a key, then put it in apiKey. The field is stored as a secret.
Why did my run end with status: "api_unreachable"?
The actor could not get any answer from api.data.gov.in: the connection was refused, timed out or the name did not resolve. That happened to every run we made from 30 Sep 2026, from Apify's servers and through Apify's datacenter proxy. Nothing is charged except the start fee. Run it again later; your key does not change this.
How do I find a resource ID?
Open the dataset on data.gov.in, go to its API tab, and copy the 36-character ID (the part after /resource/ in the API address). Pasting the whole API address into resourceId also works.
Which fields will I get?
Whatever the dataset publishes, unchanged, plus _resource_id and _scraped_at on every record. Different datasets have different columns.
How do I know what to filter on?
Run once without filters and with a small maxResults, read the column names in the output, then filter on those exact names. If your filters match nothing, the run ends with status: "no_records" and no charge.
How many records can I pull in one run? Up to 100,000, fetched 100 per request. For large pulls, raise the run timeout above the 300 second default; a run that reaches its limit keeps and charges only the records it saved.
What happens if the resource ID is wrong?
If it has no resource ID in it at all, the run stops at once with status: "invalid_input", before any request. If it looks like an ID but data.gov.in has no such dataset, the run ends with status: "resource_not_found" or "api_error" and the API's own message. Nothing is charged in either case.
Does it use a proxy? No. It calls the official API directly from Apify's servers, with no browser and no cookies.
Can I run it on a schedule?
Yes. Save your input as a task and add a schedule in Apify Console (Schedules, then Create new). The summary row's status lets an integration tell a good day from a refused one. Set monitorMode to true and each run delivers, and charges, only records not delivered by an earlier run with the same input.
How do I export the data? From the run's Storage tab as JSON, CSV, Excel, XML or HTML, or through the Apify API.
Can I use it from Claude, ChatGPT or another AI assistant?
- Connector URL:
https://mcp.apify.com/?tools=themineworks/india-data-gov-scraper. - Claude: Settings > Connectors > Add custom connector, paste the URL, sign in with Apify.
- ChatGPT: developer mode, add an MCP connector with the URL, sign in with Apify.
- Cursor or VS Code: add it as an HTTP MCP server with that URL.
- Claude Code:
claude mcp add -t http india-data-gov-scraper "https://mcp.apify.com/?tools=themineworks/india-data-gov-scraper".
Is this data free to use? data.gov.in publishes datasets under the Government Open Data License, India. Check the licence notes on each dataset's page for attribution and reuse terms. The actor reads only public open data; you are responsible for how you use it, including data protection law such as GDPR and CCPA if a dataset holds personal information.
Integrations
- Google Sheets: send each run's records to a sheet with Apify's Google Sheets integration, for example a daily mandi price log.
- Make, Zapier and n8n: start a run on a schedule and route the records, and check the summary row's
statusbefore you use them. - Webhooks: have Apify call your system when a run finishes.
- API: run it and read the dataset over HTTP, or with the Python and JavaScript clients.
- MCP clients: Claude, ChatGPT and Cursor can call the actor through
https://mcp.apify.com.
More from The Mine Works
Science, health and government data
- CourtListener Scraper
- Socrata Open Data Scraper
- Academic Research MCP
- OpenAlex Scraper
- FDA 510(k) Clearances Scraper
- Crossref Scraper
- PubMed Scraper
- arXiv Scraper
- NPI Registry Scraper
- OpenCitations Scraper
- CMS Hospital Quality Scraper
- FDA Recalls Scraper
Social media and video
Leads and business directories
Marketing, SEO and reviews
Real estate
Jobs and hiring
E-commerce and marketplaces
Company and business data
Food and local services
Developer and AI tools
More tools
- Tennis Match & Player Data Scraper
- Google Hotels Prices Scraper
- LandWatch Scraper
- Taobao Products Scraper 淘宝 天猫
Support
Questions and bug reports: the Issues tab. New data sources: dmineworks@gmail.com.
data.gov.in Scraper pulls any Indian open government dataset by resource ID as flat rows, up to 100,000 a run from $1.00 per 1,000 records, and charges nothing for records when data.gov.in refuses the request.

