NCI Genomic Data Commons Projects Scraper
Pricing
from $3.40 / 1,000 result items
NCI Genomic Data Commons Projects Scraper
Pull project-level metadata from the NCI Genomic Data Commons public API. Retrieves project id, name, disease type, primary site, dbGaP accession number, state, and releasable status. Ideal for cancer researchers compiling study inventories.
Pricing
from $3.40 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
NCI GDC Projects Scraper
Scrape NCI Genomic Data Commons project listings. Every project comes with its ID, disease type, primary site, and release state. No login or API key. Export to CSV, JSON, Excel, or XML.
The NCI Genomic Data Commons web portal is built for browsing, not bulk export. Researchers who need the full project catalog for meta-analysis or grant writing have to click through pages or write their own API scripts. This Actor reads the public GDC API directly, filters by keyword, and returns every matching project in one flat dataset.
| Who uses it | What they scrape NCI Genomic Data Commons for |
|---|---|
| Bioinformatics researchers | Catalog available genomic studies for a specific cancer type before requesting access. |
| Grant writers | Compile a list of funded projects and their disease focus to support a funding application. |
| Data librarians | Build a searchable index of GDC projects for their institution's internal data portal. |
| Cancer epidemiologists | Identify all projects linked to a particular primary site for a cross-study analysis. |
What it does
This Actor collects NCI GDC project metadata from the public API and returns each project as a flat row with its ID, name, disease type, primary site, and release status.
- š Keyword search: filter projects by name or any text field, matching the GDC site's own search behavior.
- š Structured output: each project is a flat row with consistent columns, ready for analysis in any tool.
- ā” Bulk extraction: pull project records in a single run, bypassing the portal's page-by-page browsing.
Results export to CSV, JSON, Excel, or XML, or straight from the API.
What you can do with NCI Genomic Data Commons data
𧬠Map the cancer research landscape.
A bioinformatician scrapes all GDC projects, groups them by disease_type, and identifies which cancers have the most open-access genomic data available.
š Support a grant proposal with data.
A principal investigator extracts every project related to 'lung adenocarcinoma' and lists the dbGaP accession numbers as evidence of existing research infrastructure.
š„ Build an institutional data catalog.
A data librarian runs the Actor weekly, ingests the JSON output into a local database, and lets researchers search GDC projects from the institution's own portal.
š¬ Find projects by tissue site.
An epidemiologist searches for 'breast' across all project names and primary_site fields, then downloads the resulting project list to plan a meta-analysis.
Why choose this scraper
| What you get | |
|---|---|
| No API coding | The Actor calls the GDC API for you and handles pagination and rate limits. |
| Fixed schema | Every run returns the same columns, so your downstream scripts never break. |
| Keyword filtering | Narrow results to projects matching a disease name, tissue, or any text field. |
| Full catalog access | Retrieve all public project metadata without clicking through the web portal. |
What a NCI Genomic Data Commons record looks like
Every record returns as one flat JSON row. Here is a real one from a run:
{"id": "TCGA-LUAD","projectId": "TCGA-LUAD","name": "Lung Adenocarcinoma","releasable": true,"state": "open","released": true,"scrapedAt": "2026-09-26T13:55:44.944Z"}
Every value above comes from a real run. A field a record does not have comes back as null.
Configure the run
Drive the Actor with an optional search term to filter projects by name, and set a maximum item count to control the size of your dataset. The Input tab lists every parameter.
A first run with the defaults:
{"maxItems": 10}
A larger pull:
{"maxItems": 200}
Free users
Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect more results per run.
Run it
- Create a free Apify account.
- Set your inputs and any filters, then click Start.
- Export the results as CSV, Excel, JSON, or XML from the Dataset tab.
Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.
Use with AI agents (MCP)
Give an AI agent live access to NCI Genomic Data Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:
$undefined
Then prompt it in plain language to run the scraper and read back the results.
Troubleshooting
Why am I getting fewer results than expected?
Check the max items setting. Free users are capped at 10 items. If you are on a paid plan, make sure your search term is not too restrictive. Try broadening the keyword or leaving it empty to see all projects.
The Actor returns an empty dataset.
Your search term may not match any project name or text field. Try a different keyword, check for typos, or run without a search term to confirm the API is reachable.
I need a field that is not in the output.
The Actor returns the core project metadata from the GDC API. If you need additional fields, check the GDC API documentation to see if they are available, and contact support with the field names you require.
The run fails with a timeout error.
The GDC API may be experiencing high load. Reduce the max items count and retry, or schedule the run for a less busy time. The API is public and rate-limited by the server.
FAQ
| Question | Answer |
|---|---|
| Do I need an API key or login to scrape GDC projects? | No. The NCI GDC API is fully public and requires no authentication. The Actor calls the endpoint directly. |
| What data does each project row contain? | Each row includes the project ID, name, disease type, primary site, dbGaP accession number, release state, and whether the project is releasable. The exact fields are shown in the sample output on the Actor's page. |
| Can I filter projects by a specific cancer type? | Yes. Use the search term input to filter by any text that appears in the project name or other text fields. For example, entering 'melanoma' returns only projects that mention melanoma. |
| How many projects can I scrape in one run? | Free users can scrape up to 10 projects as a preview. Paid users can set the max items up to 1,000,000, which covers the entire GDC project catalog many times over. |
| What format is the output? | The Actor returns data as a flat dataset. You can export it to CSV, JSON, Excel, or XML from the Apify platform. |
| Does this Actor scrape the GDC web portal HTML? | No. It calls the official GDC REST API directly, which is faster, more reliable, and returns structured JSON that the Actor flattens into rows. |
| Can I get the list of all GDC projects without any filter? | Yes. Leave the search term empty and set a high max items value to retrieve the complete public project catalog. |
| Is the data from the GDC API live? | Yes. Every run calls the live GDC API, so you always get the current project listings, including newly added projects. |
Related actors
Browse the full ParseForge collection for more scrapers.
š Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.
ā ļø Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by National Cancer Institute. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.
