NCI GDC Cancer Genomics Scraper
Pricing
from $29.62 / 1,000 results
NCI GDC Cancer Genomics Scraper
Scrape projects, cases, files, and annotations from the NCI Genomic Data Commons (GDC) public API. Filter by primary site or program (TCGA / CPTAC / TARGET) and get rich summary fields like case_count, file_count, file_size, disease_type and demographics. No API key required.
Pricing
from $29.62 / 1,000 results
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
0
Monthly active users
10 days ago
Last modified
Categories
Share

๐งฌ NCI Genomic Data Commons (GDC) Cancer Scraper
๐ Export cancer genomics metadata in seconds. Pull TCGA, CPTAC, TARGET, HCMI, BEATAML, and 20+ NCI programs across projects, cases, files, and annotations. No API key, no registration, no manual REST stitching.
The NCI GDC Cancer Scraper queries the NCI Genomic Data Commons public REST API and returns rich records across four entity types: projects, cases, files, and annotations. GDC is the National Cancer Institute's open data platform for cancer research, hosting harmonized genomic and clinical data from TCGA, CPTAC, TARGET, HCMI, BEATAML, MMRF, CCDI, and 20+ other landmark programs.
The catalog covers 50+ primary tumour sites (breast, lung, brain, colon, ovary, pancreas, prostate, kidney, liver, and more) and 26 NCI cancer programs, totalling thousands of projects, hundreds of thousands of cases, and millions of harmonized data files. This Actor exposes both site and program filters at the API level, so disease-specific or program-specific exports are fast.
| ๐ฏ Target Audience | ๐ก Primary Use Cases |
|---|---|
| Cancer biologists, computational oncologists, clinical bioinformaticians, biostatisticians, pharma R&D teams, journalists, regulatory analysts, ML researchers | Cohort discovery, program-level metadata audits, file inventory, annotation tracking, ML training datasets, demographic surveys, cross-program comparisons |
๐ What the NCI GDC Scraper does
Four entity modes in a single Actor:
- ๐ฅ Projects. Project IDs (TCGA-BRCA, CPTAC-3, TARGET-AML, etc.), program affiliation, primary sites, disease types, dbGaP accession, releasable / released state, and a summary block with file count, case count, and total file size in bytes.
- ๐ค Cases. Case IDs, submitter IDs, primary site, disease type, project context, demographic block, diagnoses (with stage, vital status, age at diagnosis), exposures, index date, and timestamps.
- ๐ Files. File IDs, file names, data category (Sequencing Reads, Transcriptome Profiling, etc.), data format (BAM, VCF, TSV, etc.), data type, experimental strategy (WGS, WXS, RNA-Seq, etc.), file size, MD5 sum, access (open / controlled), state, associated cases, and analysis workflow.
- ๐ Annotations. Annotation IDs, entity ID and type, submitter ID, category, classification, notes, status, project context, and timestamps.
Filter any mode by primary site (50 options) or NCI program (26 options) to scope your export. Filters are pushed to the GDC API server-side.
๐ก Why it matters: TCGA and CPTAC underpin most modern cancer-genomics research. Building your own GDC filter compiler and paginator means days of plumbing; this Actor returns ready-joined records on every run.
๐ Data fields
Each record includes: dbgap_accession_number, disease_type, name, primary_site, program, project_id, releasable, released, scrapedAt, state, summary, url. All 12 field names come from a real production run, so what you see here is what lands in your dataset.
๐ How to use
- ๐ Sign up. Create a free account with $5 credit (takes 2 minutes).
- ๐ Open the Actor. Go to the NCI GDC Cancer Scraper page on the Apify Store.
- ๐ฏ Set input. Pick an entity, optionally filter by primary site or program, and set
maxItems. - ๐ Run it. Click Start and let the Actor collect your data.
- ๐ฅ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.
โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.
๐ Recommended Actors
- ๐ค Hugging Face Model Scraper - Model metadata, downloads, and benchmarks
- ๐ฅ FINRA BrokerCheck Scraper - U.S. broker and firm regulatory disclosures
- ๐จ Greatschools Scraper - U.S. school ratings and demographics
- ๐ Smart Apify Actor Scraper - Apify Store actor metadata and quality signals
๐ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.
โ ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the NCI Genomic Data Commons, the National Cancer Institute, or the National Institutes of Health. All trademarks mentioned are the property of their respective owners. Only publicly available open cancer-genomics metadata is collected.
๐ Need Help?
If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.
For faster answers, join our Discord. It's the best place to get support and suggest new actors.