NCI GDC Cancer Genomics Scraper avatar

NCI GDC Cancer Genomics Scraper

Pricing

from $29.62 / 1,000 results

Go to Apify Store
NCI GDC Cancer Genomics Scraper

NCI GDC Cancer Genomics Scraper

Scrape projects, cases, files, and annotations from the NCI Genomic Data Commons (GDC) public API. Filter by primary site or program (TCGA / CPTAC / TARGET) and get rich summary fields like case_count, file_count, file_size, disease_type and demographics. No API key required.

Pricing

from $29.62 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

0

Monthly active users

10 days ago

Last modified

Share

ParseForge Banner

๐Ÿงฌ NCI Genomic Data Commons (GDC) Cancer Scraper

๐Ÿš€ Export cancer genomics metadata in seconds. Pull TCGA, CPTAC, TARGET, HCMI, BEATAML, and 20+ NCI programs across projects, cases, files, and annotations. No API key, no registration, no manual REST stitching.

The NCI GDC Cancer Scraper queries the NCI Genomic Data Commons public REST API and returns rich records across four entity types: projects, cases, files, and annotations. GDC is the National Cancer Institute's open data platform for cancer research, hosting harmonized genomic and clinical data from TCGA, CPTAC, TARGET, HCMI, BEATAML, MMRF, CCDI, and 20+ other landmark programs.

The catalog covers 50+ primary tumour sites (breast, lung, brain, colon, ovary, pancreas, prostate, kidney, liver, and more) and 26 NCI cancer programs, totalling thousands of projects, hundreds of thousands of cases, and millions of harmonized data files. This Actor exposes both site and program filters at the API level, so disease-specific or program-specific exports are fast.

๐ŸŽฏ Target Audience๐Ÿ’ก Primary Use Cases
Cancer biologists, computational oncologists, clinical bioinformaticians, biostatisticians, pharma R&D teams, journalists, regulatory analysts, ML researchersCohort discovery, program-level metadata audits, file inventory, annotation tracking, ML training datasets, demographic surveys, cross-program comparisons

๐Ÿ“‹ What the NCI GDC Scraper does

Four entity modes in a single Actor:

  • ๐Ÿฅ Projects. Project IDs (TCGA-BRCA, CPTAC-3, TARGET-AML, etc.), program affiliation, primary sites, disease types, dbGaP accession, releasable / released state, and a summary block with file count, case count, and total file size in bytes.
  • ๐Ÿ‘ค Cases. Case IDs, submitter IDs, primary site, disease type, project context, demographic block, diagnoses (with stage, vital status, age at diagnosis), exposures, index date, and timestamps.
  • ๐Ÿ“ Files. File IDs, file names, data category (Sequencing Reads, Transcriptome Profiling, etc.), data format (BAM, VCF, TSV, etc.), data type, experimental strategy (WGS, WXS, RNA-Seq, etc.), file size, MD5 sum, access (open / controlled), state, associated cases, and analysis workflow.
  • ๐Ÿ“ Annotations. Annotation IDs, entity ID and type, submitter ID, category, classification, notes, status, project context, and timestamps.

Filter any mode by primary site (50 options) or NCI program (26 options) to scope your export. Filters are pushed to the GDC API server-side.

๐Ÿ’ก Why it matters: TCGA and CPTAC underpin most modern cancer-genomics research. Building your own GDC filter compiler and paginator means days of plumbing; this Actor returns ready-joined records on every run.

๐Ÿ“Š Data fields

Each record includes: dbgap_accession_number, disease_type, name, primary_site, program, project_id, releasable, released, scrapedAt, state, summary, url. All 12 field names come from a real production run, so what you see here is what lands in your dataset.

๐Ÿš€ How to use

  1. ๐Ÿ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. ๐ŸŒ Open the Actor. Go to the NCI GDC Cancer Scraper page on the Apify Store.
  3. ๐ŸŽฏ Set input. Pick an entity, optionally filter by primary site or program, and set maxItems.
  4. ๐Ÿš€ Run it. Click Start and let the Actor collect your data.
  5. ๐Ÿ“ฅ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.

โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

โš ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the NCI Genomic Data Commons, the National Cancer Institute, or the National Institutes of Health. All trademarks mentioned are the property of their respective owners. Only publicly available open cancer-genomics metadata is collected.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.