# NCI GDC Cases Scraper (`acquistion-automation/nci-gdc-cases-scraper`) Actor

Collects cancer case records from the NCI Genomic Data Commons API, filtered by project ID or primary site, and returns each case as a flat row with clinical identifiers, demographics, and biospecimen counts.

- **URL**: https://apify.com/acquistion-automation/nci-gdc-cases-scraper.md
- **Developed by:** [Acquisition Automation Co.](https://apify.com/acquistion-automation) (community)
- **Categories:** Automation, Integrations, Education
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $7.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

![Acquisition Automation Co. Search less. Close more.](https://api.apify.com/v2/key-value-stores/AOdPHdOpeDpzEPS5f/records/banner.jpg)

## 🧬 NCI GDC Cases Scraper

> **Export cancer case records from the National Cancer Institute's Genomic Data Commons as flat rows: case ID, submitter ID, project, primary site, disease type, race, age at diagnosis and how many diagnoses the case carries.** No API key, no registration, no login.

The GDC holds the case records behind TCGA, TARGET and the other NCI programs. The portal shows them as a faceted browser and the API answers in nested JSON, so counting cases by project or primary site means either clicking through facets or writing a client. This Actor queries the API, flattens each case to one row, and writes a dataset you can open in a spreadsheet.

| Who uses it | What they use the cases for |
|---|---|
| 🔬 Cancer researchers | Pulling a cohort by project or primary site before requesting the underlying files |
| 🧪 Bioinformatics teams | Building a case index to join against GDC file and mutation data |
| 🏥 Clinical trial feasibility analysts | Counting how many cases a site or disease type has on record |
| 💼 Biotech corporate development | Sizing the public evidence base behind an oncology asset or a target company's indication |

### 📋 What it does

> 💡 **Why it matters:** the first question about any cohort is how many cases there are and how they split. That answer is one run here, and a morning of facet clicking in the portal.

- 🎯 **Filter by project.** `projectId` takes a GDC project ID such as `TCGA-LUAD`.
- 🫁 **Or by primary site.** `primarySite` takes the site as the GDC writes it, for example `Bronchus and lung`.
- 🆔 **Both identifiers on every row.** The GDC `case_id` UUID and the program's own `submitter_id`, for example `TCGA-44-3918`, so rows join to anything else you hold.
- 📊 **Disease type, race and age at diagnosis** as the GDC publishes them, with no recoding.
- 🔗 **A direct portal link** per case, so any row can be opened and checked.
- 💾 **Exports to CSV, Excel, JSON or XML**, from the run page or the API.

### 📊 Output

Every case is one flat row with 12 fields. `null` means the GDC publishes no value for that case.

| Field | Type | Description |
|---|---|---|
| 🆔 `case_id` | string | GDC case UUID |
| 🔖 `submitter_id` | string | The submitting program's own case identifier, for example `TCGA-44-3918` |
| 📁 `project_id` | string | GDC project the case belongs to, for example `TCGA-LUAD` |
| 🫁 `primary_site` | string | Primary site as the GDC writes it, for example `Bronchus and lung` |
| 🧫 `disease_type` | string | Disease type, for example `Adenomas and Adenocarcinomas` |
| 🚻 `gender` | string | Demographic gender where the GDC publishes it. `null` on every row in the runs checked |
| 🧑 `race` | string | Race as recorded, for example `white` |
| 📅 `ageAtDiagnosis` | integer | Age at diagnosis in days, the GDC's own unit. Divide by 365.25 for years |
| 🔢 `diagnosisCount` | integer | How many diagnosis records the case carries |
| 🔗 `url` | string | Direct link to the case on the GDC portal |
| 🕒 `scrapedAt` | string | ISO timestamp of collection |
| ⚠️ `error` | string | `null` on a normal row |

#### Example rows

```json
{
  "case_id": "6e3b6b72-142d-4b8d-a462-28a205796e41",
  "submitter_id": "TCGA-44-3918",
  "project_id": "TCGA-LUAD",
  "primary_site": "Bronchus and lung",
  "disease_type": "Adenomas and Adenocarcinomas",
  "gender": null,
  "race": "white",
  "ageAtDiagnosis": 22236,
  "diagnosisCount": 3,
  "url": "https://portal.gdc.cancer.gov/cases/6e3b6b72-142d-4b8d-a462-28a205796e41",
  "scrapedAt": "2026-09-14T17:23:10.118Z",
  "error": null
}
```

```json
{
  "case_id": "0c0b610e-fe4c-406d-a5ed-5cc3b11dabf5",
  "submitter_id": "TCGA-44-6146",
  "project_id": "TCGA-LUAD",
  "primary_site": "Bronchus and lung",
  "disease_type": "Cystic, Mucinous and Serous Neoplasms",
  "gender": null,
  "race": "white",
  "ageAtDiagnosis": 24009,
  "diagnosisCount": 3,
  "url": "https://portal.gdc.cancer.gov/cases/0c0b610e-fe4c-406d-a5ed-5cc3b11dabf5",
  "scrapedAt": "2026-09-14T17:23:10.210Z",
  "error": null
}
```

### ✨ Why choose this Actor

| | What you get |
|---|---|
| **The NCI's own record** | Rows come from the Genomic Data Commons API, not from a secondary database that mirrors it. |
| **Flat instead of nested** | The API returns nested case objects. Here it is one row, 12 columns, ready for a pivot table. |
| **No credentials** | Open GDC case metadata is public. No API key, no account, no data access request. |
| **Both identifier systems** | The GDC UUID and the program submitter ID on the same row, so cohorts join cleanly. |
| **You pay per row** | No subscription. A query that returns nothing costs nothing. |

### 🚀 How to use it

1. [Create a free Apify account](https://console.apify.com/sign-up). New accounts start with $5 of credit.
2. Open the Actor and select **Try for free**.
3. Set `projectId`, or `primarySite`, or leave both empty for a general pull.
4. Set `maxItems` to cap the run.
5. Select **Start**, then export from the **Dataset** tab as CSV, Excel, JSON or XML.

A first run:

```json
{
  "maxItems": 10
}
```

One project:

```json
{
  "projectId": "TCGA-LUAD",
  "maxItems": 500
}
```

One primary site across projects:

```json
{
  "primarySite": "Bronchus and lung",
  "maxItems": 1000
}
```

### ⚙️ Input

| Field | Required | Description |
|---|---|---|
| `projectId` | No | GDC project ID, for example `TCGA-LUAD` |
| `primarySite` | No | Primary site as the GDC writes it, for example `Bronchus and lung` |
| `maxItems` | No | How many cases to collect per run |

### 💰 Pricing

Pay per result. No subscription, and no Apify platform usage on top.

| Apify plan | Free | Bronze | Silver | Gold | Platinum | Diamond |
|---|---|---|---|---|---|---|
| Per case row | $0.0085 | $0.00817 | $0.00783 | $0.0075 | $0.0075 | $0.0075 |

| Rows collected | Cost on the Free plan |
|---|---|
| 100 | $0.85 |
| 1,000 | $8.50 |
| 10,000 | $85.00 |

**Free plan runs** return a preview. Any paid Apify plan lifts the ceiling for a full run.

### 🔌 Integrate with any app

The dataset is available through the Apify API as soon as the run finishes. Use `run-sync-get-dataset-items` for a one-shot call, webhooks to trigger what happens next, or the Make, Zapier, Airbyte and LangChain integrations listed on the Actor page.

### 🤖 Use with an AI agent

Give an agent live access to the case index over the Model Context Protocol:

```bash
claude mcp add --transport http apify "https://mcp.apify.com?tools=acquistion-automation/nci-gdc-cases-scraper"
```

Then ask it in plain language how many cases a project holds and have it read the rows back.

### ❓ Frequently asked questions

**Does this return genomic files, mutations or biospecimen records?**
No. It returns the case-level record: identifiers, project, primary site, disease type, the demographic fields above, age at diagnosis and a diagnosis count. Files, mutations and sample records are separate GDC endpoints that this Actor does not query.

**Why is `ageAtDiagnosis` a five-digit number?**
The GDC publishes age at diagnosis in days. `22236` is roughly 60.9 years. Divide by 365.25 after export if you want years.

**Why is `gender` empty?**
It came back `null` on every row in the runs checked. The field is present in the row shape, but the demographic block does not always carry it for a case.

**Is any of this controlled-access data?**
No. These are open case metadata records, the same ones the GDC portal shows without a login. Controlled-access genomic files require dbGaP authorisation and are not reachable here.

**What project IDs can I use?**
Any GDC project ID, written the way the portal writes it, for example `TCGA-LUAD` or `TARGET-AML`. If a filter returns nothing, check the spelling against the portal.

**What can I export?**
CSV, Excel, JSON and XML from the run page, or JSON straight from the API.

### 🔗 More from Acquisition Automation Co.

- [DrugBank Open Data Scraper](https://apify.com/acquistion-automation/drugbank-open-scraper)
- [IRS Exempt Organizations Scraper](https://apify.com/acquistion-automation/irs-eo-master-file-scraper)
- [SAM.gov Contract Opportunities Scraper](https://apify.com/acquistion-automation/sam-gov-contracts-scraper)
- [Clutch Agencies Scraper](https://apify.com/acquistion-automation/clutch-agencies-scraper)
- [USCG PSIX Vessel Registry Scraper](https://apify.com/acquistion-automation/uscg-psix-vessel-incidents-scraper)

### About Acquisition Automation Co.

We build automation for people buying businesses. The repetitive part of an acquisition search, checking listings, pulling public records, tracking owners and assets, is work a machine should do, so the buyer's time goes into judging deals instead of collecting them.

We add new Actors regularly. If there is a source you need and do not see here, tell us.

### 🆘 Support

Open an issue in the **Issues** tab of this Actor with your run ID, the input you used, and what you expected to get back.

### ⚠️ Disclaimer

This Actor is independent and is not affiliated with, endorsed by, or sponsored by the National Cancer Institute, the National Institutes of Health or any government agency. It collects only publicly available data. You are responsible for using that data in compliance with the source's terms of service and applicable law. Nothing here is medical advice.

# Actor input Schema

## `projectId` (type: `string`):

Filter by project ID.

## `primarySite` (type: `string`):

Filter by primary site.

## `maxItems` (type: `integer`):

How many cases to collect per run.

## Actor input object example

```json
{
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("acquistion-automation/nci-gdc-cases-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxItems": 10 }

# Run the Actor and wait for it to finish
run = client.actor("acquistion-automation/nci-gdc-cases-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10
}' |
apify call acquistion-automation/nci-gdc-cases-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,acquistion-automation/nci-gdc-cases-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/EAddbU6fiKqaTQzON/builds/aFqA6kb4h5TlQ5FkG/openapi.json
