# NCI Genomic Data Commons Projects Scraper (`parseforge/nci-gdc-projects-scraper`) Actor

Pull project-level metadata from the NCI Genomic Data Commons public API. Retrieves project id, name, disease type, primary site, dbGaP accession number, state, and releasable status. Ideal for cancer researchers compiling study inventories.

- **URL**: https://apify.com/parseforge/nci-gdc-projects-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** Other, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.40 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner-v4.webp)](https://apify.com/parseforge?fpr=vmoqkp)

### NCI GDC Projects Scraper

**Scrape NCI Genomic Data Commons project listings.** Every project comes with its ID, disease type, primary site, and release state. No login or API key. Export to CSV, JSON, Excel, or XML.

The NCI Genomic Data Commons web portal is built for browsing, not bulk export. Researchers who need the full project catalog for meta-analysis or grant writing have to click through pages or write their own API scripts. This Actor reads the public GDC API directly, filters by keyword, and returns every matching project in one flat dataset.

| Who uses it | What they scrape NCI Genomic Data Commons for |
|---|---|
| Bioinformatics researchers | Catalog available genomic studies for a specific cancer type before requesting access. |
| Grant writers | Compile a list of funded projects and their disease focus to support a funding application. |
| Data librarians | Build a searchable index of GDC projects for their institution's internal data portal. |
| Cancer epidemiologists | Identify all projects linked to a particular primary site for a cross-study analysis. |

### What it does

This Actor collects NCI GDC project metadata from the public API and returns each project as a flat row with its ID, name, disease type, primary site, and release status.

- 🔍 **Keyword search:** filter projects by name or any text field, matching the GDC site's own search behavior.
- 📊 **Structured output:** each project is a flat row with consistent columns, ready for analysis in any tool.
- ⚡ **Bulk extraction:** pull project records in a single run, bypassing the portal's page-by-page browsing.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with NCI Genomic Data Commons data

**🧬 Map the cancer research landscape.**

A bioinformatician scrapes all GDC projects, groups them by disease\_type, and identifies which cancers have the most open-access genomic data available.

**📝 Support a grant proposal with data.**

A principal investigator extracts every project related to 'lung adenocarcinoma' and lists the dbGaP accession numbers as evidence of existing research infrastructure.

**🏥 Build an institutional data catalog.**

A data librarian runs the Actor weekly, ingests the JSON output into a local database, and lets researchers search GDC projects from the institution's own portal.

**🔬 Find projects by tissue site.**

An epidemiologist searches for 'breast' across all project names and primary\_site fields, then downloads the resulting project list to plan a meta-analysis.

### Why choose this scraper

| | What you get |
|---|---|
| **No API coding** | The Actor calls the GDC API for you and handles pagination and rate limits. |
| **Fixed schema** | Every run returns the same columns, so your downstream scripts never break. |
| **Keyword filtering** | Narrow results to projects matching a disease name, tissue, or any text field. |
| **Full catalog access** | Retrieve all public project metadata without clicking through the web portal. |

### What a NCI Genomic Data Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

```json
{
 "id": "TCGA-LUAD",
 "projectId": "TCGA-LUAD",
 "name": "Lung Adenocarcinoma",
 "releasable": true,
 "state": "open",
 "released": true,
 "scrapedAt": "2026-09-26T13:55:44.944Z"
}
```

Every value above comes from a real run. A field a record does not have comes back as `null`.

### Configure the run

Drive the Actor with an optional search term to filter projects by name, and set a maximum item count to control the size of your dataset. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "maxItems": 10
}
```

A larger pull:

```json
{
 "maxItems": 200
}
```

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect more results per run.

### Run it

1. [Create a free Apify account](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Set your inputs and any filters, then click **Start**.
3. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to NCI Genomic Data Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting fewer results than expected?**

Check the max items setting. Free users are capped at 10 items. If you are on a paid plan, make sure your search term is not too restrictive. Try broadening the keyword or leaving it empty to see all projects.

**The Actor returns an empty dataset.**

Your search term may not match any project name or text field. Try a different keyword, check for typos, or run without a search term to confirm the API is reachable.

**I need a field that is not in the output.**

The Actor returns the core project metadata from the GDC API. If you need additional fields, check the GDC API documentation to see if they are available, and contact support with the field names you require.

**The run fails with a timeout error.**

The GDC API may be experiencing high load. Reduce the max items count and retry, or schedule the run for a less busy time. The API is public and rate-limited by the server.

### FAQ

| Question | Answer |
|---|---|
| Do I need an API key or login to scrape GDC projects? | No. The NCI GDC API is fully public and requires no authentication. The Actor calls the endpoint directly. |
| What data does each project row contain? | Each row includes the project ID, name, disease type, primary site, dbGaP accession number, release state, and whether the project is releasable. The exact fields are shown in the sample output on the Actor's page. |
| Can I filter projects by a specific cancer type? | Yes. Use the search term input to filter by any text that appears in the project name or other text fields. For example, entering 'melanoma' returns only projects that mention melanoma. |
| How many projects can I scrape in one run? | Free users can scrape up to 10 projects as a preview. Paid users can set the max items up to 1,000,000, which covers the entire GDC project catalog many times over. |
| What format is the output? | The Actor returns data as a flat dataset. You can export it to CSV, JSON, Excel, or XML from the Apify platform. |
| Does this Actor scrape the GDC web portal HTML? | No. It calls the official GDC REST API directly, which is faster, more reliable, and returns structured JSON that the Actor flattens into rows. |
| Can I get the list of all GDC projects without any filter? | Yes. Leave the search term empty and set a high max items value to retrieve the complete public project catalog. |
| Is the data from the GDC API live? | Yes. Every run calls the live GDC API, so you always get the current project listings, including newly added projects. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by National Cancer Institute. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `searchTerm` (type: `string`):

A keyword to filter projects by name or other text fields. Matches the site's own search behavior.

## Actor input object example

```json
{
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/nci-gdc-projects-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxItems": 10 }

# Run the Actor and wait for it to finish
run = client.actor("parseforge/nci-gdc-projects-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10
}' |
apify call parseforge/nci-gdc-projects-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/nci-gdc-projects-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Mk6Gw7dpneIzbvdZO/builds/zdkpdul0bGVC1DDwS/openapi.json
