# NCI GDC Slide Image Metadata Scraper (`acquistion-automation/nci-gdc-slide-image-metadata-scraper`) Actor

Scrapes slide image file metadata from the NCI Genomic Data Commons API. Returns each file as a flat row with file ID, submitter ID, data format, experimental strategy, and file size.

- **URL**: https://apify.com/acquistion-automation/nci-gdc-slide-image-metadata-scraper.md
- **Developed by:** [Acquisition Automation Co.](https://apify.com/acquistion-automation) (community)
- **Categories:** AI, Developer tools, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $29.62 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

[![Acquisition Automation](https://api.apify.com/v2/key-value-stores/AOdPHdOpeDpzEPS5f/records/banner.jpg)](https://apify.com/acquistion-automation)

### NCI GDC Slide Image Metadata Scraper

**Scrape slide image metadata from the NCI Genomic Data Commons API, filtered by data type and format.** Each record includes file ID, submitter ID, experimental strategy, data category, and file size. No API key required. Export to CSV, JSON, Excel, or XML.

The NCI GDC portal requires manual browsing to find whole-slide image metadata across thousands of cancer research files. This Actor queries the public GDC API directly for slide image metadata, applies your data type and format filters, and returns every matching file record in a flat, analysis-ready schema.

| Who uses it | What they scrape NCI Genomic Data Commons for |
|---|---|
| Bioinformatics researchers | Build a catalog of available digital pathology slides for a specific cancer study |
| Clinical data managers | Audit slide image submissions across projects by data format and release state |
| Machine learning engineers | Gather metadata for SVS and NDPI whole-slide images to assemble training datasets |
| Cancer registry analysts | Track the volume and types of slide images available per experimental strategy |

### What it does

This Actor collects slide image file metadata from the NCI Genomic Data Commons API and returns each file as a flat row with fields like file ID, submitter ID, data format, experimental strategy, and file size.

- 🔬 **Data type filter:** Pre-set to Slide Image, adjustable to any GDC data type like Gene Expression Quantification.
- 📁 **Format filter:** Narrow results to specific whole-slide image formats such as SVS or NDPI.
- 📊 **Flat row output:** Every file record is returned with its file ID, submitter ID, data category, experimental strategy, file size, and more.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with NCI Genomic Data Commons data

**🔬 Build a slide image inventory for a cancer type.**

A bioinformatics researcher filters slide image metadata by experimental strategy to list all available SVS files for a lung adenocarcinoma cohort.

**📈 Monitor new data releases.**

A data manager runs the Actor weekly with a data type filter to detect newly released slide images in the GDC and update internal tracking sheets.

**🤖 Prepare metadata for deep learning pipelines.**

An ML engineer extracts metadata for all NDPI slide images, then joins the output with clinical data to label training samples.

**📋 Audit submission completeness.**

A project coordinator checks that every expected submitter ID has a corresponding slide image file record in the latest GDC data release.

### Why choose this scraper

| | What you get |
|---|---|
| **No API key** | The GDC API is fully public, so you start scraping immediately without registration. |
| **Structured metadata** | Every row follows the same schema, ready for direct import into analysis tools. |
| **Format filtering** | Restrict results to SVS, NDPI, or any comma-separated list of data formats. |
| **Paginated fetching** | The Actor handles API pagination automatically up to your defined item limit. |

### What a NCI Genomic Data Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

```json
{
 "id": "18081e12-b7d5-43b4-ab26-7574b94b98f0",
 "data_format": "SVS",
 "access": "open",
 "file_name": "TCGA-05-4245-01A-01-BS1.41d3cf23-4e36-4e42-9e08-adfea139f37e.svs",
 "submitter_id": "TCGA-05-4245-01A-01-BS1_slide_image",
 "data_category": "Biospecimen",
 "type": "slide_image",
 "file_size": 85438631,
 "created_datetime": "2021-10-13T22:33:10.080104-05:00",
 "md5sum": "2724134b4ee0d567b948481d709947b4",
 "updated_datetime": "2021-10-14T18:24:32.587756-05:00",
 "file_id": "18081e12-b7d5-43b4-ab26-7574b94b98f0",
 "data_type": "Slide Image",
 "state": "released",
 "experimental_strategy": "Tissue Slide"
}
```

Every value above comes from a real run. A field a record does not have comes back as `null`.

### Configure the run

Drive the Actor with a data type and optional data format filters, and set a maximum number of items to cap the API pages fetched. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "maxItems": 10,
 "dataType": "Slide Image"
}
```

A larger pull:

```json
{
 "maxItems": 200,
 "dataType": "Slide Image"
}
```

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account with $5 in credit](https://console.apify.com/sign-up).
2. Open the [NCI GDC Slide Image Metadata Scraper](https://apify.com/acquistion-automation/nci-gdc-slide-image-metadata-scraper).
3. Set your inputs and any filters, then click **Start**.
4. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to NCI Genomic Data Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
claude mcp add --transport http apify "https://mcp.apify.com?tools=acquistion-automation/nci-gdc-slide-image-metadata-scraper"
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check that your data format filter matches the formats present in the GDC for the selected data type. Try running without a format filter first to see what is available.

**The run stops before reaching my max items limit.**

The GDC API may return fewer records than your limit if the total matching files are exhausted. Verify the total count in the API response and adjust your filters if needed.

**I see a timeout error during the run.**

The GDC API can be slow for large queries. Reduce the max items value and run multiple smaller batches, or try again during off-peak hours.

**The exported CSV contains unexpected characters.**

Some metadata fields may contain special characters. Open the CSV with UTF-8 encoding in your analysis tool to preserve all characters correctly.

### FAQ

| Question | Answer |
|---|---|
| Do I need an API key or GDC account to use this Actor? | No. The NCI GDC API is open and requires no authentication. You can start scraping slide image metadata immediately. |
| What data formats can I filter on? | You can provide any comma-separated list of GDC data format values. Common whole-slide image formats are SVS and NDPI. |
| Can I change the data type from Slide Image to something else? | Yes. The data type input defaults to Slide Image, but you can set it to any valid GDC data type such as Gene Expression Quantification or Aligned Reads. |
| How many records can I fetch in one run? | Free users are limited to 10 items for preview. Paid users can fetch up to 1,000,000 records per run by adjusting the max items setting. |
| What export formats are supported? | You can export the scraped metadata to CSV, JSON, Excel, or XML directly from the dataset tab. |
| Does this Actor download the actual slide image files? | No. It scrapes only the metadata records for slide images, not the SVS or NDPI files themselves. |
| How does pagination work with the GDC API? | The Actor automatically follows the pagination links in the API response until it reaches your max items limit or exhausts the available records. |
| Can I filter by project or case ID? | The current input schema supports data type and data format filters. For project-level or case-level filtering, you can post-process the exported dataset. |

### Related actors

Browse the full [Acquisition Automation collection](https://apify.com/acquistion-automation) for more scrapers.

🆘 **Need help?** Open an issue in the Issues tab of this Actor with your run ID, your input, and what you expected.

### Pricing

This Actor uses **pay-per-result** pricing: **$0.0395 per result** collected. You are billed only for the results you receive, so a run that returns nothing costs nothing.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by National Cancer Institute. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `dataType` (type: `string`):

The data type filter for files. Defaults to Slide Image.

## `dataFormat` (type: `string`):

Optional comma-separated list of data formats to filter by, e.g. SVS,NDPI.

## Actor input object example

```json
{
  "maxItems": 10,
  "dataType": "Slide Image"
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10,
    "dataType": "Slide Image"
};

// Run the Actor and wait for it to finish
const run = await client.actor("acquistion-automation/nci-gdc-slide-image-metadata-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxItems": 10,
    "dataType": "Slide Image",
}

# Run the Actor and wait for it to finish
run = client.actor("acquistion-automation/nci-gdc-slide-image-metadata-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10,
  "dataType": "Slide Image"
}' |
apify call acquistion-automation/nci-gdc-slide-image-metadata-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,acquistion-automation/nci-gdc-slide-image-metadata-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/y7yQhQVvdFa6nYLV2/builds/5QiwcedPM7kQjhWzs/openapi.json
