NCI GDC Slide Image Metadata Scraper avatar

NCI GDC Slide Image Metadata Scraper

Pricing

from $29.62 / 1,000 results

Go to Apify Store
NCI GDC Slide Image Metadata Scraper

NCI GDC Slide Image Metadata Scraper

Scrapes slide image file metadata from the NCI Genomic Data Commons API. Returns each file as a flat row with file ID, submitter ID, data format, experimental strategy, and file size.

Pricing

from $29.62 / 1,000 results

Rating

0.0

(0)

Developer

Acquisition Automation Co.

Acquisition Automation Co.

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Acquisition Automation

NCI GDC Slide Image Metadata Scraper

Scrape slide image metadata from the NCI Genomic Data Commons API, filtered by data type and format. Each record includes file ID, submitter ID, experimental strategy, data category, and file size. No API key required. Export to CSV, JSON, Excel, or XML.

The NCI GDC portal requires manual browsing to find whole-slide image metadata across thousands of cancer research files. This Actor queries the public GDC API directly for slide image metadata, applies your data type and format filters, and returns every matching file record in a flat, analysis-ready schema.

Who uses itWhat they scrape NCI Genomic Data Commons for
Bioinformatics researchersBuild a catalog of available digital pathology slides for a specific cancer study
Clinical data managersAudit slide image submissions across projects by data format and release state
Machine learning engineersGather metadata for SVS and NDPI whole-slide images to assemble training datasets
Cancer registry analystsTrack the volume and types of slide images available per experimental strategy

What it does

This Actor collects slide image file metadata from the NCI Genomic Data Commons API and returns each file as a flat row with fields like file ID, submitter ID, data format, experimental strategy, and file size.

  • ๐Ÿ”ฌ Data type filter: Pre-set to Slide Image, adjustable to any GDC data type like Gene Expression Quantification.
  • ๐Ÿ“ Format filter: Narrow results to specific whole-slide image formats such as SVS or NDPI.
  • ๐Ÿ“Š Flat row output: Every file record is returned with its file ID, submitter ID, data category, experimental strategy, file size, and more.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with NCI Genomic Data Commons data

๐Ÿ”ฌ Build a slide image inventory for a cancer type.

A bioinformatics researcher filters slide image metadata by experimental strategy to list all available SVS files for a lung adenocarcinoma cohort.

๐Ÿ“ˆ Monitor new data releases.

A data manager runs the Actor weekly with a data type filter to detect newly released slide images in the GDC and update internal tracking sheets.

๐Ÿค– Prepare metadata for deep learning pipelines.

An ML engineer extracts metadata for all NDPI slide images, then joins the output with clinical data to label training samples.

๐Ÿ“‹ Audit submission completeness.

A project coordinator checks that every expected submitter ID has a corresponding slide image file record in the latest GDC data release.

Why choose this scraper

What you get
No API keyThe GDC API is fully public, so you start scraping immediately without registration.
Structured metadataEvery row follows the same schema, ready for direct import into analysis tools.
Format filteringRestrict results to SVS, NDPI, or any comma-separated list of data formats.
Paginated fetchingThe Actor handles API pagination automatically up to your defined item limit.

What a NCI Genomic Data Commons record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

{
"id": "18081e12-b7d5-43b4-ab26-7574b94b98f0",
"data_format": "SVS",
"access": "open",
"file_name": "TCGA-05-4245-01A-01-BS1.41d3cf23-4e36-4e42-9e08-adfea139f37e.svs",
"submitter_id": "TCGA-05-4245-01A-01-BS1_slide_image",
"data_category": "Biospecimen",
"type": "slide_image",
"file_size": 85438631,
"created_datetime": "2021-10-13T22:33:10.080104-05:00",
"md5sum": "2724134b4ee0d567b948481d709947b4",
"updated_datetime": "2021-10-14T18:24:32.587756-05:00",
"file_id": "18081e12-b7d5-43b4-ab26-7574b94b98f0",
"data_type": "Slide Image",
"state": "released",
"experimental_strategy": "Tissue Slide"
}

Every value above comes from a real run. A field a record does not have comes back as null.

Configure the run

Drive the Actor with a data type and optional data format filters, and set a maximum number of items to cap the API pages fetched. The Input tab lists every parameter.

A first run with the defaults:

{
"maxItems": 10,
"dataType": "Slide Image"
}

A larger pull:

{
"maxItems": 200,
"dataType": "Slide Image"
}

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account with $5 in credit.
  2. Open the NCI GDC Slide Image Metadata Scraper.
  3. Set your inputs and any filters, then click Start.
  4. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to NCI Genomic Data Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

$claude mcp add --transport http apify "https://mcp.apify.com?tools=acquistion-automation/nci-gdc-slide-image-metadata-scraper"

Then prompt it in plain language to run the scraper and read back the results.

Troubleshooting

Why am I getting no results?

Check that your data format filter matches the formats present in the GDC for the selected data type. Try running without a format filter first to see what is available.

The run stops before reaching my max items limit.

The GDC API may return fewer records than your limit if the total matching files are exhausted. Verify the total count in the API response and adjust your filters if needed.

I see a timeout error during the run.

The GDC API can be slow for large queries. Reduce the max items value and run multiple smaller batches, or try again during off-peak hours.

The exported CSV contains unexpected characters.

Some metadata fields may contain special characters. Open the CSV with UTF-8 encoding in your analysis tool to preserve all characters correctly.

FAQ

QuestionAnswer
Do I need an API key or GDC account to use this Actor?No. The NCI GDC API is open and requires no authentication. You can start scraping slide image metadata immediately.
What data formats can I filter on?You can provide any comma-separated list of GDC data format values. Common whole-slide image formats are SVS and NDPI.
Can I change the data type from Slide Image to something else?Yes. The data type input defaults to Slide Image, but you can set it to any valid GDC data type such as Gene Expression Quantification or Aligned Reads.
How many records can I fetch in one run?Free users are limited to 10 items for preview. Paid users can fetch up to 1,000,000 records per run by adjusting the max items setting.
What export formats are supported?You can export the scraped metadata to CSV, JSON, Excel, or XML directly from the dataset tab.
Does this Actor download the actual slide image files?No. It scrapes only the metadata records for slide images, not the SVS or NDPI files themselves.
How does pagination work with the GDC API?The Actor automatically follows the pagination links in the API response until it reaches your max items limit or exhausts the available records.
Can I filter by project or case ID?The current input schema supports data type and data format filters. For project-level or case-level filtering, you can post-process the exported dataset.

Browse the full Acquisition Automation collection for more scrapers.

๐Ÿ†˜ Need help? Open an issue in the Issues tab of this Actor with your run ID, your input, and what you expected.

Pricing

This Actor uses pay-per-result pricing: $0.0395 per result collected. You are billed only for the results you receive, so a run that returns nothing costs nothing.

โš ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by National Cancer Institute. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.