NCI GDC Slide Image Metadata Scraper
Pricing
from $29.62 / 1,000 results
NCI GDC Slide Image Metadata Scraper
Scrapes slide image file metadata from the NCI Genomic Data Commons API. Returns each file as a flat row with file ID, submitter ID, data format, experimental strategy, and file size.
Pricing
from $29.62 / 1,000 results
Rating
0.0
(0)
Developer
Acquisition Automation Co.
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
NCI GDC Slide Image Metadata Scraper
Scrape slide image metadata from the NCI Genomic Data Commons API, filtered by data type and format. Each record includes file ID, submitter ID, experimental strategy, data category, and file size. No API key required. Export to CSV, JSON, Excel, or XML.
The NCI GDC portal requires manual browsing to find whole-slide image metadata across thousands of cancer research files. This Actor queries the public GDC API directly for slide image metadata, applies your data type and format filters, and returns every matching file record in a flat, analysis-ready schema.
| Who uses it | What they scrape NCI Genomic Data Commons for |
|---|---|
| Bioinformatics researchers | Build a catalog of available digital pathology slides for a specific cancer study |
| Clinical data managers | Audit slide image submissions across projects by data format and release state |
| Machine learning engineers | Gather metadata for SVS and NDPI whole-slide images to assemble training datasets |
| Cancer registry analysts | Track the volume and types of slide images available per experimental strategy |
What it does
This Actor collects slide image file metadata from the NCI Genomic Data Commons API and returns each file as a flat row with fields like file ID, submitter ID, data format, experimental strategy, and file size.
- ๐ฌ Data type filter: Pre-set to Slide Image, adjustable to any GDC data type like Gene Expression Quantification.
- ๐ Format filter: Narrow results to specific whole-slide image formats such as SVS or NDPI.
- ๐ Flat row output: Every file record is returned with its file ID, submitter ID, data category, experimental strategy, file size, and more.
Results export to CSV, JSON, Excel, or XML, or straight from the API.
What you can do with NCI Genomic Data Commons data
๐ฌ Build a slide image inventory for a cancer type.
A bioinformatics researcher filters slide image metadata by experimental strategy to list all available SVS files for a lung adenocarcinoma cohort.
๐ Monitor new data releases.
A data manager runs the Actor weekly with a data type filter to detect newly released slide images in the GDC and update internal tracking sheets.
๐ค Prepare metadata for deep learning pipelines.
An ML engineer extracts metadata for all NDPI slide images, then joins the output with clinical data to label training samples.
๐ Audit submission completeness.
A project coordinator checks that every expected submitter ID has a corresponding slide image file record in the latest GDC data release.
Why choose this scraper
| What you get | |
|---|---|
| No API key | The GDC API is fully public, so you start scraping immediately without registration. |
| Structured metadata | Every row follows the same schema, ready for direct import into analysis tools. |
| Format filtering | Restrict results to SVS, NDPI, or any comma-separated list of data formats. |
| Paginated fetching | The Actor handles API pagination automatically up to your defined item limit. |
What a NCI Genomic Data Commons record looks like
Every record returns as one flat JSON row. Here is a real one from a run:
{"id": "18081e12-b7d5-43b4-ab26-7574b94b98f0","data_format": "SVS","access": "open","file_name": "TCGA-05-4245-01A-01-BS1.41d3cf23-4e36-4e42-9e08-adfea139f37e.svs","submitter_id": "TCGA-05-4245-01A-01-BS1_slide_image","data_category": "Biospecimen","type": "slide_image","file_size": 85438631,"created_datetime": "2021-10-13T22:33:10.080104-05:00","md5sum": "2724134b4ee0d567b948481d709947b4","updated_datetime": "2021-10-14T18:24:32.587756-05:00","file_id": "18081e12-b7d5-43b4-ab26-7574b94b98f0","data_type": "Slide Image","state": "released","experimental_strategy": "Tissue Slide"}
Every value above comes from a real run. A field a record does not have comes back as null.
Configure the run
Drive the Actor with a data type and optional data format filters, and set a maximum number of items to cap the API pages fetched. The Input tab lists every parameter.
A first run with the defaults:
{"maxItems": 10,"dataType": "Slide Image"}
A larger pull:
{"maxItems": 200,"dataType": "Slide Image"}
Free users
Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.
Run it
- Create a free Apify account with $5 in credit.
- Open the NCI GDC Slide Image Metadata Scraper.
- Set your inputs and any filters, then click Start.
- Export the results as CSV, Excel, JSON, or XML from the Dataset tab.
Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.
Use with AI agents (MCP)
Give an AI agent live access to NCI Genomic Data Commons through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:
$claude mcp add --transport http apify "https://mcp.apify.com?tools=acquistion-automation/nci-gdc-slide-image-metadata-scraper"
Then prompt it in plain language to run the scraper and read back the results.
Troubleshooting
Why am I getting no results?
Check that your data format filter matches the formats present in the GDC for the selected data type. Try running without a format filter first to see what is available.
The run stops before reaching my max items limit.
The GDC API may return fewer records than your limit if the total matching files are exhausted. Verify the total count in the API response and adjust your filters if needed.
I see a timeout error during the run.
The GDC API can be slow for large queries. Reduce the max items value and run multiple smaller batches, or try again during off-peak hours.
The exported CSV contains unexpected characters.
Some metadata fields may contain special characters. Open the CSV with UTF-8 encoding in your analysis tool to preserve all characters correctly.
FAQ
| Question | Answer |
|---|---|
| Do I need an API key or GDC account to use this Actor? | No. The NCI GDC API is open and requires no authentication. You can start scraping slide image metadata immediately. |
| What data formats can I filter on? | You can provide any comma-separated list of GDC data format values. Common whole-slide image formats are SVS and NDPI. |
| Can I change the data type from Slide Image to something else? | Yes. The data type input defaults to Slide Image, but you can set it to any valid GDC data type such as Gene Expression Quantification or Aligned Reads. |
| How many records can I fetch in one run? | Free users are limited to 10 items for preview. Paid users can fetch up to 1,000,000 records per run by adjusting the max items setting. |
| What export formats are supported? | You can export the scraped metadata to CSV, JSON, Excel, or XML directly from the dataset tab. |
| Does this Actor download the actual slide image files? | No. It scrapes only the metadata records for slide images, not the SVS or NDPI files themselves. |
| How does pagination work with the GDC API? | The Actor automatically follows the pagination links in the API response until it reaches your max items limit or exhausts the available records. |
| Can I filter by project or case ID? | The current input schema supports data type and data format filters. For project-level or case-level filtering, you can post-process the exported dataset. |
Related actors
Browse the full Acquisition Automation collection for more scrapers.
๐ Need help? Open an issue in the Issues tab of this Actor with your run ID, your input, and what you expected.
Pricing
This Actor uses pay-per-result pricing: $0.0395 per result collected. You are billed only for the results you receive, so a run that returns nothing costs nothing.
โ ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by National Cancer Institute. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.
