# ClinicalTrials.gov Scraper (`aurenic/clinicaltrials-scraper`) Actor

Extract clinical trial data from the official ClinicalTrials.gov API v2. 585K+ studies: conditions, interventions, sponsors, phases, eligibility, outcomes, and locations. No API key, no browser, no proxy.

- **URL**: https://apify.com/aurenic/clinicaltrials-scraper.md
- **Developed by:** [Aurenic](https://apify.com/aurenic) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ClinicalTrials.gov Scraper

Extract clinical trial data from the official ClinicalTrials.gov API v2. 585K+ studies: conditions, interventions, sponsors, phases, eligibility, outcomes, and locations. No API key, no browser, no proxy.

### What does ClinicalTrials.gov Scraper do?

Scrape the U.S. National Library of Medicine's clinical trial registry in four modes:

- **Search studies** — filter by condition/disease, intervention/drug, sponsor, location, recruitment status, phase, and study type. Returns full flattened records with outcomes, eligibility, and site locations.
- **Fetch by NCT ID** — pass one or more NCT identifiers, get the same complete record.
- **Database size stats** — current total study count.
- **Field value enums** — the canonical value lists for status, phase, study type, and other controlled fields.

The ClinicalTrials.gov v2 API is **fully public** — no API key, no registration, no license barrier. Rate limit is ~50 requests/minute.

### Output fields

| Field | Description |
|---|---|
| nctId | ClinicalTrials.gov identifier |
| briefTitle / officialTitle / acronym | Title fields |
| orgStudyId / organization | Study identifiers and organization |
| overallStatus | Recruiting / Completed / Terminated / etc. |
| startDate / completionDate / lastUpdatePostDate | Date fields |
| whyStopped | Reason if stopped early |
| leadSponsor / leadSponsorClass | Primary sponsor |
| collaborators | Array of collaborator names |
| briefSummary / detailedDescription | Study description |
| conditions | Array of conditions/diseases |
| keywords | Study keywords |
| studyType | Interventional / Observational / Expanded Access |
| phases | Phase array (Phase 1, 2, 3, 4) |
| allocation / interventionModel / primaryPurpose / maskingInfo | Design |
| enrollmentCount / enrollmentType | Study size |
| interventions | Array of interventions with type, name, description |
| primaryOutcomes / secondaryOutcomes | Outcome measure arrays |
| eligibilityCriteria | Full eligibility text |
| healthyVolunteers / sex / minimumAge / maximumAge / stdAges | Eligibility |
| locations | Array of facility, city, state, country, status |
| ipdSharing | IPD sharing statement |
| hasResults | Whether the study has posted results |
| studyUrl | Direct ClinicalTrials.gov URL |

### Who is it for?

- **Pharmaceutical and biotech teams** tracking competitor trial pipelines
- **Clinical research organizations (CROs)** building site and investigator databases
- **Medical researchers** sourcing trial populations and outcomes for meta-analysis
- **Healthcare investors** monitoring pipeline progress by condition and phase
- **Patient advocacy groups** building trial-finder tools for specific diseases
- **AI and RAG builders** ingesting clinical trial corpora

### Pricing

**$1.50 per 1,000 results.** No subscription.

| Results | Cost |
|---|---|
| 100 | $0.15 |
| 1,000 | $1.50 |
| 10,000 | $15.00 |

### How to use it

1. Pick a **Mode**.
2. Enter a **Condition**, **Intervention**, **Sponsor**, or **Search Term**.
3. Optionally filter by **Overall Status**, **Phase**, and **Study Type**.
4. Set **Max Items** (default 500).
5. Click **Start**.

### Output example

```json
{
  "recordType": "study",
  "nctId": "NCT04280705",
  "briefTitle": "Adaptive COVID-19 Treatment Trial (ACTT-1)",
  "officialTitle": "A Multicenter, Adaptive, Randomized Blinded Controlled Trial of the Safety and Efficacy of Investigational Therapeutics...",
  "overallStatus": "COMPLETED",
  "startDate": "2020-02-21",
  "completionDate": "2021-04-01",
  "lastUpdatePostDate": "2023-05-15",
  "leadSponsor": "National Institute of Allergy and Infectious Diseases (NIAID)",
  "leadSponsorClass": "NIH",
  "conditions": ["COVID-19"],
  "studyType": "INTERVENTIONAL",
  "phases": ["PHASE3"],
  "enrollmentCount": 1062,
  "enrollmentType": "ACTUAL",
  "interventionCount": 2,
  "interventions": [
    { "type": "DRUG", "name": "Remdesivir", "description": "..." },
    { "type": "DRUG", "name": "Placebo", "description": "..." }
  ],
  "primaryOutcomeCount": 1,
  "primaryOutcomes": ["Time to Recovery"],
  "eligibilityCriteria": "Inclusion Criteria: ... Exclusion Criteria: ...",
  "sex": "ALL",
  "minimumAge": "18 Years",
  "stdAges": ["ADULT", "OLDER_ADULT"],
  "locationCount": 55,
  "locations": [
    { "facility": "Massachusetts General Hospital", "city": "Boston", "state": "Massachusetts", "country": "United States", "status": "COMPLETED" }
  ],
  "hasResults": true,
  "studyUrl": "https://clinicaltrials.gov/study/NCT04280705",
  "scrapedAt": "2026-09-24T12:00:00.000Z"
}
```

### Technical details

- **Official ClinicalTrials.gov API v2** — `https://clinicaltrials.gov/api/v2`. Fully public, no key, no registration.
- **Rate limit:** ~50 requests/minute (2 requests/second). The actor throttles to 600ms between requests by default.
- **Cursor pagination** via `pageToken` — up to 200 studies per page.
- **Four modes** — search, direct NCT lookup, stats, and enum reference.
- **No browser, no proxy** — pure REST JSON.
- **Flattened output** — deeply nested protocol sections (identification, status, design, eligibility, outcomes, locations) are flattened to a single wide row.

### Known limits

- **Rate limit is 2 requests/second.** Large runs (10,000+ studies) take proportional time. 10,000 studies ≈ 50 pages ≈ 60 seconds.
- **Eligibility criteria is free text** truncated at 5,000 characters.
- **Outcome measures are truncated to 20 each** (primary and secondary). Full outcome data requires per-study fetches.
- **Location lists can be very long** for multi-site trials — the full array is included.
- **`hasResults`** indicates whether the sponsor has posted results; it does not mean results are complete.

### FAQ

**Do I need an API key?** No. The ClinicalTrials.gov v2 API is fully public with no authentication.

**Do I need a proxy?** No. Datacenter IPs are accepted at up to ~50 requests/minute.

**What's the difference between condition, intervention, and term?** `condition` matches disease/condition fields; `intervention` matches drug/device names; `term` searches across all fields.

**How do I find NCT IDs?** Run search mode first — the `nctId` field is in every output record. Pass those IDs to study mode for direct fetches.

**How do I export data?** After a run, go to Storage → Export as JSON, CSV, Excel.

### Support

Open an issue on the Actor's page for bugs or feature requests.

# Actor input Schema

## `mode` (type: `string`):

What to fetch.

## `queryTerm` (type: `string`):

Free-text search across all study fields.

## `condition` (type: `string`):

Condition or disease (e.g. diabetes, breast cancer, Alzheimer).

## `intervention` (type: `string`):

Drug, device, or intervention name (e.g. pembrolizumab, insulin).

## `sponsor` (type: `string`):

Sponsor or collaborator (e.g. National Cancer Institute, Pfizer).

## `location` (type: `string`):

Geographic search (e.g. Boston, California, Germany).

## `overallStatus` (type: `string`):

Filter by recruitment status.

## `phase` (type: `string`):

Filter by trial phase.

## `studyType` (type: `string`):

Filter by study type.

## `nctIds` (type: `array`):

Direct NCT identifiers to fetch (e.g. NCT04280705). Used in study mode.

## `maxItems` (type: `integer`):

Hard cap on studies per run.

## `requestDelayMs` (type: `integer`):

Delay between requests. ClinicalTrials.gov allows ~50 req/min (2 req/sec). Default 600ms stays safely under.

## Actor input object example

```json
{
  "mode": "search",
  "queryTerm": "",
  "condition": "diabetes",
  "intervention": "",
  "sponsor": "",
  "location": "",
  "overallStatus": "",
  "phase": "",
  "studyType": "",
  "nctIds": [],
  "maxItems": 500,
  "requestDelayMs": 600
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "nctIds": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("aurenic/clinicaltrials-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "nctIds": [] }

# Run the Actor and wait for it to finish
run = client.actor("aurenic/clinicaltrials-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "nctIds": []
}' |
apify call aurenic/clinicaltrials-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,aurenic/clinicaltrials-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/SN90Cj62C5DESICj4/builds/rIEcace6GJyRPCURq/openapi.json
