# ClinicalTrials.gov Scraper — Trials, Sponsors & Contacts (`haketa/clinicaltrials-scraper`) Actor

Scrape ClinicalTrials.gov: NCT ID, title, status, phase, sponsor, interventions, conditions, eligibility, enrollment, central contacts (name/email/phone) and study sites. Search by condition, sponsor, drug or location. For pharma, CRO, research and lead gen. Not affiliated with ClinicalTrials.gov.

- **URL**: https://apify.com/haketa/clinicaltrials-scraper.md
- **Developed by:** [Haketa](https://apify.com/haketa) (community)
- **Categories:** Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.75 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ClinicalTrials.gov Scraper — Trials, Sponsors & Contacts

> **Search and scrape ClinicalTrials.gov: NCT ID, title, status, phase, study type, sponsor, interventions, conditions, eligibility, enrollment, brief summary — plus central contacts (name, email, phone) and every study site with its contacts.** Search by condition, sponsor, drug, location or NCT ID. Clean JSON/CSV/Excel in seconds. Built for pharma, CRO, research and lead generation.

[![ClinicalTrials.gov](https://img.shields.io/badge/ClinicalTrials.gov-API%20v2-0b5394)]()
[![Sponsors + Contacts](https://img.shields.io/badge/Sponsors%20%2B%20Contacts-2da44e)]()
[![Pharma / CRO Leads](https://img.shields.io/badge/Pharma%20%2F%20CRO%20Leads-8250df)]()
[![Export](https://img.shields.io/badge/Export-JSON%20%2F%20CSV%20%2F%20Excel-fb8500)]()

***

### What This Actor Does

Search the official ClinicalTrials.gov registry and get a clean record per study:

- **Identity** — NCT ID, brief & official title, acronym, study URL
- **Status & design** — recruitment status, study type, phases, start/completion dates, enrollment
- **Sponsor** — lead sponsor and its class (**Industry / NIH / Other**) plus collaborators
- **Clinical** — conditions, interventions (drugs/devices), brief summary
- **Eligibility** — sex, min/max age, healthy-volunteers flag
- **Contacts** — central contact **name, email and phone**, plus every **study site** (facility, city, state, country, status) and its site contacts

Search by **condition, sponsor, drug/intervention, location or NCT ID**, and filter by recruitment status.

***

### Why Use This

- **Reach the right people.** Recruiting trials expose central contacts with **email and phone** — the people running the study.
- **Map the industry.** Pull every trial for a sponsor, drug or condition with phases, enrollment and sites.
- **Site & investigator lists.** Each study lists its sites and their contacts — ready for outreach or feasibility.
- **Official & free.** Reads the ClinicalTrials.gov API directly — no key, no anti-bot, no browser.

***

### Quick Start

#### Run it in the console (no code)

1. Add **conditions**, **sponsors**, **interventions**, **locations** and/or **NCT IDs**.
2. Optionally set a **status** filter (e.g. *Recruiting*).
3. Set **Max studies**, click **Start**, export as **JSON, CSV, Excel or HTML**.

#### Build a recruiting-trial contact list (Python)

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run_input = {"conditions": ["diabetes"], "status": ["RECRUITING"], "maxItems": 1000}

run = client.actor("YOUR_USERNAME/clinicaltrials-scraper").call(run_input=run_input)

for s in client.dataset(run["defaultDatasetId"]).iterate_items():
    if s["primaryContactEmail"]:
        print(s["nctId"], "·", s["leadSponsor"], "·", s["primaryContactEmail"], "·", s["primaryContactPhone"])
```

#### Pull every trial for a sponsor (Python)

```python
run = client.actor("YOUR_USERNAME/clinicaltrials-scraper").call(run_input={
    "sponsors": ["Novo Nordisk A/S"], "maxItems": 2000,
})
```

***

### Input Parameters

| Field | Type | Description |
|---|---|---|
| `conditions` | array | Conditions/diseases. Each runs as a separate search. |
| `searchTerms` | array | Free-text terms (title, keywords). |
| `sponsors` | array | Sponsor/organization names. |
| `interventions` | array | Drug, device or intervention names. |
| `locations` | array | Location terms (city, state, country). |
| `status` | array | Recruitment-status filter (applies to all searches). |
| `nctIds` | array | Specific NCT IDs to look up directly. |
| `includeLocations` | boolean | Include the full list of sites and their contacts (default on). |
| `maxItems` | integer | Max studies across all searches. `0` = no limit. |
| `maxPagesPerQuery` | integer | Pagination cap per search (100 per page). |
| `proxyConfiguration` | object | Apify Proxy. Datacenter is enough (public API). |

***

### Output

Each study is one record:

```json
{
  "nctId": "NCT07668388",
  "title": "A Research Study Comparing How Well …",
  "url": "https://clinicaltrials.gov/study/NCT07668388",
  "status": "RECRUITING",
  "studyType": "INTERVENTIONAL",
  "phases": ["PHASE3"],
  "enrollmentCount": 1200,
  "leadSponsor": "Novo Nordisk A/S", "leadSponsorClass": "INDUSTRY",
  "collaborators": [],
  "conditions": ["Type 2 Diabetes"],
  "interventionNames": ["Semaglutide", "Placebo"],
  "sex": "ALL", "minimumAge": "18 Years", "maximumAge": "N/A",
  "primaryContactName": "Clinical Trials", "primaryContactEmail": "clinicaltrials@novonordisk.com", "primaryContactPhone": "+1 ...",
  "locationsCount": 85, "countries": ["United States", "Germany"],
  "locations": [{"facility": "Research Site", "city": "Boston", "country": "United States", "status": "RECRUITING", "contacts": [{"name": "…", "email": "…"}]}]
}
```

**About contacts:** central and site contacts (name/email/phone) are published for **recruiting and not-yet-recruiting** trials — so filtering by `status: ["RECRUITING"]` yields the highest contact coverage. Completed trials often have contacts removed. Study details (sponsor, phase, conditions, sites) are available for all studies.

***

### Use Cases

#### 1. Pharma & CRO lead generation

Build contact lists of trials actively recruiting by condition, drug or region — with sponsor, site and contact details.

#### 2. Competitive intelligence

Track a sponsor's or a drug's entire pipeline: phases, enrollment, status and sites over time.

#### 3. Site feasibility & recruitment

Find study sites and investigators by location and condition for feasibility and patient recruitment.

#### 4. Research & market analysis

Analyse trial volume, phases, sponsors and geography across a therapeutic area.

***

### Tips

- **`status: ["RECRUITING"]`** gives the best contact coverage (active trials expose email/phone).
- **`sponsors`** pulls a company's whole pipeline; **`interventions`** pulls everything for a drug.
- **`primaryContactEmail` / `primaryContactPhone`** are the flat central-contact fields for quick lead lists; `centralContacts` and `locations[].contacts` have the rest.
- **`leadSponsorClass`** separates Industry from NIH/academic sponsors.
- **Schedule it** with Apify Schedules to monitor new trials in your area.

***

### Frequently Asked Questions

**Do I need an account or key?**
No. The ClinicalTrials.gov API is public and free — no login, key or anti-bot.

**Which trials include contact emails?**
Recruiting and not-yet-recruiting trials expose central and site contacts. Filter by `status` for the best coverage.

**Can I get all sites of a trial?**
Yes — each study includes its full `locations` list with per-site contacts (toggle with `includeLocations`).

**What is an NCT ID?**
It is the unique ClinicalTrials.gov identifier for a study (e.g. `NCT07668388`).

**What export formats are supported?**
JSON, CSV, Excel, HTML, or via API — plus Google Sheets, webhooks, Make and Zapier.

***

### Legal & Responsible Use

This Actor is an independent tool and is **not affiliated with, endorsed by, or sponsored by ClinicalTrials.gov, the U.S. National Library of Medicine or the NIH.** It reads only public registry data. Contact details are published for study recruitment — use them responsibly and in line with applicable terms and data-protection laws.

# Actor input Schema

## `conditions` (type: `array`):

Conditions or diseases to search (e.g. "diabetes", "breast cancer"). Each runs as a separate search.

## `searchTerms` (type: `array`):

Free-text terms (title, keywords). Each runs as a separate search.

## `sponsors` (type: `array`):

Sponsor/organization names (e.g. "Pfizer", "NIH"). Each runs as a separate search.

## `interventions` (type: `array`):

Drug, device or intervention names (e.g. "semaglutide").

## `locations` (type: `array`):

Location terms (city, state, country) to search sites by.

## `status` (type: `array`):

Filter by recruitment status (applies to all searches).

## `nctIds` (type: `array`):

Specific NCT IDs to look up (e.g. NCT07505745).

## `includeLocations` (type: `boolean`):

Include the full list of study locations and their contacts.

## `maxItems` (type: `integer`):

Maximum studies across all searches. 0 = no limit.

## `maxPagesPerQuery` (type: `integer`):

Pagination cap per search (100 per page).

## `proxyConfiguration` (type: `object`):

Apify Proxy. The API is public — datacenter is enough and enabled by default.

## Actor input object example

```json
{
  "conditions": [
    "diabetes"
  ],
  "status": [
    "RECRUITING"
  ],
  "includeLocations": true,
  "maxItems": 200,
  "maxPagesPerQuery": 20,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `nctId` (type: `string`):

NCT identifier

## `title` (type: `string`):

Brief title

## `officialTitle` (type: `string`):

Official title

## `url` (type: `string`):

Study URL

## `status` (type: `string`):

Recruitment status

## `studyType` (type: `string`):

Interventional/etc.

## `phases` (type: `string`):

Trial phases

## `startDate` (type: `string`):

Start date

## `completionDate` (type: `string`):

Completion date

## `enrollmentCount` (type: `string`):

Enrollment count

## `leadSponsor` (type: `string`):

Lead sponsor

## `leadSponsorClass` (type: `string`):

INDUSTRY/NIH/OTHER

## `collaborators` (type: `string`):

Collaborators

## `conditions` (type: `string`):

Conditions

## `interventionNames` (type: `string`):

Intervention names

## `sex` (type: `string`):

Eligible sex

## `minimumAge` (type: `string`):

Minimum age

## `maximumAge` (type: `string`):

Maximum age

## `briefSummary` (type: `string`):

Brief summary

## `primaryContactName` (type: `string`):

Central contact name

## `primaryContactEmail` (type: `string`):

Central contact email

## `primaryContactPhone` (type: `string`):

Central contact phone

## `centralContacts` (type: `string`):

All central contacts

## `locationsCount` (type: `string`):

Number of sites

## `countries` (type: `string`):

Site countries

## `locations` (type: `string`):

Study sites + contacts

## `hasResults` (type: `string`):

Results posted

## `scrapedAt` (type: `string`):

ISO timestamp

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "conditions": [
        "diabetes"
    ],
    "status": [
        "RECRUITING"
    ],
    "maxItems": 200,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("haketa/clinicaltrials-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "conditions": ["diabetes"],
    "status": ["RECRUITING"],
    "maxItems": 200,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("haketa/clinicaltrials-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "conditions": [
    "diabetes"
  ],
  "status": [
    "RECRUITING"
  ],
  "maxItems": 200,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call haketa/clinicaltrials-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,haketa/clinicaltrials-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9aeIMpU0o3p7HK31X/builds/VGYB3IfqMHq15iwpU/openapi.json
