# ClinicalTrials.gov Scraper — Trials, Sponsors, Sites, Contacts (`chorelet/clinical-trials-scraper`) Actor

Search ClinicalTrials.gov by condition, drug, sponsor, place, phase, status and date: one row per trial with enrollment, sponsor, interventions, dates, eligibility and contacts — plus one row per study site with its own address and contact. Official API v2, no key.

- **URL**: https://apify.com/chorelet/clinical-trials-scraper.md
- **Developed by:** [Chorelet](https://apify.com/chorelet) (community)
- **Categories:** Business, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 trials

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## ClinicalTrials.gov Scraper — Trials, Sponsors, Sites, Contacts

Search the **US National Library of Medicine's trial registry** the way you actually think about it — by condition, drug, sponsor, country, phase, status, start date or last update — and get a clean row per trial: title, status, phase, enrollment, sponsor and collaborators, conditions, interventions, all four key dates, eligibility, outcome measure and contact details. As JSON, CSV or Excel, or through the API.

Turn on **one row per study site** and every hospital running the trial comes with it: facility name, city, state, country, coordinates, its own recruiting status, the site contact's name, email and phone, and the principal investigator. One phase-3 oncology trial in a test run expanded into 229 sites; eight trials produced 1,189 site rows.

Data comes from the official **API v2** — no key, no account, no scraping of HTML, and 200 trials per request.

### Why this Actor

- **Watch a pipeline, not a page.** `Updated since 7 days` plus a sponsor or a drug returns only the trials whose record actually changed — the query a competitive-intelligence run is built on.
- **Sites and their contacts, not just trials.** Every participating centre with city, country, coordinates, its own recruiting status, contact email and principal investigator. Eight phase-3 trials expanded into 1,189 site rows in a test run.
- **The registry's own API, not its HTML.** Official API v2: no key, no account, 200 trials per request, and filters the website itself uses — phase, status, study type, start date, last update.
- **Site rows cost a third of a trial row**, so expanding a search into hundreds of centres stays cheap.

### Sample output

One item of the dataset (long values shortened):

```json
{
  "nctId": "NCT07671092",
  "title": "Imaging Study in Metastatic UC, HR+ and HER2- Breast Cancer, TNBC, and NSCLC",
  "status": "RECRUITING",
  "phase": "EARLY_PHASE1",
  "enrollment": 40,
  "sponsor": "RayzeBio, Inc.",
  "conditions": [
    "Metastatic Urothelial Carcinoma",
    "Breast Cancer",
    "TNBC - Triple-Negative Breast Cancer",
    "…"
  ],
  "startDate": "2026-06-15",
  "lastUpdateDate": "2026-09-29",
  "url": "https://clinicaltrials.gov/study/NCT07671092"
}
```

### What you get

- **The whole search surface of the registry**: conditions, drugs and other interventions, sponsors, places, free text, recruitment status, phase, study type, "updated since" and "started since"
- **`Updated since 7 days` is the competitive-intelligence query** — it returns exactly the trials whose record changed, which is how you watch a rival's pipeline without reading the registry by hand
- **Sites with contacts**: the address, coordinates, status and named contact of every participating centre — the list CROs and site-selection teams build by hand
- **One schema for both row types**, so trials and sites land in the same spreadsheet and a `type` column tells them apart
- Sort by last update, newest, start date, enrollment size or relevance; cap sites per trial to keep a big search predictable
- Every row carries its NCT number and registry link

### Input example

```json
{
  "conditions": [
    "breast cancer"
  ],
  "statuses": [
    "RECRUITING"
  ],
  "sortBy": "last update",
  "includeSites": false,
  "maxSitesPerStudy": 0,
  "maxResults": 100
}
```

### How much does it cost?

Pay per trial — no subscription, no minimum, no charge for platform usage.

| Volume | Price |
|---|---|
| 1,000 trials | $1.00 (+ $0.30 with `site`) |
| 10,000 trials | $10.00 (+ $3.00 with `site`) |
| 100,000 trials | $100.00 (+ $30.00 with `site`) |

The Apify **free plan includes $5 of usage every month** — about 5,000 trials with this Actor, no card needed. Nothing else is charged: platform usage is included in the price, and Apify Bronze, Silver and Gold subscribers get 10%, 20% and 30% off these prices.

### Use it from code, n8n, Make, Zapier or an AI agent

Run the Actor and download the dataset in one call (JSON by default; add `&format=csv` or `xlsx`):

```bash
curl -X POST "https://api.apify.com/v2/acts/chorelet~clinical-trials-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"conditions": ["breast cancer"], "statuses": ["RECRUITING"], "sortBy": "last update", "includeSites": false, "maxSitesPerStudy": 0, "maxResults": 100}'
```

Python:

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("chorelet/clinical-trials-scraper").call(run_input={"conditions": ["breast cancer"], "statuses": ["RECRUITING"], "sortBy": "last update", "includeSites": false, "maxSitesPerStudy": 0, "maxResults": 100})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

- **n8n, Make, Zapier** — use the Apify node/module: run the Actor, then "get dataset items".
- **Google Sheets, Slack, webhooks** — add an integration on the run's *Integrations* tab.
- **AI agents** — the Actor is available as a tool through the Apify MCP server; the dataset schema describes every field for the model.
- **Schedules** — run it hourly, daily or weekly from the *Schedules* tab.

### FAQ

**Do I need an API key?**

No. ClinicalTrials.gov publishes its API v2 openly — no key, no account, no proxy.

**How do I watch a competitor's trials?**

Put the company in *Sponsors*, set *Updated since* to `7 days` and schedule the run daily. Every row that comes back is a record that changed since the last run.

**What counts as a site row?**

One participating hospital or clinic of a trial, with its address, coordinates, recruiting status, contact and investigator. Turn on *Add one row per study site*; a large phase-3 trial can list several hundred, so *Max sites per trial* caps them.

**Can I search by drug?**

Yes — *Drugs or interventions* searches the intervention fields, so `semaglutide` returns the trials testing it rather than every record that mentions it.

**How far can one run go?**

The registry returns 200 trials per request and the Actor pages through them; a 50,000-trial search is a few minutes of requests.

**Is the data free to reuse?**

ClinicalTrials.gov is a US government service and its records are in the public domain; check the registry's terms for how to credit it in a publication.

### Support

Questions, missing fields or a source that changed? Open an issue on the *Issues* tab or write to support@chorelet.app — problems are usually fixed within a day, and the Actor is checked every morning by an automated test run. If the Actor saved you time, a short review on its Store page helps other people find it.

# Actor input Schema

## `conditions` (type: `array`):

`breast cancer`, `type 2 diabetes`, `long covid`. Several are OR-ed.

## `interventions` (type: `array`):

`semaglutide`, `pembrolizumab`, `CBT`. Several are OR-ed.

## `sponsors` (type: `array`):

Lead sponsor or collaborator: `Pfizer`, `Novo Nordisk`, `Mayo Clinic`.

## `locations` (type: `array`):

Country, state or city where the trial runs: `Germany`, `Texas`, `Tokyo`.

## `terms` (type: `array`):

Free text searched across the whole record — an NCT number, a gene, a device name.

## `statuses` (type: `array`):

Empty = any status.

## `phases` (type: `array`):

Empty = any phase.

## `studyTypes` (type: `array`):

Empty = any type.

## `updatedSince` (type: `string`):

`30 days`, `last 2 weeks` or a date such as `2026-09-01` — the way to watch a field for changes.

## `startedSince` (type: `string`):

Same formats; filters on the trial's start date instead.

## `sortBy` (type: `string`):

What comes first.

## `includeSites` (type: `boolean`):

Every hospital or clinic running the trial, with its city, country, coordinates, recruiting status and its own contact. A large phase-3 trial can list several hundred.

## `maxSitesPerStudy` (type: `integer`):

0 = every site. Use it to keep the row count predictable on big trials.

## `maxResults` (type: `integer`):

Site rows do not count towards this cap.

## Actor input object example

```json
{
  "conditions": [
    "breast cancer"
  ],
  "interventions": [],
  "sponsors": [],
  "locations": [],
  "terms": [],
  "statuses": [
    "RECRUITING"
  ],
  "phases": [],
  "studyTypes": [],
  "sortBy": "last update",
  "includeSites": false,
  "maxSitesPerStudy": 0,
  "maxResults": 100
}
```

# Actor output Schema

## `trials` (type: `string`):

Everything scraped — items of the default dataset. Use ?format=csv or xlsx on this URL for spreadsheets.

## `summary` (type: `string`):

The search that ran, how many trials matched, how many rows were saved.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "conditions": [
        "breast cancer"
    ],
    "interventions": [],
    "sponsors": [],
    "locations": [],
    "terms": [],
    "statuses": [
        "RECRUITING"
    ],
    "phases": [],
    "studyTypes": [],
    "updatedSince": "",
    "startedSince": "",
    "sortBy": "last update",
    "includeSites": false,
    "maxSitesPerStudy": 0,
    "maxResults": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("chorelet/clinical-trials-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "conditions": ["breast cancer"],
    "interventions": [],
    "sponsors": [],
    "locations": [],
    "terms": [],
    "statuses": ["RECRUITING"],
    "phases": [],
    "studyTypes": [],
    "updatedSince": "",
    "startedSince": "",
    "sortBy": "last update",
    "includeSites": False,
    "maxSitesPerStudy": 0,
    "maxResults": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("chorelet/clinical-trials-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "conditions": [
    "breast cancer"
  ],
  "interventions": [],
  "sponsors": [],
  "locations": [],
  "terms": [],
  "statuses": [
    "RECRUITING"
  ],
  "phases": [],
  "studyTypes": [],
  "updatedSince": "",
  "startedSince": "",
  "sortBy": "last update",
  "includeSites": false,
  "maxSitesPerStudy": 0,
  "maxResults": 100
}' |
apify call chorelet/clinical-trials-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,chorelet/clinical-trials-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/dSLthTRcum1uVZVe5/builds/bKV2h4B1yJHfWHNNL/openapi.json
