ClinicalTrials.gov Scraper — Trials & Sponsors
Pricing
from $11.54 / 1,000 studies
ClinicalTrials.gov Scraper — Trials & Sponsors
Scrape the ClinicalTrials.gov registry — trials with status, phase, sponsor, conditions, interventions, enrollment, locations and contacts. Filter by condition, sponsor and status. Official keyless API. Schedule as a pharma/biotech trial monitor. Pay per result.
Pricing
from $11.54 / 1,000 studies
Rating
0.0
(0)
Developer
Vitalii Bondarev
Maintained by CommunityActor stats
0
Bookmarked
5
Total users
4
Monthly active users
8 days ago
Last modified
Share
ClinicalTrials.gov Scraper — Trials, Sponsors & Status
Extract structured clinical-trial records from the official ClinicalTrials.gov registry — the U.S. National Library of Medicine database of public and private clinical studies run around the world. Filter by condition, sponsor and status, and get clean JSON for every matching trial: status, phase, sponsor, conditions, interventions, enrollment, eligibility, locations and contacts.
Built on the official keyless ClinicalTrials.gov API v2 — no login, no API key, no fragile HTML scraping. Stable, fast, and complete.
What you get (per study)
| Field | Description |
|---|---|
nct_id | The trial's NCT registry identifier |
brief_title / official_title | Short and full study titles |
overall_status | Recruiting, Completed, Terminated, … |
phases | Trial phase(s) — Phase 1/2/3/4 |
study_type | Interventional / Observational |
conditions | Diseases / conditions studied |
interventions | Drugs, devices, procedures tested |
lead_sponsor / lead_sponsor_class | Sponsoring organisation + class (Industry / NIH / Other) |
collaborators | Collaborating organisations |
enrollment_count | Target / actual participant count |
start_date / primary_completion_date / completion_date | Key milestone dates |
eligibility_min_age / eligibility_max_age / eligibility_sex | Participant eligibility |
central_contact_name / central_contact_email | Study contact |
locations_count / first_location_country | Site footprint |
study_url | Direct link to the trial page |
Input
condition— disease / condition to search (e.g.cancer,diabetes).term— any other free-text search term (drug, NCT id, keyword).sponsor— filter by sponsoring organisation.location— filter by trial location.statuses— list of recruitment statuses to include (optional).maxItems— cap the number of studies returned (0= all matches).
Use cases
- Pharma & biotech competitive intelligence — track who is running trials for a given condition or drug class, and at what phase.
- Patient-recruitment & site selection — find recruiting trials by condition, geography and eligibility.
- Market & investment research — monitor a sponsor's pipeline and milestone dates.
- Academic & systematic reviews — pull structured trial metadata at scale.
Pricing
Pay-per-result: $0.0119 per study (study event). You pay only for the
records actually returned — no subscription, no minimum. Platform compute runs on
your own Apify account.
Notes
Data comes from the public ClinicalTrials.gov registry and is provided as-is for research and intelligence use. The actor reads the official API only.
Usage statistics
This Actor creates a small, content-free summary at the end of each run. It is used only to monitor reliability and improve this Actor. A copy is saved as USAGE_STATS in your own Apify key-value store, so you can see the exact record created for your run.
Set disableUsageStats to true in the input to opt out. Nothing is sent then; your USAGE_STATS record only says that statistics were disabled.
Only these fields are recorded:
- schema version, Actor name and build number;
- UTC start and finish hour (not a precise timestamp);
- run duration, number of results and time to the first result, each as a coarse range;
- whether the result was empty, the end status, and an error type from a fixed list;
- memory setting and counts of charged events;
- names of the input fields you set, never their values;
- the selected option for input fields that offer a fixed list of choices (for example a sort order).
We do not collect input text, search terms, URLs, domains, usernames, email addresses, names, proxy credentials, tokens, scraped records, output items, raw error messages, stack traces, or hashes of any of those values. Records are kept for no longer than 13 months, used only as aggregated operational statistics, and never sold or shared.
Additional fields (Phase 2)
This Actor also records your Apify user ID, whether Apify marks the account as paying, the size range of list inputs, the selected country when the input offers a fixed list of countries, and one category from a fixed Actor taxonomy. We use these fields only for aggregate reliability, repeat-use and cross-Actor analysis; reports suppress any cell with fewer than five distinct users.
The same disableUsageStats: true input flag turns these fields off too. The user ID is removed after 13 months; we do not export, sell, share, or attempt to re-identify this data.
Run-outcome signals (v2)
To learn whether a run did what it was asked to do, the record also holds a few more coarse ranges and yes/no flags. None of them contains content:
- the result limit you asked for (a range, when the input has one) and what share of it was delivered;
- results delivered per input item you listed (a range);
- output quality as ranges: how fully the result fields were filled, the share of rows that look like errors, the share of duplicate rows, and how many different fields appeared. These are counted in memory while results are saved; no result content is kept;
- how the run was started (console, API, schedule, webhook, another Actor);
- how it ended: stopped by you, timed out, reached the requested limit, stopped by the charge limit, and how many times the platform moved the run;
- if this Actor reports it: how many items to process worked or failed (ranges) and one failure reason from a fixed list;
- a short code made from the names of the input fields you set, never their values.
Repeat-run fingerprint (v2)
When your Apify user ID is recorded (see above), the record also holds an 8-character one-way code made from your input (proxy settings left out) and this Actor's name. It only lets us see that the same account ran the same input again soon after an unsatisfying run; we never see the input itself. It is stored only in the database, never published, and reports use it in aggregate with the same five-user minimum. It is the one exception to the statement above that no hashes are collected, and disableUsageStats: true turns it off.