ClinicalTrials Sponsor Scraper | FDA Cross-Ref, 12 Fields avatar

ClinicalTrials Sponsor Scraper | FDA Cross-Ref, 12 Fields

Pricing

from $0.60 / 1,000 trial scrapeds

Go to Apify Store
ClinicalTrials Sponsor Scraper | FDA Cross-Ref, 12 Fields

ClinicalTrials Sponsor Scraper | FDA Cross-Ref, 12 Fields

Scrape ClinicalTrials.gov trials by sponsor or condition with phase, status & completion date, then auto-flag FDA drug approvals to build pharma pipeline intelligence. No API key. Use it as an MCP server in Claude, ChatGPT & AI agents.

Pricing

from $0.60 / 1,000 trial scrapeds

Rating

0.0

(0)

Developer

The Mine Works

The Mine Works

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

a day ago

Last modified

Share

ClinicalTrials Sponsor Scraper: FDA Cross-Ref, 12 Fields

Pay only for results delivered. Browse all Actors.

💰 From $0.60 / 1,000 results.

Guide and FAQs: ClinicalTrials Sponsor Intel on themineworks.com. Tutorial: Pharma Pipeline Tracking: Trial Sponsors vs FDA Approvals.

Track a pharma sponsor's clinical pipeline and cross-reference every intervention against FDA's Drugs@FDA approval database in the same run. 13 fields per trial: NCT ID, title, phase, overall status, start date, completion date, results-first-posted date, lead sponsor, interventions, an fda_approved flag, the matching FDA approval date, the study URL and a capture timestamp.

Two official APIs, pure HTTP. No API key, no login, no browser, up to 2,000 trials per run.

Why use this ClinicalTrials sponsor scraper

  • The FDA cross-reference is the point. For each trial the Actor takes the first drug or biological intervention that is not a placebo arm and checks it against api.fda.gov/drug/drugsfda.json, by generic (INN) name first and brand name second, returning fda_approved: true, false or null plus the approval date on a match. That turns a trial list into a pipeline picture.
  • Results-first-posted date. results_first_posted tells you whether a completed trial actually reported results, which is the difference between a finished study and a published one.
  • Sponsor and condition compose. Search a sponsor's whole book, or one sponsor within one disease area.
  • Completed and terminated by default. The default status filter is COMPLETED,TERMINATED, the retrospective view a pipeline analysis wants. Override it for a forward-looking one.
  • The cross-reference is optional. Set crossRefFDA: false to skip the FDA calls and run faster when you only need the trial list.

Map a sponsor's completed pipeline

{
"sponsor": "Pfizer",
"status": "COMPLETED,TERMINATED",
"crossRefFDA": true,
"maxResults": 500
}

Completed and terminated trials with an approval flag per intervention is the standard read on what a sponsor has actually converted.

Watch a sponsor's active programmes

{
"sponsor": "Moderna",
"status": "RECRUITING,ACTIVE_NOT_RECRUITING",
"phase": "PHASE2,PHASE3",
"maxResults": 300
}

Late-phase and active is the forward-looking view: what is still in flight and how far along.

Compare sponsors within one disease area

{
"condition": "multiple myeloma",
"status": "COMPLETED",
"crossRefFDA": true,
"maxResults": 1000
}

With no sponsor filter, lead_sponsor on every row lets you group by company and compare approval conversion across a therapeutic area.

Run a fast trial list without the FDA lookup

{
"sponsor": "AstraZeneca",
"condition": "asthma",
"crossRefFDA": false,
"maxResults": 2000
}

Turning off the cross-reference removes the FDA requests and the 200 ms pause between them, which materially shortens a large run.

What data does the sponsor scraper return

{
"nct_id": "NCT04368728",
"title": "Study to Evaluate the Efficacy and Safety of an Example Vaccine",
"phase": "PHASE3",
"overall_status": "COMPLETED",
"start_date": "2020-07-27",
"completion_date": "2023-06-30",
"results_first_posted": "2023-11-14",
"lead_sponsor": "Example Pharmaceuticals",
"interventions": "BIOLOGICAL:Example Vaccine; OTHER:Placebo",
"fda_approved": true,
"fda_approval_date": "2004-12-30",
"url": "https://clinicaltrials.gov/study/NCT04368728",
"scraped_at": "2026-08-02T09:00:00.000Z"
}
FieldDescription
nct_id, title, urlStudy identity and link
phase, overall_statusWhere the study sits and how it ended
start_date, completion_dateStudy timeline
results_first_postedWhen results were first posted, if ever
lead_sponsorSponsor name as registered
interventionsSemicolon-joined TYPE:Name list
fda_approvedDrugs@FDA match on the trial's first non-placebo drug or biological intervention
fda_approval_dateEarliest original approval date among the matching FDA applications
scraped_atCapture timestamp

Read fda_approved for what it is. It is a name match of the trial's first drug or biological intervention against Drugs@FDA, generic name first and brand name second. Placebo and vehicle arms are stepped over, so a trial that lists its control first is still matched on the drug under study. It is a strong screening signal, not a regulatory determination: placebo arms, compounds still carrying a development code such as PF-04531083, and interventions named differently in the trial than on the label will not match. null means the check could not be completed or there was no drug name to check, which is not the same as false.

fda_approval_date is the earliest original approval among the applications that matched the name, not necessarily the drug's first ever US approval. A generic name can match later applications only, so read it as the date attached to this match rather than a regulatory milestone.

How the scraper works without an API key

Both APIs are open. The Actor queries ClinicalTrials.gov API v2 for the sponsor, condition and status you specify, requesting only the protocol-section modules it maps, and follows the page token. For each trial it then pauses 200 ms and queries Drugs@FDA for the first non-placebo drug or biological intervention, trying openfda.generic_name before openfda.brand_name. Names already looked up in the run are reused rather than queried again. No key, no login, no browser on either side.

What can you build with sponsor pipeline data

Pharma competitive intelligence. A sponsor's trials by phase and status, with an approval signal per programme.

Approval-conversion analysis. How many of a sponsor's completed phase 3 programmes carry an approved intervention.

Therapeutic-area mapping. Every sponsor active in one condition, grouped and ranked.

Investment diligence. Termination rates and results-reporting behaviour as quality signals on a sponsor.

How much does it cost to scrape sponsor pipelines

Pay per trial delivered: $0.001 on the Apify Free plan, $0.0006 on Gold and above ($0.0009 Bronze, $0.00075 Silver). Every run also charges the standard Actor start event of $0.005. A sponsor or condition that matches nothing delivers no trials, so nothing past the start event is charged, and the FDA cross-reference costs nothing extra.

How do I use sponsor intelligence in Claude or ChatGPT

https://mcp.apify.com/?tools=themineworks/clinicaltrials-sponsor-intelligence
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('themineworks/clinicaltrials-sponsor-intelligence').call({
sponsor: 'Pfizer',
status: 'COMPLETED,TERMINATED',
maxResults: 200,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();

Do I need an API key? No. ClinicalTrials.gov API v2 and openFDA are both open.

How is fda_approved determined? By searching the trial's first non-placebo drug or biological intervention against Drugs@FDA, generic name first and brand name second, and confirming the application really carries that name. Treat it as a screening signal rather than a regulatory fact.

What does fda_approved: null mean? The check could not be completed, usually no drug name to search (a device or an arm label rather than a drug), or an FDA request that failed. It is not a negative result.

Can I search by condition instead of sponsor? Yes. Both are optional and both can be supplied together.

How do I pass several phases or statuses? As a comma-separated string, for example PHASE2,PHASE3.

Complete your pharma intelligence pipeline

Found a bug or want a field added? Open an issue on the Actor's Apify Console page.

Last verified: 2026-08