ChEMBL Molecules Scraper avatar

ChEMBL Molecules Scraper

Pricing

from $28.50 / 1,000 results

Go to Apify Store
ChEMBL Molecules Scraper

ChEMBL Molecules Scraper

Scrape molecules from EBI ChEMBL public API including SMILES, InChI, molecular properties (MW, logP, HBA, HBD, PSA, RTB), max phase, ATC classifications, oral/parenteral/topical flags, first approval, black box warning, prodrug and withdrawn flag. No API key required.

Pricing

from $28.50 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

0

Monthly active users

6 days ago

Last modified

Share

ParseForge Banner

🧪 ChEMBL Bioactive Molecules Scraper

🚀 Export ChEMBL drug discovery data in seconds. Pull 2.5 million+ bioactive molecules with SMILES, InChI, ATC codes, clinical phase, and approval status. No API key, no registration, no manual REST stitching.

The ChEMBL Molecules Scraper queries the EBI ChEMBL public REST API and returns 17 fields per molecule, including the canonical ChEMBL ID, preferred name, molecule type, max clinical phase, full structure descriptors (canonical SMILES, InChI, InChI Key), calculated molecular properties (molecular weight, LogP, hydrogen-bond donors and acceptors, polar surface area, rotatable bonds, Lipinski Rule of Five violations), ATC classifications, route of administration flags, first-approval year, and withdrawn status. ChEMBL is maintained by the European Bioinformatics Institute and is one of the largest manually curated databases of bioactive molecules in drug discovery.

The catalog covers small molecules, antibodies, enzymes, proteins, oligonucleotides, oligosaccharides, cells, genes, and unknowns, totalling more than 2.5 million entries. This Actor makes the data downloadable as CSV, Excel, JSON, or XML in under a minute. The molecule type filter runs server-side, so antibody-only or small-molecule-only exports are fast.

🎯 Target Audience💡 Primary Use Cases
Cheminformaticians, drug discovery scientists, computational chemists, pharma data teams, ML researchers, bioinformaticians, academic labs, regulatory analystsQSAR datasets, virtual screening libraries, ADMET feature tables, ATC mapping, clinical-phase tracking, approved-drug audits, withdrawn-drug watchlists

📋 What the ChEMBL Molecules Scraper does

Two filtering workflows in a single run:

  • 🔎 Full-text query. Substring match across molecule names and synonyms (e.g. aspirin, imatinib, bevacizumab).
  • 🧬 Type filter. Server-side filter on molecule_type. Pick from small molecule, antibody, enzyme, protein, oligonucleotide, oligosaccharide, cell, gene, or unknown.
  • 📜 Paginated catalog dump. Leave both filters empty to walk the entire ChEMBL catalog by offset.

Each record returns the canonical ChEMBL ID, the public explorer URL, the structure block (SMILES, InChI, InChI Key, molfile) when present, the property block (MW, LogP, HBA, HBD, PSA, RTB, full MWT, Rule-of-Five violations), the molecule hierarchy (active / parent / salt), the ATC classifications array, administration route flags (oral, parenteral, topical), the black-box-warning flag, the first-approval year, the withdrawn flag, and the prodrug flag.

💡 Why it matters: ChEMBL underpins most modern drug discovery pipelines. Building your own REST pagination, retry logic, and field selection means a week of plumbing. This Actor returns clean, joined records on every run.

📊 Data fields

Each record includes: atc_classifications, black_box_warning, first_approval, max_phase, molecule_chembl_id, molecule_hierarchy, molecule_properties, molecule_structures, molecule_type, oral, parenteral, pref_name, prodrug, scrapedAt, topical, url, withdrawn_flag. All 17 field names come from a real production run, so what you see here is what lands in your dataset.

🚀 How to use

  1. 📝 Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. 🌐 Open the Actor. Go to the ChEMBL Bioactive Molecules Scraper page on the Apify Store.
  3. 🎯 Set input. Pick a molecule type, enter a text query, and set maxItems.
  4. 🚀 Run it. Click Start and let the Actor collect your data.
  5. 📥 Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.

⏱️ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

💡 Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

⚠️ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by ChEMBL, the European Bioinformatics Institute, or EMBL-EBI. All trademarks mentioned are the property of their respective owners. Only publicly available open ChEMBL data is collected.

🆘 Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.