RCSB PDB Protein Structure Scraper avatar

RCSB PDB Protein Structure Scraper

Pricing

from $28.87 / 1,000 results

Go to Apify Store
RCSB PDB Protein Structure Scraper

RCSB PDB Protein Structure Scraper

Scrape protein structure entries from the RCSB Protein Data Bank including title, authors, citation, experimental method (X-ray, EM, NMR), resolution, cell parameters, symmetry, polymer entities, keywords and entry metadata. No API key required.

Pricing

from $28.87 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

0

Monthly active users

a day ago

Last modified

Share

ParseForge Banner

๐Ÿงฌ RCSB Protein Data Bank Scraper

๐Ÿš€ Export 3D macromolecular structure metadata in seconds. Pull 220,000+ PDB entries with resolution, experimental method, unit cell, primary citation, and deposit history. No API key, no registration, no manual REST stitching.

The RCSB PDB Scraper queries the RCSB Search API and Data API and returns 22 fields per structure, including the 4-character PDB ID, title and descriptor, classification keywords, experimental method (X-ray, cryo-EM, NMR, neutron, fiber, powder, scattering), combined resolution, unit-cell dimensions and crystal symmetry (for X-ray entries), deposit and release dates, polymer composition and atom count, the audit-author list, and the full primary citation (title, journal, year, authors, DOI, PubMed ID). The Protein Data Bank has been the global archive of 3D biological macromolecular structures since 1971.

The catalog covers proteins, nucleic acids, complexes, viruses, ribosomes, membrane proteins, and small-molecule ligands across X-ray diffraction, electron microscopy (cryo-EM), solution and solid-state NMR, neutron diffraction, fiber, powder, electron crystallography, and solution scattering. This Actor makes the data downloadable as CSV, Excel, JSON, or XML in under a minute. Crystallographic fields (unit cell, space group, resolution refinement) are surfaced only when relevant to the experiment.

๐ŸŽฏ Target Audience๐Ÿ’ก Primary Use Cases
Structural biologists, cryo-EM researchers, computational chemists, drug discovery teams, bioinformaticians, journal editors, citation analysts, ML researchersStructure browsing, citation graphs, method benchmarking, drug-target validation, training sets for AI structure prediction, deposition tracking, journal scientometrics

๐Ÿ“‹ What the RCSB PDB Scraper does

Two retrieval modes in a single run:

  • ๐Ÿ”Ž Full-text search. Query the RCSB search API for any text (e.g. hemoglobin, SARS-CoV-2 spike, kinase inhibitor).
  • ๐Ÿ†” Explicit IDs. Pass a list of 4-character PDB entry IDs (e.g. ["3GOU", "1HHO"]) to fetch full metadata directly.
  • ๐Ÿ”ฌ Method filter. Restrict by experimental method (X-ray, cryo-EM, NMR, neutron, fiber, powder, scattering, electron crystallography).

Each record returns the PDB ID, RCSB explorer URL, structure title and descriptor, classification keywords, experimental method, combined resolution, unit-cell dimensions and space group (for X-ray only), refinement resolution, deposit and release dates, polymer entity count, atom and monomer counts, the audit-author list, and the full primary citation block.

๐Ÿ’ก Why it matters: PDB structures are the bedrock of structural biology, drug discovery, and the AlphaFold era. The RCSB API surfaces fields across multiple endpoints; this Actor joins them into a single, denormalized row per entry, complete with citation metadata.

๐Ÿ“Š Data fields

Each record includes: audit_authors, branched_entity_count, cell, crystals_number, deposit_date, deposited_atom_count, deposited_polymer_monomer_count, experimental_method, keyword_text, keywords, ls_d_res_high, polymer_composition, polymer_entity_count, primary_citation, rcsb_id, release_date, resolution_combined, revision_date, scrapedAt, symmetry, title, url. All 22 field names come from a real production run, so what you see here is what lands in your dataset.

๐Ÿš€ How to use

  1. ๐Ÿ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. ๐ŸŒ Open the Actor. Go to the RCSB Protein Data Bank Scraper page on the Apify Store.
  3. ๐ŸŽฏ Set input. Enter a search query or paste a list of PDB IDs, optionally filter by method.
  4. ๐Ÿš€ Run it. Click Start and let the Actor collect your data.
  5. ๐Ÿ“ฅ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.

โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

โš ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by RCSB PDB, the wwPDB, or any of its partner sites. All trademarks mentioned are the property of their respective owners. Only publicly available open structural-biology data is collected.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.