UniProt Protein Sequence & Annotation Scraper avatar

UniProt Protein Sequence & Annotation Scraper

Pricing

from $28.12 / 1,000 results

Go to Apify Store
UniProt Protein Sequence & Annotation Scraper

UniProt Protein Sequence & Annotation Scraper

Export UniProt Knowledgebase entries โ€” search Swiss-Prot by organism, keyword, gene, or any UniProt query, or fetch a single accession. Returns names, genes, organism, sequence length & molecular weight, keywords, comments, features, and PDB/RefSeq/Ensembl/KEGG cross-refs.

Pricing

from $28.12 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

0

Monthly active users

a day ago

Last modified

Share

ParseForge Banner

๐Ÿงฌ UniProt Protein Sequence & Annotation Scraper

๐Ÿš€ Export UniProt Knowledgebase entries in seconds. Query Swiss-Prot and TrEMBL by organism, gene, keyword, subcellular location, length range, or any UniProt field, or fetch a single accession with full annotations. No API key, no SPARQL, no XML parsing.

The UniProt Protein Scraper queries the official UniProt REST API and returns standardized protein records from the world's largest protein-sequence knowledgebase. Each entry carries the primary accession, UniProtKB ID, entry type (reviewed Swiss-Prot vs unreviewed TrEMBL), protein name, alternative names, gene names, organism (scientific + common + taxon ID + lineage), evidence level, annotation score, sequence length, molecular weight, CRC64 / MD5 sequence hashes, keywords (with categories), curated comments (function, subunit, subcellular location, etc.), structural features, reference counts, last-update date, entry version, and the canonical UniProt URL.

UniProt is maintained jointly by EMBL-EBI, SIB, and PIR and is the de facto reference for protein biology in research, pharma, and bioinformatics. Coverage spans 250 million+ entries across 2.7 million+ species in TrEMBL, with ~570,000 manually curated entries in Swiss-Prot. This Actor flattens UniProt's nested JSON into rows that drop into pandas, R, or any warehouse.

๐ŸŽฏ Target Audience๐Ÿ’ก Primary Use Cases
Bioinformatics teams, computational biologists, pharma research, structural biologists, drug-discovery startups, science journalistsProteome exports, gene-to-protein mapping, target dossier builds, organism-level annotation, sequence + feature retrieval, cross-database joining

๐Ÿ“‹ What the UniProt Scraper does

Two lookup modes in one Actor:

  • ๐Ÿ” Query mode. Pass any UniProt query (reviewed:true AND organism_id:9606, keyword:KW-0181, gene:BRCA1, cc_subcellular_location:nucleus, existence:1, taxonomy_id:10090 AND length:[100 TO 500]).
  • ๐Ÿ†” Accession mode. Set accession (e.g. P00533) for a single full-entry pull. Skips the search query entirely.

Each record carries identifiers (primary accession, UniProtKB ID, entry type), names (protein name, alternative names, gene names), taxonomy (scientific + common organism, taxon ID, lineage), evidence (protein existence, annotation score), sequence facts (length, molecular weight, CRC64, MD5, plus optional full sequence string), curated annotations (keywords, comments, features), reference + feature counts, last-updated date, version, and the canonical UniProt URL.

๐Ÿ’ก Why it matters: UniProt's REST API is rich but verbose. Researchers and engineering teams spend days writing parsers for keywords, comments, and features. This Actor flattens the response into 25 spreadsheet-ready fields so target dossiers, comparative proteomics, and dataset prep land in one query.

๐Ÿ“Š Data fields

Each record includes: alternativeNames, annotationScore, comments, crossReferences, ecNumbers, entryType, entryVersion, featureCount, features, geneNames, geneSynonyms, keywords, lastUpdated, organismCommon, organismLineage, organismScientific, primaryAccession, proteinExistence, proteinName, referenceCount, scrapedAt, sequenceCrc64, sequenceLength, sequenceMd5, sequenceMolWeight, taxonId, uniProtkbId, url. All 28 field names come from a real production run, so what you see here is what lands in your dataset.

๐Ÿš€ How to use

  1. ๐Ÿ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. ๐ŸŒ Open the Actor. Go to the UniProt Protein Scraper page on the Apify Store.
  3. ๐ŸŽฏ Set input. Pick a query (reviewed:true AND organism_id:9606 is a great starter) or an accession.
  4. ๐Ÿš€ Run it. Click Start and let the Actor walk the UniProt API.
  5. ๐Ÿ“ฅ Download. Grab results in the Dataset tab as CSV, Excel, JSON, or XML.

โฑ๏ธ Total time from signup to a downloaded proteome slice: 3-5 minutes. No coding required.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

โš ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by EMBL-EBI, the SIB Swiss Institute of Bioinformatics, the Protein Information Resource (PIR), the UniProt Consortium, or any of their funding agencies. All trademarks mentioned are the property of their respective owners. Only publicly available UniProtKB data is collected. Please cite UniProt as required by their CC BY 4.0 license.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.