NCBI Gene Database Scraper avatar

NCBI Gene Database Scraper

Pricing

from $7.50 / 1,000 results

Go to Apify Store
NCBI Gene Database Scraper

NCBI Gene Database Scraper

Scrapes NCBI Gene records by Entrez query and returns each gene as a flat row with identifier, symbol, name, location, aliases, OMIM ID, organism, and functional summary.

Pricing

from $7.50 / 1,000 results

Rating

0.0

(0)

Developer

Acquisition Automation Co.

Acquisition Automation Co.

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Acquisition Automation Co. Search less. Close more.

๐Ÿงฌ NCBI Gene Database Scraper

Run an Entrez query against the NCBI Gene database and get each gene back as a flat row: Gene ID, symbol, description, organism, chromosome, map location, aliases, protein synonyms, the RefSeq summary and the record URL. No API key, no registration, no login.

NCBI Gene is the reference record for a gene: the identifier everything else cites. The web interface answers one query at a time and its export options are built for downstream NCBI tools rather than for a spreadsheet. This Actor runs the same Entrez query you would type into the search box and writes one row per gene into a dataset you can open anywhere.

Who uses itWhat they use Gene records for
๐Ÿ”ฌ Researchers and bioinformaticiansResolving a list of symbols to stable Gene IDs before joining it to any other dataset
๐Ÿ“š Curators and database buildersPulling official symbol, description and aliases for a panel in one pass
๐Ÿงพ Diligence analysts in life sciencesChecking that the genes named in a target's pipeline or patent claims resolve to real records
๐Ÿ–Š Science writers and educatorsGetting the RefSeq summary and map location for a gene without opening ten tabs

๐Ÿ“‹ What it does

๐Ÿ’ก Why it matters: gene symbols get renamed, reused and abbreviated differently by every group that touches them. The Gene ID does not. Every row here carries one.

  • ๐Ÿ”Ž Takes a real Entrez query. Anything the NCBI search box accepts, for example BRCA1[gene] AND human[orgn].
  • ๐Ÿ†” Returns the Gene ID alongside the symbol, so a list of symbols becomes a list of stable identifiers.
  • ๐Ÿงญ Chromosome and map location, for example chromosome 17 at 17q21.31.
  • ๐Ÿ”ค Aliases and protein synonyms, both as delimited strings, which is where an old symbol in a legacy spreadsheet turns up.
  • ๐Ÿ“ The RefSeq summary, the curated paragraph describing what the gene does.
  • ๐Ÿ”— A direct link to the NCBI record for every row.
  • ๐Ÿ’พ Exports to CSV, Excel, JSON or XML, from the run page or the API.

๐Ÿ“Š Output

Every gene is one flat row with 12 fields.

FieldTypeDescription
๐Ÿ†” gene_idstringNCBI Gene ID, the stable identifier for the record
๐Ÿ”ค symbolstringOfficial gene symbol, for example BRCA1
๐Ÿ“„ descriptionstringOfficial full name of the gene
๐Ÿ organismstringScientific name of the organism, for example Homo sapiens
๐Ÿงฌ chromosomestringChromosome the gene sits on
๐Ÿ“ maptypestringCytogenetic map location, for example 17q21.31
๐Ÿ“ summarystringThe curated RefSeq summary of the gene's function, as one paragraph
๐Ÿ” aliasesstringAlternative symbols, comma separated
๐Ÿงช synonymsstringProtein and long-form names, separated by a pipe character
๐Ÿ”— ncbi_urlstringLink to the gene's page on NCBI
๐Ÿ•’ scrapedAtstringISO timestamp of collection
โš ๏ธ errorstringnull on a normal row

Example row

This is the single row returned by a run with the default input, maxItems set to 10 and no query.

{
"gene_id": "672",
"symbol": "BRCA1",
"description": "BRCA1 DNA repair associated",
"organism": "Homo sapiens",
"chromosome": "17",
"maptype": "17q21.31",
"summary": "This gene encodes a 190 kD nuclear phosphoprotein that plays a role in maintaining genomic stability, and it also acts as a tumor suppressor. The BRCA1 gene contains 22 exons spanning about 110 kb of DNA. The encoded protein combines with other tumor suppressors, DNA damage sensors, and signal transducers to form a large multi-subunit protein complex known as the BRCA1-associated genome surveillance complex (BASC). This gene product associates with RNA polymerase II, and through the C-terminal domain, also interacts with histone deacetylase complexes. This protein thus plays a role in transcription, DNA repair of double-stranded breaks, and recombination. Mutations in this gene are responsible for approximately 40% of inherited breast cancers and more than 80% of inherited breast and ovarian cancers. Alternative splicing plays a role in modulating the subcellular localization and physiological function of this gene. Many alternatively spliced transcript variants, some of which are disease-associated mutations, have been described for this gene, but the full-length natures of only some of these variants has been described. A related pseudogene, which is also located on chromosome 17, has been identified. [provided by RefSeq, May 2020]",
"aliases": "BRCAI, BRCC1, BROVCA1, FANCS, IRIS, PNCA4, PPP1R53, PSCP, RNF53",
"synonyms": "breast cancer type 1 susceptibility protein|BRCA1/BRCA2-containing complex, subunit 1|Fanconi anemia, complementation group S|RING finger protein 53|breast and ovarian cancer susceptibility protein 1|breast cancer 1, early onset|early onset breast cancer 1|protein phosphatase 1, regulatory subunit 53",
"ncbi_url": "https://www.ncbi.nlm.nih.gov/gene/672",
"scrapedAt": "2026-09-14T16:29:57.894Z",
"error": null
}

Write a query to get more than one row. A broad Entrez query returns as many genes as maxItems allows.

โœจ Why choose this Actor

What you get
The reference recordData comes from NCBI Gene itself, not from a mirror or a summary site.
Entrez syntax, unchangedThe query you already know from the NCBI search box works here as written.
Identifiers, not just namesgene_id is what makes the export joinable to anything else.
The same 12 fields every runAppend several queries into one sheet without reconciling columns.
You pay per rowNo subscription. A query that matches nothing costs nothing.

๐Ÿš€ How to use it

  1. Create a free Apify account. New accounts start with $5 of credit.
  2. Open the Actor and select Try for free.
  3. Write an Entrez query in query, for example BRCA1[gene] AND human[orgn].
  4. Set maxItems to cap the run.
  5. Select Start, then export from the Dataset tab as CSV, Excel, JSON or XML.

One gene in one organism:

{
"query": "BRCA1[gene] AND human[orgn]",
"maxItems": 10
}

A broader pull:

{
"query": "DNA repair[All Fields] AND human[orgn]",
"maxItems": 500
}

โš™๏ธ Input

FieldRequiredDescription
queryNoNCBI Gene Entrez query, for example BRCA1[gene] AND human[orgn]
maxItemsNoHow many genes to collect per run. Default 10

๐Ÿ’ฐ Pricing

Pay per result. No subscription, and no Apify platform usage on top.

Apify planFreeBronzeSilverGoldPlatinumDiamond
Per gene row$0.0085$0.0082$0.0078$0.0075$0.0075$0.0075
Rows collectedCost on the Free plan
100$0.85
1,000$8.50
10,000$85.00

Free plan runs return up to 10 rows as a preview. Any paid Apify plan lifts that to 1,000,000 per run.

๐Ÿ”Œ Integrate with any app

The dataset is available through the Apify API as soon as the run finishes. Use run-sync-get-dataset-items for a one-shot call, webhooks to trigger what happens next, or the Make, Zapier, Airbyte and LangChain integrations listed on the Actor page.

๐Ÿค– Use with an AI agent

Give an agent live access to NCBI Gene over the Model Context Protocol:

$claude mcp add --transport http apify "https://mcp.apify.com?tools=acquistion-automation/ncbi-eutils-gene-scraper"

Then ask it in plain language what a gene does and have it read the record back.

โ“ Frequently asked questions

Do I need an NCBI API key? No. The Actor reads the public Entrez service. You only need an Apify account to run it.

Which query syntax does query take? Entrez syntax, the same as the NCBI Gene search box. Field tags such as [gene], [orgn] and [All Fields] work, joined with AND, OR and NOT.

Does this return OMIM IDs, sequences or transcript variants? No. The row carries the fields listed in the Output table: identifier, symbol, description, organism, chromosome, map location, summary, aliases, synonyms and the record URL. Sequence and variant data live in other NCBI databases this Actor does not read.

Why did I only get one row? A narrow query matches one gene. Broaden the query or raise maxItems.

How do I read the synonyms field? It is one string with entries separated by a pipe character. Split on | in your spreadsheet or script.

Can I query non-human organisms? Yes. Add an organism tag, for example AND mouse[orgn], and organism comes back on every row so you can check what you got.

๐Ÿ”— More from Acquisition Automation Co.

About Acquisition Automation Co.

We build automation for people buying businesses. The repetitive part of an acquisition search, checking listings, pulling public records, tracking owners and assets, is work a machine should do, so the buyer's time goes into judging deals instead of collecting them.

We add new Actors regularly. If there is a source you need and do not see here, tell us.

๐Ÿ†˜ Support

Open an issue in the Issues tab of this Actor with your run ID, the input you used, and what you expected to get back.

โš ๏ธ Disclaimer

This Actor is independent and is not affiliated with, endorsed by, or sponsored by the National Center for Biotechnology Information, the National Library of Medicine or any government agency. It collects only publicly available data, and it is not medical advice. You are responsible for using that data in compliance with the source's terms of service and applicable law.