NCBI Gene Database Scraper
Pricing
from $7.50 / 1,000 results
NCBI Gene Database Scraper
Scrapes NCBI Gene records by Entrez query and returns each gene as a flat row with identifier, symbol, name, location, aliases, OMIM ID, organism, and functional summary.
Pricing
from $7.50 / 1,000 results
Rating
0.0
(0)
Developer
Acquisition Automation Co.
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share

๐งฌ NCBI Gene Database Scraper
Run an Entrez query against the NCBI Gene database and get each gene back as a flat row: Gene ID, symbol, description, organism, chromosome, map location, aliases, protein synonyms, the RefSeq summary and the record URL. No API key, no registration, no login.
NCBI Gene is the reference record for a gene: the identifier everything else cites. The web interface answers one query at a time and its export options are built for downstream NCBI tools rather than for a spreadsheet. This Actor runs the same Entrez query you would type into the search box and writes one row per gene into a dataset you can open anywhere.
| Who uses it | What they use Gene records for |
|---|---|
| ๐ฌ Researchers and bioinformaticians | Resolving a list of symbols to stable Gene IDs before joining it to any other dataset |
| ๐ Curators and database builders | Pulling official symbol, description and aliases for a panel in one pass |
| ๐งพ Diligence analysts in life sciences | Checking that the genes named in a target's pipeline or patent claims resolve to real records |
| ๐ Science writers and educators | Getting the RefSeq summary and map location for a gene without opening ten tabs |
๐ What it does
๐ก Why it matters: gene symbols get renamed, reused and abbreviated differently by every group that touches them. The Gene ID does not. Every row here carries one.
- ๐ Takes a real Entrez query. Anything the NCBI search box accepts, for example
BRCA1[gene] AND human[orgn]. - ๐ Returns the Gene ID alongside the symbol, so a list of symbols becomes a list of stable identifiers.
- ๐งญ Chromosome and map location, for example chromosome
17at17q21.31. - ๐ค Aliases and protein synonyms, both as delimited strings, which is where an old symbol in a legacy spreadsheet turns up.
- ๐ The RefSeq summary, the curated paragraph describing what the gene does.
- ๐ A direct link to the NCBI record for every row.
- ๐พ Exports to CSV, Excel, JSON or XML, from the run page or the API.
๐ Output
Every gene is one flat row with 12 fields.
| Field | Type | Description |
|---|---|---|
๐ gene_id | string | NCBI Gene ID, the stable identifier for the record |
๐ค symbol | string | Official gene symbol, for example BRCA1 |
๐ description | string | Official full name of the gene |
๐ organism | string | Scientific name of the organism, for example Homo sapiens |
๐งฌ chromosome | string | Chromosome the gene sits on |
๐ maptype | string | Cytogenetic map location, for example 17q21.31 |
๐ summary | string | The curated RefSeq summary of the gene's function, as one paragraph |
๐ aliases | string | Alternative symbols, comma separated |
๐งช synonyms | string | Protein and long-form names, separated by a pipe character |
๐ ncbi_url | string | Link to the gene's page on NCBI |
๐ scrapedAt | string | ISO timestamp of collection |
โ ๏ธ error | string | null on a normal row |
Example row
This is the single row returned by a run with the default input, maxItems set to 10 and no query.
{"gene_id": "672","symbol": "BRCA1","description": "BRCA1 DNA repair associated","organism": "Homo sapiens","chromosome": "17","maptype": "17q21.31","summary": "This gene encodes a 190 kD nuclear phosphoprotein that plays a role in maintaining genomic stability, and it also acts as a tumor suppressor. The BRCA1 gene contains 22 exons spanning about 110 kb of DNA. The encoded protein combines with other tumor suppressors, DNA damage sensors, and signal transducers to form a large multi-subunit protein complex known as the BRCA1-associated genome surveillance complex (BASC). This gene product associates with RNA polymerase II, and through the C-terminal domain, also interacts with histone deacetylase complexes. This protein thus plays a role in transcription, DNA repair of double-stranded breaks, and recombination. Mutations in this gene are responsible for approximately 40% of inherited breast cancers and more than 80% of inherited breast and ovarian cancers. Alternative splicing plays a role in modulating the subcellular localization and physiological function of this gene. Many alternatively spliced transcript variants, some of which are disease-associated mutations, have been described for this gene, but the full-length natures of only some of these variants has been described. A related pseudogene, which is also located on chromosome 17, has been identified. [provided by RefSeq, May 2020]","aliases": "BRCAI, BRCC1, BROVCA1, FANCS, IRIS, PNCA4, PPP1R53, PSCP, RNF53","synonyms": "breast cancer type 1 susceptibility protein|BRCA1/BRCA2-containing complex, subunit 1|Fanconi anemia, complementation group S|RING finger protein 53|breast and ovarian cancer susceptibility protein 1|breast cancer 1, early onset|early onset breast cancer 1|protein phosphatase 1, regulatory subunit 53","ncbi_url": "https://www.ncbi.nlm.nih.gov/gene/672","scrapedAt": "2026-09-14T16:29:57.894Z","error": null}
Write a query to get more than one row. A broad Entrez query returns as many genes as maxItems allows.
โจ Why choose this Actor
| What you get | |
|---|---|
| The reference record | Data comes from NCBI Gene itself, not from a mirror or a summary site. |
| Entrez syntax, unchanged | The query you already know from the NCBI search box works here as written. |
| Identifiers, not just names | gene_id is what makes the export joinable to anything else. |
| The same 12 fields every run | Append several queries into one sheet without reconciling columns. |
| You pay per row | No subscription. A query that matches nothing costs nothing. |
๐ How to use it
- Create a free Apify account. New accounts start with $5 of credit.
- Open the Actor and select Try for free.
- Write an Entrez query in
query, for exampleBRCA1[gene] AND human[orgn]. - Set
maxItemsto cap the run. - Select Start, then export from the Dataset tab as CSV, Excel, JSON or XML.
One gene in one organism:
{"query": "BRCA1[gene] AND human[orgn]","maxItems": 10}
A broader pull:
{"query": "DNA repair[All Fields] AND human[orgn]","maxItems": 500}
โ๏ธ Input
| Field | Required | Description |
|---|---|---|
query | No | NCBI Gene Entrez query, for example BRCA1[gene] AND human[orgn] |
maxItems | No | How many genes to collect per run. Default 10 |
๐ฐ Pricing
Pay per result. No subscription, and no Apify platform usage on top.
| Apify plan | Free | Bronze | Silver | Gold | Platinum | Diamond |
|---|---|---|---|---|---|---|
| Per gene row | $0.0085 | $0.0082 | $0.0078 | $0.0075 | $0.0075 | $0.0075 |
| Rows collected | Cost on the Free plan |
|---|---|
| 100 | $0.85 |
| 1,000 | $8.50 |
| 10,000 | $85.00 |
Free plan runs return up to 10 rows as a preview. Any paid Apify plan lifts that to 1,000,000 per run.
๐ Integrate with any app
The dataset is available through the Apify API as soon as the run finishes. Use run-sync-get-dataset-items for a one-shot call, webhooks to trigger what happens next, or the Make, Zapier, Airbyte and LangChain integrations listed on the Actor page.
๐ค Use with an AI agent
Give an agent live access to NCBI Gene over the Model Context Protocol:
$claude mcp add --transport http apify "https://mcp.apify.com?tools=acquistion-automation/ncbi-eutils-gene-scraper"
Then ask it in plain language what a gene does and have it read the record back.
โ Frequently asked questions
Do I need an NCBI API key? No. The Actor reads the public Entrez service. You only need an Apify account to run it.
Which query syntax does query take?
Entrez syntax, the same as the NCBI Gene search box. Field tags such as [gene], [orgn] and [All Fields] work, joined with AND, OR and NOT.
Does this return OMIM IDs, sequences or transcript variants? No. The row carries the fields listed in the Output table: identifier, symbol, description, organism, chromosome, map location, summary, aliases, synonyms and the record URL. Sequence and variant data live in other NCBI databases this Actor does not read.
Why did I only get one row?
A narrow query matches one gene. Broaden the query or raise maxItems.
How do I read the synonyms field?
It is one string with entries separated by a pipe character. Split on | in your spreadsheet or script.
Can I query non-human organisms?
Yes. Add an organism tag, for example AND mouse[orgn], and organism comes back on every row so you can check what you got.
๐ More from Acquisition Automation Co.
- DrugBank Open Scraper
- IRS Exempt Organizations Scraper
- SAM.gov Contract Opportunities Scraper
- USASpending Contracts Scraper
- DANE Colombia Statistics Scraper
About Acquisition Automation Co.
We build automation for people buying businesses. The repetitive part of an acquisition search, checking listings, pulling public records, tracking owners and assets, is work a machine should do, so the buyer's time goes into judging deals instead of collecting them.
We add new Actors regularly. If there is a source you need and do not see here, tell us.
๐ Support
Open an issue in the Issues tab of this Actor with your run ID, the input you used, and what you expected to get back.
โ ๏ธ Disclaimer
This Actor is independent and is not affiliated with, endorsed by, or sponsored by the National Center for Biotechnology Information, the National Library of Medicine or any government agency. It collects only publicly available data, and it is not medical advice. You are responsible for using that data in compliance with the source's terms of service and applicable law.