# NCBI Gene Database Scraper (`acquistion-automation/ncbi-eutils-gene-scraper`) Actor

Scrapes NCBI Gene records by Entrez query and returns each gene as a flat row with identifier, symbol, name, location, aliases, OMIM ID, organism, and functional summary.

- **URL**: https://apify.com/acquistion-automation/ncbi-eutils-gene-scraper.md
- **Developed by:** [Acquisition Automation Co.](https://apify.com/acquistion-automation) (community)
- **Categories:** Education, Automation, Integrations
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $7.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

![Acquisition Automation Co. Search less. Close more.](https://api.apify.com/v2/key-value-stores/AOdPHdOpeDpzEPS5f/records/banner.jpg)

## 🧬 NCBI Gene Database Scraper

> **Run an Entrez query against the NCBI Gene database and get each gene back as a flat row: Gene ID, symbol, description, organism, chromosome, map location, aliases, protein synonyms, the RefSeq summary and the record URL.** No API key, no registration, no login.

NCBI Gene is the reference record for a gene: the identifier everything else cites. The web interface answers one query at a time and its export options are built for downstream NCBI tools rather than for a spreadsheet. This Actor runs the same Entrez query you would type into the search box and writes one row per gene into a dataset you can open anywhere.

| Who uses it | What they use Gene records for |
|---|---|
| 🔬 Researchers and bioinformaticians | Resolving a list of symbols to stable Gene IDs before joining it to any other dataset |
| 📚 Curators and database builders | Pulling official symbol, description and aliases for a panel in one pass |
| 🧾 Diligence analysts in life sciences | Checking that the genes named in a target's pipeline or patent claims resolve to real records |
| 🖊 Science writers and educators | Getting the RefSeq summary and map location for a gene without opening ten tabs |

### 📋 What it does

> 💡 **Why it matters:** gene symbols get renamed, reused and abbreviated differently by every group that touches them. The Gene ID does not. Every row here carries one.

- 🔎 **Takes a real Entrez query.** Anything the NCBI search box accepts, for example `BRCA1[gene] AND human[orgn]`.
- 🆔 **Returns the Gene ID** alongside the symbol, so a list of symbols becomes a list of stable identifiers.
- 🧭 **Chromosome and map location**, for example chromosome `17` at `17q21.31`.
- 🔤 **Aliases and protein synonyms**, both as delimited strings, which is where an old symbol in a legacy spreadsheet turns up.
- 📝 **The RefSeq summary**, the curated paragraph describing what the gene does.
- 🔗 **A direct link** to the NCBI record for every row.
- 💾 **Exports to CSV, Excel, JSON or XML**, from the run page or the API.

### 📊 Output

Every gene is one flat row with 12 fields.

| Field | Type | Description |
|---|---|---|
| 🆔 `gene_id` | string | NCBI Gene ID, the stable identifier for the record |
| 🔤 `symbol` | string | Official gene symbol, for example `BRCA1` |
| 📄 `description` | string | Official full name of the gene |
| 🐁 `organism` | string | Scientific name of the organism, for example `Homo sapiens` |
| 🧬 `chromosome` | string | Chromosome the gene sits on |
| 📍 `maptype` | string | Cytogenetic map location, for example `17q21.31` |
| 📝 `summary` | string | The curated RefSeq summary of the gene's function, as one paragraph |
| 🔁 `aliases` | string | Alternative symbols, comma separated |
| 🧪 `synonyms` | string | Protein and long-form names, separated by a pipe character |
| 🔗 `ncbi_url` | string | Link to the gene's page on NCBI |
| 🕒 `scrapedAt` | string | ISO timestamp of collection |
| ⚠️ `error` | string | `null` on a normal row |

#### Example row

This is the single row returned by a run with the default input, `maxItems` set to 10 and no query.

```json
{
  "gene_id": "672",
  "symbol": "BRCA1",
  "description": "BRCA1 DNA repair associated",
  "organism": "Homo sapiens",
  "chromosome": "17",
  "maptype": "17q21.31",
  "summary": "This gene encodes a 190 kD nuclear phosphoprotein that plays a role in maintaining genomic stability, and it also acts as a tumor suppressor. The BRCA1 gene contains 22 exons spanning about 110 kb of DNA. The encoded protein combines with other tumor suppressors, DNA damage sensors, and signal transducers to form a large multi-subunit protein complex known as the BRCA1-associated genome surveillance complex (BASC). This gene product associates with RNA polymerase II, and through the C-terminal domain, also interacts with histone deacetylase complexes. This protein thus plays a role in transcription, DNA repair of double-stranded breaks, and recombination. Mutations in this gene are responsible for approximately 40% of inherited breast cancers and more than 80% of inherited breast and ovarian cancers. Alternative splicing plays a role in modulating the subcellular localization and physiological function of this gene. Many alternatively spliced transcript variants, some of which are disease-associated mutations, have been described for this gene, but the full-length natures of only some of these variants has been described. A related pseudogene, which is also located on chromosome 17, has been identified. [provided by RefSeq, May 2020]",
  "aliases": "BRCAI, BRCC1, BROVCA1, FANCS, IRIS, PNCA4, PPP1R53, PSCP, RNF53",
  "synonyms": "breast cancer type 1 susceptibility protein|BRCA1/BRCA2-containing complex, subunit 1|Fanconi anemia, complementation group S|RING finger protein 53|breast and ovarian cancer susceptibility protein 1|breast cancer 1, early onset|early onset breast cancer 1|protein phosphatase 1, regulatory subunit 53",
  "ncbi_url": "https://www.ncbi.nlm.nih.gov/gene/672",
  "scrapedAt": "2026-09-14T16:29:57.894Z",
  "error": null
}
```

Write a query to get more than one row. A broad Entrez query returns as many genes as `maxItems` allows.

### ✨ Why choose this Actor

| | What you get |
|---|---|
| **The reference record** | Data comes from NCBI Gene itself, not from a mirror or a summary site. |
| **Entrez syntax, unchanged** | The query you already know from the NCBI search box works here as written. |
| **Identifiers, not just names** | `gene_id` is what makes the export joinable to anything else. |
| **The same 12 fields every run** | Append several queries into one sheet without reconciling columns. |
| **You pay per row** | No subscription. A query that matches nothing costs nothing. |

### 🚀 How to use it

1. [Create a free Apify account](https://console.apify.com/sign-up). New accounts start with $5 of credit.
2. Open the Actor and select **Try for free**.
3. Write an Entrez query in `query`, for example `BRCA1[gene] AND human[orgn]`.
4. Set `maxItems` to cap the run.
5. Select **Start**, then export from the **Dataset** tab as CSV, Excel, JSON or XML.

One gene in one organism:

```json
{
  "query": "BRCA1[gene] AND human[orgn]",
  "maxItems": 10
}
```

A broader pull:

```json
{
  "query": "DNA repair[All Fields] AND human[orgn]",
  "maxItems": 500
}
```

### ⚙️ Input

| Field | Required | Description |
|---|---|---|
| `query` | No | NCBI Gene Entrez query, for example `BRCA1[gene] AND human[orgn]` |
| `maxItems` | No | How many genes to collect per run. Default 10 |

### 💰 Pricing

Pay per result. No subscription, and no Apify platform usage on top.

| Apify plan | Free | Bronze | Silver | Gold | Platinum | Diamond |
|---|---|---|---|---|---|---|
| Per gene row | $0.0085 | $0.0082 | $0.0078 | $0.0075 | $0.0075 | $0.0075 |

| Rows collected | Cost on the Free plan |
|---|---|
| 100 | $0.85 |
| 1,000 | $8.50 |
| 10,000 | $85.00 |

**Free plan runs** return up to 10 rows as a preview. Any paid Apify plan lifts that to 1,000,000 per run.

### 🔌 Integrate with any app

The dataset is available through the Apify API as soon as the run finishes. Use `run-sync-get-dataset-items` for a one-shot call, webhooks to trigger what happens next, or the Make, Zapier, Airbyte and LangChain integrations listed on the Actor page.

### 🤖 Use with an AI agent

Give an agent live access to NCBI Gene over the Model Context Protocol:

```bash
claude mcp add --transport http apify "https://mcp.apify.com?tools=acquistion-automation/ncbi-eutils-gene-scraper"
```

Then ask it in plain language what a gene does and have it read the record back.

### ❓ Frequently asked questions

**Do I need an NCBI API key?**
No. The Actor reads the public Entrez service. You only need an Apify account to run it.

**Which query syntax does `query` take?**
Entrez syntax, the same as the NCBI Gene search box. Field tags such as `[gene]`, `[orgn]` and `[All Fields]` work, joined with `AND`, `OR` and `NOT`.

**Does this return OMIM IDs, sequences or transcript variants?**
No. The row carries the fields listed in the Output table: identifier, symbol, description, organism, chromosome, map location, summary, aliases, synonyms and the record URL. Sequence and variant data live in other NCBI databases this Actor does not read.

**Why did I only get one row?**
A narrow query matches one gene. Broaden the query or raise `maxItems`.

**How do I read the `synonyms` field?**
It is one string with entries separated by a pipe character. Split on `|` in your spreadsheet or script.

**Can I query non-human organisms?**
Yes. Add an organism tag, for example `AND mouse[orgn]`, and `organism` comes back on every row so you can check what you got.

### 🔗 More from Acquisition Automation Co.

- [DrugBank Open Scraper](https://apify.com/acquistion-automation/drugbank-open-scraper)
- [IRS Exempt Organizations Scraper](https://apify.com/acquistion-automation/irs-eo-master-file-scraper)
- [SAM.gov Contract Opportunities Scraper](https://apify.com/acquistion-automation/sam-gov-contracts-scraper)
- [USASpending Contracts Scraper](https://apify.com/acquistion-automation/usaspending-contracts-scraper)
- [DANE Colombia Statistics Scraper](https://apify.com/acquistion-automation/dane-colombia-statistics-scraper)

### About Acquisition Automation Co.

We build automation for people buying businesses. The repetitive part of an acquisition search, checking listings, pulling public records, tracking owners and assets, is work a machine should do, so the buyer's time goes into judging deals instead of collecting them.

We add new Actors regularly. If there is a source you need and do not see here, tell us.

### 🆘 Support

Open an issue in the **Issues** tab of this Actor with your run ID, the input you used, and what you expected to get back.

### ⚠️ Disclaimer

This Actor is independent and is not affiliated with, endorsed by, or sponsored by the National Center for Biotechnology Information, the National Library of Medicine or any government agency. It collects only publicly available data, and it is not medical advice. You are responsible for using that data in compliance with the source's terms of service and applicable law.

# Actor input Schema

## `query` (type: `string`):

NCBI Gene Entrez query. Example: BRCA1\[gene] AND human\[orgn].

## `maxItems` (type: `integer`):

How many genes to collect per run.

## Actor input object example

```json
{
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("acquistion-automation/ncbi-eutils-gene-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "maxItems": 10 }

# Run the Actor and wait for it to finish
run = client.actor("acquistion-automation/ncbi-eutils-gene-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10
}' |
apify call acquistion-automation/ncbi-eutils-gene-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,acquistion-automation/ncbi-eutils-gene-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BmkDYR4eyrVeyn6G5/builds/iYUss9rxRzmNdENO0/openapi.json
