# Scholar Author Metrics — Citations, h-index, Papers (`s-r/google-scholar-author-metrics`) Actor

Look up any researcher and get their citation metrics: total citations, h-index, i10-index, per-year citation history, most-cited publications, and top co-authors. By name, ORCID, or author id. Powered by OpenAlex — no rate limits.

- **URL**: https://apify.com/s-r/google-scholar-author-metrics.md
- **Developed by:** [SR](https://apify.com/s-r) (community)
- **Categories:** Developer tools, Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 author profiles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google Scholar Author Metrics

Turn any Google Scholar profile into clean, structured data. Give this Actor an
author's profile ID (or profile URL) and it returns their citation record
straight from **Google Scholar**: total citations, recent citations, **h-index**,
**i10-index**, the year-by-year citation histogram, the complete publication
list, and the co-authors shown on the profile. One tidy row per author, ready
for a spreadsheet, a dashboard, or a database.

If you have ever copied these numbers by hand for a grant report, a hiring
committee, a literature review, or a league table, this Actor does it for you in
seconds and keeps doing it on a schedule.

### What you get

For every author you submit, one dataset record with:

- **Profile** — name, affiliation, verified email domain, homepage link,
  research interests, and profile photo.
- **Headline metrics** — citations (all-time and since the recent cut-off year
  Scholar shows), h-index (all-time and recent), and i10-index (all-time and
  recent).
- **Citation histogram** — citations per year, exactly as the bar chart on the
  profile shows them, as a simple `[{ "year": 2019, "count": 700 }, ...]` list
  you can plot directly.
- **Publications** — every paper on the profile: title, link, authors, venue,
  year, and its cited-by count. Paginated automatically, so you get the whole
  list, not just the first page.
- **Co-authors** — the co-authors listed on the profile, each with their name,
  affiliation, and their own profile link, so you can walk an entire research
  group.

### Example output

```json
{
  "author_id": "JicYPdAAAAAJ",
  "name": "Geoffrey Hinton",
  "affiliation": "Emeritus Professor of Computer Science, University of Toronto",
  "verified_email_domain": "cs.toronto.edu",
  "interests": ["machine learning", "neural networks", "artificial intelligence"],
  "citations_all": 900000,
  "citations_since": 500000,
  "h_index_all": 186,
  "h_index_since": 150,
  "i10_index_all": 483,
  "i10_index_since": 420,
  "since_year": 2019,
  "citations_per_year": [
    { "year": 2019, "count": 70000 }
  ],
  "publication_count": 100,
  "publications": [
    {
      "title": "Deep learning",
      "url": "https://scholar.google.com/citations?view_op=view_citation&user=...",
      "authors": "Y LeCun, Y Bengio, G Hinton",
      "venue": "nature 521 (7553), 436-444, 2015",
      "year": 2015,
      "cited_by_count": 75000
    }
  ],
  "coauthors": [
    {
      "name": "Yann LeCun",
      "author_id": "WLN3QrAAAAAJ",
      "affiliation": "Chief AI Scientist, Facebook & Courant Institute, NYU",
      "profile_url": "https://scholar.google.com/citations?user=WLN3QrAAAAAJ"
    }
  ],
  "profile_url": "https://scholar.google.com/citations?user=JicYPdAAAAAJ&hl=en"
}
```

### Input

| Field | Type | Description |
|---|---|---|
| `author_ids` | array of strings | **Required.** Google Scholar author IDs — the `user=` value in a profile URL (for example `JicYPdAAAAAJ`). Add as many as you like; you get one row per author. |
| `profile_urls` | array of strings | Optional. Full `scholar.google.com/citations?user=...` URLs, used instead of or alongside `author_ids`. The ID is pulled out for you. |
| `include_publications` | boolean | Include the publication list. Default `true`. Turn off for a faster, metrics-only run. |
| `max_publications` | integer | Cap the publication list per author. `0` means all of them (paginated). Default `100`. |
| `include_coauthors` | boolean | Include the co-authors listed on the profile. Default `true`. |
| `hl` | string | Interface language (`en`, `de`, `fr`, `es`, …). Affects labels, not the numbers. Default `en`. |

#### Where do I find an author's ID?

Open the author's Google Scholar profile and look at the address bar. The ID is
the value after `user=`:

```
https://scholar.google.com/citations?user=JicYPdAAAAAJ&hl=en
                                          ^^^^^^^^^^^^
```

You can paste the whole URL into `profile_urls` and skip the copying.

### Common uses

- **Research office reporting** — pull citations, h-index and i10-index for a
  whole department on a monthly schedule and drop them into a report.
- **Hiring and tenure review** — collect a candidate's metrics and full
  publication record without visiting a dozen profiles by hand.
- **Collaborator mapping** — start from one author, read the co-authors, then
  feed those IDs back in to map an entire group.
- **Bibliometrics and rankings** — build league tables or track how a metric
  moves over time by re-running on a schedule and storing each snapshot.
- **Literature reviews** — export an author's complete publication list with
  citation counts as a starting bibliography.

### How it is billed

Pay per result: you are charged once for each author profile successfully
returned. A run that finds nothing (a bad ID, or a temporary busy period)
returns no billable results. There is a small fixed charge when a run starts.
See the Pricing tab for the exact amounts.

### Notes and limits

- **One row per author.** Submitting the same ID twice still returns it once.
- **Publications are paginated for you.** Large profiles are walked page by page
  up to your `max_publications` cap. A cap keeps big runs fast; set it to `0`
  when you truly need every paper.
- **Numbers match the profile**, including Scholar's "since" cut-off year, which
  Scholar moves forward over time; the `since_year` field tells you which year
  the recent columns refer to.
- **Public profiles only.** Only authors who have made their Scholar profile
  public are available; a private or non-existent profile is reported in the
  run's error record rather than silently dropped.
- **Busy periods.** Google Scholar sometimes throttles heavy traffic; when that
  happens the affected author is reported with a clear message and the rest of
  the run still completes. Keep batches sensible and schedule large jobs.

### Tips

- Keep each run to a sensible batch of authors and schedule it, rather than
  submitting thousands at once.
- If you only need the headline numbers, set `include_publications` to `false`
  for a noticeably faster run.
- Store the whole output over time: with the citation histogram and the "since"
  year on every run, you can chart how an author's impact grows.

# Actor input Schema

## `author_ids` (type: `array`):

Google Scholar author IDs — the `user=` value in a profile URL (e.g. JicYPdAAAAAJ). One row per author is returned with full metrics.

## `profile_urls` (type: `array`):

Full Scholar citations URLs (https://scholar.google.com/citations?user=...). Accepted instead of, or alongside, author\_ids — the user id is extracted.

## `include_publications` (type: `boolean`):

Include the author's publication list (title, authors, venue, year, cited-by count). Turn off for a metrics-only, faster run.

## `max_publications` (type: `integer`):

Cap the publication list per author (0 = all, paginated). Each extra page of 100 is another request to Scholar.

## `include_coauthors` (type: `boolean`):

Include the co-authors shown on the profile (name, affiliation, profile link).

## `hl` (type: `string`):

Interface language code (en, de, fr, es…). Affects labels, not the metrics.

## `use_proxy` (type: `boolean`):

Default ON. Routes requests through the rotating proxy with automatic IP-rotation retry, which Scholar's per-address limit needs for volume. Disable to fetch directly from the Apify datacenter IP.

## Actor input object example

```json
{
  "author_ids": [
    "JicYPdAAAAAJ",
    "kukA0LcAAAAJ"
  ],
  "profile_urls": [
    "https://scholar.google.com/citations?user=JicYPdAAAAAJ&hl=en"
  ],
  "include_publications": true,
  "max_publications": 100,
  "include_coauthors": true,
  "hl": "en",
  "use_proxy": true
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset rows produced by this run, one per author.

## `output` (type: `string`):

OUTPUT record with the run's counts and status flags.

## `errors` (type: `string`):

Failures with a code and message. Absent when the run had none.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "author_ids": [
        "JicYPdAAAAAJ"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("s-r/google-scholar-author-metrics").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "author_ids": ["JicYPdAAAAAJ"] }

# Run the Actor and wait for it to finish
run = client.actor("s-r/google-scholar-author-metrics").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "author_ids": [
    "JicYPdAAAAAJ"
  ]
}' |
apify call s-r/google-scholar-author-metrics --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,s-r/google-scholar-author-metrics"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Wq8iKXuamUgQoBG9u/builds/XOcQSFnEiUivwnYBl/openapi.json
