# GitHub Repository Search, Stats and Releases (`pistachio_implementation/github-repo-search-stats`) Actor

Search GitHub repositories by keyword, language, stars and dates, or look up a list of repositories. Returns stars, forks, topics, license, language, activity dates, latest release, release notes and optional README text, from GitHub's official API. Works without a key.

- **URL**: https://apify.com/pistachio\_implementation/github-repo-search-stats.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## GitHub Repository Search, Stats and Releases

Search GitHub repositories the way you would on github.com, or give a list of repositories, and get clean rows with the numbers that matter: stars, forks, open issues, language, topics, license, homepage, creation and last push dates. Optionally add the latest release, recent release notes with download counts, and the README text in Markdown, ready for an LLM.

The actor uses GitHub's official REST API. No page scraping, no browser, no proxies. It works without any key; add a free GitHub token only if you need large runs.

### What you can use it for

- **AI builders and agents:** find the best libraries for a task ("vector database", "rag framework" with more than 1,000 stars) and read their READMEs in the same call.
- **Developer tool market research:** track the stars, forks and release pace of competitors and adjacent projects.
- **New and rising projects:** search repositories created after a date, sorted by stars.
- **Release monitoring:** get the latest releases of your dependencies, with notes and download counts, on a schedule.
- **Open source due diligence:** license, archive status, last push and open issues for a list of repositories.

### Input

| Field | What it does | Default |
|---|---|---|
| Search queries | Keywords plus any GitHub qualifier (`topic:rag`, `org:vercel`, `license:mit`, `in:readme`) | none |
| Repositories to look up | `owner/name` or a github.com link | none |
| Sort search results by | Best match, most stars, most forks, recently updated | Best match |
| Language | Main language, for example Python | none |
| Minimum stars | Skip smaller repositories | 0 |
| Created on or after, Last code push on or after | Dates such as 2026-01-01 | none |
| Include forks, Include archived | Search filters | off, on |
| Maximum repositories per query | Up to 1,000 (GitHub's own cap for one search) | 100 |
| Add latest release | Latest release tag, name, date and link | off |
| Release rows per repository | Recent releases as separate rows with notes and downloads | 0 |
| Add README text | README in Markdown | off |
| Maximum README length | Characters kept per README | 20,000 |
| GitHub token | Optional, for large runs (see Limits) | none |
| Maximum rows in total | Repositories plus releases | 5,000 |

Example input:

```json
{
  "searchQueries": ["rag framework"],
  "repositories": ["apify/crawlee", "https://github.com/vercel/next.js"],
  "sort": "stars",
  "language": "Python",
  "minStars": 1000,
  "maxReposPerQuery": 50,
  "includeLatestRelease": true,
  "includeReadme": true
}
```

### Output

Repository row:

```json
{
  "type": "repository",
  "source": "search",
  "query": "rag framework",
  "rank": 1,
  "fullName": "langchain-ai/langchain",
  "owner": "langchain-ai",
  "ownerType": "Organization",
  "url": "https://github.com/langchain-ai/langchain",
  "description": "The agent engineering platform.",
  "homepage": "https://docs.langchain.com/langchain/",
  "language": "Python",
  "topics": ["agents", "ai", "llm", "rag"],
  "stars": 147136,
  "forks": 24626,
  "openIssues": 568,
  "license": "MIT",
  "isFork": false,
  "isArchived": false,
  "defaultBranch": "master",
  "createdAt": "2022-10-17T02:58:36Z",
  "pushedAt": "2026-09-26T08:55:23Z",
  "latestRelease": { "tag": "langchain-fireworks==1.6.3", "publishedAt": "2026-09-25T12:57:38Z", "url": "https://github.com/langchain-ai/langchain/releases/tag/langchain-fireworks%3D%3D1.6.3" },
  "readme": "# LangChain ...",
  "scrapedAt": "2026-09-27T07:24:58.271Z"
}
```

Release row (when "Release rows per repository" is above 0):

```json
{
  "type": "release",
  "fullName": "langchain-ai/langchain",
  "tag": "langchain-fireworks==1.6.3",
  "url": "https://github.com/langchain-ai/langchain/releases/tag/langchain-fireworks%3D%3D1.6.3",
  "isPrerelease": false,
  "publishedAt": "2026-09-25T12:57:38Z",
  "body": "Changes since langchain-fireworks==1.6.2 ...",
  "assetCount": 2,
  "downloadCount": 4
}
```

`openIssues` is GitHub's own count and includes open pull requests. `watchers` is filled only for repositories looked up by name. Anything that returns nothing (a misspelled repository, a search with no matches) is listed in `RUN_SUMMARY` in the run's key value store and costs nothing.

### Pricing

Pay per event. No start fee, no subscription, no platform usage charged on top.

| Event | Price |
|---|---|
| Repository saved | $0.001 (one dollar per 1,000 repositories) |
| Release saved | $0.0003 (30 cents per 1,000 releases) |

The latest release and the README come with the repository row at no extra charge. Example: the top 100 Python RAG repositories with READMEs cost $0.10.

### Limits

- **GitHub's own limits apply.** Without a token GitHub allows 10 searches a minute and 60 other calls an hour from one address. Search results need no other calls, so plain searches of up to 1,000 repositories work without a token. The latest release and release rows use one extra call per repository, so without a token they cover up to about 50 repositories an hour; once GitHub's hourly allowance is used, the remaining repositories are still saved without release data, and `RUN_SUMMARY` says so (`coreLimitHit`). Release rows that were not fetched are not charged. A free personal access token with no scopes raises the limits to 30 searches a minute and 5,000 calls an hour.
- One search returns at most 1,000 repositories (GitHub's cap). Split big searches by language, date or star range.
- Public repositories only. No user profiles, emails or contributor lists are collected.
- READMEs are read from GitHub's raw file host and cut to your maximum length.

### FAQ

**Do I need a GitHub token?** No for searches and small lookups. Yes if you want releases for hundreds of repositories in one run. Create one at github.com under Settings, Developer settings, Personal access tokens, with no scopes. It is stored as a secret input and never written to the output.

**Can I use GitHub search syntax?** Yes. Anything that works in the github.com search box for repositories works in a query, for example `topic:mcp stars:>500 pushed:>2026-06-01`.

**How do I find trending new projects?** Set "Created on or after" to a recent date and sort by stars.

**Can an AI agent call it?** Yes. Narrow inputs, predictable rows and pay per event pricing with no start fee.

# Actor input Schema

## `searchQueries` (type: `array`):

One per line. Keywords (llm agent framework) and any GitHub search qualifier (topic:rag, org:vercel, license:mit, in:readme) work.

## `repositories` (type: `array`):

One per line: owner/name (apify/crawlee) or a github.com link.

## `sort` (type: `string`):

Order of search results. Best match is GitHub's own relevance order.

## `language` (type: `string`):

Optional. Only repositories whose main language is this, for example Python or TypeScript.

## `minStars` (type: `integer`):

Optional. Skip repositories with fewer stars.

## `createdAfter` (type: `string`):

Optional date such as 2026-01-01. With sort by stars this finds new and rising repositories.

## `pushedAfter` (type: `string`):

Optional date such as 2026-06-01. Keeps actively maintained repositories only.

## `includeForks` (type: `boolean`):

Forks are left out of search results unless this is on.

## `includeArchived` (type: `boolean`):

Archived (read only) repositories are included unless this is off.

## `maxReposPerQuery` (type: `integer`):

GitHub returns at most 1,000 results for one search.

## `includeLatestRelease` (type: `boolean`):

Adds the latest release tag, name, date and link to each repository. Uses one extra GitHub call per repository.

## `releasesPerRepo` (type: `integer`):

Also save up to this many recent releases per repository as separate rows, with release notes and download counts. 0 for none.

## `includeReadme` (type: `boolean`):

Adds the README in its original Markdown, ready for an LLM or a RAG index.

## `maxReadmeChars` (type: `integer`):

Cut each README to this many characters. 0 means no limit.

## `githubToken` (type: `string`):

Optional. Without a token GitHub allows 10 searches a minute and 60 other calls an hour from one address. A free personal access token with no scopes raises that to 30 searches a minute and 5,000 calls an hour. Stored encrypted and never written to the output.

## `maxItems` (type: `integer`):

Stop after saving this many rows (repositories plus releases).

## Actor input object example

```json
{
  "searchQueries": [
    "llm agent framework"
  ],
  "repositories": [
    "apify/crawlee"
  ],
  "sort": "best-match",
  "minStars": 0,
  "includeForks": false,
  "includeArchived": true,
  "maxReposPerQuery": 100,
  "includeLatestRelease": false,
  "releasesPerRepo": 0,
  "includeReadme": false,
  "maxReadmeChars": 20000,
  "maxItems": 5000
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

## `summary` (type: `string`):

The RUN\_SUMMARY record: counts and problems for the whole run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "llm agent framework"
    ],
    "repositories": [
        "apify/crawlee"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/github-repo-search-stats").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": ["llm agent framework"],
    "repositories": ["apify/crawlee"],
}

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/github-repo-search-stats").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "llm agent framework"
  ],
  "repositories": [
    "apify/crawlee"
  ]
}' |
apify call pistachio_implementation/github-repo-search-stats --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/github-repo-search-stats"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/OOQhxY3dllSulNBzq/builds/mItNd2LGFFHtfOnIL/openapi.json
