# Hugging Face Scraper: Models, Datasets & Spaces (`arman-bd/huggingface-models-scraper`) Actor

Scrape Hugging Face models, datasets and spaces: downloads, likes, tags, pipeline task, licence, library and last-modified. No key, no proxy, no browser.

- **URL**: https://apify.com/arman-bd/huggingface-models-scraper.md
- **Developed by:** [Arman Hossain](https://apify.com/arman-bd) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.37 / 1,000 repository scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hugging Face Scraper: Models, Datasets & Spaces

![Hugging Face Scraper: Models, datasets and spaces ranked by downloads, likes or trending score, with licence and gating status](https://api.apify.com/v2/key-value-stores/ZQOcNAOHrIgTacAmy/records/huggingface-models-scraper.jpg)

Lists models, datasets and spaces from the Hugging Face Hub. You get downloads, likes, trending score, tags, pipeline task, library, licence, gating status and both repository timestamps.

The Hub publishes its whole index through a public read API at `huggingface.co/api`, intended for programmatic use. This Actor reads that index directly, so there's no token, no browser and no proxies. A 500-model sweep is five HTTP requests and finishes in a couple of seconds.

**Agent skill: [SKILL.md](https://api.apify.com/v2/key-value-stores/t7YoTxpZEJOWvw4Ug/records/huggingface-models-scraper.md)**

```
https://api.apify.com/v2/key-value-stores/t7YoTxpZEJOWvw4Ug/records/huggingface-models-scraper.md
```

### What you get

| Field | What it holds |
|---|---|
| `id`, `author` | Repository ID (`meta-llama/Llama-3.2-1B-Instruct`) and the owning user or organisation |
| `resourceType` | `models`, `datasets` or `spaces` |
| `url` | Direct link to the repository page on the Hub |
| `downloads`, `downloadsAllTime` | Downloads in the last 30 days, and since the repo was created |
| `likes`, `trendingScore` | Hub likes and the current trending score |
| `pipelineTag`, `libraryName` | Task (`text-generation`) and library (`transformers`), models only |
| `sdk` | Space runtime (`gradio`, `streamlit`, `docker`), spaces only |
| `license` | Licence identifier, extracted from the repo's `license:` tag |
| `tags` | Full tag list: languages, datasets, arXiv IDs, base models, regions |
| `gated`, `gatedType` | Whether access is restricted, and whether approval is `auto` or `manual` |
| `private`, `createdAt`, `lastModified` | Visibility flag and both repository timestamps |
| `scrapedAt` | When the run happened |

`RUN_SUMMARY` in the key-value store holds per-run counts, the filters you used, and any query that failed.

### Use cases

- **Model adoption tracking.** Snapshot downloads and likes on a schedule, then diff over time.
- **Dataset discovery.** Filter datasets by search term and read the size and modality tags.
- **Release watching.** Search by organisation name and sort by `createdAt` to see what a lab just shipped.
- **Licence and SBOM audits.** The `license` and `gated` columns tell you what you may actually ship.
- **Trend reporting.** Sort by `trendingScore` for what the community is picking up this week.

### Quick start

Top downloaded models for a term:

```json
{
 "searchQueries": ["llama"],
 "maxResults": 100
}
```

Speech models only, most-liked first:

```json
{
 "searchQueries": ["whisper", "wav2vec"],
 "resourceType": "models",
 "pipelineTag": "automatic-speech-recognition",
 "sortBy": "likes",
 "maxResults": 50
}
```

The 200 most recently updated datasets on the Hub, no search filter:

```json
{
 "resourceType": "datasets",
 "sortBy": "lastModified",
 "maxResults": 200
}
```

### Input

| Field | Type | Default | Notes |
|---|---|---|---|
| `searchQueries` | array | `[]` | Free-text terms matched against repo names and metadata. Each term is a separate pass. Empty means walk the hub in sort order. |
| `resourceType` | string | `models` | `models`, `datasets` or `spaces`. |
| `pipelineTag` | string | `""` | Task filter, e.g. `text-generation`. Models only: it is ignored with a warning for datasets and spaces. |
| `sortBy` | string | `downloads` | `downloads`, `likes`, `lastModified`, `createdAt` or `trendingScore`. Always descending. |
| `maxResults` | integer | `100` | Cap per search query, or on the single listing pass when no query is given. |

`searchQueries` and `pipelineTag` combine with AND: a model has to match the term and carry the task. `sortBy: "downloads"` with `resourceType: "spaces"` falls back to likes, because spaces have no download counter.

### Output example

```json
{
 "id": "meta-llama/Llama-3.2-1B-Instruct",
 "resourceType": "models",
 "author": "meta-llama",
 "url": "https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct",
 "downloads": 10016506,
 "downloadsAllTime": 93318314,
 "likes": 1554,
 "trendingScore": 2,
 "pipelineTag": "text-generation",
 "libraryName": "transformers",
 "sdk": null,
 "license": "llama3.2",
 "tags": ["transformers", "safetensors", "llama", "text-generation", "conversational", "license:llama3.2"],
 "gated": true,
 "gatedType": "manual",
 "private": false,
 "createdAt": "2024-09-18T15:12:47.000Z",
 "lastModified": "2024-10-24T15:07:51.000Z",
 "scrapedAt": "2026-08-06T11:35:15.076Z"
}
```

### Finding a pipeline task

Open any model page on the Hub and look at the badge under the title. That string is the pipeline tag verbatim. The common ones:

| Task | `pipelineTag` |
|---|---|
| Chat and completion models | `text-generation` |
| Embeddings | `sentence-similarity` or `feature-extraction` |
| Speech to text | `automatic-speech-recognition` |
| Image generation | `text-to-image` |
| Classification | `text-classification`, `image-classification` |

The Hub's own model list at `huggingface.co/models` shows the full task tree in its left sidebar.

### API example

```bash
curl -X POST "https://api.apify.com/v2/acts/arman-bd~huggingface-models-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
 -H "Content-Type: application/json" \
 -d '{
 "searchQueries": ["llama"],
 "pipelineTag": "text-generation",
 "sortBy": "downloads",
 "maxResults": 50
 }'
```

### JavaScript example

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('arman-bd/huggingface-models-scraper').call({
 searchQueries: ['whisper'],
 resourceType: 'models',
 sortBy: 'likes',
 maxResults: 25,
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const repo of items) console.log(`${repo.id} - ${repo.downloads} downloads, ${repo.license ?? 'no licence'}`);
```

### Notes

- Gated repos are kept, not dropped. A gated model still returns full public metadata, flagged with `gated: true` and `gatedType: "manual"` so your counts stay honest. Only the weights sit behind the approval form, not the record.
- Pagination is cursor-based. The Hub returns a `Link` header with `rel="next"` and no page numbers, so the Actor follows that cursor 100 records at a time until your cap is reached.
- Licence comes from the tags. There is no licence field on the wire. It is encoded as `license:apache-2.0` in the tag list and lifted out into its own column here.
- One bad query won't kill the run. Failures are recorded in `RUN_SUMMARY.failures`, and the Actor only errors out if every query fails.
- Rate limits are generous. The Hub advertises 500 requests per 5-minute window on an anonymous IP, and a 10,000-record sweep uses 100 of them.

### FAQ

**Do I need a Hugging Face token?** No. The listing API is public. A token only matters for downloading gated weights, which this Actor never does.

**Do I need a proxy?** No. Proxy configuration is not required to run this Actor.

**Can I get the model card text?** Not from this Actor. The listing index carries metadata only, and the card is a separate file on the repo. Everything you get here is one request per 100 records.

**Why is `downloads` null on spaces?** Spaces are apps, not artefacts. The Hub tracks likes for them but not downloads, so `sdk` is the field that carries useful signal there instead.

**How far back can I list?** As far as you like. Sort by `createdAt` and raise `maxResults`, and the cursor keeps walking until the index is exhausted.

**Can I plug it into something else?** Yes. Apify API, the client libraries, webhooks, scheduled runs, dataset exports to JSON, CSV or Excel, or MCP. The output is structured JSON.

# Actor input Schema

## `searchQueries` (type: `array`):

Free-text terms matched against repository names and metadata. Each term is a separate pass, so 'llama' and 'whisper' return both sets. Leave empty to list the hub in sort order without filtering.

## `resourceType` (type: `string`):

Which hub index to read. Models carry a pipeline task and a library, datasets carry neither, and spaces carry an SDK but no download counter.

## `pipelineTag` (type: `string`):

Keep only models for one task, e.g. text-generation, automatic-speech-recognition, image-classification. Models only: it is ignored for datasets and spaces. Leave empty for every task.

## `sortBy` (type: `string`):

Ordering of the hub listing, always descending. Downloads counts the last 30 days. Spaces have no downloads, so that choice falls back to likes.

## `maxResults` (type: `integer`):

Cap the records saved for each search query. With no search queries it caps the single listing pass. Results arrive 100 per request, so 500 costs five requests.

## Actor input object example

```json
{
  "searchQueries": [
    "llama",
    "whisper"
  ],
  "resourceType": "models",
  "pipelineTag": "text-generation",
  "sortBy": "downloads",
  "maxResults": 100
}
```

# Actor output Schema

## `items` (type: `string`):

Every record the run produced.

## `runsummary` (type: `string`):

The RUN_SUMMARY record from the run's key-value store.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "llama"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("arman-bd/huggingface-models-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchQueries": ["llama"] }

# Run the Actor and wait for it to finish
run = client.actor("arman-bd/huggingface-models-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "llama"
  ]
}' |
apify call arman-bd/huggingface-models-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,arman-bd/huggingface-models-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Myd5JT3SIpYLCCKd8/builds/LB7TZDEQOrvlQz5X1/openapi.json
