# Hugging Face Scraper (`dtrungtin/huggingface-scraper`) Actor

Scrape Hugging Face models, datasets, Spaces & papers: downloads, likes, parameters, license, tasks, inference providers and model cards. Search or paste any URL.

- **URL**: https://apify.com/dtrungtin/huggingface-scraper.md
- **Developed by:** [Tin](https://apify.com/dtrungtin) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Hugging Face Scraper do?

**Hugging Face Scraper** extracts structured data about **models, datasets, Spaces (AI apps) and research papers** from the [Hugging Face Hub](https://huggingface.co), the largest open-source AI platform with millions of repositories.

Search with filters (task, library, license, language, model size, author...) or **paste any huggingface.co URL**, such as a filtered listing page, a model page, an organization profile, a collection or the Daily Papers page. You get clean JSON with **downloads, likes, trending score, parameter count, license, base model, inference providers and their prices, README model cards** and more.

It uses Hugging Face's official public API instead of a browser, so it is **fast and cheap**: thousands of results in seconds. Running on the Apify platform also gives you API access, scheduling, integrations (Google Sheets, Slack, Zapier, Make, webhooks), monitoring and storage out of the box.

### Why use Hugging Face Scraper?

- **AI market research and competitive intelligence**: track which models, organizations and tasks are gaining traction, and compare download and like counts over time.
- **Model selection**: shortlist open models by task, size, license and available inference providers, including their price per million tokens, context length and speed.
- **Trend monitoring**: schedule a daily run with *Created after = 1 day* to get every new model, dataset or paper in your niche, straight into Slack or a spreadsheet.
- **Lead generation and ecosystem mapping**: list every model, dataset and Space published by a company or research lab.
- **Research and datasets for ML**: build catalogs of datasets by task, modality, language and size, or collect README model cards for LLM and RAG pipelines.
- **Academic tracking**: collect Daily Papers with upvotes, AI summaries, keywords, GitHub repos and stars.

### How to scrape Hugging Face data

1. Click **Try for free** (or **Start**) to open the Actor.
2. Choose **What to scrape**: Models, Datasets, Spaces or Papers.
3. Optionally enter **search terms** and **filters**, such as task *Text Generation*, library *gguf* and license *apache-2.0*.
4. Or paste one or more **Hugging Face URLs**, e.g. `https://huggingface.co/models?pipeline_tag=text-generation&sort=trending`.
5. Set **Max results per search or URL** and click **Start**.
6. When the run finishes, download your data as JSON, CSV, Excel or HTML, or use the API.

### Input

All options are on the **Input** tab. The most important ones:

| Field                                      | Description                                                                                |
| ------------------------------------------ | ------------------------------------------------------------------------------------------ |
| What to scrape                             | `models`, `datasets`, `spaces` or `papers`                                                 |
| Search terms                               | One or more keywords; each one is scraped separately                                       |
| Sort by                                    | Trending, most downloads, most likes, recently created or recently updated                 |
| Max results per search or URL              | Limit per keyword or URL (0 = no limit)                                                    |
| Hugging Face URLs                          | Any huggingface.co listing, repo, profile, collection or papers URL                        |
| Filters                                    | Author, task, library, language, license, tags, min/max parameters, Space SDK, papers date |
| Min downloads / Min likes / Created after  | Keep only popular or recent results; *Created after* also accepts `7 days`                 |
| Include README / file list / card metadata | Add model card text, repository files or raw card YAML                                     |
| Hugging Face access token                  | Optional, for higher rate limits and your own private repos                                |

Example input for the 100 most downloaded small GGUF text-generation models:

```json
{
    "resourceType": "models",
    "task": "text-generation",
    "library": "gguf",
    "maxParameters": "8B",
    "sort": "downloads",
    "maxItems": 100
}
```

Example input using URLs:

```json
{
    "startUrls": [
        { "url": "https://huggingface.co/models?pipeline_tag=text-to-image&sort=likes" },
        { "url": "https://huggingface.co/meta-llama" },
        { "url": "https://huggingface.co/papers/trending" }
    ],
    "maxItems": 50
}
```

### Output

Each result is one item in the dataset. The Output tab has a table view for each type (Models, Datasets, Spaces, Papers). You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

Example model (shortened):

```json
{
    "type": "model",
    "id": "Qwen/Qwen3-8B",
    "author": "Qwen",
    "url": "https://huggingface.co/Qwen/Qwen3-8B",
    "task": "text-generation",
    "library": "transformers",
    "downloads": 12522822,
    "downloadsAllTime": 125437423,
    "likes": 2039,
    "trendingScore": 20,
    "parameters": 8190735360,
    "parameterSize": "8.19B",
    "license": "apache-2.0",
    "baseModels": ["Qwen/Qwen3-8B-Base"],
    "baseModelRelation": "finetune",
    "arxivIds": ["2309.00071", "2505.09388"],
    "gated": false,
    "inferenceProviders": [
        {
            "provider": "nscale",
            "status": "live",
            "contextLength": 40960,
            "inputPricePer1M": 0.07,
            "outputPricePer1M": 0.18,
            "tokensPerSecond": 126.5
        }
    ],
    "createdAt": "2025-04-27T03:42:21.000Z",
    "lastModified": "2025-07-26T03:49:13.000Z"
}
```

Example paper (shortened):

```json
{
    "type": "paper",
    "id": "2505.09388",
    "title": "Qwen3 Technical Report",
    "url": "https://huggingface.co/papers/2505.09388",
    "pdfUrl": "https://arxiv.org/pdf/2505.09388",
    "aiSummary": "Qwen3, a unified series of large language models, integrates thinking and non-thinking modes...",
    "authors": ["An Yang", "Anfeng Li", "Baosong Yang"],
    "upvotes": 347,
    "githubRepo": "https://github.com/QwenLM/Qwen3",
    "githubStars": 27658,
    "publishedAt": "2025-05-14T13:41:34.000Z"
}
```

### Data fields

| Type         | Main fields                                                                                                                                                                                                                                                                                                                                                                             |
| ------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Models**   | id, author, url, task, library, downloads (30 days), downloadsAllTime, likes, trendingScore, parameters, parameterSize, tensorTypes, ggufContextLength, license, languages, baseModels, baseModelRelation, trainingDatasets, arxivIds, gated, inferenceProviders (provider, status, context length, input/output price per 1M tokens, tokens per second), tags, createdAt, lastModified |
| **Datasets** | id, author, url, description, downloads, downloadsAllTime, likes, trendingScore, tasks, taskIds, modalities, formats, libraries, sizeCategory, sizeBytes, languages, license, arxivIds, citation, gated, tags, createdAt, lastModified                                                                                                                                                  |
| **Spaces**   | id, author, url, title, emoji, shortDescription, sdk, sdkVersion, appUrl, likes, trendingScore, status (running, paused...), hardware, models, datasets, license, tags, createdAt, lastModified                                                                                                                                                                                         |
| **Papers**   | id (arXiv), title, url, arxivUrl, pdfUrl, summary, aiSummary, aiKeywords, authors, publishedAt, upvotes, numComments, githubRepo, githubStars, projectPage, organization                                                                                                                                                                                                                |
| Optional     | readme (Markdown model/dataset card), files (repository file list), cardData (raw card metadata)                                                                                                                                                                                                                                                                                        |

### How much does it cost to scrape Hugging Face?

This Actor uses **pay-per-event** pricing: you pay a small fee when a run starts and a fixed price per result saved. Check the **Pricing** tab for current prices. At $2 per 1,000 results, the top 1,000 trending models cost about $2.

Because it calls the API directly instead of rendering web pages, runs are very light and finish quickly. You can set a **maximum cost per run** in the run options, and the Actor stops cleanly when it is reached.

### Tips and advanced options

- **Filtering is fastest when it matches the sort order.** With *Sort by = Most downloads* and *Min downloads = 10000*, the scraper stops as soon as results drop below 10,000. With another sort order, it has to scan further to find matching items.
- **Monitor new releases**: sort by *Recently created*, set *Created after* to `1 day` and schedule the Actor daily.
- **Use URLs for complex filters.** Set up filters on huggingface.co, then copy the URL from your browser into *Hugging Face URLs*.
- **README and dataset file lists** need one extra request per result, so runs with those options are slower. Leave them off when you only need metrics.
- **Rate limits**: Hugging Face limits anonymous API usage. The Actor waits and retries automatically, but a free [Hugging Face token](https://huggingface.co/settings/tokens) raises the limit for very large runs.
- **Duplicates are removed**: a repository found by several search terms is saved (and charged) only once per run.

### FAQ, disclaimers, and support

**Is it legal to scrape Hugging Face?** This Actor only collects publicly available metadata through Hugging Face's official public API, the same data shown on the website. It does not collect personal data beyond public usernames and author names. Always review the [Hugging Face Terms of Service](https://huggingface.co/terms-of-service) and the licenses of the models and datasets you use, and consult a lawyer if unsure.

**Why is the README empty for some models?** Gated models, such as Meta Llama, only share files with accounts that accepted their license. Add a token from an account that has access to include their READMEs.

**Why are inference prices missing for a single model URL?** Hugging Face includes provider prices only in list results. Scrape the model through a search or listing URL to get pricing.

**Does it download model weights or dataset files?** No. It collects metadata and, optionally, README text and file names.

**Found a bug or need a feature?** Open an issue on the **Issues** tab, and it will be looked at quickly. Need a custom solution, such as tracking metrics over time or scraping other parts of Hugging Face? Get in touch through the Issues tab.

# Actor input Schema

## `resourceType` (type: `string`):

Which part of the Hugging Face Hub to search. Ignored when you provide URLs below.

## `searchTerms` (type: `array`):

Full-text search, e.g. <code>llama</code>, <code>whisper</code> or <code>sentiment</code>. Each term is scraped separately. Leave empty to browse everything that matches the filters. For papers, searches titles and abstracts; leave empty for the Daily Papers feed.

## `sort` (type: `string`):

Order of results. Spaces cannot be sorted by downloads. Papers support Trending; any other value lists the newest papers.

## `maxItems` (type: `integer`):

Maximum number of results saved for each search term or URL. Set to 0 for no limit (a full crawl of all models is over 2 million results).

## `startUrls` (type: `array`):

Optional. Paste any huggingface.co URL and its data is scraped instead of the search above. Supported: listing pages with filters (e.g. <code>https://huggingface.co/models?pipeline\_tag=text-generation\&sort=trending</code>), model / dataset / Space pages, user or organization profiles, collections, and papers pages (daily, trending, a date, or a single paper).

## `author` (type: `string`):

Only results published by this user or organization, e.g. <code>meta-llama</code>, <code>google</code>, <code>Qwen</code>.

## `task` (type: `string`):

Models: pipeline task. Datasets: task category. Not available for Spaces.

## `library` (type: `string`):

Models: e.g. <code>transformers</code>, <code>gguf</code>, <code>diffusers</code>, <code>mlx</code>, <code>onnx</code>. Datasets: e.g. <code>datasets</code>, <code>pandas</code>, <code>polars</code>.

## `language` (type: `string`):

ISO language code, e.g. <code>en</code>, <code>vi</code>, <code>fr</code>, <code>zh</code>.

## `license` (type: `string`):

License id, e.g. <code>apache-2.0</code>, <code>mit</code>, <code>cc-by-4.0</code>, <code>llama3.1</code>.

## `tags` (type: `array`):

Any Hugging Face tags that results must have, e.g. <code>conversational</code>, <code>base\_model:finetune:Qwen/Qwen3-8B</code>, <code>size\_categories:1K\<n<10K</code>, <code>mcp-server</code>.

## `minParameters` (type: `string`):

Minimum model size, e.g. <code>500M</code>, <code>7B</code>.

## `maxParameters` (type: `string`):

Maximum model size, e.g. <code>3B</code>, <code>70B</code>.

## `spaceSdk` (type: `string`):

Only Spaces built with this SDK.

## `papersDate` (type: `string`):

Daily Papers for a day (<code>2026-09-22</code>), ISO week (<code>2026-W38</code>) or month (<code>2026-09</code>). Leave empty for the latest papers.

## `minDownloads` (type: `integer`):

Skip results with fewer downloads. Fastest when sorting by Most downloads.

## `minLikes` (type: `integer`):

Skip results with fewer likes (upvotes for papers). Fastest when sorting by Most likes.

## `createdAfter` (type: `string`):

Only results created (papers: published) after this date. Also accepts a relative period such as <code>7 days</code>, which is useful for scheduled monitoring runs. Fastest when sorting by Recently created.

## `includeReadme` (type: `boolean`):

Adds the full README (model card, dataset card or Space description) as Markdown. Great for LLM/RAG pipelines. Makes runs slower (one extra request per result).

## `includeFiles` (type: `boolean`):

Adds the list of files in each repository (weights, configs, notebooks...). For datasets this needs one extra request per result.

## `includeCardData` (type: `boolean`):

Adds the raw YAML metadata of the model/dataset card (base model, datasets, metrics, widget examples...). Can be large.

## `hfToken` (type: `string`):

Optional. A free <a href='https://huggingface.co/settings/tokens' target='_blank'>read token</a> gives you a higher rate limit and access to your own private or gated repos. Stored encrypted.

## `proxyConfiguration` (type: `object`):

Not needed in most cases: the Hugging Face API is public. Enable only if you hit rate limits on very large runs without a token.

## Actor input object example

```json
{
  "resourceType": "models",
  "searchTerms": [
    "llama"
  ],
  "sort": "trending",
  "maxItems": 50,
  "includeReadme": false,
  "includeFiles": false,
  "includeCardData": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "llama"
    ],
    "maxItems": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("dtrungtin/huggingface-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["llama"],
    "maxItems": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("dtrungtin/huggingface-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "llama"
  ],
  "maxItems": 50
}' |
apify call dtrungtin/huggingface-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dtrungtin/huggingface-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/NWlqQETHHj0rMDLSr/builds/M1bsg2Fdjr7pXNGZJ/openapi.json
