# Hugging Face Scraper: Models, Datasets & Spaces (`punkrecordsdata/huggingface-hub-scraper`) Actor

Scrape Hugging Face Hub models, datasets and Spaces: downloads, likes, trending score, license, gated status, parameter counts and file lists. Export CSV, Excel, JSON, XML.

- **URL**: https://apify.com/punkrecordsdata/huggingface-hub-scraper.md
- **Developed by:** [PunkRecordsData](https://apify.com/punkrecordsdata) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $12.27 / 1,000 model records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<p align="center">
  <img src="https://api.apify.com/v2/key-value-stores/AAm3a1h3Z9nYfrvh9/records/banner?v=2" alt="PunkRecordsData" width="100%" />
</p>

## 🤗 Hugging Face Scraper - Models, Datasets & Spaces - PunkRecordsData

> 🚀 **Export Hugging Face Hub data in seconds.** Search AI models, datasets and Spaces and get structured rows with downloads, likes, trending score, task, license, gated status, parameter count and file lists. A single "llama" search returns models with 7.3M monthly downloads ranked for you. Export to CSV, Excel, JSON or XML.

The Hugging Face Scraper reads the Hub's own public API, so every number is exactly what the site shows: real download counts, likes, trending scores and per-model metadata across the largest open AI model repository. Filter by task (38 pipeline tags), author or keyword, sort by trending, downloads, likes or dates, and optionally pull full model details including license, parameter count and every file in the repo.

| 🎯 Target Audience | 💡 Primary Use Cases |
|---|---|
| AI engineers and MLOps teams | Shortlist models by task with real download and license data |
| Analysts and VCs tracking AI | Measure model ecosystem momentum with trending scores |
| Compliance and legal teams | Audit licenses and gated status across the models you use |
| Researchers | Track datasets and Spaces around a research area |

### 📋 What the Hugging Face Hub Scraper does

- **Model records**: id, author, task (pipeline tag), library, downloads, likes, trending score, gated status, tags, created and modified dates.
- **Full model details** (optional, per model): license, parameter count from safetensors, base model, file count and file list.
- **Dataset records**: the same search across Hub datasets, with downloads and likes.
- **Space records**: demo apps with SDK and likes.
- **Filters that map 1:1 to the Hub**: 38 task tags as a dropdown, author/organization, keyword search, five sort orders.

> 💡 **Why it matters:** model selection decisions need license, gating and adoption data side by side. The Hub shows them one page at a time; this actor gives you the whole comparison in one spreadsheet.

### 📊 Output of the Hugging Face model search

Real sample from a live run (22 fields per model with details on):

```json
{
  "recordType": "model",
  "id": "meta-llama/Llama-3.1-8B-Instruct",
  "author": "meta-llama",
  "url": "https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct",
  "likes": 8002,
  "downloads": 6079804,
  "trendingScore": 108,
  "gated": "manual",
  "pipelineTag": "text-generation",
  "libraryName": "transformers",
  "license": "llama3.1",
  "parameterCount": 8030261248,
  "baseModel": "meta-llama/Meta-Llama-3.1-8B",
  "fileCount": 17,
  "error": null
}
```

### ✨ Why choose this Hugging Face scraper

- **4 billable events, each switchable**: models, datasets, Spaces and per-model details. No measured alternative offers more than one.
- **License + parameter count + gated status in one row**, the three fields every "can we use this model" decision needs.
- **Real Hub numbers**, read from the same API the site uses, not estimates.
- **All three Hub surfaces in one actor** instead of one actor per surface.
- **Honest billing**: details are charged only when the extra call actually returns; failures cost nothing.

### 📈 How this Hugging Face scraper compares to alternatives

Measured against the Hugging Face actors on the Apify Store (September 2026):

| | This actor | Most-used alternative | Other HF actors |
|---|---|---|---|
| Billable data events | 4 (models, datasets, spaces, details) | 1 | 1 |
| License / parameters / files per model | Yes, details event | No | No |
| Datasets and Spaces | Yes | Models only | Varies |
| Task filter | 38-tag dropdown | Varies | Varies |
| Price per 1,000 models | $5.50 | $5.00 | $1.00 to $10.00 |

### 🚀 How to use the Hugging Face Hub Scraper

1. Create a free Apify account (with $5 of platform credit) at console.apify.com.
2. Open this actor's page and click **Try for free**.
3. Type search terms or an author (e.g. mistralai), pick a task if you want one.
4. Toggle datasets, Spaces or full model details as needed and press **Start**.
5. Download the dataset as CSV, Excel, JSON or XML.

### 💼 Business use cases

#### Model selection with legal sign-off

One run gives engineering the downloads and trending data and legal the license and gated status, for every candidate model at once.

#### AI market intelligence

Track which organizations ship models weekly and which tasks are heating up, using trending scores instead of anecdotes.

#### Dependency audits

List every file and the parameter count of the models your stack pulls, before they change under you.

#### Dataset discovery

Find datasets around your domain with adoption numbers, not just names.

### 🔌 Automating the Hugging Face Scraper

Connect to **Make**, **Zapier**, **Slack**, **Airbyte**, **GitHub** or **Google Drive** through Apify integrations: weekly trending sweeps to Slack, license-change monitoring, or warehouse syncs of your model watchlist.

### 🌟 Beyond business use cases

- **Research:** ecosystem studies of model releases by task and organization.
- **Personal:** watch your own models' downloads and likes over time.
- **Non-profit:** track open-license models suitable for public-interest work.
- **Experimentation:** a zero-setup Hub API playground.

### 🤖 Ask an AI assistant about this scraper

> "I need a spreadsheet of Hugging Face models for a task, with downloads, license, gated status and parameter counts. Would the Hugging Face Scraper on Apify (apify.com/punkrecordsdata/huggingface-hub-scraper) do this on a schedule?"

### ❓ Frequently Asked Questions

#### 🤖 How do I export a list of Hugging Face models to CSV or Excel?

Enter a search term or author, click Start, and download the dataset from the Storage tab in CSV, Excel, JSON or XML.

#### 🔥 Can I get the trending models on Hugging Face?

Yes. Leave sort on "Trending" and the rows come back in the Hub's own trending order with the score included.

#### 📜 How do I check the license of many models at once?

Enable "Full model details" and each model row carries its license, parameter count, base model and file list.

#### 🔒 What does the gated field mean?

"No" means downloadable directly; "manual" or "auto" mirror the Hub's gated-access setting, where you must request access first.

#### 📦 Does it cover datasets and Spaces too?

Yes, both are separate toggles and separate billable events, using the same search terms.

#### 👤 Can I list everything from one organization?

Yes, set the author field (e.g. google) and leave search terms empty.

#### 🧮 Are download numbers real?

They come from the Hub's own API, the same numbers the website renders.

#### 💵 Do I pay for details on models where the call fails?

No. The details event is charged only when the extra call returns data.

#### 📊 How many models can one run return?

Up to 1,000,000 on paid plans, following the Hub's own pagination. Free users get a 10-record preview.

#### ⚙️ Does it need a Hugging Face API token?

No. It reads only public Hub data, no token required.

#### 🗂 Can I filter models by task?

Yes, 38 official pipeline tags (text-generation, text-to-image, ASR and more) as a dropdown.

### 🔌 Integrate with any app

Datasets are available via the Apify API in JSON, CSV, Excel or XML for Python, Node.js, Sheets or BI tools, with webhooks on run completion.

### 🔗 Recommended Actors

- [arXiv Research Papers Scraper](https://apify.com/punkrecordsdata/arxiv-research-papers-scraper) - the papers behind the models, with citation metrics
- [OpenCV Image Analyzer](https://apify.com/punkrecordsdata/opencv-image-analyzer) - deterministic computer vision on any image URL
- [Steam Games Player Stats](https://apify.com/punkrecordsdata/steam-games-player-stats) - another live-API data product
- [Discord Scraper](https://apify.com/punkrecordsdata/discord-scraper) - community signals around AI projects

> 💡 **Pro Tip:** browse the complete [PunkRecordsData collection](https://apify.com/punkrecordsdata) for more data tools.

**🆘 Need Help?** contact.punkrecordsdata@gmail.com

> **⚠️ Disclaimer:** independent tool, not affiliated with Hugging Face; only publicly available data.

# Actor input Schema

## `searchTerms` (type: `array`):

One Hub search per term (e.g. "llama", "sentiment analysis", "stable diffusion"). Leave empty to browse by author or task only.

## `author` (type: `string`):

Limit results to one author or org (e.g. meta-llama, google, mistralai).

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `includeModels` (type: `boolean`):

Scrape model records (downloads, likes, trending score, pipeline task, tags).

## `includeDatasets` (type: `boolean`):

Scrape dataset records for the same search terms.

## `includeSpaces` (type: `boolean`):

Scrape Space records (demo apps) for the same search terms.

## `includeModelDetails` (type: `boolean`):

One extra API call per model: license, gated status, parameter count, file list and base model. Billed as a separate event.

## `pipelineTag` (type: `string`):

Only models for this task.

## `sortBy` (type: `string`):

Order of results from the Hub.

## Actor input object example

```json
{
  "searchTerms": [
    "llama"
  ],
  "maxItems": 10,
  "includeModels": true,
  "includeDatasets": false,
  "includeSpaces": false,
  "includeModelDetails": false,
  "pipelineTag": "",
  "sortBy": "trendingScore"
}
```

# Actor output Schema

## `overview` (type: `string`):

Key fields per record

## `fullData` (type: `string`):

Complete dataset with all fields

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "llama"
    ],
    "author": "",
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("punkrecordsdata/huggingface-hub-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["llama"],
    "author": "",
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("punkrecordsdata/huggingface-hub-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "llama"
  ],
  "author": "",
  "maxItems": 10
}' |
apify call punkrecordsdata/huggingface-hub-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,punkrecordsdata/huggingface-hub-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qGhPNeewg6quUPuLh/builds/lPbrLTdjvCa6PDWZr/openapi.json
