# HuggingFace Datasets Scraper (`parseforge/huggingface-datasets-scraper`) Actor

Discover and collect dataset listings from the HuggingFace Hub API. Pulls dataset name, description, downloads, tags including task categories and languages, author, likes, trending score, and creation date. Ideal for curating training data catalogs or monitoring new and trending datasets.

- **URL**: https://apify.com/parseforge/huggingface-datasets-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.70 / 1,000 result items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

[![ParseForge](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner-v4.webp)](https://apify.com/parseforge?fpr=vmoqkp)

### HuggingFace Datasets Scraper

**Scrape HuggingFace dataset metadata from search results, tags, or trending lists, up to a million records per run.** Every dataset comes with its downloads, likes, trending score, description, author, and tags. No API key required. Export to CSV, JSON, Excel, or XML.

HuggingFace's official Datasets Server needs a token and rate-limits you. This reads the public datasets API directly, filtered by search term, tags, or sort order, and returns each match in one fixed schema. No login, no token, no browser needed.

| Who uses it | What they scrape HuggingFace for |
|---|---|
| ML engineers | Find the most downloaded text or image datasets for their next model fine-tune. |
| Data analysts | Track which dataset categories and languages are growing fastest each month. |
| AI product managers | Monitor trending datasets to spot emerging model capabilities and community interests. |
| Academic researchers | Build a corpus of dataset metadata for a survey on open-source data availability. |

### What it does

This Actor collects HuggingFace dataset metadata by search query, tag filter, or sort order, and returns each dataset as a flat row.

- 🔍 **Search by keyword:** Filter datasets by any term, such as 'text', 'image', or 'multilingual'.
- 🏷️ **Filter by tags:** Narrow results by task categories, languages, size, or license tags.
- 📊 **Sort by popularity:** Order results by downloads, likes, trending score, or last modified date.
- 📦 **Bulk export:** Dump up to a million dataset records into CSV, JSON, Excel, or XML.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

### What you can do with HuggingFace data

**📈 Monitor trending datasets.**

An ML engineer runs the Actor weekly with sort set to 'trending' to catch new datasets before they become mainstream.

**🏷️ Find datasets by task and language.**

A researcher filters by task\_categories:text-classification and language:de to locate German NLP datasets for a benchmark.

**📊 Audit dataset popularity.**

A product manager scrapes all datasets sorted by downloads to prioritize integrations with the most-used community resources.

**📦 Build a dataset catalog.**

A data platform team exports the full metadata dump to CSV and loads it into their internal search tool for offline browsing.

### Why choose this scraper

| | What you get |
|---|---|
| **No API key** | Reads the public datasets API endpoint with zero authentication. |
| **Fixed flat schema** | Every dataset row has the same columns, ready for pandas or a database. |
| **Tag-aware filtering** | Include or exclude datasets by task, language, license, or size tags. |
| **Sort by real signals** | Rank by downloads, likes, trending score, or last modified date. |

### What a HuggingFace record looks like

Every record returns as one flat JSON row. Here is a real one from a run:

```json
{
 "id": "textmachinelab/quail",
 "author": "textmachinelab",
 "description": "\n\t\n\t\t\n\t\tDataset Card for \"quail\"\n\t\n\n\n\t\n\t\t\n\t\tDataset Summary\n\t\n\nQuAIL is a reading comprehension dataset. QuAIL contains 15K multi-choice questions in texts 300-350 tokens long 4 domains (news, user stories, fiction, blogs).QuAIL is balanced and annotated for question types.\n\n\t\n\t\t\n\t\tSupported Tasks and Leaderboards\n\t\n\nMore Information Needed\n\n\t\n\t\t\n\t\tLanguages\n\t\n\nMore Information Needed\n\n\t\n\t\t\n\t\tDataset Structure\n\t\n\n\n\t\n\t\t\n\t\tData Instances\n\t\n\n\n\t\n\t\t\n\t\tquail\n\t\n\n\nSize of downloaded dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/textmachinelab/quail.",
 "downloads": 71086,
 "likes": 8,
 "trendingScore": 0,
 "gated": false,
 "private": false,
 "disabled": false,
 "sha": "2bd9d7f90a532fe1a910b70972cef5fda341c8fe",
 "createdAt": "2022-03-02T23:29:22.000Z",
 "lastModified": "2024-01-04T16:18:32.000Z",
 "scrapedAt": "2026-09-24T12:04:00.022Z"
}
```

Every value above comes from a real run. A field a record does not have comes back as `null`.

### Configure the run

Drive the Actor from a search term, tag filters, and a sort order, alone or together, and filters run as each dataset is read so only matches reach your dataset. The Input tab lists every parameter.

A first run with the defaults:

```json
{
 "maxItems": 10,
 "search": "text",
 "sort": "downloads"
}
```

A larger pull:

```json
{
 "maxItems": 200,
 "search": "text",
 "sort": "downloads"
}
```

### Free users

Free-plan runs return up to 10 results as a preview. [Upgrade your Apify plan](https://console.apify.com/sign-up?fpr=vmoqkp) to collect up to 1,000,000 results per run.

### Run it

1. [Create a free Apify account](https://console.apify.com/sign-up?fpr=vmoqkp).
2. Set your inputs and any filters, then click **Start**.
3. Export the results as CSV, Excel, JSON, or XML from the **Dataset** tab.

Run it programmatically through the [Apify API](https://docs.apify.com/api/v2) (`run-sync-get-dataset-items`) or the [ApifyClient](https://docs.apify.com/api/client/js) for JavaScript and Python.

### Use with AI agents (MCP)

Give an AI agent live access to HuggingFace through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

```bash
```

Then prompt it in plain language to run the scraper and read back the results.

### Troubleshooting

**Why am I getting no results?**

Check that your search term or tags filter is not too restrictive. Try removing filters one by one to see which one eliminates all matches. Also confirm that the sort order is set to a valid value.

**Why does the run stop at 10 items?**

Free users are capped at 10 items as a preview. Upgrade to a paid plan and set the maxItems input to a higher number to scrape more datasets.

**Some datasets are missing from the results.**

The API only returns public, non-disabled datasets. Gated or private datasets are not included. If a dataset was recently created, wait a few minutes and retry.

**The Actor returns an error or empty dataset.**

The HuggingFace API may be temporarily slow. Wait a minute and retry. If the problem persists, check the Actor's log for HTTP status codes and report them to Apify support.

### FAQ

| Question | Answer |
|---|---|
| Do I need a HuggingFace API key or token? | No. This Actor reads the public, unauthenticated datasets API endpoint. No login, no token, and no HuggingFace account are required. |
| What data fields does the Actor return? | Each row includes the dataset ID, author, description, downloads, likes, trending score, tags, creation date, last modified date, and more. The exact field list is shown in the sample output on the Actor's page. |
| Can I filter datasets by task category or language? | Yes. Use the tags filter to include only datasets that match specific task\_categories, language, size\_categories, or license tags. |
| How many datasets can I scrape in one run? | Free users are limited to 10 items as a preview. Paid users can set maxItems up to 1,000,000 and pull the full catalog. |
| Can I search for a specific dataset name? | Yes. The search input accepts a keyword and returns datasets whose ID or description matches, exactly like the search bar on the HuggingFace website. |
| What export formats are supported? | You can export the results to CSV, JSON, Excel, or XML directly from the dataset tab after the run finishes. |
| Does this Actor download the actual dataset files? | No. It scrapes only the metadata (name, description, stats, tags). To download the dataset files themselves, use the HuggingFace datasets Python library. |
| How do I sort results by the most downloaded datasets? | Set the sort input to 'downloads'. You can also sort by 'likes', 'trending', or 'lastModified'. |
| Is this Actor affected by HuggingFace rate limits? | The public API has generous limits. The Actor uses polite defaults, but if you need very high throughput, contact support for guidance. |
| Can I combine multiple filters in one run? | Yes. You can set a search term, one or more tags, and a sort order all at once. Only datasets that match every filter are returned. |

### Related actors

Browse the full [ParseForge collection](https://apify.com/parseforge?fpr=vmoqkp) for more scrapers.

🆘 **Need help?** Email parseforge@protonmail.com with your run ID, your input, and what you expected.

⚠️ **Disclaimer.** This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Hugging Face, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.

# Actor input Schema

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `search` (type: `string`):

Filter datasets by a keyword search (e.g., text, image, multilingual). Leave empty to list all datasets.

## `sort` (type: `string`):

Sort datasets by downloads, likes, trending score, or last modified. Default is trending.

## `tags` (type: `array`):

Filter datasets by one or more tags (e.g., task\_categories:image-classification, language:en).

## Actor input object example

```json
{
  "maxItems": 10,
  "search": "text",
  "sort": "downloads"
}
```

# Actor output Schema

## `results` (type: `string`):

Complete dataset of all scraped records.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10,
    "search": "text",
    "sort": "downloads"
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/huggingface-datasets-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxItems": 10,
    "search": "text",
    "sort": "downloads",
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/huggingface-datasets-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10,
  "search": "text",
  "sort": "downloads"
}' |
apify call parseforge/huggingface-datasets-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/huggingface-datasets-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/MC85MvoZahdaew4xT/builds/sggB2uBnHi1v587Dh/openapi.json
