HuggingFace Datasets Scraper
Pricing
from $1.70 / 1,000 result items
HuggingFace Datasets Scraper
Discover and collect dataset listings from the HuggingFace Hub API. Pulls dataset name, description, downloads, tags including task categories and languages, author, likes, trending score, and creation date. Ideal for curating training data catalogs or monitoring new and trending datasets.
Pricing
from $1.70 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
HuggingFace Datasets Scraper
Scrape HuggingFace dataset metadata from search results, tags, or trending lists, up to a million records per run. Every dataset comes with its downloads, likes, trending score, description, author, and tags. No API key required. Export to CSV, JSON, Excel, or XML.
HuggingFace's official Datasets Server needs a token and rate-limits you. This reads the public datasets API directly, filtered by search term, tags, or sort order, and returns each match in one fixed schema. No login, no token, no browser needed.
| Who uses it | What they scrape HuggingFace for |
|---|---|
| ML engineers | Find the most downloaded text or image datasets for their next model fine-tune. |
| Data analysts | Track which dataset categories and languages are growing fastest each month. |
| AI product managers | Monitor trending datasets to spot emerging model capabilities and community interests. |
| Academic researchers | Build a corpus of dataset metadata for a survey on open-source data availability. |
What it does
This Actor collects HuggingFace dataset metadata by search query, tag filter, or sort order, and returns each dataset as a flat row.
- ๐ Search by keyword: Filter datasets by any term, such as 'text', 'image', or 'multilingual'.
- ๐ท๏ธ Filter by tags: Narrow results by task categories, languages, size, or license tags.
- ๐ Sort by popularity: Order results by downloads, likes, trending score, or last modified date.
- ๐ฆ Bulk export: Dump up to a million dataset records into CSV, JSON, Excel, or XML.
Results export to CSV, JSON, Excel, or XML, or straight from the API.
What you can do with HuggingFace data
๐ Monitor trending datasets.
An ML engineer runs the Actor weekly with sort set to 'trending' to catch new datasets before they become mainstream.
๐ท๏ธ Find datasets by task and language.
A researcher filters by task_categories:text-classification and language:de to locate German NLP datasets for a benchmark.
๐ Audit dataset popularity.
A product manager scrapes all datasets sorted by downloads to prioritize integrations with the most-used community resources.
๐ฆ Build a dataset catalog.
A data platform team exports the full metadata dump to CSV and loads it into their internal search tool for offline browsing.
Why choose this scraper
| What you get | |
|---|---|
| No API key | Reads the public datasets API endpoint with zero authentication. |
| Fixed flat schema | Every dataset row has the same columns, ready for pandas or a database. |
| Tag-aware filtering | Include or exclude datasets by task, language, license, or size tags. |
| Sort by real signals | Rank by downloads, likes, trending score, or last modified date. |
What a HuggingFace record looks like
Every record returns as one flat JSON row. Here is a real one from a run:
{"id": "textmachinelab/quail","author": "textmachinelab","description": "\n\t\n\t\t\n\t\tDataset Card for \"quail\"\n\t\n\n\n\t\n\t\t\n\t\tDataset Summary\n\t\n\nQuAIL is a reading comprehension dataset. QuAIL contains 15K multi-choice questions in texts 300-350 tokens long 4 domains (news, user stories, fiction, blogs).QuAIL is balanced and annotated for question types.\n\n\t\n\t\t\n\t\tSupported Tasks and Leaderboards\n\t\n\nMore Information Needed\n\n\t\n\t\t\n\t\tLanguages\n\t\n\nMore Information Needed\n\n\t\n\t\t\n\t\tDataset Structure\n\t\n\n\n\t\n\t\t\n\t\tData Instances\n\t\n\n\n\t\n\t\t\n\t\tquail\n\t\n\n\nSize of downloaded dataset files:โฆ See the full description on the dataset page: https://huggingface.co/datasets/textmachinelab/quail.","downloads": 71086,"likes": 8,"trendingScore": 0,"gated": false,"private": false,"disabled": false,"sha": "2bd9d7f90a532fe1a910b70972cef5fda341c8fe","createdAt": "2022-03-02T23:29:22.000Z","lastModified": "2024-01-04T16:18:32.000Z","scrapedAt": "2026-09-24T12:04:00.022Z"}
Every value above comes from a real run. A field a record does not have comes back as null.
Configure the run
Drive the Actor from a search term, tag filters, and a sort order, alone or together, and filters run as each dataset is read so only matches reach your dataset. The Input tab lists every parameter.
A first run with the defaults:
{"maxItems": 10,"search": "text","sort": "downloads"}
A larger pull:
{"maxItems": 200,"search": "text","sort": "downloads"}
Free users
Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.
Run it
- Create a free Apify account.
- Set your inputs and any filters, then click Start.
- Export the results as CSV, Excel, JSON, or XML from the Dataset tab.
Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.
Use with AI agents (MCP)
Give an AI agent live access to HuggingFace through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:
$undefined
Then prompt it in plain language to run the scraper and read back the results.
Troubleshooting
Why am I getting no results?
Check that your search term or tags filter is not too restrictive. Try removing filters one by one to see which one eliminates all matches. Also confirm that the sort order is set to a valid value.
Why does the run stop at 10 items?
Free users are capped at 10 items as a preview. Upgrade to a paid plan and set the maxItems input to a higher number to scrape more datasets.
Some datasets are missing from the results.
The API only returns public, non-disabled datasets. Gated or private datasets are not included. If a dataset was recently created, wait a few minutes and retry.
The Actor returns an error or empty dataset.
The HuggingFace API may be temporarily slow. Wait a minute and retry. If the problem persists, check the Actor's log for HTTP status codes and report them to Apify support.
FAQ
| Question | Answer |
|---|---|
| Do I need a HuggingFace API key or token? | No. This Actor reads the public, unauthenticated datasets API endpoint. No login, no token, and no HuggingFace account are required. |
| What data fields does the Actor return? | Each row includes the dataset ID, author, description, downloads, likes, trending score, tags, creation date, last modified date, and more. The exact field list is shown in the sample output on the Actor's page. |
| Can I filter datasets by task category or language? | Yes. Use the tags filter to include only datasets that match specific task_categories, language, size_categories, or license tags. |
| How many datasets can I scrape in one run? | Free users are limited to 10 items as a preview. Paid users can set maxItems up to 1,000,000 and pull the full catalog. |
| Can I search for a specific dataset name? | Yes. The search input accepts a keyword and returns datasets whose ID or description matches, exactly like the search bar on the HuggingFace website. |
| What export formats are supported? | You can export the results to CSV, JSON, Excel, or XML directly from the dataset tab after the run finishes. |
| Does this Actor download the actual dataset files? | No. It scrapes only the metadata (name, description, stats, tags). To download the dataset files themselves, use the HuggingFace datasets Python library. |
| How do I sort results by the most downloaded datasets? | Set the sort input to 'downloads'. You can also sort by 'likes', 'trending', or 'lastModified'. |
| Is this Actor affected by HuggingFace rate limits? | The public API has generous limits. The Actor uses polite defaults, but if you need very high throughput, contact support for guidance. |
| Can I combine multiple filters in one run? | Yes. You can set a search term, one or more tags, and a sort order all at once. Only datasets that match every filter are returned. |
Related actors
Browse the full ParseForge collection for more scrapers.
๐ Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.
โ ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by Hugging Face, Inc. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.
