Hugging Face Models and Datasets: search, downloads and likes
Pricing
from $0.65 / 1,000 result listeds
Hugging Face Models and Datasets: search, downloads and likes
Models and datasets on the Hugging Face Hub by search term, task, library, author and tag, sorted by downloads, likes or date: id, author, task, library, downloads, likes, licence, tags, gated flag, created and modified dates, from the official Hub API. Up to 50 searches a run. Pay per result.
Pricing
from $0.65 / 1,000 result listeds
Rating
0.0
(0)
Developer
Steadydata Team
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 hours ago
Last modified
Categories
Share
Models and datasets on the Hugging Face Hub by search term, task, library, author and tag, sorted by downloads, likes or date: id, author, task, library, downloads, likes, licence, tags, gated flag, created and modified dates, from the official Hub API. Up to 50 searches a run. Pay per result.
Why this scraper
- Only delivered results are charged. Inputs that fail come back as clear error records at no cost.
- Straight from the Hub's own API, the same index behind huggingface.co's search, with cursor paging so a search is not capped at the first page. Measured on the platform: 240 models for two search terms in 9 seconds for a third of a cent.
- One row per repository with what a model choice needs: id and author, the task (text-classification, text-generation, automatic-speech-recognition), the library, downloads over the last 30 days, likes, the declared licence and languages lifted out of the tags, the gated flag (access on request), created and last-modified dates and the commit hash.
- Filters that map to the Hub's own: task, library and author, plus sorting on downloads, likes, trending, created or modified, most first; datasets are one switch away.
- Licence and languages come as their own columns (
apache-2.0,["en"]) instead of buried in a tag list, so a compliance filter is a column filter.
Who this is for
Put search terms in queries (up to 50 per run; * with a task, library or author filter lists everything that matches), choose kind (models or datasets), optionally task, library, author, sortBy and maxResultsPerQuery (default 100). Built for ML engineers shortlisting models for a task, AI product teams tracking what competitors publish, licence and compliance reviews of open models in use, researchers measuring adoption by downloads, and anyone building a catalogue of open models and datasets.
Who this is not for
The Hub lists what publishers upload and tag; a model without a pipeline tag has no task, one without a model card has no licence or languages, and download counts cover the last 30 days only. Private repositories never appear. The row is the repository's catalogue entry, not its weights, files or model card text (the url has those). The Hub allows a few hundred requests per five minutes without a token; the actor spaces its requests and identifies itself.
Input fields
| Field | Type | Required or default | What it does |
|---|---|---|---|
queries | list of text | required | One search per row, up to 50: words in the model or dataset id (sentiment, llama, whisper). An asterisk lists everything that matches the other filters. |
kind | text (models, datasets) | models | models or datasets. |
task | text | Only models for this pipeline task, for example text-classification, text-generation, automatic-speech-recognition. Empty means every task. | |
library | text | Only models for this library, for example transformers, diffusers, sentence-transformers, gguf. Empty means every library. | |
author | text | Only models or datasets published by this organisation or account, for example openai, meta-llama. Empty means every author. | |
sortBy | text (downloads, likes, trending, created, modified) | downloads | downloads (last 30 days), likes, trending, created or modified, most first. |
maxResultsPerQuery | number | 100 | Cost ceiling per search. |
Input example
{"queries": ["sentiment","whisper"],"kind": "models","sortBy": "downloads","maxResultsPerQuery": 100}
Output example
| Field | Type | What it holds |
|---|---|---|
id | text | The repository id on the Hub, author/name, the key to the model or dataset page. |
kind | text | model or dataset, the repository type the search was run over. |
author | text | The organisation or account that publishes the repository, the first part of the id. |
name | text | The repository name without the author. |
task | text | The pipeline task of a model as tagged on the Hub, for example text-classification; empty for datasets and untagged models. |
library | text | The library the model is built for, for example transformers or diffusers; empty for datasets. |
downloads | number | Downloads in the last 30 days as the Hub counts them. |
likes | number | How many users liked the repository. |
license | text | The licence declared in the model card, for example apache-2.0 or mit; empty when none is declared. |
languages | list | Language codes tagged on the repository, for example en, nl. |
tags | list | The remaining tags: frameworks, datasets used, arXiv references and free tags. |
isGated | true/false | Whether access requires accepting the author's conditions first. |
isPrivate | true/false | Whether the repository is private; public search never returns private ones, so false. |
createdAt | text | When the repository was created on the Hub. |
lastModified | text | When it was last updated. |
sha | text | The commit hash of the current revision. |
url | text | The model or dataset page on huggingface.co. |
query | text | The search term this row was found with. |
Error codes: INVALID_QUERY, NO_RESULTS, BLOCKED.
One delivered row looks like this:
{"id": "cardiffnlp/twitter-roberta-base-sentiment-latest","kind": "model","author": "cardiffnlp","name": "twitter-roberta-base-sentiment-latest","task": "text-classification","library": "transformers","downloads": 3077037,"likes": 832,"license": "cc-by-4.0","languages": ["en"],"tags": ["transformers","pytorch","tf","roberta","text-classification","en","endpoints_compatible"],"isGated": false,"isPrivate": false,"createdAt": "2022-03-15T01:21:58.000Z","lastModified": "2025-08-04T07:58:29.000Z","sha": "3216a57f2a0d9c45a2e6c20157c20c49fb4bf9c7","url": "https://huggingface.co/cardiffnlp/twitter-roberta-base-sentiment-latest","query": "sentiment","status": "ok"}
Related actors from steadydata
- arxiv-papers: the papers behind the models
- github-repo-details: stars, forks and activity of the code repositories
- hacker-news-search: what was said about a model on Hacker News
Pricing
Pay per event: one result-listed event per delivered result. No charge for inputs
that fail, no separate platform-usage surcharge.
Free Apify plan: this actor delivers up to 25 rows per run for accounts on the Apify free plan, and then stops with a message. That limit is set by us, not by Apify. It exists so the actor keeps paying for itself for the people who do pay. Any paid Apify plan runs it at full size, billed per delivered row, with failed rows never charged.
Reviews: if this actor saves you time, a short review on this page is the one thing that helps most. Ratings are what other buyers look at first, and we have no other way to ask.
FAQ
Is personal data collected?
No. author is the organisation or account name that publishes the repository, as shown on every Hub page; no profile data is read.
How do I list every model for a task?
Set task (for example text-generation), queries to * and sortBy to downloads; raise maxResultsPerQuery for the size of the list.
How do I find models I am allowed to use commercially?
Filter the rows on license: apache-2.0, mit and bsd-3-clause are permissive; cc-by-nc-4.0 is not. A model without a licence in its card comes back with an empty license, which is worth treating as unknown, not as free.
What does isGated mean?
The publisher requires users to accept conditions (or request access) before downloading; the model page is public, the weights are not until access is granted.
Are downloads all-time?
No: the Hub reports downloads over the last 30 days, which is why the numbers move. likes are all-time.
What does a run cost when a search finds nothing?
Nothing. NO_RESULTS, INVALID_QUERY and BLOCKED rows are free; only delivered results are charged.
What happens when the source changes? Sources change from time to time; that is the nature of this work. The actor is monitored daily and fixed fast, and while it is broken you are not charged, because only delivered results cost anything.