Hugging Face Models and Datasets: search, downloads and likes avatar

Hugging Face Models and Datasets: search, downloads and likes

Pricing

from $0.65 / 1,000 result listeds

Go to Apify Store
Hugging Face Models and Datasets: search, downloads and likes

Hugging Face Models and Datasets: search, downloads and likes

Models and datasets on the Hugging Face Hub by search term, task, library, author and tag, sorted by downloads, likes or date: id, author, task, library, downloads, likes, licence, tags, gated flag, created and modified dates, from the official Hub API. Up to 50 searches a run. Pay per result.

Pricing

from $0.65 / 1,000 result listeds

Rating

0.0

(0)

Developer

Steadydata Team

Steadydata Team

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

Models and datasets on the Hugging Face Hub by search term, task, library, author and tag, sorted by downloads, likes or date: id, author, task, library, downloads, likes, licence, tags, gated flag, created and modified dates, from the official Hub API. Up to 50 searches a run. Pay per result.

Why this scraper

  • Only delivered results are charged. Inputs that fail come back as clear error records at no cost.
  • Straight from the Hub's own API, the same index behind huggingface.co's search, with cursor paging so a search is not capped at the first page. Measured on the platform: 240 models for two search terms in 9 seconds for a third of a cent.
  • One row per repository with what a model choice needs: id and author, the task (text-classification, text-generation, automatic-speech-recognition), the library, downloads over the last 30 days, likes, the declared licence and languages lifted out of the tags, the gated flag (access on request), created and last-modified dates and the commit hash.
  • Filters that map to the Hub's own: task, library and author, plus sorting on downloads, likes, trending, created or modified, most first; datasets are one switch away.
  • Licence and languages come as their own columns (apache-2.0, ["en"]) instead of buried in a tag list, so a compliance filter is a column filter.

Who this is for

Put search terms in queries (up to 50 per run; * with a task, library or author filter lists everything that matches), choose kind (models or datasets), optionally task, library, author, sortBy and maxResultsPerQuery (default 100). Built for ML engineers shortlisting models for a task, AI product teams tracking what competitors publish, licence and compliance reviews of open models in use, researchers measuring adoption by downloads, and anyone building a catalogue of open models and datasets.

Who this is not for

The Hub lists what publishers upload and tag; a model without a pipeline tag has no task, one without a model card has no licence or languages, and download counts cover the last 30 days only. Private repositories never appear. The row is the repository's catalogue entry, not its weights, files or model card text (the url has those). The Hub allows a few hundred requests per five minutes without a token; the actor spaces its requests and identifies itself.

Input fields

FieldTypeRequired or defaultWhat it does
querieslist of textrequiredOne search per row, up to 50: words in the model or dataset id (sentiment, llama, whisper). An asterisk lists everything that matches the other filters.
kindtext (models, datasets)modelsmodels or datasets.
tasktextOnly models for this pipeline task, for example text-classification, text-generation, automatic-speech-recognition. Empty means every task.
librarytextOnly models for this library, for example transformers, diffusers, sentence-transformers, gguf. Empty means every library.
authortextOnly models or datasets published by this organisation or account, for example openai, meta-llama. Empty means every author.
sortBytext (downloads, likes, trending, created, modified)downloadsdownloads (last 30 days), likes, trending, created or modified, most first.
maxResultsPerQuerynumber100Cost ceiling per search.

Input example

{
"queries": [
"sentiment",
"whisper"
],
"kind": "models",
"sortBy": "downloads",
"maxResultsPerQuery": 100
}

Output example

FieldTypeWhat it holds
idtextThe repository id on the Hub, author/name, the key to the model or dataset page.
kindtextmodel or dataset, the repository type the search was run over.
authortextThe organisation or account that publishes the repository, the first part of the id.
nametextThe repository name without the author.
tasktextThe pipeline task of a model as tagged on the Hub, for example text-classification; empty for datasets and untagged models.
librarytextThe library the model is built for, for example transformers or diffusers; empty for datasets.
downloadsnumberDownloads in the last 30 days as the Hub counts them.
likesnumberHow many users liked the repository.
licensetextThe licence declared in the model card, for example apache-2.0 or mit; empty when none is declared.
languageslistLanguage codes tagged on the repository, for example en, nl.
tagslistThe remaining tags: frameworks, datasets used, arXiv references and free tags.
isGatedtrue/falseWhether access requires accepting the author's conditions first.
isPrivatetrue/falseWhether the repository is private; public search never returns private ones, so false.
createdAttextWhen the repository was created on the Hub.
lastModifiedtextWhen it was last updated.
shatextThe commit hash of the current revision.
urltextThe model or dataset page on huggingface.co.
querytextThe search term this row was found with.

Error codes: INVALID_QUERY, NO_RESULTS, BLOCKED.

One delivered row looks like this:

{
"id": "cardiffnlp/twitter-roberta-base-sentiment-latest",
"kind": "model",
"author": "cardiffnlp",
"name": "twitter-roberta-base-sentiment-latest",
"task": "text-classification",
"library": "transformers",
"downloads": 3077037,
"likes": 832,
"license": "cc-by-4.0",
"languages": [
"en"
],
"tags": [
"transformers",
"pytorch",
"tf",
"roberta",
"text-classification",
"en",
"endpoints_compatible"
],
"isGated": false,
"isPrivate": false,
"createdAt": "2022-03-15T01:21:58.000Z",
"lastModified": "2025-08-04T07:58:29.000Z",
"sha": "3216a57f2a0d9c45a2e6c20157c20c49fb4bf9c7",
"url": "https://huggingface.co/cardiffnlp/twitter-roberta-base-sentiment-latest",
"query": "sentiment",
"status": "ok"
}

Pricing

Pay per event: one result-listed event per delivered result. No charge for inputs that fail, no separate platform-usage surcharge.

Free Apify plan: this actor delivers up to 25 rows per run for accounts on the Apify free plan, and then stops with a message. That limit is set by us, not by Apify. It exists so the actor keeps paying for itself for the people who do pay. Any paid Apify plan runs it at full size, billed per delivered row, with failed rows never charged.

Reviews: if this actor saves you time, a short review on this page is the one thing that helps most. Ratings are what other buyers look at first, and we have no other way to ask.

FAQ

Is personal data collected? No. author is the organisation or account name that publishes the repository, as shown on every Hub page; no profile data is read.

How do I list every model for a task? Set task (for example text-generation), queries to * and sortBy to downloads; raise maxResultsPerQuery for the size of the list.

How do I find models I am allowed to use commercially? Filter the rows on license: apache-2.0, mit and bsd-3-clause are permissive; cc-by-nc-4.0 is not. A model without a licence in its card comes back with an empty license, which is worth treating as unknown, not as free.

What does isGated mean? The publisher requires users to accept conditions (or request access) before downloading; the model page is public, the weights are not until access is granted.

Are downloads all-time? No: the Hub reports downloads over the last 30 days, which is why the numbers move. likes are all-time.

What does a run cost when a search finds nothing? Nothing. NO_RESULTS, INVALID_QUERY and BLOCKED rows are free; only delivered results are charged.

What happens when the source changes? Sources change from time to time; that is the nature of this work. The actor is monitored daily and fixed fast, and while it is broken you are not charged, because only delivered results cost anything.