# Hugging Face Models and Datasets: search, downloads and likes (`steadydata/huggingface-models`) Actor

Models and datasets on the Hugging Face Hub by search term, task, library, author and tag, sorted by downloads, likes or date: id, author, task, library, downloads, likes, licence, tags, gated flag, created and modified dates, from the official Hub API. Up to 50 searches a run. Pay per result.

- **URL**: https://apify.com/steadydata/huggingface-models.md
- **Developed by:** [Steadydata Team](https://apify.com/steadydata) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.65 / 1,000 result listeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hugging Face Models and Datasets: search, downloads and likes

Models and datasets on the Hugging Face Hub by search term, task, library, author and tag, sorted by downloads, likes or date: id, author, task, library, downloads, likes, licence, tags, gated flag, created and modified dates, from the official Hub API. Up to 50 searches a run. Pay per result.

### Why this scraper

- **Only delivered results are charged.** Inputs that fail come back as clear error
  records at no cost.
- Straight from the Hub's own API, the same index behind huggingface.co's search, with cursor paging so a search is not capped at the first page. Measured on the platform: 240 models for two search terms in 9 seconds for a third of a cent.
- One row per repository with what a model choice needs: id and author, the task (text-classification, text-generation, automatic-speech-recognition), the library, downloads over the last 30 days, likes, the declared licence and languages lifted out of the tags, the gated flag (access on request), created and last-modified dates and the commit hash.
- Filters that map to the Hub's own: task, library and author, plus sorting on downloads, likes, trending, created or modified, most first; datasets are one switch away.
- Licence and languages come as their own columns (`apache-2.0`, `["en"]`) instead of buried in a tag list, so a compliance filter is a column filter.

### Who this is for

Put search terms in `queries` (up to 50 per run; `*` with a task, library or author filter lists everything that matches), choose `kind` (models or datasets), optionally `task`, `library`, `author`, `sortBy` and `maxResultsPerQuery` (default 100). Built for ML engineers shortlisting models for a task, AI product teams tracking what competitors publish, licence and compliance reviews of open models in use, researchers measuring adoption by downloads, and anyone building a catalogue of open models and datasets.

### Who this is not for

The Hub lists what publishers upload and tag; a model without a pipeline tag has no `task`, one without a model card has no licence or languages, and download counts cover the last 30 days only. Private repositories never appear. The row is the repository's catalogue entry, not its weights, files or model card text (the `url` has those). The Hub allows a few hundred requests per five minutes without a token; the actor spaces its requests and identifies itself.

### Input fields

| Field | Type | Required or default | What it does |
|---|---|---|---|
| `queries` | list of text | required | One search per row, up to 50: words in the model or dataset id (sentiment, llama, whisper). An asterisk lists everything that matches the other filters. |
| `kind` | text (models, datasets) | models | models or datasets. |
| `task` | text |  | Only models for this pipeline task, for example text-classification, text-generation, automatic-speech-recognition. Empty means every task. |
| `library` | text |  | Only models for this library, for example transformers, diffusers, sentence-transformers, gguf. Empty means every library. |
| `author` | text |  | Only models or datasets published by this organisation or account, for example openai, meta-llama. Empty means every author. |
| `sortBy` | text (downloads, likes, trending, created, modified) | downloads | downloads (last 30 days), likes, trending, created or modified, most first. |
| `maxResultsPerQuery` | number | 100 | Cost ceiling per search. |

### Input example

```json
{
    "queries": [
        "sentiment",
        "whisper"
    ],
    "kind": "models",
    "sortBy": "downloads",
    "maxResultsPerQuery": 100
}
```

### Output example

| Field | Type | What it holds |
|---|---|---|
| `id` | text | The repository id on the Hub, author/name, the key to the model or dataset page. |
| `kind` | text | model or dataset, the repository type the search was run over. |
| `author` | text | The organisation or account that publishes the repository, the first part of the id. |
| `name` | text | The repository name without the author. |
| `task` | text | The pipeline task of a model as tagged on the Hub, for example text-classification; empty for datasets and untagged models. |
| `library` | text | The library the model is built for, for example transformers or diffusers; empty for datasets. |
| `downloads` | number | Downloads in the last 30 days as the Hub counts them. |
| `likes` | number | How many users liked the repository. |
| `license` | text | The licence declared in the model card, for example apache-2.0 or mit; empty when none is declared. |
| `languages` | list | Language codes tagged on the repository, for example en, nl. |
| `tags` | list | The remaining tags: frameworks, datasets used, arXiv references and free tags. |
| `isGated` | true/false | Whether access requires accepting the author's conditions first. |
| `isPrivate` | true/false | Whether the repository is private; public search never returns private ones, so false. |
| `createdAt` | text | When the repository was created on the Hub. |
| `lastModified` | text | When it was last updated. |
| `sha` | text | The commit hash of the current revision. |
| `url` | text | The model or dataset page on huggingface.co. |
| `query` | text | The search term this row was found with. |

Error codes: `INVALID_QUERY`, `NO_RESULTS`, `BLOCKED`.

One delivered row looks like this:

```json
{
  "id": "cardiffnlp/twitter-roberta-base-sentiment-latest",
  "kind": "model",
  "author": "cardiffnlp",
  "name": "twitter-roberta-base-sentiment-latest",
  "task": "text-classification",
  "library": "transformers",
  "downloads": 3077037,
  "likes": 832,
  "license": "cc-by-4.0",
  "languages": [
    "en"
  ],
  "tags": [
    "transformers",
    "pytorch",
    "tf",
    "roberta",
    "text-classification",
    "en",
    "endpoints_compatible"
  ],
  "isGated": false,
  "isPrivate": false,
  "createdAt": "2022-03-15T01:21:58.000Z",
  "lastModified": "2025-08-04T07:58:29.000Z",
  "sha": "3216a57f2a0d9c45a2e6c20157c20c49fb4bf9c7",
  "url": "https://huggingface.co/cardiffnlp/twitter-roberta-base-sentiment-latest",
  "query": "sentiment",
  "status": "ok"
}
```

### Related actors from steadydata

- [arxiv-papers](https://apify.com/steadydata/arxiv-papers): the papers behind the models
- [github-repo-details](https://apify.com/steadydata/github-repo-details): stars, forks and activity of the code repositories
- [hacker-news-search](https://apify.com/steadydata/hacker-news-search): what was said about a model on Hacker News

### Pricing

Pay per event: one `result-listed` event per delivered result. No charge for inputs
that fail, no separate platform-usage surcharge.

**Free Apify plan:** this actor delivers up to 25 rows per run for accounts on the Apify free
plan, and then stops with a message. That limit is set by us, not by Apify. It exists so the
actor keeps paying for itself for the people who do pay. Any paid Apify plan runs it at full
size, billed per delivered row, with failed rows never charged.

**Reviews:** if this actor saves you time, a short review on this page is the one thing that
helps most. Ratings are what other buyers look at first, and we have no other way to ask.

### FAQ

**Is personal data collected?**
No. `author` is the organisation or account name that publishes the repository, as shown on every Hub page; no profile data is read.

**How do I list every model for a task?**
Set `task` (for example `text-generation`), `queries` to `*` and `sortBy` to downloads; raise `maxResultsPerQuery` for the size of the list.

**How do I find models I am allowed to use commercially?**
Filter the rows on `license`: `apache-2.0`, `mit` and `bsd-3-clause` are permissive; `cc-by-nc-4.0` is not. A model without a licence in its card comes back with an empty `license`, which is worth treating as unknown, not as free.

**What does `isGated` mean?**
The publisher requires users to accept conditions (or request access) before downloading; the model page is public, the weights are not until access is granted.

**Are downloads all-time?**
No: the Hub reports downloads over the last 30 days, which is why the numbers move. `likes` are all-time.

**What does a run cost when a search finds nothing?**
Nothing. `NO_RESULTS`, `INVALID_QUERY` and `BLOCKED` rows are free; only delivered results are charged.

**What happens when the source changes?**
Sources change from time to time; that is the nature of this work. The actor is
monitored daily and fixed fast, and while it is broken you are not charged, because
only delivered results cost anything.

# Changelog

This Actor's version history is a separate document: https://apify.com/steadydata/huggingface-models/changelog.md

# Actor input Schema

## `queries` (type: `array`):

One search per row, up to 50: words in the model or dataset id (sentiment, llama, whisper). An asterisk lists everything that matches the other filters.

## `kind` (type: `string`):

models or datasets.

## `task` (type: `string`):

Only models for this pipeline task, for example text-classification, text-generation, automatic-speech-recognition. Empty means every task.

## `library` (type: `string`):

Only models for this library, for example transformers, diffusers, sentence-transformers, gguf. Empty means every library.

## `author` (type: `string`):

Only models or datasets published by this organisation or account, for example openai, meta-llama. Empty means every author.

## `sortBy` (type: `string`):

downloads (last 30 days), likes, trending, created or modified, most first.

## `maxResultsPerQuery` (type: `integer`):

Cost ceiling per search.

## Actor input object example

```json
{
  "queries": [
    "sentiment",
    "whisper"
  ],
  "kind": "models",
  "sortBy": "downloads",
  "maxResultsPerQuery": 100
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "sentiment",
        "whisper"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("steadydata/huggingface-models").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": [
        "sentiment",
        "whisper",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("steadydata/huggingface-models").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "sentiment",
    "whisper"
  ]
}' |
apify call steadydata/huggingface-models --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,steadydata/huggingface-models"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WhDYfpxdumjUn3Kgj/builds/TrV0RD9epNAZgyOVc/openapi.json
