Hugging Face Scraper: Models, Datasets & Spaces avatar

Hugging Face Scraper: Models, Datasets & Spaces

Pricing

from $0.37 / 1,000 repository scrapeds

Go to Apify Store
Hugging Face Scraper: Models, Datasets & Spaces

Hugging Face Scraper: Models, Datasets & Spaces

Scrape Hugging Face models, datasets and spaces: downloads, likes, tags, pipeline task, licence, library and last-modified. No key, no proxy, no browser.

Pricing

from $0.37 / 1,000 repository scrapeds

Rating

0.0

(0)

Developer

Arman Hossain

Arman Hossain

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

Hugging Face Scraper: Models, datasets and spaces ranked by downloads, likes or trending score, with licence and gating status

Lists models, datasets and spaces from the Hugging Face Hub. You get downloads, likes, trending score, tags, pipeline task, library, licence, gating status and both repository timestamps.

The Hub publishes its whole index through a public read API at huggingface.co/api, intended for programmatic use. This Actor reads that index directly, so there's no token, no browser and no proxies. A 500-model sweep is five HTTP requests and finishes in a couple of seconds.

Agent skill: SKILL.md

https://api.apify.com/v2/key-value-stores/t7YoTxpZEJOWvw4Ug/records/huggingface-models-scraper.md

What you get

FieldWhat it holds
id, authorRepository ID (meta-llama/Llama-3.2-1B-Instruct) and the owning user or organisation
resourceTypemodels, datasets or spaces
urlDirect link to the repository page on the Hub
downloads, downloadsAllTimeDownloads in the last 30 days, and since the repo was created
likes, trendingScoreHub likes and the current trending score
pipelineTag, libraryNameTask (text-generation) and library (transformers), models only
sdkSpace runtime (gradio, streamlit, docker), spaces only
licenseLicence identifier, extracted from the repo's license: tag
tagsFull tag list: languages, datasets, arXiv IDs, base models, regions
gated, gatedTypeWhether access is restricted, and whether approval is auto or manual
private, createdAt, lastModifiedVisibility flag and both repository timestamps
scrapedAtWhen the run happened

RUN_SUMMARY in the key-value store holds per-run counts, the filters you used, and any query that failed.

Use cases

  • Model adoption tracking. Snapshot downloads and likes on a schedule, then diff over time.
  • Dataset discovery. Filter datasets by search term and read the size and modality tags.
  • Release watching. Search by organisation name and sort by createdAt to see what a lab just shipped.
  • Licence and SBOM audits. The license and gated columns tell you what you may actually ship.
  • Trend reporting. Sort by trendingScore for what the community is picking up this week.

Quick start

Top downloaded models for a term:

{
"searchQueries": ["llama"],
"maxResults": 100
}

Speech models only, most-liked first:

{
"searchQueries": ["whisper", "wav2vec"],
"resourceType": "models",
"pipelineTag": "automatic-speech-recognition",
"sortBy": "likes",
"maxResults": 50
}

The 200 most recently updated datasets on the Hub, no search filter:

{
"resourceType": "datasets",
"sortBy": "lastModified",
"maxResults": 200
}

Input

FieldTypeDefaultNotes
searchQueriesarray[]Free-text terms matched against repo names and metadata. Each term is a separate pass. Empty means walk the hub in sort order.
resourceTypestringmodelsmodels, datasets or spaces.
pipelineTagstring""Task filter, e.g. text-generation. Models only: it is ignored with a warning for datasets and spaces.
sortBystringdownloadsdownloads, likes, lastModified, createdAt or trendingScore. Always descending.
maxResultsinteger100Cap per search query, or on the single listing pass when no query is given.

searchQueries and pipelineTag combine with AND: a model has to match the term and carry the task. sortBy: "downloads" with resourceType: "spaces" falls back to likes, because spaces have no download counter.

Output example

{
"id": "meta-llama/Llama-3.2-1B-Instruct",
"resourceType": "models",
"author": "meta-llama",
"url": "https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct",
"downloads": 10016506,
"downloadsAllTime": 93318314,
"likes": 1554,
"trendingScore": 2,
"pipelineTag": "text-generation",
"libraryName": "transformers",
"sdk": null,
"license": "llama3.2",
"tags": ["transformers", "safetensors", "llama", "text-generation", "conversational", "license:llama3.2"],
"gated": true,
"gatedType": "manual",
"private": false,
"createdAt": "2024-09-18T15:12:47.000Z",
"lastModified": "2024-10-24T15:07:51.000Z",
"scrapedAt": "2026-08-06T11:35:15.076Z"
}

Finding a pipeline task

Open any model page on the Hub and look at the badge under the title. That string is the pipeline tag verbatim. The common ones:

TaskpipelineTag
Chat and completion modelstext-generation
Embeddingssentence-similarity or feature-extraction
Speech to textautomatic-speech-recognition
Image generationtext-to-image
Classificationtext-classification, image-classification

The Hub's own model list at huggingface.co/models shows the full task tree in its left sidebar.

API example

curl -X POST "https://api.apify.com/v2/acts/arman-bd~huggingface-models-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"searchQueries": ["llama"],
"pipelineTag": "text-generation",
"sortBy": "downloads",
"maxResults": 50
}'

JavaScript example

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('arman-bd/huggingface-models-scraper').call({
searchQueries: ['whisper'],
resourceType: 'models',
sortBy: 'likes',
maxResults: 25,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const repo of items) console.log(`${repo.id} - ${repo.downloads} downloads, ${repo.license ?? 'no licence'}`);

Notes

  • Gated repos are kept, not dropped. A gated model still returns full public metadata, flagged with gated: true and gatedType: "manual" so your counts stay honest. Only the weights sit behind the approval form, not the record.
  • Pagination is cursor-based. The Hub returns a Link header with rel="next" and no page numbers, so the Actor follows that cursor 100 records at a time until your cap is reached.
  • Licence comes from the tags. There is no licence field on the wire. It is encoded as license:apache-2.0 in the tag list and lifted out into its own column here.
  • One bad query won't kill the run. Failures are recorded in RUN_SUMMARY.failures, and the Actor only errors out if every query fails.
  • Rate limits are generous. The Hub advertises 500 requests per 5-minute window on an anonymous IP, and a 10,000-record sweep uses 100 of them.

FAQ

Do I need a Hugging Face token? No. The listing API is public. A token only matters for downloading gated weights, which this Actor never does.

Do I need a proxy? No. Proxy configuration is not required to run this Actor.

Can I get the model card text? Not from this Actor. The listing index carries metadata only, and the card is a separate file on the repo. Everything you get here is one request per 100 records.

Why is downloads null on spaces? Spaces are apps, not artefacts. The Hub tracks likes for them but not downloads, so sdk is the field that carries useful signal there instead.

How far back can I list? As far as you like. Sort by createdAt and raise maxResults, and the cursor keeps walking until the index is exhausted.

Can I plug it into something else? Yes. Apify API, the client libraries, webhooks, scheduled runs, dataset exports to JSON, CSV or Excel, or MCP. The output is structured JSON.