Hugging Face Scraper: Models, Datasets & Spaces
Pricing
from $0.37 / 1,000 repository scrapeds
Hugging Face Scraper: Models, Datasets & Spaces
Scrape Hugging Face models, datasets and spaces: downloads, likes, tags, pipeline task, licence, library and last-modified. No key, no proxy, no browser.
Pricing
from $0.37 / 1,000 repository scrapeds
Rating
0.0
(0)
Developer
Arman Hossain
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share

Lists models, datasets and spaces from the Hugging Face Hub. You get downloads, likes, trending score, tags, pipeline task, library, licence, gating status and both repository timestamps.
The Hub publishes its whole index through a public read API at huggingface.co/api, intended for programmatic use. This Actor reads that index directly, so there's no token, no browser and no proxies. A 500-model sweep is five HTTP requests and finishes in a couple of seconds.
Agent skill: SKILL.md
https://api.apify.com/v2/key-value-stores/t7YoTxpZEJOWvw4Ug/records/huggingface-models-scraper.md
What you get
| Field | What it holds |
|---|---|
id, author | Repository ID (meta-llama/Llama-3.2-1B-Instruct) and the owning user or organisation |
resourceType | models, datasets or spaces |
url | Direct link to the repository page on the Hub |
downloads, downloadsAllTime | Downloads in the last 30 days, and since the repo was created |
likes, trendingScore | Hub likes and the current trending score |
pipelineTag, libraryName | Task (text-generation) and library (transformers), models only |
sdk | Space runtime (gradio, streamlit, docker), spaces only |
license | Licence identifier, extracted from the repo's license: tag |
tags | Full tag list: languages, datasets, arXiv IDs, base models, regions |
gated, gatedType | Whether access is restricted, and whether approval is auto or manual |
private, createdAt, lastModified | Visibility flag and both repository timestamps |
scrapedAt | When the run happened |
RUN_SUMMARY in the key-value store holds per-run counts, the filters you used, and any query that failed.
Use cases
- Model adoption tracking. Snapshot downloads and likes on a schedule, then diff over time.
- Dataset discovery. Filter datasets by search term and read the size and modality tags.
- Release watching. Search by organisation name and sort by
createdAtto see what a lab just shipped. - Licence and SBOM audits. The
licenseandgatedcolumns tell you what you may actually ship. - Trend reporting. Sort by
trendingScorefor what the community is picking up this week.
Quick start
Top downloaded models for a term:
{"searchQueries": ["llama"],"maxResults": 100}
Speech models only, most-liked first:
{"searchQueries": ["whisper", "wav2vec"],"resourceType": "models","pipelineTag": "automatic-speech-recognition","sortBy": "likes","maxResults": 50}
The 200 most recently updated datasets on the Hub, no search filter:
{"resourceType": "datasets","sortBy": "lastModified","maxResults": 200}
Input
| Field | Type | Default | Notes |
|---|---|---|---|
searchQueries | array | [] | Free-text terms matched against repo names and metadata. Each term is a separate pass. Empty means walk the hub in sort order. |
resourceType | string | models | models, datasets or spaces. |
pipelineTag | string | "" | Task filter, e.g. text-generation. Models only: it is ignored with a warning for datasets and spaces. |
sortBy | string | downloads | downloads, likes, lastModified, createdAt or trendingScore. Always descending. |
maxResults | integer | 100 | Cap per search query, or on the single listing pass when no query is given. |
searchQueries and pipelineTag combine with AND: a model has to match the term and carry the task. sortBy: "downloads" with resourceType: "spaces" falls back to likes, because spaces have no download counter.
Output example
{"id": "meta-llama/Llama-3.2-1B-Instruct","resourceType": "models","author": "meta-llama","url": "https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct","downloads": 10016506,"downloadsAllTime": 93318314,"likes": 1554,"trendingScore": 2,"pipelineTag": "text-generation","libraryName": "transformers","sdk": null,"license": "llama3.2","tags": ["transformers", "safetensors", "llama", "text-generation", "conversational", "license:llama3.2"],"gated": true,"gatedType": "manual","private": false,"createdAt": "2024-09-18T15:12:47.000Z","lastModified": "2024-10-24T15:07:51.000Z","scrapedAt": "2026-08-06T11:35:15.076Z"}
Finding a pipeline task
Open any model page on the Hub and look at the badge under the title. That string is the pipeline tag verbatim. The common ones:
| Task | pipelineTag |
|---|---|
| Chat and completion models | text-generation |
| Embeddings | sentence-similarity or feature-extraction |
| Speech to text | automatic-speech-recognition |
| Image generation | text-to-image |
| Classification | text-classification, image-classification |
The Hub's own model list at huggingface.co/models shows the full task tree in its left sidebar.
API example
curl -X POST "https://api.apify.com/v2/acts/arman-bd~huggingface-models-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \-H "Content-Type: application/json" \-d '{"searchQueries": ["llama"],"pipelineTag": "text-generation","sortBy": "downloads","maxResults": 50}'
JavaScript example
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_TOKEN' });const run = await client.actor('arman-bd/huggingface-models-scraper').call({searchQueries: ['whisper'],resourceType: 'models',sortBy: 'likes',maxResults: 25,});const { items } = await client.dataset(run.defaultDatasetId).listItems();for (const repo of items) console.log(`${repo.id} - ${repo.downloads} downloads, ${repo.license ?? 'no licence'}`);
Notes
- Gated repos are kept, not dropped. A gated model still returns full public metadata, flagged with
gated: trueandgatedType: "manual"so your counts stay honest. Only the weights sit behind the approval form, not the record. - Pagination is cursor-based. The Hub returns a
Linkheader withrel="next"and no page numbers, so the Actor follows that cursor 100 records at a time until your cap is reached. - Licence comes from the tags. There is no licence field on the wire. It is encoded as
license:apache-2.0in the tag list and lifted out into its own column here. - One bad query won't kill the run. Failures are recorded in
RUN_SUMMARY.failures, and the Actor only errors out if every query fails. - Rate limits are generous. The Hub advertises 500 requests per 5-minute window on an anonymous IP, and a 10,000-record sweep uses 100 of them.
FAQ
Do I need a Hugging Face token? No. The listing API is public. A token only matters for downloading gated weights, which this Actor never does.
Do I need a proxy? No. Proxy configuration is not required to run this Actor.
Can I get the model card text? Not from this Actor. The listing index carries metadata only, and the card is a separate file on the repo. Everything you get here is one request per 100 records.
Why is downloads null on spaces? Spaces are apps, not artefacts. The Hub tracks likes for them but not downloads, so sdk is the field that carries useful signal there instead.
How far back can I list? As far as you like. Sort by createdAt and raise maxResults, and the cursor keeps walking until the index is exhausted.
Can I plug it into something else? Yes. Apify API, the client libraries, webhooks, scheduled runs, dataset exports to JSON, CSV or Excel, or MCP. The output is structured JSON.