Hugging Face Scraper
Pricing
from $5.00 / 1,000 results
Hugging Face Scraper
Scrape Hugging Face models, datasets, Spaces & papers: downloads, likes, parameters, license, tasks, inference providers and model cards. Search or paste any URL.
Pricing
from $5.00 / 1,000 results
Rating
0.0
(0)
Developer
Tin
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 days ago
Last modified
Categories
Share
What does Hugging Face Scraper do?
Hugging Face Scraper extracts structured data about models, datasets, Spaces (AI apps) and research papers from the Hugging Face Hub, the largest open-source AI platform with millions of repositories.
Search with filters (task, library, license, language, model size, author...) or paste any huggingface.co URL, such as a filtered listing page, a model page, an organization profile, a collection or the Daily Papers page. You get clean JSON with downloads, likes, trending score, parameter count, license, base model, inference providers and their prices, README model cards and more.
It uses Hugging Face's official public API instead of a browser, so it is fast and cheap: thousands of results in seconds. Running on the Apify platform also gives you API access, scheduling, integrations (Google Sheets, Slack, Zapier, Make, webhooks), monitoring and storage out of the box.
Why use Hugging Face Scraper?
- AI market research and competitive intelligence: track which models, organizations and tasks are gaining traction, and compare download and like counts over time.
- Model selection: shortlist open models by task, size, license and available inference providers, including their price per million tokens, context length and speed.
- Trend monitoring: schedule a daily run with Created after = 1 day to get every new model, dataset or paper in your niche, straight into Slack or a spreadsheet.
- Lead generation and ecosystem mapping: list every model, dataset and Space published by a company or research lab.
- Research and datasets for ML: build catalogs of datasets by task, modality, language and size, or collect README model cards for LLM and RAG pipelines.
- Academic tracking: collect Daily Papers with upvotes, AI summaries, keywords, GitHub repos and stars.
How to scrape Hugging Face data
- Click Try for free (or Start) to open the Actor.
- Choose What to scrape: Models, Datasets, Spaces or Papers.
- Optionally enter search terms and filters, such as task Text Generation, library gguf and license apache-2.0.
- Or paste one or more Hugging Face URLs, e.g.
https://huggingface.co/models?pipeline_tag=text-generation&sort=trending. - Set Max results per search or URL and click Start.
- When the run finishes, download your data as JSON, CSV, Excel or HTML, or use the API.
Input
All options are on the Input tab. The most important ones:
| Field | Description |
|---|---|
| What to scrape | models, datasets, spaces or papers |
| Search terms | One or more keywords; each one is scraped separately |
| Sort by | Trending, most downloads, most likes, recently created or recently updated |
| Max results per search or URL | Limit per keyword or URL (0 = no limit) |
| Hugging Face URLs | Any huggingface.co listing, repo, profile, collection or papers URL |
| Filters | Author, task, library, language, license, tags, min/max parameters, Space SDK, papers date |
| Min downloads / Min likes / Created after | Keep only popular or recent results; Created after also accepts 7 days |
| Include README / file list / card metadata | Add model card text, repository files or raw card YAML |
| Hugging Face access token | Optional, for higher rate limits and your own private repos |
Example input for the 100 most downloaded small GGUF text-generation models:
{"resourceType": "models","task": "text-generation","library": "gguf","maxParameters": "8B","sort": "downloads","maxItems": 100}
Example input using URLs:
{"startUrls": [{ "url": "https://huggingface.co/models?pipeline_tag=text-to-image&sort=likes" },{ "url": "https://huggingface.co/meta-llama" },{ "url": "https://huggingface.co/papers/trending" }],"maxItems": 50}
Output
Each result is one item in the dataset. The Output tab has a table view for each type (Models, Datasets, Spaces, Papers). You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.
Example model (shortened):
{"type": "model","id": "Qwen/Qwen3-8B","author": "Qwen","url": "https://huggingface.co/Qwen/Qwen3-8B","task": "text-generation","library": "transformers","downloads": 12522822,"downloadsAllTime": 125437423,"likes": 2039,"trendingScore": 20,"parameters": 8190735360,"parameterSize": "8.19B","license": "apache-2.0","baseModels": ["Qwen/Qwen3-8B-Base"],"baseModelRelation": "finetune","arxivIds": ["2309.00071", "2505.09388"],"gated": false,"inferenceProviders": [{"provider": "nscale","status": "live","contextLength": 40960,"inputPricePer1M": 0.07,"outputPricePer1M": 0.18,"tokensPerSecond": 126.5}],"createdAt": "2025-04-27T03:42:21.000Z","lastModified": "2025-07-26T03:49:13.000Z"}
Example paper (shortened):
{"type": "paper","id": "2505.09388","title": "Qwen3 Technical Report","url": "https://huggingface.co/papers/2505.09388","pdfUrl": "https://arxiv.org/pdf/2505.09388","aiSummary": "Qwen3, a unified series of large language models, integrates thinking and non-thinking modes...","authors": ["An Yang", "Anfeng Li", "Baosong Yang"],"upvotes": 347,"githubRepo": "https://github.com/QwenLM/Qwen3","githubStars": 27658,"publishedAt": "2025-05-14T13:41:34.000Z"}
Data fields
| Type | Main fields |
|---|---|
| Models | id, author, url, task, library, downloads (30 days), downloadsAllTime, likes, trendingScore, parameters, parameterSize, tensorTypes, ggufContextLength, license, languages, baseModels, baseModelRelation, trainingDatasets, arxivIds, gated, inferenceProviders (provider, status, context length, input/output price per 1M tokens, tokens per second), tags, createdAt, lastModified |
| Datasets | id, author, url, description, downloads, downloadsAllTime, likes, trendingScore, tasks, taskIds, modalities, formats, libraries, sizeCategory, sizeBytes, languages, license, arxivIds, citation, gated, tags, createdAt, lastModified |
| Spaces | id, author, url, title, emoji, shortDescription, sdk, sdkVersion, appUrl, likes, trendingScore, status (running, paused...), hardware, models, datasets, license, tags, createdAt, lastModified |
| Papers | id (arXiv), title, url, arxivUrl, pdfUrl, summary, aiSummary, aiKeywords, authors, publishedAt, upvotes, numComments, githubRepo, githubStars, projectPage, organization |
| Optional | readme (Markdown model/dataset card), files (repository file list), cardData (raw card metadata) |
How much does it cost to scrape Hugging Face?
This Actor uses pay-per-event pricing: you pay a small fee when a run starts and a fixed price per result saved. Check the Pricing tab for current prices. At $2 per 1,000 results, the top 1,000 trending models cost about $2.
Because it calls the API directly instead of rendering web pages, runs are very light and finish quickly. You can set a maximum cost per run in the run options, and the Actor stops cleanly when it is reached.
Tips and advanced options
- Filtering is fastest when it matches the sort order. With Sort by = Most downloads and Min downloads = 10000, the scraper stops as soon as results drop below 10,000. With another sort order, it has to scan further to find matching items.
- Monitor new releases: sort by Recently created, set Created after to
1 dayand schedule the Actor daily. - Use URLs for complex filters. Set up filters on huggingface.co, then copy the URL from your browser into Hugging Face URLs.
- README and dataset file lists need one extra request per result, so runs with those options are slower. Leave them off when you only need metrics.
- Rate limits: Hugging Face limits anonymous API usage. The Actor waits and retries automatically, but a free Hugging Face token raises the limit for very large runs.
- Duplicates are removed: a repository found by several search terms is saved (and charged) only once per run.
FAQ, disclaimers, and support
Is it legal to scrape Hugging Face? This Actor only collects publicly available metadata through Hugging Face's official public API, the same data shown on the website. It does not collect personal data beyond public usernames and author names. Always review the Hugging Face Terms of Service and the licenses of the models and datasets you use, and consult a lawyer if unsure.
Why is the README empty for some models? Gated models, such as Meta Llama, only share files with accounts that accepted their license. Add a token from an account that has access to include their READMEs.
Why are inference prices missing for a single model URL? Hugging Face includes provider prices only in list results. Scrape the model through a search or listing URL to get pricing.
Does it download model weights or dataset files? No. It collects metadata and, optionally, README text and file names.
Found a bug or need a feature? Open an issue on the Issues tab, and it will be looked at quickly. Need a custom solution, such as tracking metrics over time or scraping other parts of Hugging Face? Get in touch through the Issues tab.