Hugging Face Scraper avatar

Hugging Face Scraper

Pricing

from $5.00 / 1,000 results

Go to Apify Store
Hugging Face Scraper

Hugging Face Scraper

Scrape Hugging Face models, datasets, Spaces & papers: downloads, likes, parameters, license, tasks, inference providers and model cards. Search or paste any URL.

Pricing

from $5.00 / 1,000 results

Rating

0.0

(0)

Developer

Tin

Tin

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Categories

Share

What does Hugging Face Scraper do?

Hugging Face Scraper extracts structured data about models, datasets, Spaces (AI apps) and research papers from the Hugging Face Hub, the largest open-source AI platform with millions of repositories.

Search with filters (task, library, license, language, model size, author...) or paste any huggingface.co URL, such as a filtered listing page, a model page, an organization profile, a collection or the Daily Papers page. You get clean JSON with downloads, likes, trending score, parameter count, license, base model, inference providers and their prices, README model cards and more.

It uses Hugging Face's official public API instead of a browser, so it is fast and cheap: thousands of results in seconds. Running on the Apify platform also gives you API access, scheduling, integrations (Google Sheets, Slack, Zapier, Make, webhooks), monitoring and storage out of the box.

Why use Hugging Face Scraper?

  • AI market research and competitive intelligence: track which models, organizations and tasks are gaining traction, and compare download and like counts over time.
  • Model selection: shortlist open models by task, size, license and available inference providers, including their price per million tokens, context length and speed.
  • Trend monitoring: schedule a daily run with Created after = 1 day to get every new model, dataset or paper in your niche, straight into Slack or a spreadsheet.
  • Lead generation and ecosystem mapping: list every model, dataset and Space published by a company or research lab.
  • Research and datasets for ML: build catalogs of datasets by task, modality, language and size, or collect README model cards for LLM and RAG pipelines.
  • Academic tracking: collect Daily Papers with upvotes, AI summaries, keywords, GitHub repos and stars.

How to scrape Hugging Face data

  1. Click Try for free (or Start) to open the Actor.
  2. Choose What to scrape: Models, Datasets, Spaces or Papers.
  3. Optionally enter search terms and filters, such as task Text Generation, library gguf and license apache-2.0.
  4. Or paste one or more Hugging Face URLs, e.g. https://huggingface.co/models?pipeline_tag=text-generation&sort=trending.
  5. Set Max results per search or URL and click Start.
  6. When the run finishes, download your data as JSON, CSV, Excel or HTML, or use the API.

Input

All options are on the Input tab. The most important ones:

FieldDescription
What to scrapemodels, datasets, spaces or papers
Search termsOne or more keywords; each one is scraped separately
Sort byTrending, most downloads, most likes, recently created or recently updated
Max results per search or URLLimit per keyword or URL (0 = no limit)
Hugging Face URLsAny huggingface.co listing, repo, profile, collection or papers URL
FiltersAuthor, task, library, language, license, tags, min/max parameters, Space SDK, papers date
Min downloads / Min likes / Created afterKeep only popular or recent results; Created after also accepts 7 days
Include README / file list / card metadataAdd model card text, repository files or raw card YAML
Hugging Face access tokenOptional, for higher rate limits and your own private repos

Example input for the 100 most downloaded small GGUF text-generation models:

{
"resourceType": "models",
"task": "text-generation",
"library": "gguf",
"maxParameters": "8B",
"sort": "downloads",
"maxItems": 100
}

Example input using URLs:

{
"startUrls": [
{ "url": "https://huggingface.co/models?pipeline_tag=text-to-image&sort=likes" },
{ "url": "https://huggingface.co/meta-llama" },
{ "url": "https://huggingface.co/papers/trending" }
],
"maxItems": 50
}

Output

Each result is one item in the dataset. The Output tab has a table view for each type (Models, Datasets, Spaces, Papers). You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

Example model (shortened):

{
"type": "model",
"id": "Qwen/Qwen3-8B",
"author": "Qwen",
"url": "https://huggingface.co/Qwen/Qwen3-8B",
"task": "text-generation",
"library": "transformers",
"downloads": 12522822,
"downloadsAllTime": 125437423,
"likes": 2039,
"trendingScore": 20,
"parameters": 8190735360,
"parameterSize": "8.19B",
"license": "apache-2.0",
"baseModels": ["Qwen/Qwen3-8B-Base"],
"baseModelRelation": "finetune",
"arxivIds": ["2309.00071", "2505.09388"],
"gated": false,
"inferenceProviders": [
{
"provider": "nscale",
"status": "live",
"contextLength": 40960,
"inputPricePer1M": 0.07,
"outputPricePer1M": 0.18,
"tokensPerSecond": 126.5
}
],
"createdAt": "2025-04-27T03:42:21.000Z",
"lastModified": "2025-07-26T03:49:13.000Z"
}

Example paper (shortened):

{
"type": "paper",
"id": "2505.09388",
"title": "Qwen3 Technical Report",
"url": "https://huggingface.co/papers/2505.09388",
"pdfUrl": "https://arxiv.org/pdf/2505.09388",
"aiSummary": "Qwen3, a unified series of large language models, integrates thinking and non-thinking modes...",
"authors": ["An Yang", "Anfeng Li", "Baosong Yang"],
"upvotes": 347,
"githubRepo": "https://github.com/QwenLM/Qwen3",
"githubStars": 27658,
"publishedAt": "2025-05-14T13:41:34.000Z"
}

Data fields

TypeMain fields
Modelsid, author, url, task, library, downloads (30 days), downloadsAllTime, likes, trendingScore, parameters, parameterSize, tensorTypes, ggufContextLength, license, languages, baseModels, baseModelRelation, trainingDatasets, arxivIds, gated, inferenceProviders (provider, status, context length, input/output price per 1M tokens, tokens per second), tags, createdAt, lastModified
Datasetsid, author, url, description, downloads, downloadsAllTime, likes, trendingScore, tasks, taskIds, modalities, formats, libraries, sizeCategory, sizeBytes, languages, license, arxivIds, citation, gated, tags, createdAt, lastModified
Spacesid, author, url, title, emoji, shortDescription, sdk, sdkVersion, appUrl, likes, trendingScore, status (running, paused...), hardware, models, datasets, license, tags, createdAt, lastModified
Papersid (arXiv), title, url, arxivUrl, pdfUrl, summary, aiSummary, aiKeywords, authors, publishedAt, upvotes, numComments, githubRepo, githubStars, projectPage, organization
Optionalreadme (Markdown model/dataset card), files (repository file list), cardData (raw card metadata)

How much does it cost to scrape Hugging Face?

This Actor uses pay-per-event pricing: you pay a small fee when a run starts and a fixed price per result saved. Check the Pricing tab for current prices. At $2 per 1,000 results, the top 1,000 trending models cost about $2.

Because it calls the API directly instead of rendering web pages, runs are very light and finish quickly. You can set a maximum cost per run in the run options, and the Actor stops cleanly when it is reached.

Tips and advanced options

  • Filtering is fastest when it matches the sort order. With Sort by = Most downloads and Min downloads = 10000, the scraper stops as soon as results drop below 10,000. With another sort order, it has to scan further to find matching items.
  • Monitor new releases: sort by Recently created, set Created after to 1 day and schedule the Actor daily.
  • Use URLs for complex filters. Set up filters on huggingface.co, then copy the URL from your browser into Hugging Face URLs.
  • README and dataset file lists need one extra request per result, so runs with those options are slower. Leave them off when you only need metrics.
  • Rate limits: Hugging Face limits anonymous API usage. The Actor waits and retries automatically, but a free Hugging Face token raises the limit for very large runs.
  • Duplicates are removed: a repository found by several search terms is saved (and charged) only once per run.

FAQ, disclaimers, and support

Is it legal to scrape Hugging Face? This Actor only collects publicly available metadata through Hugging Face's official public API, the same data shown on the website. It does not collect personal data beyond public usernames and author names. Always review the Hugging Face Terms of Service and the licenses of the models and datasets you use, and consult a lawyer if unsure.

Why is the README empty for some models? Gated models, such as Meta Llama, only share files with accounts that accepted their license. Add a token from an account that has access to include their READMEs.

Why are inference prices missing for a single model URL? Hugging Face includes provider prices only in list results. Scrape the model through a search or listing URL to get pricing.

Does it download model weights or dataset files? No. It collects metadata and, optionally, README text and file names.

Found a bug or need a feature? Open an issue on the Issues tab, and it will be looked at quickly. Need a custom solution, such as tracking metrics over time or scraping other parts of Hugging Face? Get in touch through the Issues tab.