Steam Workshop Collection & Mod Metadata
Pricing
from $3.00 / 1,000 results
Steam Workshop Collection & Mod Metadata
Extract structured metadata, creators, tags, engagement counts, app IDs, dates, and collection membership from public Steam Workshop collections and mod pages. Built for mod discovery, catalog analysis, compatibility research, and scheduled change monitoring without login or private APIs.
Pricing
from $3.00 / 1,000 results
Rating
0.0
(0)
Developer
coolinbex
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
Steam Workshop Collection and Mod Intelligence Scraper
Build a structured catalog of public Steam Workshop content from collection pages and individual mod/item URLs. This Actor is designed for mod discovery, game-content research, creator analysis, collection auditing, compatibility work, and scheduled monitoring of public Workshop metadata.
What it extracts
For each public Workshop item, the dataset can include:
- Published file ID and canonical public URL
- Mod/item title and creator
- Steam app/game ID
- Workshop tags
- Description excerpt
- Published and updated timestamps when exposed
- Subscriptions, favorites, positive votes, and negative votes
- The collection URLs containing the item
- Explicit
recordTypevalues foritemanderror - Retrieval timestamp and actionable error text for failed pages
Counts, dates, creator names, and tags are best-effort public-page fields. Steam can change its markup or withhold fields from anonymous visitors; unavailable fields are returned as null rather than guessed.
Input
Provide at least one collectionUrls or itemUrls entry. Collections are paginated with a bounded maxPages; item IDs are deduplicated before detail requests. Filtering is applied before item records are written.
{"collectionUrls": ["https://steamcommunity.com/sharedfiles/filedetails/?id=123456789"],"itemUrls": ["https://steamcommunity.com/sharedfiles/filedetails/?id=987654321"],"appId": "730","tags": ["Gameplay"],"titleContains": "weapon","creatorContains": "studio","maxItems": 100,"maxPages": 20,"requestDelayMs": 350,"retries": 3,"timeoutMs": 20000}
Input controls
| Field | Purpose | Default / limit |
|---|---|---|
collectionUrls | Public Steam Workshop collection pages | Up to the configured run limits |
itemUrls | Direct public Workshop item pages | Deduplicated by published file ID |
appId | Keep items for one Steam app/game ID | Optional numeric string |
tags | Require every supplied tag | Optional |
titleContains | Case-insensitive title filter | Optional |
creatorContains | Case-insensitive creator filter | Optional |
maxItems | Maximum item records | 100 by default |
maxPages | Maximum pages per collection | 20 by default |
requestDelayMs | Delay between public requests | 350 ms by default |
retries | Retries for transient failures | 3 by default |
timeoutMs | Per-request timeout | 20,000 ms by default |
Output
The Actor publishes a dataset schema with a table view for file ID, title, creator, app ID, tags, subscriptions, and updated time. recordType: "item" identifies extracted Workshop metadata. recordType: "error" preserves a failed collection or item request with its URL, file ID when available, collection membership, and error message.
Reliability and responsible use
The crawler is bounded, deduplicates IDs, retries only transient failures, uses a descriptive user agent, and applies a configurable delay. It does not log in, use private Steam APIs, solve CAPTCHAs, bypass access controls, or download Workshop files. Use public pages only, respect Steam’s terms, robots guidance, rate limits, copyright, and applicable law.
Steam may return an access/protection page or an unavailable item. Those cases are retained as explicit error records instead of being silently presented as valid metadata.
Development
npm cinpm testnpm run lint
Tests use mocked HTML and fetch implementations and do not contact Steam.