# Steam Workshop Collection & Mod Metadata (`coolinbex/steam-workshop-collection-mod-scraper`) Actor

Extract structured metadata, creators, tags, engagement counts, app IDs, dates, and collection membership from public Steam Workshop collections and mod pages. Built for mod discovery, catalog analysis, compatibility research, and scheduled change monitoring without login or private APIs.

- **URL**: https://apify.com/coolinbex/steam-workshop-collection-mod-scraper.md
- **Developed by:** [coolinbex](https://apify.com/coolinbex) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Steam Workshop Collection and Mod Intelligence Scraper

Build a structured catalog of public Steam Workshop content from collection pages and individual mod/item URLs. This Actor is designed for mod discovery, game-content research, creator analysis, collection auditing, compatibility work, and scheduled monitoring of public Workshop metadata.

### What it extracts

For each public Workshop item, the dataset can include:

- Published file ID and canonical public URL
- Mod/item title and creator
- Steam app/game ID
- Workshop tags
- Description excerpt
- Published and updated timestamps when exposed
- Subscriptions, favorites, positive votes, and negative votes
- The collection URLs containing the item
- Explicit `recordType` values for `item` and `error`
- Retrieval timestamp and actionable error text for failed pages

Counts, dates, creator names, and tags are best-effort public-page fields. Steam can change its markup or withhold fields from anonymous visitors; unavailable fields are returned as `null` rather than guessed.

### Input

Provide at least one `collectionUrls` or `itemUrls` entry. Collections are paginated with a bounded `maxPages`; item IDs are deduplicated before detail requests. Filtering is applied before item records are written.

```json
{
  "collectionUrls": [
    "https://steamcommunity.com/sharedfiles/filedetails/?id=123456789"
  ],
  "itemUrls": [
    "https://steamcommunity.com/sharedfiles/filedetails/?id=987654321"
  ],
  "appId": "730",
  "tags": ["Gameplay"],
  "titleContains": "weapon",
  "creatorContains": "studio",
  "maxItems": 100,
  "maxPages": 20,
  "requestDelayMs": 350,
  "retries": 3,
  "timeoutMs": 20000
}
```

#### Input controls

| Field | Purpose | Default / limit |
|---|---|---|
| `collectionUrls` | Public Steam Workshop collection pages | Up to the configured run limits |
| `itemUrls` | Direct public Workshop item pages | Deduplicated by published file ID |
| `appId` | Keep items for one Steam app/game ID | Optional numeric string |
| `tags` | Require every supplied tag | Optional |
| `titleContains` | Case-insensitive title filter | Optional |
| `creatorContains` | Case-insensitive creator filter | Optional |
| `maxItems` | Maximum item records | 100 by default |
| `maxPages` | Maximum pages per collection | 20 by default |
| `requestDelayMs` | Delay between public requests | 350 ms by default |
| `retries` | Retries for transient failures | 3 by default |
| `timeoutMs` | Per-request timeout | 20,000 ms by default |

### Output

The Actor publishes a dataset schema with a table view for file ID, title, creator, app ID, tags, subscriptions, and updated time. `recordType: "item"` identifies extracted Workshop metadata. `recordType: "error"` preserves a failed collection or item request with its URL, file ID when available, collection membership, and error message.

### Reliability and responsible use

The crawler is bounded, deduplicates IDs, retries only transient failures, uses a descriptive user agent, and applies a configurable delay. It does not log in, use private Steam APIs, solve CAPTCHAs, bypass access controls, or download Workshop files. Use public pages only, respect Steam’s terms, robots guidance, rate limits, copyright, and applicable law.

Steam may return an access/protection page or an unavailable item. Those cases are retained as explicit error records instead of being silently presented as valid metadata.

### Development

```text
npm ci
npm test
npm run lint
```

Tests use mocked HTML and fetch implementations and do not contact Steam.

# Actor input Schema

## `collectionUrls` (type: `array`):

Public steamcommunity.com/sharedfiles/filedetails/?id=... collection pages.

## `itemUrls` (type: `array`):

Public Workshop item URLs.

## `appId` (type: `string`):

Optional app ID filter.

## `maxItems` (type: `integer`):

Maximum Workshop item records to emit.

## `tags` (type: `array`):

Only items containing every supplied tag.

## `titleContains` (type: `string`):

Case-insensitive title filter.

## `creatorContains` (type: `string`):

Case-insensitive creator filter.

## `maxPages` (type: `integer`):

Maximum public collection pages to inspect.

## `requestDelayMs` (type: `integer`):

Delay between public Workshop requests.

## `retries` (type: `integer`):

Retries for transient public page failures.

## `timeoutMs` (type: `integer`):

Timeout for each public Workshop request.

## Actor input object example

```json
{
  "maxItems": 100,
  "maxPages": 20,
  "requestDelayMs": 350,
  "retries": 3,
  "timeoutMs": 20000
}
```

# Actor output Schema

## `records` (type: `string`):

Public Workshop item metadata and explicit request-error records.

## `summary` (type: `string`):

The Actor writes a run summary to the default key-value store when configured.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("coolinbex/steam-workshop-collection-mod-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("coolinbex/steam-workshop-collection-mod-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call coolinbex/steam-workshop-collection-mod-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,coolinbex/steam-workshop-collection-mod-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KZJdLbz9GAOBeAnK7/builds/Ortl1dipWhaRAXpNT/openapi.json
