# University Course Catalog Scraper (`misty_ravel/university-course-data`) Actor

Scrapes complete course catalogs from US universities in a single run. Every course with code, title, credit hours, full description and prerequisites. Filter by university and subject.

- **URL**: https://apify.com/misty\_ravel/university-course-data.md
- **Developed by:** [Brian Webster](https://apify.com/misty_ravel) (community)
- **Categories:** Education, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## University Course Catalog Scraper

Scrapes **complete course catalogs from US universities** in a single run — every course with its code, title, credit hours, full description and prerequisites — with a filter to pick which universities you want.

These catalogs share a common underlying publishing system, which is what lets one scraper cover many of them. None of them offer a bulk download; the data exists only as HTML, one subject page at a time.

### What you get

One row per course, with `institution` and `institutionName` on every row so a multi-university dataset stays sortable.

| Field | Example |
| --- | --- |
| `institution`, `institutionName` | `mit`, `Massachusetts Institute of Technology` |
| `courseKey` | `6.1000` |
| `subject`, `catalogNumber` | `6`, `1000` |
| `title` | `Introduction to Programming and Computer Science` |
| `credits` | `4-0-8 units` |
| `description`, `descriptionWordCount` | full catalog text, `64` |
| `prerequisites` | `MATH 151 with a grade of C or better` |
| `sourceUrl` | link back to the subject page |

### Input

```json
{
  "institutions": ["mit", "northeastern"],
  "subjects": ["CS"],
  "maxItemsPerInstitution": 500
}
```

| Field | Default | Notes |
| --- | --- | --- |
| `institutions` | all | keys from the **Universities** picker in the input; an unknown key fails immediately |
| `subjects` | all | codes differ between institutions |
| `maxItemsPerInstitution` | 0 (no limit) | cap per university |
| `requestDelayMs` | 600 | politeness delay |

Leaving `institutions` empty scrapes every available catalog, which is a multi-hour run. Pick a handful, or set `maxItemsPerInstitution` for a trial.

### Uses

- **Cross-institution curriculum comparison** — the reason this exists
- **Prerequisite-chain analysis** across universities
- **Transfer-credit and articulation** mapping between institutions
- **Program benchmarking** — what peers teach that you do not
- **Catalog change tracking** between academic years

### Reliability

Each institution is harvested independently, so one failing cannot lose the others — a failure is logged and recorded in the run summary rather than aborting the run. Individual subject pages that fail are skipped with a warning. `429` and `5xx` get bounded exponential backoff. Courses are de-duplicated per institution on `courseKey`.

A per-institution breakdown is written to the key-value store under `RUN_SUMMARY`, including any institution that returned zero courses — so a layout this parser has not seen shows up as a fact in the summary rather than as a quietly short dataset.

### A note on coverage

Every catalog offered in the **Universities** picker was fetched and parsed before being listed — an institution whose pages came back empty was dropped rather than shipped as a name that quietly returns nothing. Institutions not in the picker may still publish their catalog the same way.

The course markup is broadly shared across these catalogs, but the index path is **not** — it differs per institution, and there are several distinct title layouts and course-code conventions in the wild (`CS 101`, `ZACM-1000`, `ANTH101`, NYU's `ACCT-GB 2103`, MIT's `6.1000`). This actor handles all of them. Some institutions split undergraduate and graduate catalogs across two indexes, and both are crawled.

***

Not affiliated with or endorsed by any institution listed. All data is public information published by each university.

# Actor input Schema

## `institutions` (type: `array`):

Which universities to scrape. Pick the ones you need — each row is a billed item, so a handful of universities costs a fraction of the full set. To run every one you must also tick "Scrape every university".

## `scrapeAllInstitutions` (type: `boolean`):

Run every available catalog. This is a multi-hour run producing a very large dataset; leave it off and select universities instead unless you genuinely want all of it.

## `subjects` (type: `array`):

Subject codes to keep across every selected university, e.g. \["CS","MATH"]. Leave empty for all subjects. Note that codes differ between institutions.

## `maxItemsPerInstitution` (type: `integer`):

Stop after this many courses from each university. 0 means no limit.

## `requestDelayMs` (type: `integer`):

Politeness delay between requests. The default is deliberately conservative.

## `userAgent` (type: `string`):

Sent with every request. Please keep it identifying rather than impersonating a browser.

## Actor input object example

```json
{
  "institutions": [
    "mit",
    "utaustin"
  ],
  "scrapeAllInstitutions": false,
  "subjects": [],
  "maxItemsPerInstitution": 0,
  "requestDelayMs": 600,
  "userAgent": "university-catalog-scraper/1.0 (+https://apify.com/)"
}
```

# Actor output Schema

## `courses` (type: `string`):

One row per course: code, title, credits, description, prerequisites.

## `runSummary` (type: `string`):

Per-institution course counts, plus any institution that failed or returned nothing.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "institutions": [
        "mit",
        "utaustin"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("misty_ravel/university-course-data").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "institutions": [
        "mit",
        "utaustin",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("misty_ravel/university-course-data").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "institutions": [
    "mit",
    "utaustin"
  ]
}' |
apify call misty_ravel/university-course-data --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,misty_ravel/university-course-data"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/PM22FoyxOBOxEVxK5/builds/63DYTdQv6ChgTsjT7/openapi.json
