# Coursera Courses Scraper (`automation-lab/coursera-course-catalog`) Actor

Search public Coursera courses by subject and export providers, course URLs, displayed ratings, level, duration bands and skills for recurring curriculum comparisons.

- **URL**: https://apify.com/automation-lab/coursera-course-catalog.md
- **Developed by:** [Automation Lab](https://apify.com/automation-lab) (community)
- **Categories:** Education
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.08 / 1,000 item extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Coursera Courses Scraper

Search public **Coursera courses** by subject and export course URLs, providers, displayed ratings, level, duration bands, and skills. Use the records for repeat curriculum comparisons, training-program research, and educational catalog analysis.

This Actor returns individual courses only. Specializations, professional certificates, degrees, lessons, and learner accounts are outside its scope. It does not download course videos, syllabi, reviews, or enrollment records.

### Who is it for?

- Learning and development teams comparing course providers for a skills program.
- Curriculum researchers building subject-specific course shortlists.
- Education analysts comparing course-card ratings and duration bands between periodic exports.

### Why use this Actor?

Get a consistent tabular export without manually copying search cards. Stable course IDs and canonical URLs make it practical to join snapshots. Each row includes the first matching input query and page so you can trace how it was discovered.

The Actor requests course-only results rather than mixing individual courses with certificate programs. It uses public catalog data and requires no Coursera account or cookies.

### Getting started

1. Enter one or more subjects in `queries`, such as `python` or `cybersecurity`.
2. Set `maxItems` to the total number of unique courses you want across the run.
3. Optionally adjust `maxPages` for each subject.
4. Run the Actor and inspect the Courses dataset view.
5. Download JSON, CSV, Excel, or another export format supported by Apify.

Example input:

```json
{"queries":["python"],"maxItems":10,"maxPages":10}
```

### Input parameters

| Field | Default | Meaning |
| --- | --- | --- |
| `queries` | `["python"]` | 1–10 subjects, each 1–200 characters. Trimmed; identical queries deduplicated case-sensitively. |
| `maxItems` | `10` | Global maximum unique courses, 1–1000. Earlier queries consume the limit first. |
| `maxPages` | `10` | Search pages per query, 1–100. Currently up to 12 courses per page. |

Zero and unlimited limits are not supported. Unknown input fields are rejected rather than silently ignored. Explicit course URLs, language filters, category URLs, and authentication inputs are not supported.

### Search relevance and coverage

Coursera's best-match search determines relevance; this is not a literal title filter, exact phrase match, or exhaustive subject taxonomy. A nonsense or very narrow query can still produce semantic recommendations. Inspect results before using them as a definitive curriculum shortlist.

Results follow upstream ordering. Pagination stops at the global item cap, the per-query page cap, or source exhaustion. Course IDs are deduplicated across all queries; overlapping subjects do not produce duplicate charged rows. A course found under several subjects is attributed only to its first query.

### Extracted data

| Field | Meaning |
| --- | --- |
| `courseId` | Coursera course search identifier. |
| `title` | Public course title. |
| `url` | Canonical course page URL. |
| `providers` | Array of institutions or companies. |
| `rating` | Displayed course-card rating, or null when unrated. |
| `ratingCount` | Displayed course-card rating/review count, or null. |
| `level` | Source difficulty enum, for example `BEGINNER`. |
| `duration` | Duration band such as `ONE_TO_THREE_MONTHS`, not exact hours. |
| `skills` | Skills associated with the search record. |
| `imageUrl` | Catalog thumbnail URL; image files are not downloaded. |
| `query`, `page` | First query and one-based page producing this course. |
| `scrapedAt` | UTC retrieval timestamp. |

Ratings reflect the public course card at retrieval time and can change. They are not individual review text or an endorsement of course quality.

### Output example

A representative public course record (some optional arrays shortened):

```json
{
  "courseId":"course~ejOz7RDUEei99hK0xs-tsg",
  "title":"Python for Data Science, AI & Development",
  "url":"https://www.coursera.org/learn/python-for-applied-data-science-ai",
  "providers":["IBM"],
  "rating":4.62,
  "ratingCount":43825,
  "level":"BEGINNER",
  "duration":"ONE_TO_THREE_MONTHS",
  "skills":["Python Programming","NumPy"],
  "query":"python",
  "page":1,
  "scrapedAt":"2026-10-01T14:00:00.000Z"
}
```

A separate `SUMMARY` record in the run's key-value store reports the count and requested queries. The default dataset contains course records only.

### How much does it cost to export Coursera courses?

Pay-per-event pricing includes a **$0.003 start fee** and a plan-dependent fee for each unique course saved. There is no separate charge for the summary, rejected rows, or duplicates.

The BRONZE item rate is **$0.0018 per course**. At that rate, 10 courses cost $0.021, 25 cost $0.048, and 100 cost $0.183. FREE is $0.00207, SILVER $0.001404, and GOLD/PLATINUM/DIAMOND $0.00108 per course. Spend tiers follow your qualifying aggregate monthly Apify Store spend, not this Actor's run volume. These are estimated customer charges, not payout guarantees. See the Pricing tab for your plan's current rate and set an Apify spending limit as appropriate.

### Integrations

Export CSV to a spreadsheet for provider comparisons. Use dataset JSON in a data warehouse and join snapshots on `courseId` to compare rating counts and catalog presence over time.

For recurring comparisons, schedule this Actor through Apify and store each completed run's dataset separately. Scheduling does not turn the Actor into a built-in alert service or historical database. A course missing from one capped search snapshot is not proof of removal from Coursera.

### API usage

Replace `APIFY_TOKEN` with your own token in your environment. Do not place credentials in shared documents.

```bash
curl -X POST -H "Authorization: Bearer $APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  'https://api.apify.com/v2/acts/automation-lab~coursera-course-catalog/run-sync-get-dataset-items' \
  -d '{"queries":["python"],"maxItems":10}'
```

JavaScript with the official client:

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation-lab/coursera-course-catalog')
  .call({ queries: ['python'], maxItems: 10 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

Python:

```python
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ['APIFY_TOKEN'])
run = client.actor('automation-lab/coursera-course-catalog').call(
    run_input={'queries': ['python'], 'maxItems': 10})
print(client.dataset(run['defaultDatasetId']).list_items().items)
```

### MCP use

After the Actor is available to your account, configure its scoped Apify MCP endpoint:

```bash
claude mcp add --transport http apify \
  "https://mcp.apify.com?tools=automation-lab/coursera-course-catalog"
```

For Claude Desktop, Cursor, and VS Code, use the equivalent client configuration:

```json
{"mcpServers":{"apify":{"url":"https://mcp.apify.com?tools=automation-lab/coursera-course-catalog"}}}
```

Authenticate through your client's supported Apify authentication flow. Example prompts:

- “Find ten public Coursera courses about Python and compare their providers and displayed ratings.”
- “Export twenty cybersecurity courses with their duration bands for my curriculum comparison.” Availability depends on your account access; the endpoint does not expose private learner data.

### Limits and failure behavior

The Actor accesses public English-route search pages without a proxy. Titles may still be multilingual. Catalog rankings and availability can vary by location and time. It does not promise complete inventory, exact course duration, pricing, or enrollment eligibility.

Transient network errors and 429/5xx responses have at most two HTTP retries. Challenge pages, malformed state, unexpected entity types, and broken pagination fail visibly instead of producing a misleading successful empty export. Partial output may remain if a later page fails; check run status before consuming it.

### Troubleshooting and FAQ

**Why did a narrow query return unrelated courses?**
Coursera uses semantic search and may recommend courses even for unmatched words. Review the shortlist manually or filter exported titles in your downstream workflow.

**Why did I get fewer than my item cap?**
The page cap, source exhaustion, duplicate records, or Apify spending limits can stop collection earlier. Increase `maxPages` only when more source pages exist.

**Why are some ratings null?**
Unrated courses do not have a meaningful rating. Null is not zero stars.

**Can I download videos or student information?**
No. This Actor only exports public course catalog records.

**What if a run fails?**
Inspect its log and status. Retry transient incidents later. A persistent upstream format change needs a parser repair, not a larger item limit.

### Legality and responsible use

Use public catalog data responsibly, respect source terms and applicable law, and do not assume public accessibility grants permission for every downstream use. Avoid collecting private learner data. This Actor is independent and not affiliated with Coursera or course providers.

No AI model processes your input or results. Coursera receives search terms and the runtime network address. Apify provides execution and dataset/KV/log storage; operating costs are included in the documented Actor charges, with no required paid third-party account. Run data follows your Apify account retention settings and remains until you delete it or platform retention removes it; this Actor creates no cross-run cache or external backup. Delete unwanted datasets, run logs and KV stores in Apify. Do not put confidential or personal information in queries. For help, open an issue through the Actor's Apify Support tab.

### Related Actors

For another public learning catalog, see [LinkedIn Learning Course Catalog](https://apify.com/automation-lab/linkedin-learning-course-catalog). Its source and coverage differ; it is not a substitute for Coursera course records.

# Changelog

This Actor's version history is a separate document: https://apify.com/automation-lab/coursera-course-catalog/changelog.md

# Actor input Schema

## `queries` (type: `array`):

1–10 nonempty subjects or search terms, processed in order. Whitespace is trimmed and identical queries deduplicated case-sensitively. Coursera best-match search decides relevance; this is not an exact phrase or title-only filter. Only courses are requested, not specializations or degrees.

## `maxItems` (type: `integer`):

Global maximum unique courses across all queries, in source order. Later queries may not run when earlier ones fill the limit. Default 10; 1–1000; zero/unlimited is not supported. Pagination and source coverage may yield fewer.

## `maxPages` (type: `integer`):

Maximum search pages fetched per query, starting at page 1. Default 10, range 1–100; stops earlier at source exhaustion or the global item limit. A page currently contains up to 12 courses. No unlimited catalog promise.

## Actor input object example

```json
{
  "queries": [
    "python"
  ],
  "maxItems": 10,
  "maxPages": 10
}
```

# Actor output Schema

## `dataset` (type: `string`):

Unique public Coursera course records.

## `summary` (type: `string`):

Retrieval count and input query summary.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "python"
    ],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("automation-lab/coursera-course-catalog").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["python"],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("automation-lab/coursera-course-catalog").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "python"
  ],
  "maxItems": 10
}' |
apify call automation-lab/coursera-course-catalog --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation-lab/coursera-course-catalog"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ypAKq4huMQZWFKZQz/builds/tTHiycffSGfVCcBro/openapi.json
