# Coursera Courses Catalog Scraper (`bakos_bence/coursera-courses`) Actor

Scrape Coursera course catalog via the public courses API — titles, slugs, workload, languages, and partner IDs. Optional keyword filter. Unofficial; not affiliated with Coursera.

- **URL**: https://apify.com/bakos\_bence/coursera-courses.md
- **Developed by:** [Bakos Bence](https://apify.com/bakos_bence) (community)
- **Categories:** Developer tools, Automation, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.25 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

### What does Coursera Courses do?

**Scrape the public Coursera course catalog** via the **courses.v1 JSON API** — **no browser, no login, no instructor PII**. Each dataset row is one course with title, slug, description, photo, workload, primary languages, partner IDs, and a stable learn URL. Optional keyword filter (`query`) matches name/slug/description client-side. Export **JSON or CSV** via Apify API, schedule, or MCP.

**Unofficial.** This Actor is **not affiliated with, endorsed by, or maintained by** Coursera, Inc. It reads **public** catalog endpoints only.

**Dashboards:** connect **Power BI**, **Tableau**, **Looker Studio**, or **Qlik** to the dataset API — see *Export to Power BI, Tableau, Looker Studio & Qlik* near the end. **AI agents:** paste this Actor's Store page URL into an Apify-capable agent (MCP / Claude / Cursor with Apify), create an Apify account and API token once, and the agent can run the Actor and wire up the output.

#### Why use this Coursera course scraper?

- 🎓 Full public catalog (~24k courses) with offset pagination
- 🔍 Optional keyword filter without browser search pages
- 📊 Structured rows for edtech market research and curriculum tracking
- 🛡️ Direct HTTP (`curl_cffi`) — lightweight, no Playwright

#### How to scrape Coursera courses

1. Open this Actor in [Apify Console](https://console.apify.com/).
2. Click **Start** on the prefilled `python` query (`maxItems` 5).
3. Clear `query` for a full-catalog pull, or raise `maxItems` for larger dumps.
4. Export the dataset or pull via API.

```json
{
  "query": "python",
  "maxItems": 5,
  "proxyType": "none"
}
```

#### How much does it cost?

**Pay per event (PPE):** Actor start + each course **result** row. On Free, results are about **$2.49 / 1,000**; paid plans are lower (see the Pricing tab). Mid-band reliability vs a ~$12 leader and ~$1 floor Actors — public API path, no browser.

#### Input

| Field | Description |
|-------|-------------|
| `query` | Optional keyword (name/slug/description). Empty = full catalog |
| `maxItems` | Cap on course rows (preferred) |
| `pages` | Compat alias when `maxItems` omitted (~100 courses/page) |
| `proxyType` | Optional proxy (default none) |

#### Sample output

```json
{
  "courseId": "l31la3mKEe-zFg7heHyXOQ",
  "title": "Getting started with the Vertex AI Gemini 1.5 Pro Model",
  "slug": "googlecloud-getting-started-with-the-vertex-ai-gemini-1-5-pro-model-i43mr",
  "description": "This is a self-paced lab that takes place in the Google Cloud console.",
  "photoUrl": "https://d3njjcbhbojbot.cloudfront.net/...",
  "workload": "1 hour 30 minutes",
  "primaryLanguages": ["en"],
  "partnerIds": ["443"],
  "url": "https://www.coursera.org/learn/googlecloud-getting-started-with-the-vertex-ai-gemini-1-5-pro-model-i43mr",
  "scrapedAt": "2026-09-10T10:00:00+00:00",
  "sourceUrl": "https://www.coursera.org/api/courses.v1?..."
}
```

#### FAQ

##### Is this official Coursera software?

No. It is unofficial and not affiliated with Coursera. Catalog fields can change without notice.

##### Do I need a Coursera account?

No. This Actor uses the public catalog API only — no login, reviews, or instructor personal data.

##### Can AI agents call this?

Yes — Actor ID `bakos_bence/coursera-courses` via Apify API or [MCP](https://mcp.apify.com).

##### Something broke?

Open the **Issues** tab on this Actor's page.

#### Related Actors

| Actor | Use together when… |
|-------|-------------------|
| [Public Document Finder](https://apify.com/bakos_bence/public-document-finder) | You need broader public document / research search |
| [Domain Typosquat Scanner](https://apify.com/bakos_bence/domain-typosquat-scanner) | You need brand/domain risk checks alongside course catalog work |

Profile: [bakos\_bence on Apify](https://apify.com/bakos_bence).

#### For AI agents & LLM apps

Paste this Actor's **Apify Store URL** into an agent that can use Apify (MCP, Claude, Cursor, etc.). With an Apify account and API token, the agent can set input, run the Actor, and fetch the dataset.

#### Export to Power BI, Tableau, Looker Studio & Qlik

1. Run the Actor (or open a finished run).
2. Open the **Dataset** → **API** / **Export**.
3. Copy the dataset items URL (JSON or CSV).
4. In your BI tool, add a **web / API** data source with that URL (and your Apify token as a header or query param if required).
5. Refresh on a schedule that matches your catalog monitoring cadence.

Course data is public catalog information — always verify critical decisions against Coursera’s own site.

# Actor input Schema

## `query` (type: `string`):

Optional keyword filter (case-insensitive) against course name, slug, and description. Empty string = full catalog scrape up to maxItems.

## `maxItems` (type: `integer`):

Cap course rows returned. Leave at 5 for a cheap first run. Preferred over pages.

## `pages` (type: `integer`):

Optional competitor-compat alias. Used only when maxItems is omitted; converts to maxItems ≈ pages × 100.

## `proxyType` (type: `string`):

Leave on none — Coursera courses.v1 is a public JSON API. Enable Apify Proxy only if rate-limited.

## `proxyConfiguration` (type: `object`):

Used only when Proxy is Apify Proxy.

## Actor input object example

```json
{
  "query": "python",
  "maxItems": 5,
  "proxyType": "none",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

Default dataset items.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "python",
    "maxItems": 5,
    "proxyType": "none"
};

// Run the Actor and wait for it to finish
const run = await client.actor("bakos_bence/coursera-courses").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "python",
    "maxItems": 5,
    "proxyType": "none",
}

# Run the Actor and wait for it to finish
run = client.actor("bakos_bence/coursera-courses").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "python",
  "maxItems": 5,
  "proxyType": "none"
}' |
apify call bakos_bence/coursera-courses --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bakos_bence/coursera-courses"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pWnfoCEM3cCLgCqho/builds/YzmihtYOICSER1fl7/openapi.json
