# Medium Scraper - Articles by Tag, Author, Publication, Search (`unbrowseai/medium-articles-scraper`) Actor

Scrape Medium articles by tag, author, publication or keyword search: title, subtitle, author, publication, date, reading time, claps, responses, tags, member-only flag, image and optional full article text as Markdown. Half the usual price.

- **URL**: https://apify.com/unbrowseai/medium-articles-scraper.md
- **Developed by:** [Unbrowse AI](https://apify.com/unbrowseai) (community)
- **Categories:** News, Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Medium Scraper – Articles by Tag, Author, Publication or Search

Turn Medium into a clean dataset. Give it topic tags, authors, publications, keywords or single article links, and get one row per article with the title and subtitle, author name and profile, publication, publish and update dates, reading time, word count, clap count, number of responses, tags, cover image and whether the story is behind the member paywall. Switch on **Include full article text** to also get the body as Markdown and plain text.

Feeds are paginated, so you are not limited to the ten latest stories an RSS feed shows: ask for 500 articles from a tag or an author's whole back catalogue.

### What people use it for

- **Content research and trend spotting** – see which stories in a topic collect the most claps and responses this week, month or year.
- **Competitor and creator monitoring** – track what an author or publication posts and how each story performs.
- **Datasets for AI and NLP** – collect free articles as Markdown with tags and language for training, retrieval or summarisation.
- **Newsletters and aggregators** – pull the newest stories for a set of tags on a schedule.
- **Outreach** – find active writers in a niche with their profile links.

### How to use it

1. Add one or more inputs: **Tags** (`machine-learning`), **Authors** (`@karpathy`), **Publications** (`better-programming` or a custom domain such as `levelup.gitconnected.com`), **Search keywords** or **Article URLs**.
2. For tags, choose the **feed order**: newest, trending, or top stories of the week, month, year or all time.
3. Set **Max articles per input** (default 50).
4. Tick **Include full article text** if you need the body.
5. Run, then download JSON, CSV or Excel, or read the dataset through the API.

### Output example

```json
{
  "id": "a64152b37c35",
  "url": "https://karpathy.medium.com/software-2-0-a64152b37c35",
  "title": "Software 2.0",
  "subtitle": "I sometimes see people refer to neural networks as just “another tool in your machine learning toolbox”. They have some pros and cons, they…",
  "authorName": "Andrej Karpathy",
  "authorUsername": "karpathy",
  "authorUrl": "https://medium.com/@karpathy",
  "publicationName": null,
  "publishedAt": "2017-11-11T22:18:53.751Z",
  "readingTimeMinutes": 8.8,
  "wordCount": 2146,
  "claps": 61027,
  "responses": 194,
  "tags": ["machine-learning", "artificial-intelligence", "programming", "software-development", "future"],
  "imageUrl": "https://miro.medium.com/v2/resize:fit:1400/1*CHcu2L0NmAZwCpQgmS1ByA.jpeg",
  "isMemberOnly": false,
  "language": "en",
  "markdown": "## Software 2.0\n\nI sometimes see people refer to neural networks as ...",
  "contentIsPreview": false,
  "sourceType": "author",
  "sourceValue": "@karpathy"
}
```

`markdown`, `text` and `contentIsPreview` appear only with **Include full article text**. Member-only stories return the public preview, flagged with `contentIsPreview: true`.

### Pricing

You pay per article in the dataset. Inputs that fail (unknown tag, removed article) are listed in the dataset for free.

| Apify plan | Price per 1,000 articles |
|---|---|
| Free | $5.00 |
| Starter | $5.00 |
| Scale | $5.00 |
| Business | $5.00 |

Full text is included at no extra charge.

### FAQ

**Can it read member-only articles in full?** No. It works without a login, so paywalled stories come with metadata and the free preview only.

**Why can a tag return fewer articles than I asked for?** Some feeds are short, and an article that already appeared under another input in the same run is not repeated.

**Does it work with custom-domain publications?** Yes, enter the domain in Publications or paste article links from it.

**Legal note:** this Actor collects publicly available information. You are responsible for using the data in line with Medium's terms, copyright and privacy law in your jurisdiction.

# Actor input Schema

## `tags` (type: `array`):

Medium topic tags, as a slug (machine-learning), a name (Machine Learning) or a tag URL (medium.com/tag/...).

## `tagSort` (type: `string`):

Which feed to read for each tag: newest first, trending now, or top stories of a period.

## `authors` (type: `array`):

Medium usernames (@karpathy or karpathy) or profile URLs (medium.com/@name, name.medium.com). Returns the author's stories, newest first.

## `publications` (type: `array`):

Publication slugs (better-programming), URLs (medium.com/better-programming) or custom domains (levelup.gitconnected.com).

## `searchQueries` (type: `array`):

Keywords, searched like the Medium search box (e.g. rust async, product management).

## `articleUrls` (type: `array`):

Individual Medium article URLs, on medium.com, a subdomain or a publication's own domain.

## `maxItemsPerSource` (type: `integer`):

Upper limit for each tag, author, publication and search. Feeds are paginated past the 10 items an RSS feed shows.

## `includeContent` (type: `boolean`):

Adds the article body as Markdown and plain text. Free articles come in full; member-only articles give the public preview (contentIsPreview = true).

## Actor input object example

```json
{
  "tags": [
    "machine-learning"
  ],
  "tagSort": "NEW",
  "authors": [],
  "publications": [],
  "searchQueries": [],
  "articleUrls": [],
  "maxItemsPerSource": 10,
  "includeContent": false
}
```

# Actor output Schema

## `articles` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "tags": [
        "machine-learning"
    ],
    "maxItemsPerSource": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("unbrowseai/medium-articles-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "tags": ["machine-learning"],
    "maxItemsPerSource": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("unbrowseai/medium-articles-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "tags": [
    "machine-learning"
  ],
  "maxItemsPerSource": 10
}' |
apify call unbrowseai/medium-articles-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,unbrowseai/medium-articles-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/MgHDXcDpMjQ3FjEIj/builds/7tazKMbokwvrIZc2E/openapi.json
