# Medium Scraper - Articles by Tag, Author & Publication (`kmscrape/medium-articles-scraper`) Actor

Scrape Medium articles from any tag, author or publication: titles, subtitles, authors, dates, tags, images, and full text from author and publication feeds. Fast, cheap, no login. Great for content research and AI agents.

- **URL**: https://apify.com/kmscrape/medium-articles-scraper.md
- **Developed by:** [kfir messika](https://apify.com/kmscrape) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Medium Articles Scraper — Public RSS and article metadata

Collect recent stories from public Medium author, tag, and publication RSS feeds. The actor also accepts direct article URLs and tries to read structured article metadata and public text from each page. It does not log in to Medium.

### Use cases

- Track newly published writing by author, topic, or publication.
- Build reading lists and editorial research datasets.
- Supply article feeds to AI agents / MCP workflows.

### Input

```json
{
  "authors": [],
  "tags": ["artificial-intelligence"],
  "publications": [],
  "articleUrls": [],
  "maxArticlesPerSource": 5,
  "includeFullText": true
}
```

Authors accept an `@handle` or a Medium profile URL. Tags and publications accept a slug or Medium URL. For example, the prefilled input reads five items from the latest AI tag feed. If omitted, `maxArticlesPerSource` defaults to 50; RSS feeds currently expose about 10 latest items, so increasing the cap cannot add older RSS stories. Direct article URLs are fetched individually.

### Output

Charged rows have `type: "article"`; each row is charged as `article-result`. Free `summary` rows report a source's count and status, and free `error` rows report fetch or parsing failures.

#### Output fields that may be null

| Field | When it is populated | Why it can be null |
| --- | --- | --- |
| `claps`, `responses`, `readingTimeMin` | When an accessible article page exposes the value in JSON-LD or Apollo state | RSS does not include these values; a page may be blocked or omit them |
| `isMemberOnly` | When page data explicitly identifies member-only or free access | RSS does not report access status; blocked pages and pages without an explicit access marker leave it unknown |
| `text` | From `<content:encoded>` in RSS or a publicly accessible article body, when `includeFullText` is true | Tag RSS usually has only a short excerpt; a page may be blocked, omit its body, or mark it member-only |

Feed article pages are fetched with at most three concurrent requests. A blocked page does not remove the RSS row. Page requests have a 20 second timeout and retry HTTP 429 and 5xx responses with backoff.

Trimmed row from the public tag RSS feed (page-only fields are null when RSS does not expose them):

```json
{"type":"article","url":"https://medium.com/@khalidkhan3398/how-to-use-suno-ai-in-2026-songs-v6-the-new-speech-beta-and-pricing-6fe6069fc0bf","id":"6fe6069fc0bf","title":"How to Use Suno AI in 2026: Songs, v6, the New Speech Beta and Pricing","subtitle":"Learning how to use Suno is the quickest way to turn an idea, a poem or a few lines of lyrics into a finished song, and from this week…","author":{"name":"Kristen Belly ☄️","username":"khalidkhan3398","url":"https://medium.com/@khalidkhan3398"},"publication":null,"publishedAt":"2026-10-04T10:46:23.000Z","updatedAt":"2026-10-04T10:46:23.275Z","tags":["ai","artificial-intelligence","tech","technology","ai-agent"],"claps":null,"responses":null,"readingTimeMin":null,"isMemberOnly":null,"imageUrl":"https://cdn-images-1.medium.com/max/800/1*fOXOrIWD315DfhqJcM3ZFQ.webp","text":null,"source":"tag:artificial-intelligence","scrapedAt":"2026-10-04T11:00:39.917Z"}
```

### Pricing

$1.50 per 1,000 article rows, using the `article-result` pay-per-event in Apify Console. Summary and error rows are free. Local runs work without pay-per-event configuration.

### Limits

Medium RSS currently returns about 10 recent entries per author, tag, or publication feed. The unauthenticated Medium GraphQL endpoint returned HTTP 403 during verification, so archive pagination is not implemented. A live tag feed returned HTTP 200, but the article page returned HTTP 403 with a Cloudflare challenge from this runtime. The blocked response exposed no JSON-LD, Apollo state, or article meta tags, so this runtime could not verify which page fields that live article would otherwise expose. On accessible pages the parser checks JSON-LD, `window.__APOLLO_STATE__`, and meta tags. RSS `content:encoded` bodies are converted to plain text when full text is requested. Member-only article text is omitted when page metadata marks it as not freely accessible. Requests time out after 20 seconds, retry HTTP 429 and 5xx up to three times with backoff, and feed article requests run with concurrency up to three.

### FAQ

**Do I need a Medium account?** No. The actor requests public RSS and article pages only.

**Why are claps and responses null?** RSS does not include these values. They are populated only if an accessible article page exposes them in JSON-LD or Apollo state. A Cloudflare block leaves them null.

**Can I scrape more than the latest 10 posts?** Not from the public RSS feeds verified here. GraphQL archives were blocked with HTTP 403 and are not used.

**Is full text always available?** No. Some author and publication feeds include `content:encoded`; tag feeds often contain only an excerpt. Public page bodies are used when available. Pages can be blocked, and member-only text is excluded.

# Actor input Schema

## `authors` (type: `array`):

Medium @handles or author profile URLs.

## `tags` (type: `array`):

Medium topic tags to read from public RSS feeds.

## `publications` (type: `array`):

Medium publication slugs or publication URLs.

## `articleUrls` (type: `array`):

Individual Medium article URLs to fetch directly.

## `maxArticlesPerSource` (type: `integer`):

Maximum article results to emit for each source. RSS sources expose about 10 latest articles.

## `includeFullText` (type: `boolean`):

Include article body text when a public article page exposes it.

## Actor input object example

```json
{
  "authors": [],
  "tags": [
    "artificial-intelligence"
  ],
  "publications": [],
  "articleUrls": [],
  "maxArticlesPerSource": 5,
  "includeFullText": true
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "tags": [
        "artificial-intelligence"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("kmscrape/medium-articles-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "tags": ["artificial-intelligence"] }

# Run the Actor and wait for it to finish
run = client.actor("kmscrape/medium-articles-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "tags": [
    "artificial-intelligence"
  ]
}' |
apify call kmscrape/medium-articles-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,kmscrape/medium-articles-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aFsfPJpJzxR7s9XzL/builds/ZiL1pqbc2cE6obByS/openapi.json
