# Medium Articles Scraper & Monitor (RSS, Markdown) (`dima_kadirovich/medium-articles-scraper`) Actor

Get Medium articles by tag, author, publication, or any Medium URL. Full text as clean Markdown, keyword and date filters, and an 'only new articles' mode for scheduled monitoring.

- **URL**: https://apify.com/dima\_kadirovich/medium-articles-scraper.md
- **Developed by:** [Cronexa Data Tools](https://apify.com/dima_kadirovich) (community)
- **Categories:** AI, News, MCP servers
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Medium Articles Scraper do?

Medium Articles Scraper gets articles from [Medium](https://medium.com) by **tag, author, publication, or any Medium link**, and gives you clean, structured data:

- 📝 **Full article text as clean Markdown** (and plain text) for authors and publications, which is ready for AI, RAG, and LLM pipelines
- 🔔 **"Only new articles" mode**: schedule it every hour or day, and each run returns only articles you haven't seen yet. You only pay for new articles.
- 🔍 **Keyword and date filters**: keep only articles about the topics you care about, from the last N days
- 🔗 **Paste any Medium link**: profiles, publications, tags, `someone.medium.com`, or custom-domain blogs like `netflixtechblog.com`
- 💸 **Low price**: **$3 per 1,000 articles**, with no subscription and no proxy costs

It reads Medium's official public RSS feeds, so it is **fast (seconds per run) and reliable**. It needs no login or cookies and doesn't break when Medium redesigns its website.

### What can I use it for?

- **Content monitoring**: get every new article about your company, product, or competitors, sent to Slack, email, or Google Sheets through Apify integrations
- **AI and RAG datasets**: collect expert articles from top engineering blogs as Markdown for your knowledge base
- **Newsletter curation**: find the freshest articles on a topic every morning
- **Research and trend analysis**: track which topics and tags authors write about over time
- **Author and publication tracking**: archive everything a writer or publication posts

### How do I use it?

1. Enter one or more **tags**, **authors**, **publications**, or **Medium links**.
2. (Optional) Add **keywords** or a **Published in the last N days** limit.
3. Click **Start**. Results appear in seconds.
4. For monitoring, turn on **Only new articles** and add a **Schedule** (for example, every hour).

#### Example input

```json
{
  "tags": ["artificial-intelligence", "kubernetes"],
  "users": ["@netflixtechblog"],
  "publications": ["better-programming"],
  "keywords": ["LLM", "agents"],
  "publishedWithinDays": 7,
  "onlyNew": true
}
```

### Output example

```json
{
  "id": "516d5a29b252",
  "title": "Trading a Cloud Identity for Your Own: Workload Attestation on Managed Compute",
  "url": "https://netflixtechblog.com/trading-a-cloud-identity-for-your-own-workload-attestation-on-managed-compute-516d5a29b252",
  "author": "Netflix Technology Blog",
  "publishedAt": "2026-09-25T16:01:02+00:00",
  "tags": ["identity", "aws-emr", "apache-spark", "cloud", "security"],
  "isFullContent": true,
  "contentMarkdown": "### Introduction\n\nOrganizations that have been around for a while usually run two identity systems side by side...",
  "contentText": "Introduction\n\nOrganizations that have been around for a while...",
  "excerpt": "Introduction Organizations that have been around for a while usually run two identity systems...",
  "wordCount": 1877,
  "readingTimeMinutes": 8,
  "imageUrl": "https://cdn-images-1.medium.com/max/1024/1*eIWpHw7J_6qZCuaffD0lQQ.png",
  "sourceType": "user",
  "source": "@netflixtechblog",
  "feedTitle": "Stories by Netflix Technology Blog on Medium",
  "scrapedAt": "2026-09-29T18:56:34+00:00"
}
```

You can download the data as JSON, CSV, Excel, or HTML, or get it through the Apify API.

### How much does it cost?

**$3 per 1,000 articles** ($0.003 per article), which is much cheaper than similar Medium scrapers. You pay only for articles saved to your dataset. Articles skipped by filters or by "only new" mode are free. Apify's free plan includes monthly credit, so you can try it at no cost.

### Input options

| Field | Description |
|---|---|
| `tags` | Medium topics, for example `machine-learning`. Spaces are fine. |
| `users` | Usernames (`@dhh`) or profile links |
| `publications` | Publication slugs (`better-programming`) or domains (`netflixtechblog.com`) |
| `urls` | Any Medium link. The type is detected automatically. |
| `keywords` | Keep articles containing at least one keyword (title, tags, or text) |
| `publishedWithinDays` | Keep articles from the last N days |
| `onlyNew` | Skip articles returned by previous runs with the same sources |
| `maxItems` | Max articles per run (default 1000) |
| `includeHtml` | Also save the original HTML |

### Limitations (please read)

- Medium's RSS feeds contain the **latest ~10 articles per feed**. To collect more over time, schedule the Actor with **Only new articles** turned on, and it will build a complete archive run by run.
- **Tag feeds contain a short excerpt**, not the full text (`isFullContent: false`). **Author and publication feeds contain the full text.**
- Member-only (paywalled) stories contain only what Medium puts in the RSS feed.

### FAQ

**Is it legal to scrape Medium?** This Actor reads Medium's public RSS feeds, which Medium provides so that people and apps can follow its content. It collects no private data and uses no login. Respect authors' copyright when you reuse article text.

**Something doesn't work?** Open an issue in the **Issues** tab and it will be fixed quickly.

### Use it from AI assistants (Claude, ChatGPT, Cursor)

AI agents can run this Actor as a tool through the [Apify MCP server](https://mcp.apify.com). Add this to your MCP client (Claude Desktop, Claude Code, Cursor, VS Code…) and sign in with Apify in the browser when asked:

```json
{
  "mcpServers": {
    "medium-articles-scraper": { "url": "https://mcp.apify.com?tools=dima_kadirovich/medium-articles-scraper" }
  }
}
```

Then just ask, for example:

- *"Get the 10 newest Medium articles tagged 'kubernetes' and summarize the main trends."*
- *"What has @netflixtechblog published this month? Give me the key points of each post."*

The agent fills in the input, runs the Actor and reads the results. You pay the same per-result price.

### More tools from Dima Data Tools

- [Medium Articles Scraper & Monitor](https://apify.com/dima_kadirovich/medium-articles-scraper): Medium articles by tag, author, or publication as clean Markdown, with "only new" monitoring
- [Bulk Image Downloader](https://apify.com/dima_kadirovich/bulk-image-downloader): every image from any web page, as download links or ZIP, with duplicates and icons removed
- [Website SEO Audit & Broken Link Checker](https://apify.com/dima_kadirovich/website-seo-audit): crawl a site, score every page 0–100, find broken links, and get a shareable HTML report
- [Website to Markdown Crawler for AI](https://apify.com/dima_kadirovich/website-to-markdown): any website as clean main-content Markdown, with RAG chunks, llms.txt, and a cheap "only changed pages" refresh mode
- [Website Tech Stack & Domain Lookup](https://apify.com/dima_kadirovich/tech-stack-domain-lookup): technologies, email provider, SPF/DMARC, SaaS tools, SSL expiry and WHOIS for any list of domains

# Actor input Schema

## `tags` (type: `array`):

Medium topics, for example `artificial-intelligence`, `kubernetes`, `startup`. Spaces are fine (`machine learning`). Tag feeds give a short excerpt of each article; use authors or publications for full text.

## `users` (type: `array`):

Medium usernames (`@dhh` or `dhh`) or profile links (`https://medium.com/@dhh`, `https://someone.medium.com`). Returns full article text.

## `publications` (type: `array`):

Publication slugs (`better-programming`) or domains/links (`netflixtechblog.com`). Returns full article text.

## `urls` (type: `array`):

Paste any Medium link and the Actor detects whether it is a tag, author, publication, or custom-domain blog.

## `keywords` (type: `array`):

Keep only articles whose title, tags, or text contain at least one of these words (case-insensitive). Leave empty to keep everything.

## `publishedWithinDays` (type: `integer`):

Keep only articles published within this many days. Leave empty for no date limit.

## `onlyNew` (type: `boolean`):

Remember which articles earlier runs already returned and skip them. Turn this on when you schedule the Actor (for example every hour), so each run gives you only fresh articles and you only pay for new ones.

## `maxItems` (type: `integer`):

Upper limit on articles saved in one run.

## `includeHtml` (type: `boolean`):

Also save the original article HTML next to the Markdown.

## `proxyConfiguration` (type: `object`):

Not needed in most cases. Enable only if Medium starts rate-limiting your runs.

## Actor input object example

```json
{
  "tags": [
    "artificial-intelligence"
  ],
  "onlyNew": false,
  "maxItems": 1000,
  "includeHtml": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `articles` (type: `string`):

One row per Medium article: title, author, date, tags, Markdown and plain text, URL.

## `summary` (type: `string`):

Articles saved per feed, feeds that failed, and articles skipped by filters or 'only new' mode.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "tags": [
        "artificial-intelligence"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("dima_kadirovich/medium-articles-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "tags": ["artificial-intelligence"] }

# Run the Actor and wait for it to finish
run = client.actor("dima_kadirovich/medium-articles-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "tags": [
    "artificial-intelligence"
  ]
}' |
apify call dima_kadirovich/medium-articles-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dima_kadirovich/medium-articles-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/teTw2CTA3qjL37GkQ/builds/FbcMugGf8t8H2ufMg/openapi.json
