Medium Scraper — Articles by Author, Publication or Tag
Pricing
from $2.10 / 1,000 story scrapeds
Medium Scraper — Articles by Author, Publication or Tag
Extract Medium articles with title, subtitle, author, publication, claps, responses, reading time, tags, member-only flag and optional full text of free stories. Works from any Medium profile, publication, tag or story URL.
Pricing
from $2.10 / 1,000 story scrapeds
Rating
0.0
(0)
Developer
Sergey Lutsak
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Extract Medium stories with full metadata — title, subtitle, author, publication, claps, responses, reading time, tags, member-only flag — from any author profile, publication, tag or single story URL. Optionally include the full text of free stories. No login, no cookies: only Medium's public endpoints, so runs are fast and cheap.
Use cases
- AI agents & RAG pipelines — pull the latest stories on a tag or from an author as clean JSON with plain-text content, ready for embedding or summarization.
- Content & marketing research — track what competitors, thought leaders or publications publish, how much engagement (claps, responses) each story gets and which tags they use.
- Newsletter and curation — monitor a set of authors and publications on a schedule and get new stories as they appear.
- Analytics — build datasets of publishing frequency, reading time and engagement per author or topic.
Input
| Field | Type | Default | Description |
|---|---|---|---|
startUrls | array | — | Medium URLs: https://medium.com/@username, https://username.medium.com, https://medium.com/publication-slug, https://medium.com/tag/tag-name, or a story URL |
maxItems | integer | 50 | Hard cap on stories across all start URLs — you are charged only for stories actually returned |
includeContent | boolean | false | Add content_text (plain-text paragraphs). Member-only stories return the public preview only |
tagSort | latest / top | latest | Sorting for tag URLs |
proxy | object | Apify datacenter proxy | Proxy configuration |
{"startUrls": [{ "url": "https://medium.com/@barackobama" },{ "url": "https://medium.com/tag/artificial-intelligence" }],"maxItems": 100,"includeContent": false}
Output (one item)
{"id": "5623e38aa89","url": "https://medium.com/the-atlantic/im-not-yet-ready-to-abandon-the-possibility-of-america-5623e38aa89","title": "I’m Not Yet Ready to Abandon the Possibility of America","subtitle": "I wrote my book for young people — as an invitation…","author": { "id": "9e422a605dc5", "name": "Barack Obama", "username": "barackobama", "url": "https://medium.com/@barackobama" },"publication": { "id": "3666ec396666", "name": "The Atlantic", "slug": "the-atlantic", "url": "https://medium.com/the-atlantic" },"published_at": "2020-11-16T23:24:00.730000Z","updated_at": "2020-11-16T23:24:00.730000Z","language": "en","reading_time_min": 7.17,"word_count": 1846,"claps": 10355,"responses": 68,"tags": ["politics", "america", "barack-obama", "election-2020"],"is_member_only": true,"preview_image_url": "https://miro.medium.com/v2/1*20aXAMdpczvmmMn-u6zzxw.jpeg","excerpt": "I wrote my book for young people — as an invitation…","source": "profile:barackobama","content_text": "…" // only when includeContent = true}
Field names are a stable contract: new fields may be added in minor versions; renames only in a major version (see Changelog).
Coverage & limits
- Profiles and publications are paginated — you can go back through an author's or publication's full history (up to
maxItems). - Tags are not paginated by Medium: a tag URL returns up to ~35 stories (latest or top + RSS), deduplicated.
- Publications on custom domains (e.g.
towardsdatascience.com) are not supported yet — use theirmedium.com/...URL if one exists. - Member-only stories: metadata is complete;
content_textcontains only the public preview andcontent_is_preview_onlyis set totrue. - Public data only. No login, no paywall bypass, no personal contact data.
Pricing
Pay only for results: $3.00 per 1,000 stories plus a negligible $0.00005 per run start. Platform usage is included — no compute bills, no proxy costs. A run that finds nothing costs (almost) nothing.
Discounts on paid Apify plans:
| Apify plan | Price per 1,000 stories |
|---|---|
| Free plan | $3.00 |
| Starter | $2.70 |
| Scale | $2.40 |
| Business | $2.10 |
Examples (Free plan price):
- 100 stories from an author → $0.30
- 1,000 stories on a tag over a month (scheduled) → $3.00
- 10,000 stories from 50 publications → $30.00
FAQ
Does it work with username.medium.com URLs? Yes — both medium.com/@username and subdomain profiles.
Why are some stories missing full text? They are member-only. The scraper returns the public preview and flags it; it never bypasses the paywall.
Can I get the newest stories on a tag every hour? Yes — schedule the actor with a tag URL and tagSort: latest.
How far back can I go? Profiles and publications: all the way back, subject to maxItems. Tags: ~35 most recent/top stories per run.
Is it stable? The scraper relies on Medium's public JSON endpoints, not on HTML layout, which makes it far less fragile than DOM-based scrapers.
API & integrations
Run it from code — Python:
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("vulcandata/medium-scraper").call(run_input={"startUrls": [{"url": "https://medium.com/tag/artificial-intelligence"}], "maxItems": 100})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
…or plain HTTP (returns the dataset items directly):
curl -X POST "https://api.apify.com/v2/acts/vulcandata~medium-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \-H "Content-Type: application/json" -d '{"startUrls": [{"url": "https://medium.com/tag/artificial-intelligence"}], "maxItems": 100}'
- No-code: connect to Make, Zapier, n8n, Google Sheets, Slack or webhooks from the actor's Integrations tab.
- Schedules: run daily/weekly from Schedules and get fresh data automatically.
- AI agents (MCP): add
vulcandata/medium-scraperto the Apify MCP server (mcp.apify.com) and let Claude, ChatGPT or Cursor call it as a tool. - Exports: JSON, CSV, Excel, XML, RSS — straight from the dataset.
More scrapers by VULCAN
Changelog
- 0.1 — initial release: profiles, publications, tags, single stories, optional full text.