Astral Codex Ten (ACX) Scraper — Posts & Comments
Pricing
from $3.00 / 1,000 results
Astral Codex Ten (ACX) Scraper — Posts & Comments
Scrape Astral Codex Ten (astralcodexten.com) by Scott Alexander: post title, date, full text, tags, reactions and reader comments. Cross-run caching returns only new posts. Export to JSON, CSV or Excel, or call it as an API.
Pricing
from $3.00 / 1,000 results
Rating
0.0
(0)
Developer
ActorFlow
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Scrape Astral Codex Ten (astralcodexten.com), Scott Alexander's blog and the successor to Slate Star Codex. This Astral Codex Ten scraper extracts every post's title, subtitle, author, publish date, full text, word count, tags, cover image, reactions, comment and restack counts, and can collect the famous ACX comment threads with every reply. Export to JSON, CSV or Excel, or call it as an Astral Codex Ten API from Python, JavaScript or cURL. Paste the archive URL and press Start.

🔁 Only get new posts. Give a run a
cacheProjectNameand the scraper remembers every post it has already collected. Every later run with the same name skips those posts and returns only newly published ones, so a post is never scraped or billed twice. It makes this actor a low-cost Astral Codex Ten new-post monitor you can run on a schedule.
✨ Features of this Astral Codex Ten scraper
- Cross-run caching: scrape only new posts — name a cache project and every later run skips posts already collected, so scheduled runs return only fresh posts
- Full post extraction — title, subtitle, authors, publish date, body text, word count, tags, cover image and podcast URL
- ACX comments scraper — optionally collect every reader comment and reply, with author, date, text, reactions and thread position
- Engagement metrics — reaction, comment and restack counts for every post
- Paywall detection — subscriber-only posts are flagged with
isPaywalledand can be skipped entirely - Whole-archive pagination — walks the complete Astral Codex Ten archive until your item limit is reached
- Old links work —
astralcodexten.substack.comURLs are accepted as well asastralcodexten.com - No browser required — runs on plain HTTP requests, which makes it fast and cheap
- Proxy support — optional, and switched off by default
🔁 Scrape only new Astral Codex Ten posts with cross-run caching
Most blog scrapers download the whole archive again every time they run. This one can remember what it has already scraped. Set cacheProjectName to any name, for example acx-new-posts, and the actor keeps a list of every post it collects under that name in your Apify account. The next run with the same name:
- skips every post already collected, whether it came from the archive or a direct post URL
- counts only new posts toward
maxItems, somaxItems: 10means 10 posts you have not seen before - never re-fetches comments for a cached post
- saves the list only after a run finishes successfully. If a run fails part-way through, the next run fetches those posts again rather than missing them
How to monitor Astral Codex Ten for new posts:
- Enter
https://www.astralcodexten.com/archivein Start URLs. - Set
cacheProjectNameto a name you will reuse, such asacx-new-posts. - Add an Apify Schedule to run it daily.
- Each run's dataset now holds only the posts published since the last run. Connect it to Slack, email, Google Sheets or a webhook to get alerts for every new ACX post.
Leave cacheProjectName empty to scrape everything on every run.
🚀 How to scrape Astral Codex Ten in 5 steps
- Sign up for a free Apify account — includes $5 monthly credit.
- Open the actor page and click Try for free.
- Keep the default archive URL in Start URLs, or paste specific post URLs.
- Click Start and wait for the run to complete.
- Download results from the Output tab in JSON, CSV, or Excel format.
You can also run this actor via the Apify API or integrate it directly into your workflows using Zapier, Make, or n8n.
💰 How much does it cost to scrape Astral Codex Ten?
This actor uses pay-per-result billing based on the compute units a run consumes.
- New Apify accounts include $5 of free monthly credit.
- It runs on plain HTTP requests rather than a headless browser, so it costs significantly less to run than browser-based scrapers.
- Proxies are disabled by default, which keeps runs at their cheapest.
- Caching cuts repeat-run costs. With
cacheProjectNameset, posts already collected are skipped before they are fetched, so a scheduled run only pays for new posts.
🔧 Astral Codex Ten scraper input configuration
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
startUrls | array | — | https://www.astralcodexten.com/archive | Astral Codex Ten archive, home page or single post URLs. Page type is detected automatically. |
maxItems | integer | — | 5 | Maximum posts to scrape per start URL. Set to 0 for no limit. |
includePaywalled | boolean | — | true | Keep subscriber-only posts (with an isPaywalled flag) or skip them entirely. |
cacheProjectName | string | — | — | Cross-run cache. Reuse the same name and later runs skip already-scraped posts, returning only new ones. |
scrapeComments | boolean | — | false | Also collect each post's comments and replies. |
maxCommentsPerPost | integer | — | 50 | Maximum comments per post, best first. Set to 0 for all. |
proxyConfiguration | object | — | {"useApifyProxy": false} | Proxy settings. Off by default. |
Supported URL types:
- Archive —
https://www.astralcodexten.com/archive - Home page —
https://www.astralcodexten.com - Single post —
https://www.astralcodexten.com/p/king-ludd - Old Substack address —
https://astralcodexten.substack.com/p/king-ludd
📦 Astral Codex Ten scraper output data
Each result is a JSON object with the keys url, title, subtitle, slug, authors, publishedAt, audience, isPaywalled, type, description, body, wordCount, tags, coverImage, podcastUrl, reactionCount, commentCount and restackCount, plus a comments array when comment scraping is enabled. Each comment has id, parentId, depth, author, authorHandle, date, editedAt, body, isDeleted, reactionCount and replyCount. Replies keep a parentId and depth, so threads can be rebuilt from a flat CSV.
The dataset ships with three views: Overview, a compact table of title, authors, date and paywall status; Full post details, which adds the body text, tags and engagement counts; and Comments, which lists each post's comments.
Sample output (the second post was scraped with scrapeComments on):
[{"url": "https://www.astralcodexten.com/p/does-georgism-work-five-years-later","title": "Does Georgism Work? Five Years Later","subtitle": "A guest post by Lars Doucet","slug": "does-georgism-work-five-years-later","authors": ["Scott Alexander"],"publishedAt": "2026-09-24T03:16:49.656Z","audience": "everyone","isPaywalled": false,"type": "newsletter","description": "A guest post by Lars Doucet","body": "Hi, this is Lars Doucet, author of the book review of Henry George’s Progress and Poverty that won the first ACX book review contest , as well as the three-part follow-up guest post series, “ Does Georgism Work? ” A lot has happened since then, including land value tax (LVT) enablement laws passing this year in two U.S. states and the election of LVT-friendly national leaders in the UK and South K …","wordCount": 8069,"tags": [],"coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/c40c781e-9ade-4737-9d3e-2fb83913a579_1018x511.jpeg","podcastUrl": null,"reactionCount": 424,"commentCount": 465,"restackCount": 42},{"url": "https://www.astralcodexten.com/p/king-ludd","title": "King Ludd","subtitle": "...","slug": "king-ludd","authors": ["Scott Alexander"],"publishedAt": "2026-09-14T23:56:40.626Z","audience": "everyone","isPaywalled": false,"type": "newsletter","description": "...","body": "I.\n\nIn the gods’ great scramble for human followers, Nodens would seem to have lost decisively.\n\nWe know him only from vague inscriptions at two archaeological sites near the Welsh-English border. One might have been his temple. He seems to have been a Romano-British-Celtic god of . . . hunting? rivers? . . . worshipped around the Severn estuary between 100 and 400 AD. His cult may have been dispe …","wordCount": 2736,"tags": [],"coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/017c9c94-ad3d-490e-90b4-1306c0e97122_575x339.png","podcastUrl": null,"reactionCount": 548,"commentCount": 400,"restackCount": 32,"comments": [{"id": "337392616","parentId": null,"depth": 0,"author": "Mark","authorHandle": "mark7969","date": "2026-09-15T03:18:38.274Z","editedAt": "2026-09-15T03:20:36.087Z","body": "Who, then, is the modern incarnation of of Ludd?\n\nIt would need to be somebody who started out calling new technology into the world, and later regretted that and tried to put it back in the box.\n\nClearly, the modern ava …","isDeleted": false,"reactionCount": 17,"replyCount": 1},{"id": "337419613","parentId": "337392616","depth": 1,"author": "Shaked Koplewitz","authorHandle": "shakeddown","date": "2026-09-15T04:32:58.058Z","editedAt": null,"body": "\"Eliezer\" means \"god-helper\", explaining why despite himself he seems to be helping a silicon god come to being. …","isDeleted": false,"reactionCount": 5,"replyCount": 0}]}]
🐍 How to scrape Astral Codex Ten with Python, JavaScript or the API
Run the actor programmatically with the official Apify clients. Replace <YOUR_API_TOKEN> with the token from your Apify Console.
Python (pip install apify-client):
from apify_client import ApifyClientclient = ApifyClient("<YOUR_API_TOKEN>")run = client.actor("confidential_gnat/astralcodexten-scraper").call(run_input={"startUrls": [{"url": "https://www.astralcodexten.com/archive"}],"maxItems": 20,"cacheProjectName": "acx-new-posts",})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["publishedAt"], item["title"])
JavaScript (npm install apify-client):
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });const run = await client.actor('confidential_gnat/astralcodexten-scraper').call({startUrls: [{ url: 'https://www.astralcodexten.com/archive' }],maxItems: 5,scrapeComments: true,});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
cURL — start a run and wait for the dataset:
curl -X POST "https://api.apify.com/v2/acts/confidential_gnat~astralcodexten-scraper/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \-H "Content-Type: application/json" \-d '{"startUrls": [{"url": "https://www.astralcodexten.com/archive"}], "maxItems": 5}'
💡 What you can use Astral Codex Ten data for
- New-post alerts — get every new ACX post as it is published, using cross-run caching and a schedule
- Full archive backup — keep a searchable offline copy of Astral Codex Ten posts
- Comment and community research — study one of the internet's most active long-form comment sections
- Reading lists and trackers — build a feed of book reviews, open threads and essays by date
- Engagement analysis — compare reactions, comments and restacks across posts
- Research and training datasets — collect long-form essays with structured metadata
Readers, researchers, rationalist community members and media analysts use this data for archiving, discourse research and content analysis.
⚠️ Astral Codex Ten scraping limitations
- Subscriber-only posts — paid posts return only the short public preview, not the full text. They are flagged with
isPaywalledso you can filter them. The actor does not log in or bypass paywalls. - Comments on subscriber-only posts — the site hides them from non-subscribers, so
commentscomes back empty for those posts. - One site only — this actor scrapes Astral Codex Ten. For other Substack newsletters or Substack search, use the Substack Scraper.
❓ Frequently asked questions
Is it legal to scrape Astral Codex Ten?
This actor only collects data that is already publicly visible on astralcodexten.com — no login, paywall bypass, or private content is accessed. Scraping publicly available data is generally considered lawful (see hiQ Labs v. LinkedIn as precedent). You remain responsible for complying with the site's terms and applicable copyright law when republishing content.
How do I get only new Astral Codex Ten posts?
Set cacheProjectName and reuse the same name every run. The actor remembers every post it has collected under that name and skips them next time, so each run returns only newly published posts. Pair it with Apify Schedules for daily updates.
Can I scrape the whole Astral Codex Ten archive?
Yes. Set maxItems to 0 and use the archive URL. The actor pages through every post, newest first.
Can I scrape Astral Codex Ten comments?
Yes. Turn on scrapeComments and every post gets a comments array with each comment's author, date, text, reactions and replies. Use maxCommentsPerPost to cap how many are collected, best comments first, or 0 for all of them.
Does this scraper get paid-subscriber ACX posts?
No — it collects what a logged-out visitor sees. For paid posts that is the public preview, which the actor flags with isPaywalled: true. Set includePaywalled to false to skip them.
Does it work with old Slate Star Codex or Substack links?
It accepts both astralcodexten.com and the older astralcodexten.substack.com addresses. The old Slate Star Codex blog (slatestarcodex.com) is a different site and is not supported.
How do I scrape Astral Codex Ten with Python?
Install apify-client, then call the actor with the archive URL and iterate the dataset — see the Python example above. Each post comes back as structured JSON ready for pandas or a database.
🔗 Other actors you may find useful
- ✍️ Substack Scraper — Newsletter Posts, Comments & Search with AI — Scrape any Substack newsletter or search all of Substack by keyword: posts, full text, reactions and comments with free sentiment, plus optional AI summaries and cross-run caching.
- 🏠 University Living Housing Scraper — Scrapes student housing listings and property details from universityliving.com.
- 🍷 Total Wine Scraper — Scrape Total Wine & More (totalwine.com) wine, liquor and beer prices, sizes, ratings, reviews, badges, ABV, origin and taste profile from search, category or product URLs.
- 🇩🇪 German Imprint (Impressum) Scraper with AI Extraction — Finds the Impressum page on any German website and extracts the company's decision makers, legal name, address, email addresses, phone numbers, commercial register number and VAT ID as structured data using AI.
- ⭐ Google Play Store Reviews Scraper — Scrapes user reviews from Google Play Store apps (play.google.com) including review text, star rating, author, date, replies and optional sentiment tagging.
- 📜 Google Patents Scraper — Scrapes patent data from Google Patents (patents.google.com) by keyword or URL, including title, abstract, inventors, assignee, filing and publication dates, citations, figures and PDF links.
📰 Other news and article website scrapers
- 🗞️ Google News AI Scraper — Search Google News by keyword, optionally extract full article text and AI-generated summaries, and never re-scrape the same article twice across runs.
- 🦘 Sydney Morning Herald (SMH) News Scraper — Scrape news articles from The Sydney Morning Herald (smh.com.au) — headline, author, publish date, section, keywords, images and full public article text, with a paywall flag.
- 🇮🇩 Detik News Scraper — Scrapes news articles from Detik.com, including headline, author, publish date, category, images and full article text.
💬 Support & Contact
If you encounter any issues or have questions, please open an issue
You can also find more of our actors on the Actor Flow .