# Astral Codex Ten (ACX) Scraper — Posts & Comments (`confidential_gnat/astralcodexten-scraper`) Actor

Scrape Astral Codex Ten (astralcodexten.com) by Scott Alexander: post title, date, full text, tags, reactions and reader comments. Cross-run caching returns only new posts. Export to JSON, CSV or Excel, or call it as an API.

- **URL**: https://apify.com/confidential\_gnat/astralcodexten-scraper.md
- **Developed by:** [ActorFlow](https://apify.com/confidential_gnat) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Astral Codex Ten (ACX) Scraper — Posts & Comments

**Scrape Astral Codex Ten (astralcodexten.com)**, Scott Alexander's blog and the successor to Slate Star Codex. This **Astral Codex Ten scraper** extracts every post's title, subtitle, author, publish date, full text, word count, tags, cover image, reactions, comment and restack counts, and can collect the famous **ACX comment threads** with every reply. Export to JSON, CSV or Excel, or call it as an **Astral Codex Ten API** from Python, JavaScript or cURL. Paste the archive URL and press Start.

![Astral Codex Ten scraper exporting posts and comments to CSV, JSON and XML](./assets/featureimage.png)

> 🔁 **Only get new posts.** Give a run a `cacheProjectName` and the scraper remembers every post it has already collected. Every later run with the same name **skips those posts and returns only newly published ones**, so a post is never scraped or billed twice. It makes this actor a low-cost **Astral Codex Ten new-post monitor** you can run on a schedule.

### ✨ Features of this Astral Codex Ten scraper

- **Cross-run caching: scrape only new posts** — name a cache project and every later run skips posts already collected, so scheduled runs return only fresh posts
- **Full post extraction** — title, subtitle, authors, publish date, body text, word count, tags, cover image and podcast URL
- **ACX comments scraper** — optionally collect every reader comment and reply, with author, date, text, reactions and thread position
- **Engagement metrics** — reaction, comment and restack counts for every post
- **Paywall detection** — subscriber-only posts are flagged with `isPaywalled` and can be skipped entirely
- **Whole-archive pagination** — walks the complete Astral Codex Ten archive until your item limit is reached
- **Old links work** — `astralcodexten.substack.com` URLs are accepted as well as `astralcodexten.com`
- **No browser required** — runs on plain HTTP requests, which makes it fast and cheap
- **Proxy support** — optional, and switched off by default

### 🔁 Scrape only new Astral Codex Ten posts with cross-run caching

Most blog scrapers download the whole archive again every time they run. This one can **remember what it has already scraped**. Set `cacheProjectName` to any name, for example `acx-new-posts`, and the actor keeps a list of every post it collects under that name in your Apify account. The next run with the same name:

- **skips every post already collected**, whether it came from the archive or a direct post URL
- **counts only new posts toward `maxItems`**, so `maxItems: 10` means 10 posts you have not seen before
- **never re-fetches comments** for a cached post
- **saves the list only after a run finishes successfully**. If a run fails part-way through, the next run fetches those posts again rather than missing them

**How to monitor Astral Codex Ten for new posts:**

1. Enter `https://www.astralcodexten.com/archive` in **Start URLs**.
2. Set `cacheProjectName` to a name you will reuse, such as `acx-new-posts`.
3. Add an [Apify Schedule](https://docs.apify.com/platform/schedules) to run it daily.
4. Each run's dataset now holds **only the posts published since the last run**. Connect it to Slack, email, Google Sheets or a webhook to get alerts for every new ACX post.

Leave `cacheProjectName` empty to scrape everything on every run.

### 🚀 How to scrape Astral Codex Ten in 5 steps

1. [Sign up](https://apify.com/sign-up) for a free Apify account — includes **$5 monthly credit**.
2. Open the actor page and click **Try for free**.
3. Keep the default archive URL in **Start URLs**, or paste specific post URLs.
4. Click **Start** and wait for the run to complete.
5. Download results from the **Output** tab in JSON, CSV, or Excel format.

You can also run this actor via the [Apify API](https://docs.apify.com/api/v2) or integrate it directly into your workflows using [Zapier](https://zapier.com/apps/apify), [Make](https://www.make.com/), or [n8n](https://n8n.io/).

### 💰 How much does it cost to scrape Astral Codex Ten?

This actor uses **pay-per-result** billing based on the compute units a run consumes.

- New Apify accounts include **$5 of free monthly credit**.
- It runs on plain HTTP requests rather than a headless browser, so it costs significantly less to run than browser-based scrapers.
- Proxies are disabled by default, which keeps runs at their cheapest.
- **Caching cuts repeat-run costs.** With `cacheProjectName` set, posts already collected are skipped before they are fetched, so a scheduled run only pays for new posts.

### 🔧 Astral Codex Ten scraper input configuration

| Field                | Type    | Required | Default                                  | Description                                                                                                  |
| -------------------- | ------- | -------- | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| `startUrls`          | array   | —        | `https://www.astralcodexten.com/archive` | Astral Codex Ten archive, home page or single post URLs. Page type is detected automatically.                |
| `maxItems`           | integer | —        | `5`                                      | Maximum posts to scrape **per start URL**. Set to `0` for no limit.                                          |
| `includePaywalled`   | boolean | —        | `true`                                   | Keep subscriber-only posts (with an `isPaywalled` flag) or skip them entirely.                               |
| `cacheProjectName`   | string  | —        | —                                        | **Cross-run cache.** Reuse the same name and later runs skip already-scraped posts, returning only new ones. |
| `scrapeComments`     | boolean | —        | `false`                                  | Also collect each post's comments and replies.                                                               |
| `maxCommentsPerPost` | integer | —        | `50`                                     | Maximum comments per post, best first. Set to `0` for all.                                                   |
| `proxyConfiguration` | object  | —        | `{"useApifyProxy": false}`               | Proxy settings. Off by default.                                                                              |

**Supported URL types:**

- Archive — `https://www.astralcodexten.com/archive`
- Home page — `https://www.astralcodexten.com`
- Single post — `https://www.astralcodexten.com/p/king-ludd`
- Old Substack address — `https://astralcodexten.substack.com/p/king-ludd`

### 📦 Astral Codex Ten scraper output data

Each result is a JSON object with the keys `url`, `title`, `subtitle`, `slug`, `authors`, `publishedAt`, `audience`, `isPaywalled`, `type`, `description`, `body`, `wordCount`, `tags`, `coverImage`, `podcastUrl`, `reactionCount`, `commentCount` and `restackCount`, plus a `comments` array when comment scraping is enabled. Each comment has `id`, `parentId`, `depth`, `author`, `authorHandle`, `date`, `editedAt`, `body`, `isDeleted`, `reactionCount` and `replyCount`. Replies keep a `parentId` and `depth`, so threads can be rebuilt from a flat CSV.

The dataset ships with three views: **Overview**, a compact table of title, authors, date and paywall status; **Full post details**, which adds the body text, tags and engagement counts; and **Comments**, which lists each post's comments.

**Sample output** (the second post was scraped with `scrapeComments` on):

```json
[
    {
        "url": "https://www.astralcodexten.com/p/does-georgism-work-five-years-later",
        "title": "Does Georgism Work? Five Years Later",
        "subtitle": "A guest post by Lars Doucet",
        "slug": "does-georgism-work-five-years-later",
        "authors": ["Scott Alexander"],
        "publishedAt": "2026-09-24T03:16:49.656Z",
        "audience": "everyone",
        "isPaywalled": false,
        "type": "newsletter",
        "description": "A guest post by Lars Doucet",
        "body": "Hi, this is Lars Doucet, author of the book review of Henry George’s Progress and Poverty that won the first ACX book review contest , as well as the three-part follow-up guest post series, “ Does Georgism Work? ” A lot has happened since then, including land value tax (LVT) enablement laws passing this year in two U.S. states and the election of LVT-friendly national leaders in the UK and South K …",
        "wordCount": 8069,
        "tags": [],
        "coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/c40c781e-9ade-4737-9d3e-2fb83913a579_1018x511.jpeg",
        "podcastUrl": null,
        "reactionCount": 424,
        "commentCount": 465,
        "restackCount": 42
    },
    {
        "url": "https://www.astralcodexten.com/p/king-ludd",
        "title": "King Ludd",
        "subtitle": "...",
        "slug": "king-ludd",
        "authors": ["Scott Alexander"],
        "publishedAt": "2026-09-14T23:56:40.626Z",
        "audience": "everyone",
        "isPaywalled": false,
        "type": "newsletter",
        "description": "...",
        "body": "I.\n\nIn the gods’ great scramble for human followers, Nodens would seem to have lost decisively.\n\nWe know him only from vague inscriptions at two archaeological sites near the Welsh-English border. One might have been his temple. He seems to have been a Romano-British-Celtic god of . . . hunting? rivers? . . . worshipped around the Severn estuary between 100 and 400 AD. His cult may have been dispe …",
        "wordCount": 2736,
        "tags": [],
        "coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/017c9c94-ad3d-490e-90b4-1306c0e97122_575x339.png",
        "podcastUrl": null,
        "reactionCount": 548,
        "commentCount": 400,
        "restackCount": 32,
        "comments": [
            {
                "id": "337392616",
                "parentId": null,
                "depth": 0,
                "author": "Mark",
                "authorHandle": "mark7969",
                "date": "2026-09-15T03:18:38.274Z",
                "editedAt": "2026-09-15T03:20:36.087Z",
                "body": "Who, then, is the modern incarnation of of Ludd?\n\nIt would need to be somebody who started out calling new technology into the world, and later regretted that and tried to put it back in the box.\n\nClearly, the modern ava …",
                "isDeleted": false,
                "reactionCount": 17,
                "replyCount": 1
            },
            {
                "id": "337419613",
                "parentId": "337392616",
                "depth": 1,
                "author": "Shaked Koplewitz",
                "authorHandle": "shakeddown",
                "date": "2026-09-15T04:32:58.058Z",
                "editedAt": null,
                "body": "\"Eliezer\" means \"god-helper\", explaining why despite himself he seems to be helping a silicon god come to being. …",
                "isDeleted": false,
                "reactionCount": 5,
                "replyCount": 0
            }
        ]
    }
]
```

### 🐍 How to scrape Astral Codex Ten with Python, JavaScript or the API

Run the actor programmatically with the official Apify clients. Replace `<YOUR_API_TOKEN>` with the token from your [Apify Console](https://console.apify.com/account/integrations).

**Python** (`pip install apify-client`):

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")

run = client.actor("confidential_gnat/astralcodexten-scraper").call(run_input={
    "startUrls": [{"url": "https://www.astralcodexten.com/archive"}],
    "maxItems": 20,
    "cacheProjectName": "acx-new-posts",
})

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["publishedAt"], item["title"])
```

**JavaScript** (`npm install apify-client`):

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });

const run = await client.actor('confidential_gnat/astralcodexten-scraper').call({
    startUrls: [{ url: 'https://www.astralcodexten.com/archive' }],
    maxItems: 5,
    scrapeComments: true,
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

**cURL** — start a run and wait for the dataset:

```bash
curl -X POST "https://api.apify.com/v2/acts/confidential_gnat~astralcodexten-scraper/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"startUrls": [{"url": "https://www.astralcodexten.com/archive"}], "maxItems": 5}'
```

### 💡 What you can use Astral Codex Ten data for

- **New-post alerts** — get every new ACX post as it is published, using cross-run caching and a schedule
- **Full archive backup** — keep a searchable offline copy of Astral Codex Ten posts
- **Comment and community research** — study one of the internet's most active long-form comment sections
- **Reading lists and trackers** — build a feed of book reviews, open threads and essays by date
- **Engagement analysis** — compare reactions, comments and restacks across posts
- **Research and training datasets** — collect long-form essays with structured metadata

Readers, researchers, rationalist community members and media analysts use this data for archiving, discourse research and content analysis.

### ⚠️ Astral Codex Ten scraping limitations

- **Subscriber-only posts** — paid posts return only the short public preview, not the full text. They are flagged with `isPaywalled` so you can filter them. The actor does not log in or bypass paywalls.
- **Comments on subscriber-only posts** — the site hides them from non-subscribers, so `comments` comes back empty for those posts.
- **One site only** — this actor scrapes Astral Codex Ten. For other Substack newsletters or Substack search, use the [Substack Scraper](https://apify.com/confidential_gnat/substack-newsletter-scraper).

### ❓ Frequently asked questions

#### Is it legal to scrape Astral Codex Ten?

This actor only collects data that is already publicly visible on astralcodexten.com — no login, paywall bypass, or private content is accessed. Scraping publicly available data is generally considered lawful (see *hiQ Labs v. LinkedIn* as precedent). You remain responsible for complying with the site's terms and applicable copyright law when republishing content.

#### How do I get only new Astral Codex Ten posts?

Set `cacheProjectName` and reuse the same name every run. The actor remembers every post it has collected under that name and skips them next time, so each run returns only newly published posts. Pair it with [Apify Schedules](https://docs.apify.com/platform/schedules) for daily updates.

#### Can I scrape the whole Astral Codex Ten archive?

Yes. Set `maxItems` to `0` and use the archive URL. The actor pages through every post, newest first.

#### Can I scrape Astral Codex Ten comments?

Yes. Turn on `scrapeComments` and every post gets a `comments` array with each comment's author, date, text, reactions and replies. Use `maxCommentsPerPost` to cap how many are collected, best comments first, or `0` for all of them.

#### Does this scraper get paid-subscriber ACX posts?

No — it collects what a logged-out visitor sees. For paid posts that is the public preview, which the actor flags with `isPaywalled: true`. Set `includePaywalled` to `false` to skip them.

#### Does it work with old Slate Star Codex or Substack links?

It accepts both `astralcodexten.com` and the older `astralcodexten.substack.com` addresses. The old Slate Star Codex blog (slatestarcodex.com) is a different site and is not supported.

#### How do I scrape Astral Codex Ten with Python?

Install `apify-client`, then call the actor with the archive URL and iterate the dataset — see the Python example above. Each post comes back as structured JSON ready for pandas or a database.

### 🔗 Other actors you may find useful

- ✍️ **[Substack Scraper — Newsletter Posts, Comments & Search with AI](https://apify.com/confidential_gnat/substack-newsletter-scraper)** — Scrape any Substack newsletter or search all of Substack by keyword: posts, full text, reactions and comments with free sentiment, plus optional AI summaries and cross-run caching.
- 🏠 **[University Living Housing Scraper](https://apify.com/confidential_gnat/universityliving-housing-scraper)** — Scrapes student housing listings and property details from universityliving.com.
- 🍷 **[Total Wine Scraper](https://apify.com/confidential_gnat/totalwine-scraper)** — Scrape Total Wine & More (totalwine.com) wine, liquor and beer prices, sizes, ratings, reviews, badges, ABV, origin and taste profile from search, category or product URLs.
- 🇩🇪 **[German Imprint (Impressum) Scraper with AI Extraction](https://apify.com/confidential_gnat/german-imprint-scraper)** — Finds the Impressum page on any German website and extracts the company's decision makers, legal name, address, email addresses, phone numbers, commercial register number and VAT ID as structured data using AI.
- ⭐ **[Google Play Store Reviews Scraper](https://apify.com/confidential_gnat/google-play-reviews-scraper)** — Scrapes user reviews from Google Play Store apps (play.google.com) including review text, star rating, author, date, replies and optional sentiment tagging.
- 📜 **[Google Patents Scraper](https://apify.com/confidential_gnat/google-patents-scraper)** — Scrapes patent data from Google Patents (patents.google.com) by keyword or URL, including title, abstract, inventors, assignee, filing and publication dates, citations, figures and PDF links.

### 📰 Other news and article website scrapers

- 🗞️ **[Google News AI Scraper](https://apify.com/confidential_gnat/google-news-ai-scraper)** — Search Google News by keyword, optionally extract full article text and AI-generated summaries, and never re-scrape the same article twice across runs.
- 🦘 **[Sydney Morning Herald (SMH) News Scraper](https://apify.com/confidential_gnat/smh-news-scraper)** — Scrape news articles from The Sydney Morning Herald (smh.com.au) — headline, author, publish date, section, keywords, images and full public article text, with a paywall flag.
- 🇮🇩 **[Detik News Scraper](https://apify.com/confidential_gnat/detik-news-scraper)** — Scrapes news articles from Detik.com, including headline, author, publish date, category, images and full article text.

### 💬 Support & Contact

If you encounter any issues or have questions, please [open an issue](https://apify.com/confidential_gnat/astralcodexten-scraper/issues/open)

You can also find more of our actors on the [Actor Flow ](https://apify.com/confidential_gnat).

# Actor input Schema

## `startUrls` (type: `array`):

Astral Codex Ten archive or post URLs. Accepts the archive (https://www.astralcodexten.com/archive), the home page, and single post URLs (https://www.astralcodexten.com/p/some-post). Old astralcodexten.substack.com links work too. Page type is detected automatically.

## `maxItems` (type: `integer`):

Maximum number of posts to scrape per start URL. Set to 0 for no limit.

## `includePaywalled` (type: `boolean`):

Subscriber-only posts return only a short public preview rather than the full text. Leave enabled to collect them with an isPaywalled flag, or disable to skip them entirely.

## `cacheProjectName` (type: `string`):

Scrape only new posts. Give this run a name (e.g. "acx-new-posts") and reuse it: every later run with the same name skips posts already collected and returns only newly published ones, so the same post is never scraped or billed twice. Ideal for scheduled monitoring. Leave empty to scrape everything every time.

## `scrapeComments` (type: `boolean`):

Also collect each post's reader comments and replies.

## `maxCommentsPerPost` (type: `integer`):

Maximum comments (including replies) to collect per post, best comments first. Set to 0 for all comments.

## `proxyConfiguration` (type: `object`):

Proxy settings. Astral Codex Ten is reachable without a proxy, so proxies are disabled by default to keep runs cheap.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.astralcodexten.com/archive"
    }
  ],
  "maxItems": 5,
  "includePaywalled": true,
  "scrapeComments": false,
  "maxCommentsPerPost": 5,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.astralcodexten.com/archive"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("confidential_gnat/astralcodexten-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.astralcodexten.com/archive" }] }

# Run the Actor and wait for it to finish
run = client.actor("confidential_gnat/astralcodexten-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.astralcodexten.com/archive"
    }
  ]
}' |
apify call confidential_gnat/astralcodexten-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,confidential_gnat/astralcodexten-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KsWT3NRSdQvP2WpkH/builds/Z2ZhEq01dKeVEivgV/openapi.json
