# Bluesky Post Detail Scraper (`xtracto/bluesky-post-detail-scraper`) Actor

Fetch a Bluesky post with its thread (replies + parents) by AT URI or web URL. HTTP-only, no login required.

- **URL**: https://apify.com/xtracto/bluesky-post-detail-scraper.md
- **Developed by:** [Farhan Febrian Nauval](https://apify.com/xtracto) (community)
- **Categories:** Social media, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bluesky Post Detail Scraper

Extract a single Bluesky post together with its full conversation — author, text, media, counters, the parent chain, and the nested reply tree — in one clean structured JSON record.

### Why use this actor

- **No account / no login required** — paste a public bsky.app URL and run.
- **No API key needed** — get the same data the Bluesky web app shows, in JSON.
- **Whole conversation in one call** — parent chain plus the nested reply tree are returned together, no manual stitching.
- **Flexible input** — accepts `https://bsky.app/profile/<handle>/post/<rkey>` URLs, `at://...` identifiers, or `<handle>/<rkey>` shorthand; handles are resolved to stable IDs for you.
- **Configurable depth** — control how deep the reply tree goes and how many parents above the post you want to capture.
- **Stable JSON output** suitable for pipelines, spreadsheets, and databases — every row carries `_input`, `_source`, `_scrapedAt` envelope fields so you can join results back to your input list.

### How it works

1. You provide a list of Bluesky posts (URL, AT URI, or `<handle>/<rkey>`).
2. The actor fetches the post and walks the conversation up (parents) and down (replies) to the depth you set.
3. Each post in the thread arrives with author info, full text, media references, and engagement counters.
4. Results stream into your dataset, ready to download as JSON, CSV, or Excel.

You do not need to manage scrapers, browsers, or rotating IPs — all handled internally.

### Input

```json
{
  "posts": [
    "https://bsky.app/profile/bsky.app/post/3mlvllqtnmk2g",
    "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mlr7prm6g22c"
  ],
  "depth": 10,
  "parentHeight": 80,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": ["DATACENTER"]
  }
}
```

| Field | Type | Description |
|---|---|---|
| `posts` | array | List of Bluesky posts. Each entry can be a `bsky.app` URL, an `at://...` URI, or a `<handle>/<rkey>` shorthand. Handles are auto-resolved. |
| `depth` | integer | How many levels deep to follow the reply tree (0–1000). Default `10`. |
| `parentHeight` | integer | How many parent posts above the target to include (0–1000). Default `80`. |
| `proxyConfiguration` | object | Apify Proxy settings. Datacenter proxy works for most cases. |

### Output

Input: `at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mlvllqtnmk2g`

```json
{
  "_input": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mlvllqtnmk2g",
  "_uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mlvllqtnmk2g",
  "_source": "S1-xrpc",
  "_scrapedAt": "2026-05-18T10:35:57.147239+00:00",
  "thread": {
    "$type": "app.bsky.feed.defs#threadViewPost",
    "post": {
      "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mlvllqtnmk2g",
      "cid": "bafyreidkkomfz72otqfbyj4y2zjetrn456hko6pxzj3yhozyoxfx3s6eyy",
      "author": {
        "did": "did:plc:z72i7hdynmk6r22z27h6tvur",
        "handle": "bsky.app",
        "displayName": "Bluesky",
        "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreihwihm6kpd6zuwhhlro75p5qks5qtrcu55jp3gddbfjsieiv7wuka",
        "createdAt": "2023-04-12T04:53:57.057Z",
        "verification": {
          "verifications": [],
          "verifiedStatus": "none",
          "trustedVerifierStatus": "valid"
        },
        "labels": []
      },
      "record": {
        "$type": "app.bsky.feed.post",
        "createdAt": "2026-05-15T14:54:22.994Z",
        "langs": ["en"],
        "text": "THE INTERNET FOR $600: This social media app now boasts at least 200 \"Jeopardy!\" contestants and champions (including one host).",
        "embed": {
          "$type": "app.bsky.embed.record",
          "record": {
            "cid": "bafyreida3vtrwyjlngev73bmkjmxzi2ccstksiuvirkjhf55larupdvkaa",
            "uri": "at://did:plc:qrllvid7s54k4hnwtqxwetrf/app.bsky.feed.post/3kzwdmzvnf223"
          }
        }
      },
      "replyCount": 71,
      "repostCount": 141,
      "likeCount": 1920,
      "quoteCount": 28,
      "bookmarkCount": 39,
      "indexedAt": "2026-05-15T14:54:23.265Z",
      "labels": []
    },
    "replies": [
      {
        "$type": "app.bsky.feed.defs#threadViewPost",
        "post": {
          "uri": "at://did:plc:.../app.bsky.feed.post/...",
          "author": {
            "handle": "solidstalemate.bsky.social",
            "displayName": "Solid Stalemate",
            "did": "did:plc:..."
          },
          "record": {
            "text": "What is \"MySpace\"?",
            "createdAt": "2026-05-15T18:38:09.987Z",
            "langs": ["en"]
          },
          "replyCount": 0,
          "repostCount": 0,
          "likeCount": 1,
          "quoteCount": 0,
          "indexedAt": "2026-05-15T18:38:10.568Z"
        }
      },
      {
        "$type": "app.bsky.feed.defs#threadViewPost",
        "post": {
          "author": { "handle": "halfpintptbo.bsky.social" },
          "record": {
            "text": "We r an app of intelligent people unlike that other place full of idiot magats. Congrats to all!!",
            "createdAt": "2026-05-17T16:09:24.812Z"
          },
          "likeCount": 1,
          "replyCount": 0,
          "indexedAt": "2026-05-17T16:09:25.168Z"
        }
      },
      "... 66 more"
    ]
  },
  "threadgate": {
    "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.threadgate/3mlvllqtnmk2g",
    "cid": "bafyreibv64kb2qju5y7khjnvg4d2x2kwdgpdvhted2yhcgh3cxa5lzc7ty",
    "record": {
      "$type": "app.bsky.feed.threadgate",
      "createdAt": "2026-05-15T15:48:09.490Z",
      "hiddenReplies": [
        "at://did:plc:3r2ydi2f75r2hct527i637l5/app.bsky.feed.post/3mlvlot3nzc2i",
        "... 4 more"
      ],
      "post": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mlvllqtnmk2g"
    },
    "lists": []
  }
}
```

#### Envelope fields

| Field | Type | Description |
|---|---|---|
| `_input` | string | The post identifier exactly as you supplied it. Use this to join results back to your input list. |
| `_uri` | string | The canonical `at://` URI of the post (with the handle resolved to a stable DID). |
| `_source` | string | Internal tag for the path used to fetch the record. `S1-xrpc` means the primary path was used; values starting with `S2-` indicate a fallback was used. |
| `_scrapedAt` | string | ISO-8601 UTC timestamp when the record was scraped. |
| `_error` | string | Present only on failure. One of `not_found`, `resolve_failed`, `invalid_request`. |
| `_errorDetail` | string | Present only on failure. Human-readable detail explaining the error. |

#### `thread` — the post and its conversation

| Field | Type | Description |
|---|---|---|
| `thread.post` | object | The target post itself (see fields below). |
| `thread.parent` | object | Present if `parentHeight > 0` and the post is a reply. Same shape as `thread`, recurses upward. |
| `thread.replies` | array | Nested reply tree. Each entry has the same shape as `thread` (`post` + optional `replies`). Empty if the post has no replies or `depth = 0`. |
| `thread.$type` | string | Bluesky's record type. `app.bsky.feed.defs#threadViewPost` for a regular post, `#notFoundPost` or `#blockedPost` for hidden replies. |

#### `thread.post` — the post record

| Field | Type | Description |
|---|---|---|
| `uri` | string | The post's `at://` URI. Stable across handle changes. |
| `cid` | string | Content-addressed ID of this revision of the record. |
| `author` | object | Author profile: `did`, `handle`, `displayName`, `avatar`, `createdAt`, `verification`, `labels`. |
| `record.text` | string | The post body text (plain UTF-8, may contain emojis). |
| `record.createdAt` | string | ISO-8601 timestamp the author published the post. |
| `record.langs` | array | Language tags the author declared (e.g. `["en"]`). |
| `record.embed` | object | Attached media — quoted post, images, video, or external link card. `$type` indicates which. |
| `record.reply` | object | Present if the post is itself a reply. Contains `parent.uri` and `root.uri`. |
| `record.facets` | array | Inline mentions, hashtags, and links with byte offsets. |
| `embed` | object | The same embed, fully resolved (the quoted post is expanded, image URLs are populated). |
| `replyCount` | integer | Direct reply count. |
| `repostCount` | integer | Number of reposts. |
| `likeCount` | integer | Number of likes. |
| `quoteCount` | integer | Number of quote posts. |
| `bookmarkCount` | integer | Number of bookmarks. |
| `indexedAt` | string | ISO-8601 timestamp the post was indexed by the AppView. |
| `labels` | array | Moderation labels applied to the post. |
| `threadgate` | object | Author's reply restrictions, if any. |

#### `threadgate` — author's reply policy (top level)

| Field | Type | Description |
|---|---|---|
| `record.allow` | array | Who is allowed to reply (e.g. mentioned users, followers). Absent = anyone. |
| `record.hiddenReplies` | array | URIs of replies the author chose to hide from the default view. |
| `lists` | array | Lists referenced by the threadgate (for list-scoped reply permissions). |

### Error envelope

Posts that don't exist, are deleted, or fail to resolve return a structured error instead of crashing the run:

```json
{
  "_input": "https://bsky.app/profile/bsky.app/post/does-not-exist",
  "_uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/does-not-exist",
  "_source": "S1-xrpc",
  "_scrapedAt": "2026-05-18T10:36:02.487122+00:00",
  "_error": "not_found",
  "_errorDetail": "NotFound app.bsky.feed.getPostThread {'uri': 'at://...'}"
}
```

Filter on `_error` to triage failed rows.

### Pricing

This actor is billed per result: **$1.50 per 1,000 results**. Each successfully scraped post (with its full thread) counts as one result. Errors (`not_found`, `resolve_failed`) are not billed.

### Other Sosmed Actors

| Platform | Actor | Best for |
|---|---|---|
| Bluesky | [Bluesky Account Scraper](https://apify.com/xtracto/bluesky-account-scraper) | Profile, followers, bio for any handle |
| Bluesky | [Bluesky Account Posts Scraper](https://apify.com/xtracto/bluesky-account-posts-scraper) | Paginate every post by an author |
| Bluesky | [Bluesky Search Scraper](https://apify.com/xtracto/bluesky-search-scraper) | Keyword search across Bluesky posts |
| Threads | [Threads Post Detail Scraper](https://apify.com/xtracto/threads-post-detail-scraper) | Single Threads post + reply tree |
| Twitter / X | [X Post Detail Scraper](https://apify.com/xtracto/x-post-detail-scraper) | Single tweet + conversation context |
| Mastodon | [Mastodon Status Detail Scraper](https://apify.com/xtracto/mastodon-status-detail-scraper) | Single Mastodon status + replies |
| Reddit | [Reddit Post Detail Scraper](https://apify.com/xtracto/reddit-post-detail-scraper) | Reddit post + nested comments |
| Instagram | [Instagram Post Detail Scraper](https://apify.com/xtracto/instagram-post-detail-scraper) | Single Instagram post + comments |

Browse the full catalog at [apify.com/xtracto](https://apify.com/xtracto).

### Notes

- Counters (`likeCount`, `repostCount`, `replyCount`, `quoteCount`, `bookmarkCount`) are eventually-consistent and may lag the live network by a few minutes.
- Default `depth` is 10 levels of replies and `parentHeight` is 80 — generous for most analytics needs but lower them for performance on very large threads.
- Replies from blocked or labeled accounts may surface as `$type: app.bsky.feed.defs#blockedPost` or `#notFoundPost` entries inside `thread.replies` — keep this in mind when counting visible replies.
- Deleted or never-existed posts return `{"_error": "not_found", "_input": "..."}` so downstream pipelines can branch on missing data.
- The `at://` URI in `_uri` is the most reliable identifier — handles can change but the DID inside the URI does not.
- For very large threads (>5,000 replies) you may hit practical depth limits; for those, paginate replies via the Bluesky Account Posts Scraper against the original author.

# Actor input Schema

## `posts` (type: `array`):

List of Bluesky posts to fetch. Each entry can be a bsky.app post URL, an AT URI (`at://...`), or a `<handle>/<rkey>` shorthand. Handles are resolved automatically.

## `depth` (type: `integer`):

How many levels deep to follow the reply tree under the post. Range: 0 (no replies) to 1000. Default 10 covers most conversations; increase for deeply nested threads.

## `parentHeight` (type: `integer`):

How many parent posts above the target to include (for replies inside a longer chain). Range: 0 (none) to 1000. Default 80 captures the full ancestor chain in almost all cases.

## `proxyConfiguration` (type: `object`):

Apify Proxy settings. Datacenter proxy is sufficient for the public Bluesky API; switch to Residential only if you hit rate limits on very large jobs.

## Actor input object example

```json
{
  "posts": [
    "https://bsky.app/profile/bsky.app/post/3mp7sbnuyu22b",
    "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mp57vzxqlk22"
  ],
  "depth": 10,
  "parentHeight": 80,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `thread` (type: `string`):

Thread as reported by the source.

## `_input` (type: `string`):

The input value this row was produced from.

## `_uri` (type: `string`):

Uri as reported by the source.

## `_scrapedAt` (type: `string`):

UTC timestamp of the scrape, ISO 8601.

## `threadgate` (type: `string`):

Threadgate as reported by the source.

## `uri` (type: `string`):

Resource identifier at the source.

## `cid` (type: `string`):

Content identifier at the source.

## `author` (type: `string`):

Author name.

## `embed` (type: `string`):

Embed as reported by the source.

## `replyCount` (type: `string`):

Reply Count. Whole number.

## `repostCount` (type: `string`):

Repost Count. Whole number.

## `likeCount` (type: `string`):

Like count. Whole number.

## `quoteCount` (type: `string`):

Quote Count. Whole number.

## `bookmarkCount` (type: `string`):

Bookmark Count. Whole number.

## `indexedAt` (type: `string`):

Timestamp the source indexed the item, ISO 8601.

## `labels` (type: `string`):

Labels.

## `lists` (type: `string`):

Lists.

## `_source` (type: `string`):

Which extraction strategy produced the row.

## `_error` (type: `string`):

Set only on diagnostic rows - why that target produced no data.

## `_errorDetail` (type: `string`):

Extra context for the error.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "posts": [
        "https://bsky.app/profile/bsky.app/post/3mp7sbnuyu22b",
        "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mp57vzxqlk22"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("xtracto/bluesky-post-detail-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "posts": [
        "https://bsky.app/profile/bsky.app/post/3mp7sbnuyu22b",
        "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mp57vzxqlk22",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("xtracto/bluesky-post-detail-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "posts": [
    "https://bsky.app/profile/bsky.app/post/3mp7sbnuyu22b",
    "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mp57vzxqlk22"
  ]
}' |
apify call xtracto/bluesky-post-detail-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,xtracto/bluesky-post-detail-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/1030osGfqNhc9GF4M/builds/gE2pZLIRtoLQCYuPC/openapi.json
