# Bluesky Scraper: Posts, Profiles & Followers (`digital_influx/bluesky-scraper`) Actor

Scrape Bluesky posts by keyword or hashtag, profiles with their posts, followers and following, people search and post replies. Emails and links from bios, likes, reposts, media, quotes. Only-new mode for scheduled monitoring. Public API, no login.

- **URL**: https://apify.com/digital\_influx/bluesky-scraper.md
- **Developed by:** [Bruno Petrelli](https://apify.com/digital_influx) (community)
- **Categories:** Social media, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 post saveds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bluesky Scraper: Posts, Profiles & Followers

Get Bluesky data as clean rows, straight from Bluesky's public API:

- **Posts by keyword or hashtag:** text, date, likes, reposts, replies, quotes, hashtags, mentions, full links, images with alt text, video, link cards and quoted posts. Bluesky's search syntax works: `from:handle`, `mentions:handle`, `domain:example.com`, `"exact phrase"`.
- **Profiles:** bio, **emails and links written in the bio**, website, social profiles, followers, following and post counts, verification, join date, pinned post. A handle that is the account's own domain (`acme.com`) is returned as `customDomain`: Bluesky verified that the account controls it.
- **A profile's posts**, optionally with its replies and reposts.
- **Followers and following** of any account, each as a full profile row.
- **People search:** find accounts by name or bio ("marketing agency", "data engineer") and keep only those with an email or a link.
- **A post and its replies,** level by level, with the depth of each reply.

**No login, no proxies, no browser.** Bluesky's robots.txt says "Crawling the public parts of the API is allowed", and the Actor keeps to a handful of requests at a time, as it asks. In our test, 30 posts of a search plus a profile and its 10 latest posts took 2 seconds.

### Use it for

- **Brand and topic monitoring:** schedule a search for your brand, product or hashtag every hour with "Only new" and get just the new mentions.
- **Lead lists:** people search or the followers of a competitor, filtered to accounts with an email or website in the bio.
- **Influencer research:** profiles with follower counts, posting activity and their latest posts.
- **Research and social listening:** posts on a topic with language and date filters, engagement counts and links.
- **Community management:** the replies to a post, or what an account posts and reposts.
- **AI agents:** one plain JSON object per post or profile, callable through the Apify MCP server.

### Input

Search posts and read a profile (the example in the form):

```json
{
  "searchTerms": ["web scraping"],
  "maxPostsPerSearch": 30,
  "profiles": ["bsky.app"],
  "postsPerProfile": 10
}
```

Followers of an account that have an email or link in their bio:

```json
{
  "profiles": ["apify.com"],
  "postsPerProfile": 0,
  "followersPerProfile": 500,
  "onlyProfilesWithContact": true
}
```

Hourly brand monitoring in English (schedule it in Apify):

```json
{
  "searchTerms": ["\"your brand\"", "yourbrand.com"],
  "languages": ["en"],
  "onlyNew": true
}
```

Every field can be combined in one run. Profiles can be handles (`jay.bsky.team` or `@jay.bsky.team`), profile links or DIDs; posts can be bsky.app links or `at://` URIs. Filters (`postedWithinDays`, `languages`, `minLikes`) apply to searched posts, profile posts and replies; a post you list by link is always saved.

### Output

One row per post (`"type": "post"`) or profile (`"type": "profile"`). The run's page has a Posts view and a Profiles view. Real results from 2026-09-30, shortened:

```json
{
  "type": "profile",
  "url": "https://bsky.app/profile/apify.com",
  "handle": "apify.com",
  "displayName": "Apify",
  "description": "Thousands of Actors to automate your business, get real-time web data, and integrate your apps and agents.  ➡️  apify.com • mcp.apify.com • github.com/apify",
  "emails": [],
  "website": "https://apify.com",
  "customDomain": "apify.com",
  "links": ["https://apify.com", "https://mcp.apify.com", "https://github.com/apify"],
  "socialProfiles": { "github": ["https://github.com/apify"] },
  "followersCount": 117,
  "followsCount": 36,
  "postsCount": 138,
  "createdAt": "2024-11-25T17:31:53.548Z",
  "verified": false,
  "pinnedPostUrl": null,
  "did": "did:plc:vu5ic5ygdeyunpefgmlamsmw",
  "source": "profile",
  "sourceInput": "apify.com"
}
```

```json
{
  "type": "post",
  "url": "https://bsky.app/profile/apify.com/post/3mb4ki3wrd42w",
  "text": "TikTok doesn’t show follower lists - but now you can have them. …",
  "createdAt": "2025-12-29T09:45:56.843Z",
  "authorHandle": "apify.com",
  "authorName": "Apify",
  "likeCount": 2,
  "repostCount": 0,
  "replyCount": 0,
  "quoteCount": 0,
  "languages": [],
  "hashtags": [],
  "mentions": [],
  "links": ["https://apify.com/clockworks/tiktok-followers-scraper"],
  "images": [{ "url": "https://cdn.bsky.app/img/feed_fullsize/plain/did:plc:vu5ic5ygdeyunpefgmlamsmw/bafkreie2wfcjmk6hckxtadiok3jhvglar36ojo5ombk5gpqazh3jy5yhoe", "alt": "" }],
  "isReply": false,
  "quotedPostUrl": null,
  "repostedBy": null,
  "uri": "at://did:plc:vu5ic5ygdeyunpefgmlamsmw/app.bsky.feed.post/3mb4ki3wrd42w",
  "source": "profilePosts",
  "sourceInput": "apify.com"
}
```

- Posts also have `bookmarkCount`, `videoUrl`, `linkCard` (url, title, description), `replyToUrl`, `threadRootUrl`, `quotedText`, `quotedAuthorHandle`, `labels` (content labels such as `porn` or `graphic-media`), `repostedAt`, `replyDepth` (in a thread), `authorDid`, `authorAvatar`, `authorUrl`, `cid` and `indexedAt`.
- Profiles also have `trustedVerifier`, `pronouns`, `labels` (such as `bot`), `lists`, `feeds`, `starterPacks`, `isLabeler`, `acceptsMessagesFrom`, `avatar`, `banner` and `indexedAt`.
- `source` says where a row came from: `search`, `profile`, `profilePosts`, `follower`, `following`, `profileSearch` or `thread`; `sourceInput` is the line of your input that found it.
- The OUTPUT record of the run has the status of each input (found, not found, hidden, failed) and the totals, including how many rows each filter left out.

### Pricing

Pay per event: **one event per post saved** and **one per profile saved**. Posts and profiles left out by your filters, already saved by an earlier run ("Only new"), repeated within the run, or hidden at the owner's request are free. Set a maximum cost for the run and the Actor stops cleanly when it is reached.

### Good to know

- **Search returns at most 100 posts per term.** Bluesky gives apps that are not logged in only the first page of a search, and the Actor respects that. To collect more, run it on a schedule with "Only new", or use narrower searches (a language, `from:`, a hashtag). Profiles, followers, following, people search and threads have no such cap.
- **Accounts that asked not to be shown to logged-out visitors are left out.** In Bluesky's settings, people can "discourage apps from showing my account to users who are logged out". The Actor reads Bluesky logged out, so it skips those accounts and their posts and counts them as `hiddenAtUserRequest` (free). In four searches on 2026-09-30 that was between 6% and 20% of the posts (6 of 94 for "web scraping", 19 of 98 for "marketing" in English).
- **Only new:** each search, profile feed, follower list and thread has its own memory of what was saved. A profile's posts stop at the first post an earlier run saved. The profiles you list are saved on every run, so you can follow their counts over time.
- Deleted posts and accounts no longer come out of Bluesky's API, so they are not returned.

### Personal data

Posts and profiles are written by people, and bios can contain their email address. That is personal data even when it is public. Use it only with a legal basis (for example legitimate interest under the GDPR for relevant, one-to-one outreach), honor opt-outs and deletion requests, and follow the anti-spam rules of the countries you contact.

### Support

A field you need, or a result that looks wrong? Open an issue on the Actor's Issues tab with the input you used. Issues get an answer within a few days.

# Actor input Schema

## `searchTerms` (type: `array`):

Words, phrases or hashtags to find posts, one per line, e.g. 'web scraping' or '#ai'. Bluesky's search syntax works too: from:handle, mentions:handle, domain:example.com, "exact phrase".

## `maxPostsPerSearch` (type: `integer`):

1 to 100. Bluesky gives apps that are not logged in the first 100 posts of each search; run on a schedule with 'Only new' to collect more over time.

## `searchSort` (type: `string`):

Latest posts first, or Bluesky's top posts for the search.

## `profiles` (type: `array`):

Bluesky accounts, one per line: a handle (jay.bsky.team), a profile link (https://bsky.app/profile/jay.bsky.team) or a DID. You get one row per profile with its bio, emails and links in the bio, counts and verification, then its posts, followers and following as set below.

## `postsPerProfile` (type: `integer`):

Newest posts of each profile. 0 for none.

## `includeReplies` (type: `boolean`):

Also save the replies a profile wrote to other posts.

## `includeReposts` (type: `boolean`):

Also save the posts a profile reposted (each row says who reposted it and when).

## `followersPerProfile` (type: `integer`):

Newest followers of each profile, each as a full profile row (bio, emails, links, counts). 0 for none.

## `followingPerProfile` (type: `integer`):

Accounts each profile follows, as full profile rows. 0 for none.

## `searchProfiles` (type: `array`):

Find accounts by name or bio, one search per line, e.g. 'marketing agency' or 'data engineer'. Each result is a full profile row.

## `maxProfilesPerSearch` (type: `integer`):

1 to 1000.

## `onlyProfilesWithContact` (type: `boolean`):

For people searches, followers and following: skip accounts whose bio has no email address and no web link. The profiles you list above are always saved.

## `postUrls` (type: `array`):

Post links, one per line (https://bsky.app/profile/…/post/…). You get the post and its replies, level by level.

## `repliesPerPost` (type: `integer`):

0 for only the post.

## `postedWithinDays` (type: `integer`):

Only posts from the last N days, for searches, profile posts and replies. 0 for any date.

## `languages` (type: `array`):

Only posts tagged with one of these languages, as two-letter codes (en, es, pt, de, ja…). Posts without a language tag are left out when this is set.

## `minLikes` (type: `integer`):

Only posts with at least this many likes. 0 for any.

## `onlyNew` (type: `boolean`):

Save only posts and followers that earlier runs with this option did not save, so an hourly or daily schedule gives you just what is new (mentions of your brand, new posts of an account, new followers) and you pay only for those. The profiles you list are saved on every run. The memory is one record per search, profile and post, in a key-value store named "bluesky-state" in your account; entries not seen for 90 days are dropped.

## `stateKey` (type: `string`):

Give a name to keep a separate memory (for example one per client or per task), or a new name to start over. Letters, digits, dot, dash and underscore.

## Actor input object example

```json
{
  "searchTerms": [
    "web scraping"
  ],
  "maxPostsPerSearch": 30,
  "searchSort": "latest",
  "profiles": [
    "bsky.app"
  ],
  "postsPerProfile": 10,
  "includeReplies": false,
  "includeReposts": false,
  "followersPerProfile": 0,
  "followingPerProfile": 0,
  "maxProfilesPerSearch": 50,
  "onlyProfilesWithContact": false,
  "repliesPerPost": 50,
  "postedWithinDays": 0,
  "minLikes": 0,
  "onlyNew": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "web scraping"
    ],
    "maxPostsPerSearch": 30,
    "profiles": [
        "bsky.app"
    ],
    "postsPerProfile": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("digital_influx/bluesky-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["web scraping"],
    "maxPostsPerSearch": 30,
    "profiles": ["bsky.app"],
    "postsPerProfile": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("digital_influx/bluesky-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "web scraping"
  ],
  "maxPostsPerSearch": 30,
  "profiles": [
    "bsky.app"
  ],
  "postsPerProfile": 10
}' |
apify call digital_influx/bluesky-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,digital_influx/bluesky-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KbnJSrzrJs0ZfKODc/builds/AtKxXPcJp3RoP5HdW/openapi.json
