# Bluesky Profile, Follower & Creator Scraper (`haketa/bluesky-scraper`) Actor

Discover and scrape Bluesky profiles: handle, name, bio, follower/following/post counts, verification, account age and links. Search by keyword, look up handles, or scrape any account's followers/following. Extracts email, website and socials from bios. For creator discovery and lead generation.

- **URL**: https://apify.com/haketa/bluesky-scraper.md
- **Developed by:** [Haketa](https://apify.com/haketa) (community)
- **Categories:** Social media, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bluesky Profile, Follower & Creator Scraper

> **Discover and scrape Bluesky profiles at scale: handle, display name, bio, follower / following / post counts, verification, account age, avatar, banner — plus email, website and social links pulled from each bio.** Search by keyword, look up specific handles, or scrape the **followers / following of any account**. Clean JSON/CSV/Excel in seconds — built for creator discovery, influencer research and lead generation.

[![Bluesky](https://img.shields.io/badge/Bluesky-AT%20Protocol-1185fe)]()
[![Creators + Contacts](https://img.shields.io/badge/Creators%20%2B%20Contacts-2da44e)]()
[![Lead Generation](https://img.shields.io/badge/Lead%20Generation-8250df)]()
[![Export](https://img.shields.io/badge/Export-JSON%20%2F%20CSV%20%2F%20Excel-fb8500)]()

***

### What This Actor Does

This Actor pulls structured data from Bluesky's public network. For each profile it captures:

- **Identity** — handle, display name, DID, bio, avatar, banner, profile URL
- **Audience & activity** — follower count, following count, post count, account age (created date)
- **Signals** — verification status, labeler flag, and counts of lists, custom feeds and starter packs
- **Contacts (from bio)** — email, website and social links (Instagram, Twitter/X, YouTube, TikTok, LinkedIn, GitHub, Patreon, Substack)

Four ways to find profiles:

1. **Keyword search** — find creators/accounts by topic (searches names, handles and bios).
2. **Handles / DIDs** — look up specific accounts directly.
3. **Followers of** — scrape the followers of any account (build a lead list from a competitor's or influencer's audience).
4. **Following of** — scrape who any account follows.

***

### Why Use This

- **Turn any account into a lead list.** Scrape the followers of a competitor, community or influencer and enrich every one with counts and contacts.
- **Find creators by topic.** Keyword search surfaces accounts by niche — photographers, developers, marketers, artists.
- **Contacts included.** Email, website and social handles are pulled straight from each bio.
- **Clean and free.** Reads Bluesky's public API — no key, no login, no anti-bot, no browser.

***

### Quick Start

#### Run it in the console (no code)

1. Add **keyword searches**, specific **handles**, and/or accounts under **Followers of** / **Following of**.
2. Set **Max profiles**, click **Start**.
3. Export as **JSON, CSV, Excel or HTML**, or push to Google Sheets, a webhook or a database.

#### Build a lead list from an account's followers (Python)

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run_input = {"followersOf": ["jay.bsky.team"], "maxItems": 2000}

run = client.actor("YOUR_USERNAME/bluesky-scraper").call(run_input=run_input)

for p in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(p["handle"], "·", p["followersCount"], "followers ·", p["website"] or p["email"])
```

#### Discover creators by topic (Python)

```python
run = client.actor("YOUR_USERNAME/bluesky-scraper").call(run_input={
    "searchTerms": ["photographer", "illustrator"], "maxItems": 1000,
})
for p in client.dataset(run["defaultDatasetId"]).iterate_items():
    if (p.get("followersCount") or 0) > 5000:
        print(p["handle"], p["displayName"], "→", p["website"], p["instagram"])
```

***

### Input Parameters

| Field | Type | Description |
|---|---|---|
| `searchTerms` | array | Keywords to find creators/accounts. |
| `handles` | array | Specific handles or DIDs (accepts `@handle` or profile URLs). |
| `followersOf` | array | Accounts whose **followers** you want to scrape. |
| `followingOf` | array | Accounts whose **following** you want to scrape. |
| `includeContacts` | boolean | Extract email, website and socials from bios (default on). |
| `maxItems` | integer | Max profiles to return. `0` = no limit. |
| `maxPagesPerSource` | integer | Pagination cap per search term / seed account (100 per page). |
| `proxyConfiguration` | object | Apify Proxy. Datacenter is enough (public API). |

Provide at least one of `searchTerms`, `handles`, `followersOf` or `followingOf`.

***

### Output

Each profile is one record:

```json
{
  "did": "did:plc:z72i7hdynmk6r22z27h6tvur",
  "handle": "bsky.app",
  "displayName": "Bluesky",
  "bio": "official Bluesky account",
  "followersCount": 35097942,
  "followsCount": 15,
  "postsCount": 865,
  "createdAt": "2023-04-12T04:53:57.057Z",
  "verifiedStatus": "valid",
  "listsCount": 18, "feedgensCount": 7, "starterPacksCount": 15,
  "avatar": "https://cdn.bsky.app/img/avatar/...",
  "email": "", "website": "https://bsky.social",
  "instagram": "", "twitter": "", "youtube": "",
  "profileUrl": "https://bsky.app/profile/bsky.app",
  "source": "handle"
}
```

**About contacts:** website and social links are common in Bluesky bios and are extracted at a high rate. Raw email addresses are **rare** in social bios (creators tend to link a website or Linktree rather than post an email) — so the `email` field is often empty while `website` is populated. This is the nature of the platform, not a gap in coverage. Use `website` as the primary contact path and `email` where present.

***

### Use Cases

#### 1. Lead generation

Scrape the followers of any account and enrich each with counts, website and socials — a ready-made prospecting list.

#### 2. Influencer & creator discovery

Find and rank creators by topic and audience size for partnerships, PR and sponsorships.

#### 3. Competitive & audience research

Analyse who follows (or is followed by) any account — map a community or a competitor's audience.

#### 4. Social monitoring & CRM enrichment

Track accounts over time and enrich records with Bluesky handles, bios and links.

***

### Tips

- **`followersOf`** is the fastest way to a targeted lead list — point it at the right account.
- **`followersCount`** lets you rank and filter creators by reach.
- **`website`** is the most reliable contact field; **`email`** is included when present in the bio.
- **`createdAt`** gives account age — useful for spotting established vs new accounts.
- **Schedule it** with Apify Schedules to track an audience or niche over time.

***

### Frequently Asked Questions

**Do I need an account or key?**
No. Bluesky's AppView API is public — no login, key or anti-bot.

**Can I scrape a private account's followers?**
The Actor only reads public data available through Bluesky's public API.

**Why is the email field often empty?**
Social bios rarely contain raw email addresses; creators usually link a website or Linktree. `website` and social fields are populated far more often — use those as the primary contact path.

**How many profiles can I get from one account's followers?**
Pagination runs 100 per page up to your `maxItems` and `maxPagesPerSource` limits.

**What export formats are supported?**
JSON, CSV, Excel, HTML, or via API — plus Google Sheets, webhooks, Make and Zapier.

***

### Legal & Responsible Use

This Actor is an independent tool and is **not affiliated with, endorsed by, or sponsored by Bluesky (Bluesky Social PBC)**. All trademarks are the property of their respective owners. It reads only public profile data. Use the data responsibly and in line with applicable terms and data-protection laws (e.g. GDPR) when handling personal data.

# Actor input Schema

## `searchTerms` (type: `array`):

Keywords to find creators/accounts (e.g. "photographer", "crypto", "climate"). Searches names, handles and bios.

## `handles` (type: `array`):

Specific Bluesky handles or DIDs to look up (e.g. bsky.app, jay.bsky.team). Accepts @handle or profile URLs.

## `followersOf` (type: `array`):

Scrape the followers of these accounts — build an audience/lead list from any account's followers.

## `followingOf` (type: `array`):

Scrape who these accounts follow.

## `includeContacts` (type: `boolean`):

Pull email, website and social links out of each bio.

## `maxItems` (type: `integer`):

Maximum profiles to return. 0 = no limit.

## `maxPagesPerSource` (type: `integer`):

Pagination cap per search term / seed account (100 per page).

## `proxyConfiguration` (type: `object`):

Apify Proxy. The API is public — datacenter is enough and enabled by default.

## Actor input object example

```json
{
  "searchTerms": [
    "photographer"
  ],
  "includeContacts": true,
  "maxItems": 200,
  "maxPagesPerSource": 20,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `did` (type: `string`):

Decentralized ID

## `handle` (type: `string`):

Bluesky handle

## `displayName` (type: `string`):

Display name

## `bio` (type: `string`):

Profile description

## `followersCount` (type: `string`):

Follower count

## `followsCount` (type: `string`):

Following count

## `postsCount` (type: `string`):

Post count

## `createdAt` (type: `string`):

Account creation date

## `verifiedStatus` (type: `string`):

Verification status

## `isLabeler` (type: `string`):

Labeler account

## `listsCount` (type: `string`):

Lists count

## `feedgensCount` (type: `string`):

Custom feeds count

## `starterPacksCount` (type: `string`):

Starter packs count

## `avatar` (type: `string`):

Avatar image URL

## `banner` (type: `string`):

Banner image URL

## `email` (type: `string`):

Email from bio

## `website` (type: `string`):

Website from bio

## `instagram` (type: `string`):

Instagram from bio

## `twitter` (type: `string`):

Twitter/X from bio

## `youtube` (type: `string`):

YouTube from bio

## `tiktok` (type: `string`):

TikTok from bio

## `linkedin` (type: `string`):

LinkedIn from bio

## `github` (type: `string`):

GitHub from bio

## `profileUrl` (type: `string`):

Bluesky profile URL

## `source` (type: `string`):

Which query found it

## `scrapedAt` (type: `string`):

ISO timestamp

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchTerms": [
        "photographer"
    ],
    "maxItems": 200,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("haketa/bluesky-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchTerms": ["photographer"],
    "maxItems": 200,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("haketa/bluesky-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchTerms": [
    "photographer"
  ],
  "maxItems": 200,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call haketa/bluesky-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,haketa/bluesky-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/1MIV2Zl0uVSrXh6lL/builds/YNQIGphkcTnbBqV8g/openapi.json
