# Bluesky Profile Search Scraper (`scrapingmonkey/bluesky-profile-search-scraper`) Actor

Search public Bluesky profiles by keyword and collect multiple result pages. Export handles, names, biographies, avatar URLs and available verification data.

- **URL**: https://apify.com/scrapingmonkey/bluesky-profile-search-scraper.md
- **Developed by:** [ScrapingMonkey](https://apify.com/scrapingmonkey) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

Search public Bluesky accounts by keyword and export matching profiles as individual records. **Bluesky Profile Search Scraper** returns handles, names, biographies, avatar links, profile URLs and available verification data across multiple result pages.

Start with a topic, profession, organization or name to discover accounts for manual qualification, research and repeatable profile search workflows.

| At a glance | Details |
|---|---|
| 📥 Input | Add one nonblank profile-search query per item, such as python, science or art. This searches accounts, not posts. |
| 📤 Output | One unique matching profile per query |
| 📄 Pagination | Up to 25 source items per requested page |
| 🔐 Login required | No |
| ⚡ Processing | Up to 5 HTTP requests concurrently; up to 5 attempts for temporary failures |
| 💾 Delivery | One Apify dataset view with individual result rows, flattened object columns and complete nested JSON |

### What the Bluesky profile search scraper collects 🔎

The Actor uses the public account-search results shown by the website. Each successful row contains one matching profile and its source query, keeping searches easy to compare in a spreadsheet or database.

Data can include:

- Source query exactly as normalized for search
- Matching profile DID, handle, display name and canonical URL
- Biography and avatar URL, plus banner and pronouns when exposed
- Optional profile counts, dates and verification states
- Source label values when present in the public record

### How to collect profile search results from Bluesky 🚀

1. Enter one or more supported inputs.
2. Choose a mode if available and set the page count for each input.
3. Start the Actor.
4. Open the **Profiles** dataset view and review individual result rows.
5. Export the dataset or retrieve it from your application.

```json
{
  "inputList": [
    "python"
  ],
  "pagesPerSearch": 1
}
```

`pagesPerSearch` counts account-search result pages separately for each query. Each page requests up to 25 accounts. Search ranking and matching are controlled by Bluesky, and the Actor does not guarantee an exhaustive account directory.

### Bluesky profile search results output fields 📦

| Field | Type | Meaning |
|---|---|---|
| `input` | string | Original submitted input, retained on success and failure. |
| `status` | string | `success` or `failed`. |
| `did` | string or null | Stable Bluesky account identifier. |
| `handle` | string or null | Current complete account handle. |
| `display_name` | string or null | Display name returned by the source. |
| `profile_url` | string or null | Canonical account profile URL. |
| `description` | string or null | Public biography or collection description, according to row type. |
| `avatar_url` | string or null | Profile or feed avatar URL. |
| `banner_url` | string or null | Profile banner URL when exposed. |
| `pronouns` | string or null | Profile pronouns when exposed. |
| `created_at` | string or null | Creation or publication time returned by the source. |
| `indexed_at` | string or null | Time the source indexed the record. |
| `followers_count` | integer or null | Profile follower count when present in this response. |
| `following_count` | integer or null | Profile following count when present in this response. |
| `posts_count` | integer or null | Profile post count when present in this response. |
| `verified_status` | string or null | Source verification state, when supplied. |
| `trusted_verifier_status` | string or null | Source trusted-verifier state, when supplied. |
| `labels` | array or null | Source label values. |
| `query` | string or null | Normalized search query; null outside a search mode. |

The successful examples below use normalized public response data. Values are snapshots rather than promises of current content or counts; every top-level output key is included.

Complete representative successful result:

```json
{
  "input": "python",
  "status": "success",
  "query": "python",
  "did": "did:plc:sfrl4dmvaxeq4lqgaucotygo",
  "handle": "python.org",
  "profile_url": "https://bsky.app/profile/did:plc:sfrl4dmvaxeq4lqgaucotygo",
  "labels": [],
  "verified_status": "valid",
  "trusted_verifier_status": "none",
  "display_name": "Python Software Foundation",
  "description": "The nonprofit organization behind the Python programming language. For help with Python code: http://python.org/about/help/\n\nOn Mastodon: @ThePSF@fosstodon.org",
  "avatar_url": "https://cdn.bsky.app/img/avatar/plain/did:plc:sfrl4dmvaxeq4lqgaucotygo/bafkreicxrockny5ugxmk4n6ydnqb4cs5ya5da24sr4ii3dyixf2awoz7ta",
  "banner_url": null,
  "pronouns": null,
  "created_at": "2024-06-06T15:58:01.264Z",
  "indexed_at": "2026-02-25T15:35:25.157Z",
  "followers_count": null,
  "following_count": null,
  "posts_count": null
}
```

Complete failed dataset item:

```json
{
  "input": "   ",
  "status": "failed",
  "did": null,
  "handle": null,
  "display_name": null,
  "profile_url": null,
  "description": null,
  "avatar_url": null,
  "banner_url": null,
  "pronouns": null,
  "created_at": null,
  "indexed_at": null,
  "followers_count": null,
  "following_count": null,
  "posts_count": null,
  "verified_status": null,
  "trusted_verifier_status": null,
  "labels": null,
  "query": null
}
```

Each failed row preserves `input`, uses `status: failed`, and sets every other top-level field to `null`. The reason is written to the run log. Successful rows may contain null optional fields or empty arrays when the source does not supply a value. The dataset is not split into separate tables for media, authors, modes or failures.

### Input and pagination settings ⚙️

| Parameter | Type | Required | Default | Rules |
|---|---|---|---|---|
| `inputList` | array of strings | Yes | None | Add one nonblank profile-search query per item, such as python, science or art. This searches accounts, not posts. |
| `pagesPerSearch` | integer | No | `1` | Number of result pages to attempt per input. Each page requests up to 25 items; actual public rows can be fewer. Bootstrap requests do not count as pages. Stops when the cursor ends or repeats. Minimum `1`. |

Enter nonblank plain-text queries of at most 1,000 characters, such as `python`, `science` or `art`. Spaces within a phrase are preserved. URLs, account identifiers and control characters are rejected as search input.

For supported website URLs, use the exact `https://bsky.app` host and canonical path without a query string or fragment. URLs from other hosts and unsupported record types are rejected. Handles are normalized; duplicate normalized inputs in the same mode are processed once. Different source inputs retain their own results, while repeated record identifiers within one input’s pagination are removed.

The first result request counts as page 1; resolving a handle or validating source metadata does not consume a result page. Pagination stops at the requested page count, a missing cursor or a repeated cursor. A short, empty or fully filtered page can still continue when it includes a usable next cursor. No exact result total is guaranteed.

### Bluesky profile search results use cases 🎯

#### Creator and specialist discovery

Search a topic or profession and review the returned biographies and profile links. Select relevant accounts yourself before using them in further analysis.

#### Organization and name research

Collect account search matches for organization names or project terms. Keep the original query when documenting uncertain matches for human review.

#### Repeated account-search snapshots

Schedule queries and compare returned account DIDs across snapshots. The Actor does not rank leads, infer contact details or automatically contact accounts.

### Pricing and saved-result behavior 💰

See the Actor’s **Pricing** tab for the active charging model and current rate. Store settings are separate from this local implementation, so this README does not state an unverified fixed price or runtime.

Under dataset-item pricing:

- Every unique record saved as `success` is one result for that source input. Nested author, media or profile fields do not become separate rows.
- An invalid or unavailable input, an exhausted temporary failure, or an input with no publicly available results within its page budget can save one `failed` row.
- Automatic retry attempts do not create extra dataset rows by themselves.
- A normal empty continuation after earlier successes creates no extra result row.
- More requested pages can produce more saved rows. Check the active listing for how saved failed rows are billed; they are not assumed to be free.

The final total depends on source availability, duplicate removal and the chosen page budget. Start with a small run and check its actual usage before selecting a larger budget.

### Bluesky profile search results API 🔌

Replace `$ACTOR_ID` with the identifier from this Actor’s **API** tab and `$APIFY_TOKEN` with your Apify token.

```bash
curl -X POST "https://api.apify.com/v2/acts/$ACTOR_ID/runs?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"inputList":["python"],"pagesPerSearch":1}'
```

Retrieve the default dataset through the Apify API or download JSON, CSV, Excel, XML and other formats available in the Console. Schedules, completion webhooks and Apify integrations can connect results to Google Sheets, Make, Zapier, n8n, cloud storage or your own backend. These are platform connection options, not integrations preconfigured by this Actor.

### Reliability, retries, and public-data limits ⚠️

The Actor uses pure HTTP collection with up to five requests concurrently. It does not launch a browser. Invalid syntax and confirmed missing, removed or inaccessible targets stop without unnecessary retries. Temporary network and proxy failures, timeouts, blocks, malformed responses, throttling and server errors are retried up to five total attempts.

Earlier successful pages remain saved if a later page fails. If no publicly available rows are found before the source ends or the selected page budget is reached, one failed row records that input; this does not mean the underlying account or collection necessarily does not exist. A later empty page after successes is normal exhaustion. A later request that exhausts retries can append a failed row while preserving earlier results.

Bluesky controls public availability and returned fields. Objects restricted from unauthenticated viewing are excluded; a restricted target produces a failed row. Deleted, suspended, unavailable or otherwise restricted records may be missing. Optional counts are not inferred from an incomplete sample, and media links may change or expire.

This searches accounts, not post text, hashtags, likes or comments. Search results can change between runs and may omit optional profile detail fields. The Actor uses the returned search profiles without an additional detail request for each match.

Three consecutive result pages were verified for every paginated mode in HTTP runs of the packaged Actor on September 11, 2026. These checks demonstrate working continuation on the tested sources and do not guarantee future source availability.

One invalid string inside an otherwise valid input list does not stop other inputs. A configuration that fails the input schema logs an input warning and exits without starting requests: for example, a list containing a number instead of a string, an unsupported mode or an invalid page-count type. Infrastructure failures such as startup errors, unavailable dataset storage or an unrecoverable result-save error can still stop the whole run. Result-save failures are not retried as fresh scraping requests.

### Frequently asked questions ❓

#### Does this search post text?

No. It collects account search results. A keyword describes the profile search, not a full-text search of posts.

#### Can I submit a phrase with spaces?

Yes. Internal spaces are preserved and the query is encoded correctly. Leading and trailing whitespace is removed.

#### Are results guaranteed to be relevant?

Bluesky controls account matching and ranking. Review names, biographies and source links yourself before treating a match as a particular person or organization.

#### Does it require Bluesky login or cookies?

No Bluesky login, password, session cookie or account token is accepted or required. The Actor uses HTTP requests used by the public website and respects restrictions on unauthenticated access.

#### What happens to invalid or unavailable inputs?

Invalid individual inputs are saved as failed rows without an HTTP request. Confirmed missing, removed or restricted targets also become failed rows without unnecessary retries. Temporary failures are retried up to five total attempts. Other inputs and already saved pages remain available.

#### Can I export results or automate collection?

Yes. Use the Apify dataset to download JSON, CSV, Excel, XML or other supported formats, or retrieve records through its API. Apify schedules and webhooks can connect repeated runs to your own workflow.

### Support, responsible use, and related actors 🛟

For a reproducible problem, open an issue in the Actor’s **Issues** tab. Include the run ID, approximate time, mode if relevant, page count, safe public input, expected result and actual result. Do not share access tokens, proxy credentials or other secrets.

Use public data responsibly and follow applicable privacy, copyright, contractual and platform requirements before storing, combining or redistributing collected information.

# Actor input Schema

## `inputList` (type: `array`):

Add one nonblank profile-search query per item, such as python, science or art. This searches accounts, not posts.

## `pagesPerSearch` (type: `integer`):

Number of result pages to attempt per input. Each page requests up to 25 items; actual public rows can be fewer. Bootstrap requests do not count as pages. Stops when the cursor ends or repeats.

## Actor input object example

```json
{
  "inputList": [
    "python"
  ],
  "pagesPerSearch": 1
}
```

# Actor output Schema

## `profiles` (type: `string`):

One unique matching profile per query. Failed rows preserve the original input.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "inputList": [
        "python"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapingmonkey/bluesky-profile-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "inputList": ["python"] }

# Run the Actor and wait for it to finish
run = client.actor("scrapingmonkey/bluesky-profile-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "inputList": [
    "python"
  ]
}' |
apify call scrapingmonkey/bluesky-profile-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapingmonkey/bluesky-profile-search-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/62XN5jOJScCLgkJTr/builds/1w27M9S2pbTbljLbJ/openapi.json
