Bluesky Post Search Scraper avatar

Bluesky Post Search Scraper

Pricing

from $3.50 / 1,000 bluesky post scrapeds

Go to Apify Store
Bluesky Post Search Scraper

Bluesky Post Search Scraper

Search public Bluesky posts by keyword, author, date, language, hashtag, mention, domain, or URL. Use for monitoring and research; not profiles, followers, feeds, or private data. Returns post text, author, time, engagement, media, and URL. 0.0035 USD/post + 0.00005 USD start; usage extra.

Pricing

from $3.50 / 1,000 bluesky post scrapeds

Rating

0.0

(0)

Developer

Khadin Akbar

Khadin Akbar

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

19 days ago

Last modified

Share

Bluesky Post Search Scraper — Posts & Metrics

Bluesky Post Search Scraper is an Apify Actor for searching public Bluesky posts by keyword, author, date, language, hashtag, mention, linked domain, or linked URL. It accepts one focused query per run and returns one structured record per post with text, author identity, timestamps, engagement counts, media fields, facets, and source URLs. The output is useful when you need searchable public post data for monitoring, research, and AI-agent workflows that work with one post record at a time.

Best fit and connected workflows

If direct Bluesky search is blocked, the Actor uses Apify Unblocker with at most two requests when the first has a transient error or malformed response. Set useUnblockerFallback to false for direct requests only. Each successful proxy request adds 10 Unblocker units to platform usage; see Apify proxy pricing. The same query and filters apply on both routes. OUTPUT.unblocker and RUN_SUMMARY.unblocker report whether that fallback succeeded. If every route remains unavailable, the run reports UPSTREAM_FAILED without charging for posts.

This Actor fits workflows that start with a targeted Bluesky search and end with a structured list of matching public posts. Typical routing includes:

  • topic monitoring from a phrase or Lucene-style query
  • author-scoped searches for public accounts
  • language, hashtag, mention, domain, and linked-URL filtering
  • Use follow-up review of selected posts with Bluesky Posts Scraper - Search, Feeds & Threads when you need a downstream workflow that operates on a returned public URL or ID

Because each dataset row represents one post, the output is easy to filter, store, enrich, or hand to an agent for later steps.

Practical scenario

Maya is tracking posts about a product launch. She starts with the query open source, sets lang to en, and keeps maxResults at 25. The Actor returns rows with postUrl, text, authorHandle, createdAt, likeCount, repostCount, replyCount, quoteCount, tags, and externalUrl. Maya spots one post that links to a release note page and has strong engagement. Her next action is to open the post URL for review and pass the public URL into a downstream enrichment workflow.

Input

The Actor accepts one required field, searchQuery, plus optional server-side filters.

FieldTypePurpose
searchQuerystringFocused search text or supported Lucene-style query
sortstringlatest or top ordering from Bluesky search
maxResultsintegerMaximum persisted posts, from 1 to 100
sincestringLower bound for indexed search time
untilstringExclusive upper bound for indexed search time
langstringBCP 47 language tag
authorstringPublic Bluesky author handle or DID
mentionsstringMentioned account handle or DID
domainstringLinked domain hostname
urlstringAbsolute linked URL
tagsarray[string]Hashtag facets, up to 10 tags
includeRepliesbooleanWhether reply posts may appear in results

Focused input example:

{
"searchQuery": "open source",
"sort": "latest",
"maxResults": 25,
"since": "2026-07-01",
"until": "2026-07-15",
"lang": "en",
"author": "jay.bsky.team",
"mentions": "",
"domain": "github.com",
"url": "",
"tags": ["opensource"],
"includeReplies": true
}

Output

Each dataset row contains one validated public Bluesky post.

FieldTypePurpose
searchQuerystringExact query used for the row
sortstringRanking mode used for the search
rankintegerOne-based rank in the run
postUristringStable AT Protocol URI
postCidstringCID of the returned post version
postUrlstringHuman-readable Bluesky post URL
textstringPublic post text
authorDidstringAuthor DID
authorHandlestringAuthor handle
authorDisplayNamestring or nullPublic display name when available
authorAvatarUrlstring or nullPublic avatar URL when returned
createdAtstringPost record timestamp
indexedAtstringAppView index timestamp
langsarray[string]Declared language tags
likeCountintegerLikes at collection time
repostCountintegerReposts at collection time
replyCountintegerReplies at collection time
quoteCountintegerQuote posts at collection time
engagementCountintegerSum of engagement counts
isReplybooleanReply indicator
replyRootUristring or nullReply root AT URI
replyParentUristring or nullDirect parent AT URI
linksarray[string]Rich-text link URLs
mentionedDidsarray[string]Mention facet DIDs
tagsarray[string]Hashtags found on the post
labelsarray[string]Content labels returned on the post view
imageUrlsarray[string]Public image URLs
imageAltTextsarray[string]Image alt text aligned by index
videoPlaylistUrlstring or nullHLS playlist URL for video embeds
videoThumbnailUrlstring or nullVideo thumbnail URL
externalUrlstring or nullExternal card URL
externalTitlestring or nullExternal card title
externalDescriptionstring or nullExternal card description
externalThumbUrlstring or nullExternal card thumbnail
quotedPostUristring or nullQuoted post AT URI
collectedAtstringDataset normalization timestamp
sourceEndpointstringBluesky AppView endpoint used
_notestring or nullAgent-readable continuation note

Illustrative output record:

{
"searchQuery": "open source",
"sort": "latest",
"rank": 1,
"postUri": "at://did:plc:example/app.bsky.feed.post/3abc",
"postCid": "bafyreiexample",
"postUrl": "https://bsky.app/profile/alice.bsky.social/post/3abc",
"text": "A new open-source release is available today.",
"authorDid": "did:plc:example",
"authorHandle": "alice.bsky.social",
"authorDisplayName": "Alice",
"authorAvatarUrl": null,
"createdAt": "2026-07-15T10:00:00.000Z",
"indexedAt": "2026-07-15T10:00:01.000Z",
"langs": ["en"],
"likeCount": 42,
"repostCount": 7,
"replyCount": 5,
"quoteCount": 2,
"engagementCount": 56,
"isReply": false,
"replyRootUri": null,
"replyParentUri": null,
"links": ["https://example.com/release"],
"mentionedDids": [],
"tags": ["opensource"],
"labels": [],
"imageUrls": [],
"imageAltTexts": [],
"videoPlaylistUrl": null,
"videoThumbnailUrl": null,
"externalUrl": "https://example.com/release",
"externalTitle": "Release notes",
"externalDescription": "What's new in this version.",
"externalThumbUrl": null,
"quotedPostUri": null,
"collectedAt": "2026-07-15T11:00:00.000Z",
"sourceEndpoint": "https://api.bsky.app/xrpc/app.bsky.feed.searchPosts",
"_note": null
}

The Actor also writes compact OUTPUT and diagnostic RUN_SUMMARY records to the key-value store.

How it works

This Actor searches public Bluesky posts through Bluesky's public AppView API. One execution executes one focused query with optional server-side filters and a hard result cap. The implementation uses the public app.bsky.feed.searchPosts route and normalizes the returned post view into dataset rows. The dataset schema keeps all listed fields explicit, with null or empty arrays where applicable, so downstream tools can read the data consistently.

Pricing

This Actor uses Pay per event plus Apify platform usage.

  • Actor start: charged once per run, scaled by the run's memory allocation
  • Bluesky post scraped: charged for each schema-valid post persisted to the default dataset
  • Apify platform usage: billed separately by Apify

A run that persists one hundred posts would charge one start event plus one hundred post events, and Apify platform usage is added separately. See the live Pricing tab in the Apify Console for the current billable values and totals.

Use with AI agents (MCP)

This Actor is usable through Apify MCP as a tool for public Bluesky post search. The exact Actor identity is khadinakbar/search-bluesky-posts.

The tool is suitable used with a focused query, optional filters, and a bounded maxResults so an agent can read the returned dataset and continue with a second step if needed.

Search public Bluesky posts for "open source" from the last month, keep results in English, and return up to 25 posts. Then summarize the most relevant post URLs and engagement counts.

Output interpretation:

  • dataset rows contain one post each
  • rank shows the order of persisted rows in the run
  • sourceEndpoint identifies the Bluesky-hosted AppView route
  • collectedAt records when the row was normalized
  • _note may help an agent continue after a capped execution Provenance and scope:
  • results come from public Bluesky AppView search
  • the Actor returns public posts, authors, timestamps, engagement, media, and URL fields
  • the search scope is post search, not profiles, followers, or private data

Pagination and cost guidance:

  • maxResults can be set from 1 to 100
  • the public search route currently returns one page per query
  • tighter filters can make follow-up runs more specific
  • each persisted post creates a post-scraped event charge

JavaScript example

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('khadinakbar/search-bluesky-posts').call({
searchQuery: 'open source',
sort: 'latest',
maxResults: 25,
lang: 'en'
});
const dataset = await client.dataset(run.defaultDatasetId).listItems();
console.log(dataset.items);

Best results and outcome guidance

Start with one focused query and then narrow by author, language, date, hashtag, mention, domain, or URL as needed. If you are building an agent workflow, keep the first run small, inspect text, authorHandle, createdAt, and postUrl, then decide whether a second run should target a more specific slice of the search space. Use includeReplies based on whether reply posts belong in the review set.

Continue the workflow

Design note

I found that the dataset contract keeps searchQuery, sort, rank, postUrl, text, authorHandle, createdAt, and engagement counts in every row, which makes each record self-describing for later analysis.

FAQ

When should I use author instead of searchQuery?

Use author when you want public posts from one Bluesky account and searchQuery when you are tracking a topic, phrase, or query syntax pattern.

When should I use mentions?

Use mentions when the post must include a structured mention facet for a specific account handle or DID.

When should I use domain or url?

Use domain when you want posts linking to a hostname, and url when you need the exact linked web address.

How do tags work?

tags targets structured hashtags. Multiple tags are matched together, and the tags are entered without the # prefix.

What does includeReplies change?

includeReplies controls whether reply posts may remain in the final dataset after search results are returned.

How should I choose between latest and top?

Use latest for newest indexed posts and top for relevance-ranked results.

How can I use this output in a downstream workflow?

Use the returned postUrl, postUri, or authorHandle as the starting point for a second Apify workflow or for agent-side enrichment.

Responsible use

This Actor searches public Bluesky posts only. Use the returned data in line with applicable law, Bluesky and AT Protocol terms, and your own privacy and research obligations. Keep collection proportional to the task, store data carefully, and apply special care when combining public social data with other sources.