Bluesky Post Search Scraper
Pricing
from $3.50 / 1,000 bluesky post scrapeds
Bluesky Post Search Scraper
Search public Bluesky posts by keyword, author, date, language, hashtag, mention, domain, or URL. Use for monitoring and research; not profiles, followers, feeds, or private data. Returns post text, author, time, engagement, media, and URL. 0.0035 USD/post + 0.00005 USD start; usage extra.
Pricing
from $3.50 / 1,000 bluesky post scrapeds
Rating
0.0
(0)
Developer
Khadin Akbar
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
19 days ago
Last modified
Categories
Share
Bluesky Post Search Scraper — Posts & Metrics
Bluesky Post Search Scraper is an Apify Actor for searching public Bluesky posts by keyword, author, date, language, hashtag, mention, linked domain, or linked URL. It accepts one focused query per run and returns one structured record per post with text, author identity, timestamps, engagement counts, media fields, facets, and source URLs. The output is useful when you need searchable public post data for monitoring, research, and AI-agent workflows that work with one post record at a time.
Best fit and connected workflows
If direct Bluesky search is blocked, the Actor uses Apify Unblocker with at most two requests when the first has a transient error or malformed response. Set useUnblockerFallback to false for direct requests only. Each successful proxy request adds 10 Unblocker units to platform usage; see Apify proxy pricing. The same query and filters apply on both routes. OUTPUT.unblocker and RUN_SUMMARY.unblocker report whether that fallback succeeded. If every route remains unavailable, the run reports UPSTREAM_FAILED without charging for posts.
This Actor fits workflows that start with a targeted Bluesky search and end with a structured list of matching public posts. Typical routing includes:
- topic monitoring from a phrase or Lucene-style query
- author-scoped searches for public accounts
- language, hashtag, mention, domain, and linked-URL filtering
- Use follow-up review of selected posts with Bluesky Posts Scraper - Search, Feeds & Threads when you need a downstream workflow that operates on a returned public URL or ID
Because each dataset row represents one post, the output is easy to filter, store, enrich, or hand to an agent for later steps.
Practical scenario
Maya is tracking posts about a product launch. She starts with the query open source, sets lang to en, and keeps maxResults at 25. The Actor returns rows with postUrl, text, authorHandle, createdAt, likeCount, repostCount, replyCount, quoteCount, tags, and externalUrl. Maya spots one post that links to a release note page and has strong engagement. Her next action is to open the post URL for review and pass the public URL into a downstream enrichment workflow.
Input
The Actor accepts one required field, searchQuery, plus optional server-side filters.
| Field | Type | Purpose |
|---|---|---|
searchQuery | string | Focused search text or supported Lucene-style query |
sort | string | latest or top ordering from Bluesky search |
maxResults | integer | Maximum persisted posts, from 1 to 100 |
since | string | Lower bound for indexed search time |
until | string | Exclusive upper bound for indexed search time |
lang | string | BCP 47 language tag |
author | string | Public Bluesky author handle or DID |
mentions | string | Mentioned account handle or DID |
domain | string | Linked domain hostname |
url | string | Absolute linked URL |
tags | array[string] | Hashtag facets, up to 10 tags |
includeReplies | boolean | Whether reply posts may appear in results |
Focused input example:
{"searchQuery": "open source","sort": "latest","maxResults": 25,"since": "2026-07-01","until": "2026-07-15","lang": "en","author": "jay.bsky.team","mentions": "","domain": "github.com","url": "","tags": ["opensource"],"includeReplies": true}
Output
Each dataset row contains one validated public Bluesky post.
| Field | Type | Purpose |
|---|---|---|
searchQuery | string | Exact query used for the row |
sort | string | Ranking mode used for the search |
rank | integer | One-based rank in the run |
postUri | string | Stable AT Protocol URI |
postCid | string | CID of the returned post version |
postUrl | string | Human-readable Bluesky post URL |
text | string | Public post text |
authorDid | string | Author DID |
authorHandle | string | Author handle |
authorDisplayName | string or null | Public display name when available |
authorAvatarUrl | string or null | Public avatar URL when returned |
createdAt | string | Post record timestamp |
indexedAt | string | AppView index timestamp |
langs | array[string] | Declared language tags |
likeCount | integer | Likes at collection time |
repostCount | integer | Reposts at collection time |
replyCount | integer | Replies at collection time |
quoteCount | integer | Quote posts at collection time |
engagementCount | integer | Sum of engagement counts |
isReply | boolean | Reply indicator |
replyRootUri | string or null | Reply root AT URI |
replyParentUri | string or null | Direct parent AT URI |
links | array[string] | Rich-text link URLs |
mentionedDids | array[string] | Mention facet DIDs |
tags | array[string] | Hashtags found on the post |
labels | array[string] | Content labels returned on the post view |
imageUrls | array[string] | Public image URLs |
imageAltTexts | array[string] | Image alt text aligned by index |
videoPlaylistUrl | string or null | HLS playlist URL for video embeds |
videoThumbnailUrl | string or null | Video thumbnail URL |
externalUrl | string or null | External card URL |
externalTitle | string or null | External card title |
externalDescription | string or null | External card description |
externalThumbUrl | string or null | External card thumbnail |
quotedPostUri | string or null | Quoted post AT URI |
collectedAt | string | Dataset normalization timestamp |
sourceEndpoint | string | Bluesky AppView endpoint used |
_note | string or null | Agent-readable continuation note |
Illustrative output record:
{"searchQuery": "open source","sort": "latest","rank": 1,"postUri": "at://did:plc:example/app.bsky.feed.post/3abc","postCid": "bafyreiexample","postUrl": "https://bsky.app/profile/alice.bsky.social/post/3abc","text": "A new open-source release is available today.","authorDid": "did:plc:example","authorHandle": "alice.bsky.social","authorDisplayName": "Alice","authorAvatarUrl": null,"createdAt": "2026-07-15T10:00:00.000Z","indexedAt": "2026-07-15T10:00:01.000Z","langs": ["en"],"likeCount": 42,"repostCount": 7,"replyCount": 5,"quoteCount": 2,"engagementCount": 56,"isReply": false,"replyRootUri": null,"replyParentUri": null,"links": ["https://example.com/release"],"mentionedDids": [],"tags": ["opensource"],"labels": [],"imageUrls": [],"imageAltTexts": [],"videoPlaylistUrl": null,"videoThumbnailUrl": null,"externalUrl": "https://example.com/release","externalTitle": "Release notes","externalDescription": "What's new in this version.","externalThumbUrl": null,"quotedPostUri": null,"collectedAt": "2026-07-15T11:00:00.000Z","sourceEndpoint": "https://api.bsky.app/xrpc/app.bsky.feed.searchPosts","_note": null}
The Actor also writes compact OUTPUT and diagnostic RUN_SUMMARY records to the key-value store.
How it works
This Actor searches public Bluesky posts through Bluesky's public AppView API. One execution executes one focused query with optional server-side filters and a hard result cap. The implementation uses the public app.bsky.feed.searchPosts route and normalizes the returned post view into dataset rows. The dataset schema keeps all listed fields explicit, with null or empty arrays where applicable, so downstream tools can read the data consistently.
Pricing
This Actor uses Pay per event plus Apify platform usage.
- Actor start: charged once per run, scaled by the run's memory allocation
- Bluesky post scraped: charged for each schema-valid post persisted to the default dataset
- Apify platform usage: billed separately by Apify
A run that persists one hundred posts would charge one start event plus one hundred post events, and Apify platform usage is added separately. See the live Pricing tab in the Apify Console for the current billable values and totals.
Use with AI agents (MCP)
This Actor is usable through Apify MCP as a tool for public Bluesky post search. The exact Actor identity is khadinakbar/search-bluesky-posts.
The tool is suitable used with a focused query, optional filters, and a bounded maxResults so an agent can read the returned dataset and continue with a second step if needed.
Search public Bluesky posts for "open source" from the last month, keep results in English, and return up to 25 posts. Then summarize the most relevant post URLs and engagement counts.
Output interpretation:
- dataset rows contain one post each
rankshows the order of persisted rows in the runsourceEndpointidentifies the Bluesky-hosted AppView routecollectedAtrecords when the row was normalized_notemay help an agent continue after a capped execution Provenance and scope:- results come from public Bluesky AppView search
- the Actor returns public posts, authors, timestamps, engagement, media, and URL fields
- the search scope is post search, not profiles, followers, or private data
Pagination and cost guidance:
maxResultscan be set from 1 to 100- the public search route currently returns one page per query
- tighter filters can make follow-up runs more specific
- each persisted post creates a post-scraped event charge
JavaScript example
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('khadinakbar/search-bluesky-posts').call({searchQuery: 'open source',sort: 'latest',maxResults: 25,lang: 'en'});const dataset = await client.dataset(run.defaultDatasetId).listItems();console.log(dataset.items);
Best results and outcome guidance
Start with one focused query and then narrow by author, language, date, hashtag, mention, domain, or URL as needed. If you are building an agent workflow, keep the first run small, inspect text, authorHandle, createdAt, and postUrl, then decide whether a second run should target a more specific slice of the search space. Use includeReplies based on whether reply posts belong in the review set.
Continue the workflow
- Then use Bluesky Posts Scraper — Search, Feeds & Threads to continue from Bluesky Post Search Scraper — Posts & Metrics discovery into content data for the selected records.
- Then use Skool Community Scraper to extend Bluesky Post Search Scraper — Posts & Metrics with a neighboring social-media research source when the brief calls for Skool data.
Design note
I found that the dataset contract keeps searchQuery, sort, rank, postUrl, text, authorHandle, createdAt, and engagement counts in every row, which makes each record self-describing for later analysis.
FAQ
When should I use author instead of searchQuery?
Use author when you want public posts from one Bluesky account and searchQuery when you are tracking a topic, phrase, or query syntax pattern.
When should I use mentions?
Use mentions when the post must include a structured mention facet for a specific account handle or DID.
When should I use domain or url?
Use domain when you want posts linking to a hostname, and url when you need the exact linked web address.
How do tags work?
tags targets structured hashtags. Multiple tags are matched together, and the tags are entered without the # prefix.
What does includeReplies change?
includeReplies controls whether reply posts may remain in the final dataset after search results are returned.
How should I choose between latest and top?
Use latest for newest indexed posts and top for relevance-ranked results.
How can I use this output in a downstream workflow?
Use the returned postUrl, postUri, or authorHandle as the starting point for a second Apify workflow or for agent-side enrichment.
Responsible use
This Actor searches public Bluesky posts only. Use the returned data in line with applicable law, Bluesky and AT Protocol terms, and your own privacy and research obligations. Keep collection proportional to the task, store data carefully, and apply special care when combining public social data with other sources.