Reddit Posts Scraper Buyer Question Detection
Pricing
Pay per usage
Reddit Posts Scraper Buyer Question Detection
Reddit Posts Scraper Buyer Question Detection: Extract structured Reddit post data at scale. This actor gathers titles, scores, authors, dates, and engagement stats from any subreddit. Ideal for analysts, marketers, and developers who need reliable Reddit insights for monitoring or automation.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
API Empire
Maintained by CommunityActor stats
0
Bookmarked
8
Total users
1
Monthly active users
10 hours ago
Last modified
Categories
Share
Reddit Scraper — Extract Posts, Comments and Buyer-Intent Leads
Reddit Posts Scraper (Buyer Question Detection) reads subreddits, keyword searches or direct Reddit URLs and returns three things in one JSON row: the post itself, its comment thread, and a rule-based buyer-intent and question classification with a 0–100 lead score. Every response is typed JSON — no HTML, no selectors, no parsing. Filters can keep only questions, only buying-intent posts, or only threads nobody has answered yet, so a wide subreddit scan comes back as a short list you can actually act on.
What is Reddit Posts Scraper (Buyer Question Detection)?
Reddit Posts Scraper (Buyer Question Detection) is an Apify Actor that scans Reddit subreddits, keyword searches and direct post URLs, and returns every qualifying post as one dataset row carrying its full metadata, its comment thread, and a rule-based lead-detection layer: whether the post is a question, whether it shows buying intent, whether the thread is still unanswered, and a composite lead score from 0 to 100.
It runs logged out against old.reddit.com's HTML — the same page any anonymous visitor sees — so no Reddit account, login, cookie or OAuth application is required.
- Scrape posts from subreddits, keyword searches and full Reddit URLs, mixed freely in one run
- Detect questions and buying-intent phrases with a transparent, rule-based matcher you can extend with your own phrases
- Pull each qualifying post's top-level comment thread (with nested replies) to check whether it has actually been answered
- Score every post 0–100 on intent, question, freshness and unanswered signals
- Optionally add LLM intent/sentiment classification on top of the rule-based layer — off by default
- Export as JSON, CSV or Excel; no proxy account or IP rotation to manage yourself
⚠️ What to know before your first run
Two things decide whether your first run looks the way you expect.
Buyer-intent and question detection is rule-based text matching, not machine understanding — and it runs by default, whether or not you ever touch the AI settings. isQuestion triggers on a literal ? anywhere in the title, the title's first word being one of how what why when where which who whom whose can could should would does do did is are will has have any anyone anybody, or the substring does anyone, has anyone, is there, anyone or anybody appearing anywhere in the title. hasBuyingIntent triggers when any of looking for, looking to buy, in the market for, recommend, recommendation, recommendations, any suggestions, suggestions for, worth it, worth buying, should i buy, should i get, what should i get, help me choose, help me pick, which should i, best , alternative to, alternatives to, vs , budget, willing to pay, where can i buy, where to buy, how much does, how much is, pricing, price of, trying to decide, need a, need help choosing, any good, which one, purchase, planning to buy, thinking of buying or shopping for appears as a plain substring in the lowercased title and body combined — plus anything you add in intentKeywords. Because these are substring matches, not phrase-aware ones, expect real false positives in both directions: anyone/anybody flags a mid-sentence statement as a question, and short fragments like best or vs match inside unrelated text (best friend, cats vs dogs). AI enrichment (aiEnhancement) never replaces this layer — the rule-based detection always runs, on every post, whether or not AI is turned on.
proxyConfiguration defaults to no proxy. The schema ships with useApifyProxy: false prefilled, so a fresh run goes out on the container's direct IP. The Actor escalates to Apify's datacenter, then residential, proxy tiers automatically the moment a request comes back blocked — but for a large or repeated run, it's worth turning on Residential yourself from the start rather than paying for the first block.
What data does Reddit Posts Scraper (Buyer Question Detection) collect?
Every qualifying post produces one row that carries three layers of data: the post itself, its comment thread, and the computed lead-detection signals — plus a fourth, optional layer when AI enrichment is on.
| Data Type | Key Fields | JSON Field Names |
|---|---|---|
| Posts & metadata | Title, body, score, subreddit, permalink, flair, NSFW/spoiler flags | title, body, score, subreddit, permalink, linkFlairText, over_18, spoiler |
| Comment threads (nested, per post) | Top-level comments with author, text, score and nested replies | comments → array of {author, body, score, created_utc, replies[]} |
| Buyer-intent & question signals | Question flag and matched cue, buying-intent flag and matched phrase, unanswered flag, composite score | isQuestion, matchedQuestionCue, hasBuyingIntent, matchedIntentPhrase, isUnanswered, leadScore |
| AI intent/sentiment layer (optional) | LLM-classified intent, sentiment, buyer stage, topics, one-line summary | aiIntent, aiSentiment, aiBuyerStage, aiTopics, aiSummary |
Need more Reddit data?
This Actor is built for depth on a single signal — filtering wide subreddit and keyword scans down to the posts worth acting on. For breadth across many subreddits, user feeds and keyword searches without the intent filtering, Reddit Trends Scraper covers four input kinds in one flat schema. For a single user's complete post and comment history split by community, Reddit User Profile Posts And Comments Scraper By Subreddit goes deep on one account instead of wide across a topic.
How does Reddit Posts Scraper (Buyer Question Detection) differ from the official Reddit Data API?
Reddit's Data API is Reddit's own supported programmatic surface. Since Reddit moved it to paid access on 1 July 2023, it requires a registered application authenticated with OAuth 2.0 on every request, and use beyond Reddit's free non-commercial tier requires a commercial agreement with Reddit (CNBC, 1 June 2023; TechCrunch, 4 July 2023). Reddit Posts Scraper (Buyer Question Detection) reads the same public listings anonymously and returns them pre-filtered for buying intent, questions and unanswered threads.
| Feature | Reddit Data API | Reddit Posts Scraper (Buyer Question Detection) |
|---|---|---|
| Reddit account and registered app | Required | Not used |
| Authentication | OAuth 2.0 on every request | None — anonymous requests |
| Commercial use above the free tier | Requires an agreement with Reddit | Governed by Apify's terms and your own legal review |
| Buyer-intent / question filtering | Not provided — raw post and comment endpoints only | Built in, rule-based, applied to every post before it's saved |
| Rate limiting | harshmaur/reddit-scraper's Apify Store listing states Reddit's API caps free access at "600 requests/10min" (checked 25 July 2026 — not independently verified here) | Governed by your own requestDelay and proxy tier, not an OAuth quota |
| Comment access | Full comment tree and moderation actions via dedicated endpoints | Top-level comment thread (with nested replies) per qualifying post, depth-limited |
| Setup time | App registration and approval before the first call | Paste targets, press Start |
Use the Reddit Data API when you need write access, full moderation history, or a contractual basis for large-scale commercial use — it's the supported route and nothing here replaces it. Use this Actor when you want public posts pre-filtered for purchase intent and open questions, without registering an application.
Why do developers and teams scrape Reddit for buyer intent and questions?
Reddit is where purchase questions, tool comparisons and open complaints get written down in plain language before they reach anywhere else searchable. Four groups get the most out of filtering it for intent.
For sales and growth teams
hasBuyingIntent, isUnanswered and leadScore turn a subreddit scan into a warm-outreach queue: run buyingIntentOnly and unansweredOnly together against the three or four subreddits your buyers actually post in, sort new, and every row that comes back is someone who asked for a recommendation nobody has answered yet. matchedIntentPhrase shows exactly which phrase fired, so a rep can open the reply already knowing what the poster said they wanted. A daily scheduled run against the same subreddits turns this into a lead feed instead of a one-off scrape.
For AI engineers and agent builders
Every row is typed JSON with a transparent scoring rationale, which makes it usable as an agent tool without a parsing layer: an agent can call this Actor with a niche's subreddits, filter the response on leadScore above a threshold, and draft an outreach reply using title, body and the matched comments. Because isQuestion and hasBuyingIntent are rule-based rather than model-generated, they're deterministic and reproducible across runs — a cheap pre-filter before a more expensive LLM pass, or paired with the optional aiEnhancement layer for richer classification on the same rows.
For marketers and community teams
Turning on questionsOnly against a category subreddit surfaces exactly the threads worth a brand reply — genuine questions, not link shares or memes. Pair it with intentKeywords tuned to your own product category (a plugin name, a competitor, a feature term) and a weekly run becomes a community-listening feed that flags where your product gets asked about, not just mentioned. subreddit and permalink tell you exactly where to go reply.
For researchers and analysts
Everything returned is visible to an anonymous visitor — no login, no private subreddits, no member-only content — which matters for research scope and approvals. isQuestion, hasBuyingIntent and matchedIntentPhrase give a consistent, rule-based coding scheme for studying how purchase intent gets expressed in a community, without hand-labeling every post, and because the matching logic is fixed and documented, results are reproducible across a study rather than dependent on a black-box classifier.
How to scrape Reddit for buyer-intent leads (step by step)
Reddit Posts Scraper (Buyer Question Detection) runs on Apify. Start it from the Apify Console or call it through the Apify API.
- Open Reddit Posts Scraper (Buyer Question Detection) on Apify and click Try for free
- Add your subreddits, search keywords or Reddit URLs to
startUrls— this is the only required input - Turn on the filters you want —
questionsOnly,buyingIntentOnly,unansweredOnly— and setmaxCommentsToQualifyif you're using the unanswered filter - Set
maxPosts(per source),maxComments,sortOrderand, if you're sorting by Top or Rising,timeFilter - Click Start, then download the dataset as JSON, CSV or Excel, or pull it through the Apify API
What to do when Reddit changes its structure
The scraper is maintained, and its output schema is what stays stable on your end — field names and types don't change because Reddit redesigned a page. No specific turnaround time is promised for any given break.
What changed in Reddit scraping recently?
The defining change was Reddit moving its Data API to paid access on 1 July 2023, requiring OAuth-registered applications and pricing large-scale third-party use out of the free tier — several major third-party apps shut down rather than pay (CNBC, 1 June 2023; TechCrunch, 4 July 2023). Since May 2024, Reddit's Public Content Policy and an updated robots.txt (announced 25 June 2024) have formalised rate-limiting and blocking of crawlers and bots operating without a licensing agreement (TechCrunch, 25 June 2024).
On the anonymous-access side, Reddit's .json endpoints — the surface most DIY Reddit scrapers still reach for — are now hard-blocked for anonymous requests from cloud infrastructure across proxy tiers, which is why this Actor reads old.reddit.com's rendered HTML instead. For DIY scrapers this raises the floor: default HTTP clients get TLS-fingerprint-blocked, and the .json fallback most tutorials assume simply returns 403. For users of this Actor, that's already handled — public subreddit and search listings remain accessible through the HTML surface it's built against, and it's maintained against further changes there.
⬇️ Input
Fifteen parameters. Only startUrls is required.
| Parameter | Required | Type | Description | Example Value |
|---|---|---|---|---|
startUrls | Yes | array | One entry per line — a subreddit (buildapc or r/SaaS), a keyword that runs a Reddit search (crm recommendation), or a full Reddit URL. Duplicate subreddits are merged. At least one entry is required. | ["buildapc", "r/SaaS", "https://www.reddit.com/r/marketing/", "crm recommendation"] |
questionsOnly | No | boolean | Keep only posts detected as a question (title ends with ?, opens with how/what/which/should/anyone…, etc.). Default false. | true |
buyingIntentOnly | No | boolean | Keep only posts that match a purchase-intent phrase (looking for, recommend, worth it, alternative to, budget, best … for…, etc.). Default false. | true |
unansweredOnly | No | boolean | Keep only posts at or below the comment threshold below — still-open questions nobody has answered. Default false. | true |
maxCommentsToQualify | No | integer | A post counts as "unanswered" when it has this many comments or fewer. Minimum 0, maximum 1000, default 5. | 3 |
intentKeywords | No | array | Optional custom phrases (lowercased, substring match on title+body) added to the built-in buying-intent list — e.g. demo, pricing, switching from. One per line. | ["demo", "pricing", "switching from"] |
maxPosts | No | integer | Max posts to save per subreddit/keyword after filters. Minimum 1, maximum 1000, default 10. With filters on, the Actor pages deeper to find this many matches. | 25 |
maxComments | No | integer | Top-level comments to pull for each saved post. Minimum 0, maximum 1000, default 5. Set 0 to skip comments and run faster. | 5 |
sortOrder | No | string | How Reddit orders the scanned posts: hot, new, top or rising. Default "new". | "new" |
timeFilter | No | string | Recency window: hour, day, week, month, year or all. Only applies when sortOrder is top or rising — ignored for hot/new. Default "week". | "week" |
aiEnhancement | No | boolean | Adds aiIntent, aiSentiment, aiBuyerStage, aiTopics and aiSummary per post via an LLM call. Default false. The rule-based detection above runs regardless. | false |
aiModel | No | string | Model/provider for AI enrichment, auto-detected from the name prefix (claude-*=Anthropic, gpt-*/o1/o3=OpenAI, gemini-*=Google, grok-*=xAI, deepseek-*=DeepSeek, sonar*=Perplexity, mistral-*=Mistral). One of claude-haiku-4-5, claude-sonnet-5, claude-opus-4-8, gpt-4o-mini, gpt-4o, gpt-4.1-mini, o3-mini, gemini-2.0-flash-lite, gemini-2.0-flash, gemini-2.5-flash, grok-3-mini, deepseek-chat, sonar, mistral-small-latest. Default "claude-haiku-4-5". | "claude-haiku-4-5" |
aiApiKey | No | string (secret) | API key for the selected provider. Used only when aiEnhancement is on. Falls back to the matching provider env var (ANTHROPIC_API_KEY / OPENAI_API_KEY / GEMINI_API_KEY / XAI_API_KEY / DEEPSEEK_API_KEY / PERPLEXITY_API_KEY / MISTRAL_API_KEY) when left blank. | "" |
proxyConfiguration | No | object | Apify Proxy configuration. The fetch layer uses old.reddit.com HTML and auto-escalates direct → datacenter → residential on a block. Defaults to no proxy (useApifyProxy: false); Residential is recommended for large runs. | {"useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"]} |
requestDelay | No | integer | Seconds to pause between paginated requests and between sources. Minimum 0, maximum 60, default 1. | 1 |
Common pitfall: timeFilter is silently ignored unless sortOrder is top or rising — set it alongside hot or new and the Actor accepts it without error but never sends it to Reddit, so your results aren't actually windowed the way the input suggests.
Example input
{"startUrls": ["r/SaaS", "buildapc", "crm recommendation", "https://www.reddit.com/r/marketing/"],"questionsOnly": false,"buyingIntentOnly": true,"unansweredOnly": true,"maxCommentsToQualify": 3,"intentKeywords": ["demo", "pricing", "switching from"],"maxPosts": 25,"maxComments": 5,"sortOrder": "new","timeFilter": "week","aiEnhancement": false,"aiModel": "claude-haiku-4-5","aiApiKey": "","proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] },"requestDelay": 1}
⬆️ Output
Each dataset row is one qualifying post, carrying 32 fields on every row, plus 5 more when aiEnhancement is on. Export as JSON, CSV or Excel, or read the dataset through the Apify API. Rows are written as they're found, so a run you stop early still keeps what it already collected.
Scraped post & lead-detection row
{"post_id": "1n2m3k4","title": "Looking for an affordable CRM — any recommendations?","author": "founder_jane","created_utc": 1784916000,"num_comments": 2,"score": 14,"permalink": "/r/SaaS/comments/1n2m3k4/looking_for_an_affordable_crm_any_recommendations/","image_url": "","thumbnail_url": "","body": "We're a 5-person team outgrowing spreadsheets. Budget is under $50/mo. What's actually good at this price point?","comments": [{"author": "saas_sam","body": "Try Pipedrive or a lean HubSpot free tier first — both fit that budget.","score": 6,"created_utc": 1784922000,"replies": []}],"subreddit": "SaaS","success": true,"error_message": null,"isQuestion": true,"matchedQuestionCue": "?","hasBuyingIntent": true,"matchedIntentPhrase": "looking for","isUnanswered": true,"leadScore": 100,"publishedAt": "2026-07-24T18:00:00Z","postUrl": "https://www.reddit.com/r/SaaS/comments/1n2m3k4/looking_for_an_affordable_crm_any_recommendations/","authorUrl": "https://www.reddit.com/user/founder_jane","scrapedAt": "2026-07-25T09:14:02Z","over_18": false,"spoiler": false,"domain": "self.SaaS","outboundUrl": null,"linkFlairText": "Question","numCrossposts": 0,"authorFullname": "t2_9f8g7h","subredditType": "public","aiIntent": "buying_intent","aiSentiment": "neutral","aiBuyerStage": "consideration","aiTopics": ["CRM", "small business tools", "budget software"],"aiSummary": "A small SaaS team is asking for a budget-friendly CRM recommendation."}
The last five keys — aiIntent, aiSentiment, aiBuyerStage, aiTopics, aiSummary — are absent from the row entirely, not null, whenever aiEnhancement is off (the default) or a run has no valid provider key. The Actor only merges them onto the row when enrichment actually returns a result, so check for key presence rather than a null check if you're branching on them downstream.
Two more row shapes exist outside this normal flow, and both are pushed without the row_result charge (no charged_event_name, so neither bills you): a source whose very first page fails to fetch, and a source that throws an unhandled error mid-scan. Both are written as a small subset of the fields above — post_id: null, title: null, success: false, an error_message describing what happened, and scrapedAt — so a genuine block shows up in the dataset instead of silently vanishing. Filter on success === true to work only with billed, complete rows.
Search-sourced rows (from a keyword input rather than a subreddit) carry title, score, comment count and author, but Reddit's search-results markup doesn't expose a post body, thumbnail, subreddit type or NSFW flag — those come back as null, "" or false on search-sourced rows specifically, never guessed. upvote_ratio is not returned on any row: it only exists in Reddit's JSON API, which this Actor does not use, and nothing here fabricates a value in its place.
How does Reddit Posts Scraper (Buyer Question Detection) compare to other Reddit scrapers?
Claims below are as observed on each Actor's live Apify Store listing on 25 July 2026.
| Feature | Reddit Posts Scraper (Buyer Question Detection) | Generic Reddit scraper |
|---|---|---|
| Question detection | ✅ isQuestion + matched cue on every post | Not documented on signalengine/reddit-lead-finder, harshmaur/reddit-scraper or parseforge/reddit-posts-scraper's listings |
| Buying-intent matching scope | ✅ Title and body combined, plus your own intentKeywords | signalengine/reddit-lead-finder's listing states intent is "matched on the post title" only — a body-only signal is missed there |
| Unanswered-thread filter | ✅ unansweredOnly with a configurable comment threshold | Not documented on any of the three listings checked |
| Comments & engagement numbers | ✅ Nested comment thread plus real score/num_comments from the page | signalengine/reddit-lead-finder's score field "is always null — RSS doesn't expose engagement counts", and it does not return comments |
| Optional AI classification | ✅ 7 providers (Anthropic, OpenAI, Google, xAI, DeepSeek, Perplexity, Mistral), off by default | Not documented on signalengine or parseforge's listings |
If you're building an AI agent or a RAG pipeline, the output-format row is the decision-maker — parsing HTML inside an agent loop is a reliability failure mode, not a feature. harshmaur/reddit-scraper's listing documents the widest general-purpose field set of the three and MCP-connector delivery to Slack/Notion/Airtable; if broad post/comment/profile coverage without intent filtering is what you need, that's the honest recommendation over this Actor.
How many results can you scrape with Reddit Posts Scraper (Buyer Question Detection)?
maxPosts caps qualifying posts per source, not per run — minimum 1, maximum 1000, default 10. Three sources at maxPosts: 50 can return up to 150 rows total; ten sources at the maximum can return up to 10,000.
Pagination behaves differently depending on whether a filter is active. With no filters on, the Actor scans exactly maxPosts posts and saves all of them. With questionsOnly, buyingIntentOnly or unansweredOnly on, it pages deeper — scanning up to max(maxPosts × 20, 200) candidate posts per source, looking for enough matches to hit your target. Either way, a hard ceiling of 40 listing pages per source (roughly 1,000 scanned posts at old.reddit's ~25 per page) applies regardless of filters, so a narrow filter combination on a low-signal subreddit can return fewer than maxPosts even though nothing failed — the source simply ran out of listing to scan.
No benchmark run time is published here; actual duration depends on maxComments, requestDelay and how deep the Actor has to page to satisfy your filters.
Integrate Reddit Posts Scraper (Buyer Question Detection) and automate your workflow
Reddit Posts Scraper (Buyer Question Detection) works with any language or tool that can send an HTTP request.
REST API integration
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("<YOUR_USERNAME>/reddit-posts-scraper-buyer-question-detection").call(run_input={"startUrls": ["r/SaaS", "crm recommendation"],"buyingIntentOnly": True,"unansweredOnly": True,"maxPosts": 25,})for post in client.dataset(run["defaultDatasetId"]).iterate_items():if post.get("success") and post["leadScore"] >= 60:print(post["subreddit"], post["leadScore"], post["title"], post["postUrl"])
Works in Python, Node.js, Go, Ruby, cURL — any language that can make an HTTP request.
MCP for AI agents
This Actor doesn't ship a bespoke MCP server of its own, but Apify's own MCP server (mcp.apify.com) exposes any public Actor — including this one — as a callable tool for MCP-compatible clients. harshmaur/reddit-scraper's Apify Store listing (checked 25 July 2026) documents this exact integration path for Claude, ChatGPT/Codex and Cursor via mcp-remote; the same server works for this Actor by pointing the tools= value at <YOUR_USERNAME>/reddit-posts-scraper-buyer-question-detection.
Automation platforms (n8n, Make, LangChain)
In n8n, use the Apify node — or an HTTP Request node pointed at the Apify run endpoint with your token — and pass the same JSON shown above; a Filter node on leadScore then routes only high-signal posts onward. In Make, the Apify module supports run-and-wait, so a daily subreddit sweep can feed a Google Sheets, Airtable or Slack step directly without polling. For LangChain, wrap the Apify API call as a tool that takes startUrls and the three filter booleans as arguments — the flat JSON response needs no output parser before it's handed to the model.
Is it legal to scrape Reddit posts and comments?
Scraping publicly visible Reddit posts and comments is broadly treated as permissible where no authentication is bypassed, and Reddit Posts Scraper (Buyer Question Detection) collects only what an anonymous visitor already sees — no account, no cookie, no private or restricted community content.
This output is personal data. Reddit posts and comments carry usernames, author profile links and user-generated text, so GDPR, CCPA and equivalent regimes attach to it — a pseudonymous Reddit username is still personal data under GDPR, and pseudonymisation reduces risk without removing the obligation to have a lawful basis for storing and using these records. Reddit's own terms of service and its Public Content Policy are separate contractual considerations, worth reading before a sustained collection programme.
Consult legal counsel for commercial use cases involving bulk personal data.
❓ Frequently asked questions
Does Reddit Posts Scraper (Buyer Question Detection) work without a Reddit account?
Yes. No Reddit account, login, cookie or OAuth application is used anywhere in the Actor — requests are anonymous against old.reddit.com. The only credential you need is your Apify token if you're calling it through the API, plus an AI provider key only if you turn on aiEnhancement.
How often is the scraped data updated?
Every run fetches live. Nothing is cached between runs — each run requests the listing, search and comment pages fresh, so results reflect Reddit as of the moment you start it. Use Apify schedules for repeat monitoring rather than expecting a cached snapshot to refresh itself.
What happens when a subreddit scan or keyword search returns no qualifying posts?
You get zero rows for that source and the run continues to the next one — no error, no charge for a post that was never found. If the source's first page fails to fetch entirely (blocked or unreachable), an uncharged accounting row with success: false and an error_message is written instead, so a genuine block is visible in the dataset rather than silently indistinguishable from "no matches."
Can I scrape private or restricted subreddits?
No. Only publicly accessible subreddits and search results are returned — the Actor never authenticates, so private, restricted and quarantined communities are not reachable.
How is Reddit Posts Scraper (Buyer Question Detection) priced, and is there a free trial?
Pricing and any trial terms are set on this Actor's own Apify Store listing page and shown before you run it — check there for the current pay-per-event rate rather than relying on a figure here. The Actor charges the row_result event only for rows that are actually saved to the dataset; error and accounting rows are never charged.
Does Reddit Posts Scraper (Buyer Question Detection) work for AI agent workflows and LLM pipelines?
Yes. It's callable as a standard HTTP-triggered Apify Actor run, so LangChain, CrewAI, n8n or a hand-written tool definition can invoke it and get typed JSON with no parsing step — or reach it through Apify's own MCP server, as described above.
How does Reddit Posts Scraper (Buyer Question Detection) handle Reddit's anti-bot system?
Fetches use browser-impersonated HTTP clients (impit Chrome TLS impersonation first, falling back to curl_cffi then plain aiohttp) rather than a default client signature. On a 403, 429 or 503 response — or when an expected page element is silently missing — the Actor escalates a proxy ladder: no proxy → Apify datacenter proxy → Apify residential proxy, retrying residential IPs up to three times before giving up on that page.
How does Reddit Posts Scraper (Buyer Question Detection) compare to other Reddit scrapers?
Checked on the Apify Store on 25 July 2026: signalengine/reddit-lead-finder is the closest in intent but matches buying phrases on the title only, never returns comments, and always returns score: null. harshmaur/reddit-scraper and parseforge/reddit-posts-scraper both scrape broad post/comment data without any question or buying-intent filter. This Actor's difference is combining title-and-body intent matching, question detection, an unanswered-thread filter and optional multi-provider AI enrichment on the same row.
Does Reddit Posts Scraper (Buyer Question Detection) return data in a format LLMs can use directly?
Yes. Typed, normalized JSON with stable field names on every row. No HTML, no selectors, no parsing step. Pass a row straight into an LLM context window, index it into a vector store, or route it through an agent tool.
Can I use Reddit Posts Scraper (Buyer Question Detection) without managing proxies?
Yes, once you turn proxy on — select Apify Proxy in proxyConfiguration and the Actor handles escalation and retries itself. Note the default is no proxy (useApifyProxy: false), so for a large or repeated run, switch on Residential yourself rather than relying on auto-escalation to kick in only after the first blocks.
What happens when Reddit changes its structure or blocks the scraper?
The scraper is maintained, and the output schema is what stays stable on your end — field names and types don't change because Reddit changed a page layout or tightened its blocking. No specific turnaround time is promised for any individual break.
🔗 Related scrapers
| Scraper Name | What it extracts |
|---|---|
| Reddit Trends Scraper | Posts across subreddits, user feeds, keywords and URLs in one flat schema |
| Reddit User Profile Posts And Comments Scraper By Subreddit | One user's full post and comment history, split by community |
| Google Search Autocomplete Scraper (Buyer Intent Filter) | Buyer-intent signals from Google's own autocomplete suggestions |
| Quora Search Scraper (Author Lead Enrichment) | Question-intent results with author enrichment |
| Twitter Trends Scraper | Trending topics, with where each trend is also trending |
💬 Your feedback
Found a bug, or a buying-intent phrase your niche uses that isn't in the built-in list? Open an issue on the Actor's Issues tab and it will be looked at. Reports that include your exact input JSON and the post title that misbehaved are the fastest to reproduce and fix.