Reddit Scraper - Posts, Comments & Users avatar

Reddit Scraper - Posts, Comments & Users

Deprecated

Pricing

from $0.14 / 1,000 item collecteds

Go to Apify Store
Reddit Scraper - Posts, Comments & Users

Reddit Scraper - Posts, Comments & Users

Deprecated

Scrape Reddit posts, comments, subreddits, users, activity, and search through 32 API endpoints with automatic pagination and structured Dataset output. Powered by Titan Network.

Pricing

from $0.14 / 1,000 item collecteds

Rating

0.0

(0)

Developer

Titan Network

Titan Network

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

17 days ago

Last modified

Categories

Share

Dataclawd Reddit Data Collector

A reliable Actor that collects Reddit data through the Dataclawd API — a managed social-data API with API-key authentication, usage reporting and prepaid credits. This Actor is a thin client: it calls the endpoint you choose (or maps a friendly task to the right endpoint automatically), follows after pagination automatically, extracts each item, and writes one row per item into the default Dataset (each successful row is also the pay-per-event billing event).

It covers 32 endpoints: subreddit hot/new/top/controversial posts, subreddit info/rules/similar, post details & comments, global popular posts, post/user/subreddit search, and user profiles / activity / comments / posts.

Powered by Dataclawd API · Not affiliated with Reddit, Inc.

Quick start

  1. Open the Actor's Input tab.
  2. Pick a task (collection task), e.g. subredditHot — the endpoint is selected automatically.
  3. Fill subreddit, e.g. java.
  4. Optionally set limit (per-page count, max 100), maxItems / maxPages to control volume, or add multiple objects to targets to batch-collect several subreddits.
  5. Click Run and open the Dataset tab when it finishes — every collected post is one row with the raw JSON in the data field.

Input

FieldTypeDefaultDescription
taskenum—Recommended. Collection intent, auto-maps to an endpoint (see Tasks below). Leave empty only if you set endpoint explicitly.
endpointenumautoAdvanced. auto = resolve from task; or set an explicit endpoint to override.
subredditstring—Subreddit name without r/ (e.g. java). Needed by postDetail/postComments/postDuplicates and all subreddit* endpoints.
usernamestring—Reddit username without u/ (e.g. elonmusk). Needed by all user* endpoints.
articlestring—Post ID from the comments path (e.g. 1b7abcd). Needed by postDetail/postComments/postDuplicates.
qstring—Search keyword. Needed by search / searchSubreddits / searchUsers.
sortstring—Listing sort: hot/new/top/rising/best (popular) or confidence/top/new/controversial/old/random/qa/live (comments).
countrystring—ISO country code (e.g. US) for popularCountry.
restrictSrstring—Set true to restrict search to a subreddit (with subreddit).
afterstring(empty)Reddit pagination cursor (e.g. t3_1b7abcd) to start from.
limitinteger25Items per page, max 100.
tenum(empty)Time filter for top endpoints: hour / day / week / month / year / all.
targetsarray of object(empty)Batch mode: each element is a param set, collected one after another.
proxyUrlstring(empty)Proxy used by Dataclawd's server when fetching Reddit.
httpMethodenumGETDataclawd Reddit endpoints are GET-only; POST returns 501.
maxItemsinteger100Stop after this many items in total (max 100 000).
maxPagesinteger10Max pages per target (max 100).
startCursorstring(empty)Alias for after; resume from a previous cursor.
itemsPathstring(empty)Dot-path to the item list, e.g. data.children. Leave empty for auto-detection (Reddit Listing children are auto-unwrapped).
nextCursorPathstring(empty)Dot-path to the next-page cursor. Leave empty for auto-detection (data.after).
paginationParamenumautoQuery param used for paging: auto (after), cursor, or none.

Endpoints

Tasks

Pick a task and the Actor selects the endpoint for you. Fill the params below (subreddit / username / article / q …).

TaskEndpointRequired params
popular Global hotpopular—
subredditHot Subreddit hotsubredditHotsubreddit
subredditNew Subreddit newsubredditNewsubreddit
subredditTop Subreddit topsubredditTopsubreddit
postDetail Post detailpostDetailsubreddit + article
postComments Post commentspostCommentssubreddit + article
search Searchsearchq
searchSubreddits Search subredditssearchSubredditsq
searchUsers Search userssearchUsersq
userProfile User profileuserProfileusername
userPosts User postsuserPostsusername
userComments User commentsuserCommentsusername

Posts — postDetail, postComments, postDuplicates Popular / homepage — popular, popularBest, popularCountry, popularRising, popularTop Search — search, searchSubreddits, searchUsers Subreddit — subredditHot, subredditNew, subredditTop, subredditAbout, subredditComments, subredditControversial, subredditPosts, subredditRules, subredditSimilar, subredditsNew, subredditsPopular User — userAbout, userActivity, userComments, userCommentsTop, userOverview, userPosts, userPostsHot, userPostsNew, userPostsTop, userProfile

Output (Dataset — one row per collected item)

FieldDescription
platformAlways reddit.
endpointThe endpoint that produced this row.
statuscompleted (data collected) or failed (request-level failure).
targetHuman-readable target summary, e.g. subreddit=java.
pagePage number this item came from (1-based).
itemIndexSequential index within the run (1-based).
cursorCursor used to fetch that page (-1 = first page).
dataThe collected item — the raw JSON returned by Dataclawd (a post or user object; Reddit Listing wrappers are unwrapped).
errorStageWhere a failure happened: connect_failed / server_error / auth_error / rate_limited / request_error / parse_error / empty_result.
errorCodeHTTP status code (if any), e.g. 429.
errorError message on failed rows.

Run status & status message

While running, the Actor's status message is updated periodically (Collecting … → Collecting …: wrote n items → Done: …). The run succeeds (SUCCEEDED) as soon as at least one item is collected. If every target fails or returns nothing, the run ends FAILED with an explanatory message. Billing happens only for completed rows.

How to verify the results

  1. Run status: the run is SUCCEEDED when at least one item was written; the status message shows live progress.
  2. Per-item state: open the run's Dataset — every collected item has status=completed with the raw JSON in data; failed rows explain why (errorStage / errorCode / error).
  3. Spot-check against the API: re-run one target in the Dataclawd API Console and compare with the row's data.
  4. Billing: only completed rows carry the item-collected pay-per-event charge.

Developer note: for automated checks (run status, row schema, optional Dataclawd API replay), use python3 tools/verify_output.py --run <RUN_ID> [--api].

Pricing

Pay-per-event (PPE): one item-collected event per successfully collected item. Failed rows do not carry this event.

Publishing note: define an item-collected event with your price in the Actor's Monetization tab. The platform's built-in synthetic event apify-default-dataset-item is charged automatically per default-dataset item when PPE is enabled; set its price to 0 if you don't want a per-item base fee.

Configuration (publisher)

The following environment variables are configured on the Actor's version settings and are not visible to end users:

  • DATACLAWD_API_HOST — Dataclawd API base URL (e.g. https://api.dataclawd.example).
  • DATACLAWD_API_KEY — Dataclawd API key (sent as X-API-Key).

Disclaimer

This Actor is not affiliated with, endorsed by, or sponsored by Reddit, Inc. Collect only data you are authorized to collect, and comply with Reddit's Terms of Service, the Dataclawd service terms, and applicable law.