Reddit Scraper avatar

Reddit Scraper

Pricing

$1.00 / 1,000 results

Go to Apify Store
Reddit Scraper

Reddit Scraper

Collect public Reddit posts, comments, community details, and user activity from archived JSON records. Search within communities and export structured data.

Pricing

$1.00 / 1,000 results

Rating

0.0

(0)

Developer

MLG Data

MLG Data

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

2 days ago

Last modified

Share

Scrape public Reddit posts, comments, community details, and user activity into structured records. Export Reddit data to JSON, CSV, or a spreadsheet through the dataset, with community search and direct post URLs as a practical Reddit API alternative for archived public content.

This actor reads archived public Reddit records rather than authenticated account data. The archive may lag live pages, and a post that has since been edited or removed can still appear in its earlier captured state. Each row includes a dataType so posts, comments, communities, and users remain easy to separate after export.

What data can you extract from Reddit?

One dataset can contain four kinds of row. A community URL adds community metadata followed by posts. A direct post URL adds the post and its available comments. A user URL adds a profile summary and public posts; a user comments URL adds the profile summary and comments. A search term can find titles and comment text within the community you specify.

FieldDescriptionExample or availability
dataTypeRow kindpost, comment, community, user
idFull item identifiert3_1wq2vmx
parsedIdPost or comment identifier without prefix1wq2vmx
urlPublic discussion, community, or profile URLhttps://www.reddit.com/r/Python/
usernamePublic account name for a post, comment, or userPresent when captured
userIdPublic account identifierPresent when captured
titlePost headline or community titlePre-parsing change detection
communityNameCommunity with the r/ prefixr/Python
parsedCommunityNameCommunity without the prefixPython
bodyPost text or comment textEmpty after known removal markers
htmlHTML version of a text bodyOften empty in the archive
numberOfCommentsCount shown on a post0
upVotesReported up vote count1
upVoteRatioReported up vote ratio1
scoreReported score1
authorFlairAuthor flair textMay be empty
flairPost flair textDiscussion
isVideoVideo post markerfalse
isAdAdvertising markerfalse
over18Mature content markerfalse when known
imageUrlsCaptured preview, gallery, or image links[] if none
videoUrlsCaptured video fallback links[] if none
outboundUrlLinked destination when distinct from the discussion URLMay be empty
thumbnailThumbnail URLMay be empty
lockedReply lock markerfalse
stickiedPinned item markerfalse
spoilerSpoiler markerfalse
archivedArchived post markerfalse
createdAtOriginal creation time in UTC2026-09-25T17:46:25Z
retrievedAtTime the archive captured the record2026-09-25T17:46:51Z
scrapedAtTime this actor saved the rowUTC timestamp
parentIdParent post or comment ID on comment rowst3_... or t1_...
postIdParent post ID on comment rowst3_...
categoryCommunity name on comment rowsPython
numberOfRepliesReply count if the source supplies oneUsually empty
nameFull community identifiert5_...
headerImageCommunity header image URLMay be empty
iconUrlCommunity icon URLMay be empty
descriptionPublic community descriptionMay be empty
numberOfMembersSubscriber count at capture timeChanges over time
userIconUser icon URL if availableUsually empty
postKarmaArchived user post karmaAvailable on some user summaries
commentKarmaArchived user comment karmaAvailable on some user summaries

Fields are shared across row types, so a post does not contain user profile karma and a community does not contain comment text. Empty values are represented as null, except media lists, which are empty arrays. Counts and scores reflect the archived snapshot, not a promise about the current live page. The retrievedAt field helps you judge how old the snapshot is before using those numbers.

How to scrape Reddit

  1. Enter one or more public Reddit URLs in Start URLs, or add search terms and a Search community. Community, post, and user URLs are supported. Search terms are ignored when URLs are present.
  2. Set the total item limit and the per-target post and comment limits. A direct post URL can produce one post row and multiple comment rows, so leave room for both.
  3. Choose whether to include mature content, collect post comments, or include community details. For incremental collection, add a post or comment date limit.
  4. Run the actor and check the dataset row count. Download the resulting dataset as JSON or CSV, or connect it to a spreadsheet workflow.

A community URL such as https://www.reddit.com/r/Python/ is the simplest starting point. It can produce a community row and a sequence of posts. A post URL under /r/<community>/comments/<post-id>/ is useful when you need the comment tree for one discussion. A user URL under /user/<name>/ returns a user row and public posts. To collect that user's comments instead, use /user/<name>/comments/.

Input

NameTypeDefaultDescription
startUrlsURL listExample community URLPublic community, post, or user URLs. These take precedence over searches.
searchesString listEmptyTerms searched in post titles and optionally comment bodies. Requires searchCommunityName for text search.
searchCommunityNameStringEmptyCommunity for keyword search, without r/. The form suggests an example.
searchPostsBooleantrueSearch titles for each term.
searchCommentsBooleanfalseSearch comment bodies for each term.
searchCommunitiesBooleanfalseLook up an exact community name matching a term.
searchUsersBooleanfalseLook up an exact username matching a term.
skipCommentsBooleanfalseSkip comments on direct post URLs.
skipUserPostsBooleanfalseSkip public posts on user profile URLs.
skipCommunityBooleanfalseSkip community metadata on community URLs.
includeNSFWBooleanfalseInclude records marked as mature.
maxItemsInteger100Total dataset cap across targets; 0 removes the overall cap.
maxPostCountInteger100Post cap per community, user, or search term.
maxCommentsInteger100Comment cap per post, user comment URL, or search term.
maxCommunitiesCountInteger10Cap for exact community lookups.
maxUserCountInteger10Cap for exact user lookups.
postDateLimitISO dateEmptySkip posts older than this date.
commentDateLimitISO dateEmptySkip comments older than this date.
proxyConfigurationObjectPlatform proxyOptional network settings for the public source.

For a recent community sample:

{
"startUrls": [{"url": "https://www.reddit.com/r/Python/"}],
"maxItems": 35,
"maxPostCount": 35,
"skipCommunity": true,
"includeNSFW": false
}

For a focused keyword query, remove startUrls and use searches with searchCommunityName. The source requires a community for post title and comment body text search. Exact community and user lookups can be requested with their corresponding switches, but they should not be treated as broad discovery across all names. When a date limit is set, the actor reads newest records first and stops once it crosses the limit.

Output example

This shortened row came from a successful 35-item community run. The omitted fields in the actual row include media arrays, flags, and available counts. It represents one archive capture, so its scores and dates should be interpreted at that time.

{
"dataType": "post",
"id": "t3_1wq2vmx",
"parsedId": "1wq2vmx",
"url": "https://www.reddit.com/r/Python/comments/1wq2vmx/preparsing_change_detection/",
"title": "Pre-parsing change detection",
"communityName": "r/Python",
"parsedCommunityName": "Python",
"numberOfComments": 0,
"upVotes": 1,
"score": 1,
"flair": "Discussion",
"over18": false,
"createdAt": "2026-09-25T17:46:25Z",
"retrievedAt": "2026-09-25T17:46:51Z"
}

The run returned 35 rows. All 35 had dataType, id, url, title, username, communityName, and createdAt. That verifies the community post path for this example; it does not imply every optional field will be populated for every community or time period.

Use cases

  • Community monitoring: Save recent discussions from a defined community and compare new post IDs across scheduled runs. The ID lets you avoid counting repeated posts.
  • Topic research: Search titles and comments inside a community, then review discussion text and timestamps together. A specific community keeps the query focused and makes coverage easier to explain.
  • Conversation analysis: Start from a direct post URL to pair the original post with archived comments. Use postId and parentId to connect comment rows back to the discussion and their immediate parent.
  • Content inventory: Export post titles, text, links, media URLs, scores, and flair for a community. Separate post content from outbound links using url and outboundUrl.
  • Public activity review: Use a public user URL to examine archived posts or a user comments URL for comments. Karma totals may be present on the profile summary, but profile fields are thinner than post and comment fields.
  • Historical comparison: Keep createdAt, retrievedAt, and scrapedAt in downstream records. The three timestamps distinguish publication, archive capture, and your own export.

How much does it cost to scrape Reddit?

There is no published per-1,000-result price for this actor yet. One successful validation run produced 35 rows with platform usage of approximately $0.00005. That is one observed run, not a guaranteed unit price: runtime, source response time, retries, memory, proxy settings, and the number of target URLs can change usage.

For planning only, dividing that single observed usage by 35 yields about $0.00143 per 1,000 rows. At exactly that same rate, 100 rows would be about $0.00014; 1,000 rows would be about $0.00143; and 5,000 rows would be about $0.00714. These are arithmetic examples, not a Store charge or a promise about larger runs. A larger export may require more pages, and an unavailable source may spend time retrying before it returns no rows.

Set maxItems to avoid collecting more rows than you need. Set maxPostCount and maxComments as well when you have several targets: the first controls post rows per target and the second controls comment rows per target. If you run the actor regularly, keep the exported IDs and compare them with prior runs so you only process new records in your own workflow.

Tips for best results

Use a precise community URL when you know where a topic is discussed. This gives the actor a direct collection path and avoids relying on text matching. Use a post URL when comments are the main goal; a community listing alone returns posts, not all comments in every post. A user comments URL is different from a user profile URL and intentionally selects comment activity.

Start with maxItems between 30 and 100 so you can inspect coverage and field fill before raising limits. The source supports timestamp-based pagination for posts and comments; the actor moves to older records until it reaches your limit, the date cutoff, or the source stops returning rows. For longer history, run separate narrow community or user targets rather than assuming a single request will cover everything. The source can rate limit or return a temporary error, so rerun later if a source error is reported.

Use postDateLimit or commentDateLimit for recurring collection. An ISO date such as 2026-09-01 is accepted. The actor interprets a date without a timezone as UTC. Keep a small overlap between scheduled windows in your downstream process, then deduplicate by dataType and id. This protects against records arriving in the archive after their creation time.

Inspect retrievedAt before treating a count as current. Scores, subscriber counts, and comment counts are snapshots. For media exports, check both imageUrls and videoUrls; a link post may have an outboundUrl without either media list being populated. For text exports, body: null can mean that the archived record contained a removal marker, while some source records simply have no text body.

Limits

Live Reddit pages, their JSON listings, RSS feeds, and several internal routes returned access blocks or a human-verification page during discovery. The actor therefore uses a public archive of Reddit records. The archive is independent of live page availability, but it may lag, miss records, or retain an older state after edits or removals. It is unsuitable when you need a guaranteed live score, exact current member count, or a complete record of a fast-changing discussion.

Keyword text search is restricted to one community. Unrestricted global keyword search, live relevance sorting, and a complete popular-feed view are not available from the selected source. Community and user name lookups are exact matches. Post and comment pagination is time based, so very dense periods with many records at exactly the same second may require separate time windows to avoid gaps. A direct post URL can be looked up by its post ID, but comments are limited by maxComments and the archive's coverage.

Some leader-style fields are unavailable in many archive records. HTML text, a user icon, a user description, reply counts, and current profile flags are often empty. The actor leaves these values empty rather than inventing them. Public records that are deleted or removed after archive capture can still be present; evaluate them before redistribution. Private, login-only, and account-specific content is outside this actor's scope.

Use with connected agents

The input and output schemas make the actor usable through a connected workflow that can supply JSON input and read dataset rows. A useful request is: “Collect the 50 newest archived posts from r/Python, exclude mature posts, and summarize recurring questions with post URLs.” Another is: “For this public discussion URL, collect up to 100 comments and group them by parent ID.” Include the community, limits, and freshness requirement in the request so the result can be interpreted correctly.

FAQ

Is it appropriate to collect this data?

The actor handles public records. Respect the site's terms, applicable privacy law, and the expectations of participants. Do not use public usernames or comments for harassment, sensitive profiling, or other personal-data misuse. Recheck content before redistributing a historical archive copy, especially when a live post may have been edited or removed.

Do I need a login or account credentials?

No account credentials are accepted. The selected source exposes archived public records. It cannot see private communities, private messages, account settings, or content that requires a logged-in session.

Do I need a proxy?

The actor has a proxy setting and uses the platform's proxy path by default. The successful validation run used that setup. A proxy does not make live Reddit pages available here; the actor collects from the archived JSON source. Source rate limits or outages remain possible.

How fast is a run?

The 35-row validation run finished in about half a minute, including startup. A run with more pages, comments, targets, or retries can take longer. Use row caps and a date window to keep a scheduled job predictable.

Can I schedule and monitor collection?

Yes. Schedule runs at an interval that fits your use case and monitor their status and dataset row counts. Store previous IDs in your destination to identify new records. A successful run with fewer rows than expected should prompt a check of the community activity, date cutoff, and source freshness.

Can I export to a spreadsheet?

Yes. Download the dataset as CSV and open it in a spreadsheet, or pass JSON rows to an integration. Filter on dataType before building tables because the four row kinds have different fields.

Why is a field empty or a count outdated?

The archive only contains what it captured. A text body may have been removed, a media preview may not have been recorded, or a profile field may never have been in the source. Counts can change after retrievedAt. Treat nulls as missing source data, not as zero.

Can I search every community at once?

No. Text search needs searchCommunityName. This constraint keeps the query bounded and matches the selected source's search behavior. Supply a list of community URLs if your goal is to monitor several known communities.

Integrations

Use the dataset API to retrieve JSON rows after a run, download CSV for a spreadsheet, or send completed run data to a webhook. Scheduling and workflow connections can trigger repeat collection and compare IDs across runs. Keep dataType in every downstream record so fields from different row kinds are interpreted correctly.

Support

Open an issue on the Issues tab; we reply within 24h and add fields on request.