Reddit User Profile History Scraper avatar

Reddit User Profile History Scraper

Pricing

from $1.50 / 1,000 posts

Go to Apify Store
Reddit User Profile History Scraper

Reddit User Profile History Scraper

Export a Reddit account's profile, its posts and its comments. Paste one or many usernames, handles or profile URLs and get one clean row per profile, per post and per comment, with karma, community, flair, media, engagement and URL fields ready to export.

Pricing

from $1.50 / 1,000 posts

Rating

0.0

(0)

Developer

APISmith

APISmith

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Export a Reddit account's profile, its posts and its comments — one clean row per profile, per post and per comment.

Paste one or many accounts in any format you have lying around (a bare handle, a u/-prefixed one, a full profile URL) and get a dataset that exports to CSV without a column shifting: every row of a type carries the same fields in the same order, with null where the data source has no value.

What you get

dataTypeRowsFieldsWhat it is
user_profileone per account29Karma breakdown, icon, bio, verification, account age, profile URL
postup to maxPostsCount per account75Title, body, community, score, ratio, flair, media, crossposts, derived engagement
commentup to maxCommentsCount per account41Body, score, the commented post, derived length/word counts

Rows of all three types land in the same default dataset and are told apart by dataType. crawledAt is the same value on every row of a run, so a run groups by run rather than scattering across however long it took.

Every run also produces, for each account, a profile row whatever else happens — an account that has no posts, or that the data source could not serve, still gets its profile row and a line in the Run log saying so.

Input

FieldTypeDefaultNotes
usernamesstring list—Required. spez, u/spez, /user/spez and https://www.reddit.com/user/spez/ all work, mixed freely in one list. Duplicates are collapsed case-insensitively, so listing both spez and u/spez pays for the account once.
maxPostsCountinteger50Per account. 0 exports the profile and comments without fetching any posts — not "fetch a page and discard it".
maxCommentsCountinteger50Per account. 0 likewise skips the comment feed entirely.
includeNSFWbooleanfalseA post whose own NSFW flag is set is dropped unless this is on. The dropped posts are counted in the Run log.

Entries that cannot be read as an account (a comment permalink, a blank cell from a pasted spreadsheet column) are reported, not fatal: the run names them in the Run log and processes the rest. Only an input that names no usable account at all is refused.

Output columns

All results are written into a single dataset. Each row carries a dataType column (user_profile / post / comment) that tells the row's kind apart. One input account produces one user_profile row, up to maxPostsCount post rows and up to maxCommentsCount comment rows.

The three Console tabs — Profiles, Posts, Comments — are column presets, not separate result sets. Apify dataset views only pick which columns to show; they do not filter rows. So every tab renders all rows, and a row of the "wrong" type shows undefined in columns that do not apply to it (for example, the Profiles tab shows the single profile row filled in and the post/comment rows mostly blank). To read a clean type-specific table:

  • switch to All fields and read by the dataType column, or
  • fetch the dataset over the API and filter on "dataType" (e.g. "dataType": "post").

Three conventions worth knowing before you build on the data:

  • id is a fullname on profile and post rows (t2_…, t3_…) and a bare id on comment rows. parsedId is always the bare form.
  • images / galleryImages / mediaAssets / previousNames are [], never null — a pinned key set has nowhere to put "absent", and null would read as "the source failed to answer" for a post that simply has no images.
  • authorId on a comment row is the comment's author, which for an account's own comment history is the account itself.

Coverage and known limits

This Actor is explicit about what it cannot fill, because an empty column with no explanation is worse than a documented one. null means "this data source did not answer", and the fields where that is true on every row of a type are listed here.

Always null on user_profile: awardeeKarma, awarderKarma, hasVerifiedEmail, bannerImg, profileVisibility, hideFromRobots.

Always null on post: isOriginalContent, numCrossposts, numDuplicates, removedBy, bannedBy, removalReason, modReasonTitle, isRobotIndexable, totalAwardsReceived, gilded, videoUrl, media, secureMedia, mediaMetadata, galleryData.

Always null on comment: stickied, edited, editedAt, distinguished, scoreHidden, totalAwardsReceived, gilded, authorFlairText, depth, controversiality, isSubmitter, collapsed, collapsedReason.

Three limits are worth calling out on their own, because they are losses rather than absences:

  • Comment parents are only resolved when someone pays for them. parentId, parsedParentId and parentKind are null on every comment row unless the Actor is configured to look them up, because each one costs a request of its own. See Cost for the switch and what it buys.
  • Comment bodies are the first 300 characters. The comment feed serves a plain-text preview with the markdown stripped and a hard 300-character cap, so body, bodyHtml, bodyLength and wordCount on a comment row are all measured on that preview. Reaching the full text means paging whole threads to find one comment — measured at 40–60 extra requests for the handful of long comments in a typical account, roughly doubling a run's request count. That trade was declined on the record rather than made silently.
  • URLs and bodies are emitted exactly as sent. A URL that arrives with & keeps its & rather than being escaped to &, so the value in your dataset is the value you can fetch. Reddit's inline media placeholders in a post body are likewise left as the markdown that arrived, because the URLs they would expand into carry signatures that the payload does not contain.

Cost

The run's scope is logged before any data is fetched, as a bound on the number of upstream requests. Two that matter:

  • Comment parents are not looked up by default. Resolving what each comment replied to (parentId, parsedParentId, parentKind) costs one request per comment: 50 extra requests on a 50-comment account, which turns an 8-request run into a 58-request one. So it stays off, and those three columns come out null while postCommentsCount still fills for comments left on the account's own posts. (Switching it on is an Actor-environment setting on the owner's side — it is not part of a run's input, and COMMENT_PARENT_ENRICHMENT=1 is the value it reads.) Either way the opening plan line in the Run log says which one this run did.
  • maxPostsCount: 0 or maxCommentsCount: 0 genuinely removes the requests, which is what makes a profile-only run cost one request per account.

A run that hits the maximum charge for its run stops at an account boundary and is reported as a partial success, not as a failure: the rows it collected are delivered. The same label is used for a run stopped by the platform. Watch for it if you are comparing two runs' row counts.

Free plan limits

A free-plan run is capped so it cannot spend more than a few cents of paid upstream calls. The ceilings are a cost guard, not a quota — a paid plan lifts them:

  • Accounts: at most 3 usernames per run.
  • Posts: at most 20 posts per account.
  • Comments: at most 20 comments per account.

The run logs the effective scope (accounts, posts and comments it will attempt, and how many upstream requests that may take) before any data is fetched, so you always see the cap that applied. If your input asks for more than the free ceiling allows, the run delivers the capped volume and reports a partial success rather than an error.

API

The dataset is plain JSON. Every row carries dataType (user_profile / post / comment) and the same fields in the same order for its type, so you can split or pivot by dataType without a column shifting. Download it as JSON or CSV from the run's Storage tab, or pull it programmatically with the Apify API using the run's default dataset ID. crawledAt is identical across a run's rows, which makes grouping runs by run trivial.

See Output columns for the full field list of each row type.

Privacy

This Actor reads public Reddit profile, post and comment data through a third-party data provider and writes the result into your run dataset. It does not keep a copy of the scraped data, run a browser, or store cookies. Your provider credential is supplied through the Actor's environment variable on Apify and is never written into the dataset or the rows.

Errors and run status

A run ends in one of three states, reported in the Run log and the dataset summary:

  • success — every requested account was served to its limits.
  • partial_success — rows were delivered but the run did not finish its input. Causes: a reached charge limit, an account the data source could not serve (suspended, deleted, or not found), or a platform stop. The run still exits 0 and the rows it collected are kept.
  • failed — nothing usable was delivered. Causes: a missing data-source credential, an input that names no usable account, or a charge that could not be confirmed. The Run log names the cause in a single sentence.

Accounts the data source cannot serve are named in the Run log with the reason and are skipped individually — they never abort the rest of the run.

Support

Found a bug, a data gap, or a row that looks wrong? Open an issue on this Actor's Apify Store page (the Issues tab), or reach the author through the contact channel shown there. When you report, include the Run ID and the relevant lines from the Run log — they are the fastest way to reproduce what you saw.

Development

npm install
npm run build # tsc -> dist/
npm test # unit + contract + parity tests
npm run lint

Offline end-to-end verification (no credential, no requests)

npm run e2e replays the run against captured responses and asserts on what the built Actor actually wrote, in three scenarios: a run the captures can serve end to end with parent resolution turned on, a capture gap (which must fail loudly and still send zero real requests), and a bare run — no environment at all — which is the assertion that the default is the cheap one. It builds first, needs no credential, and exits non-zero on any failed expectation.

Recording and replaying a real run

cp .env.example .env.local # then fill in the data-source credential
# write the run's input to storage/key_value_stores/default/INPUT.json
npm run capture # one real run, responses recorded
UPSTREAM_REPLAY=1 UPSTREAM_REPLAY_DIR=storage/upstream-responses npm run replay

Replay answers every request from a recorded body and aborts on a request that was never recorded — it never falls through to the network, so it cannot spend money and cannot quietly produce a degraded run that looks successful. It is a local-only switch: it is unreachable on the platform, which always reads live data.

Field parity

npm run check:parity pairs this Actor's output against a saved sample of expected rows, field by field, and reports the differences in three buckets: fields this data source does not provide, differences that were chosen deliberately, and values that are live counters and therefore not comparable. It fails only on undeclared differences. The reasoning for every declared difference is in src/contract.ts and docs/.

Local run notes: the local storage root is ./storage (override with CRAWLEE_STORAGE_DIR), and a request count is a cost — an offline replay is the cheap way to iterate on anything downstream of the fetch.