Reddit User Profile History Scraper
Pricing
from $1.50 / 1,000 posts
Reddit User Profile History Scraper
Export a Reddit account's profile, its posts and its comments. Paste one or many usernames, handles or profile URLs and get one clean row per profile, per post and per comment, with karma, community, flair, media, engagement and URL fields ready to export.
Pricing
from $1.50 / 1,000 posts
Rating
0.0
(0)
Developer
APISmith
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Export a Reddit account's profile, its posts and its comments — one clean row per profile, per post and per comment.
Paste one or many accounts in any format you have lying around (a bare handle, a u/-prefixed
one, a full profile URL) and get a dataset that exports to CSV without a column shifting:
every row of a type carries the same fields in the same order, with null where the data
source has no value.
What you get
dataType | Rows | Fields | What it is |
|---|---|---|---|
user_profile | one per account | 29 | Karma breakdown, icon, bio, verification, account age, profile URL |
post | up to maxPostsCount per account | 75 | Title, body, community, score, ratio, flair, media, crossposts, derived engagement |
comment | up to maxCommentsCount per account | 41 | Body, score, the commented post, derived length/word counts |
Rows of all three types land in the same default dataset and are told apart by dataType.
crawledAt is the same value on every row of a run, so a run groups by run rather than
scattering across however long it took.
Every run also produces, for each account, a profile row whatever else happens — an account that has no posts, or that the data source could not serve, still gets its profile row and a line in the Run log saying so.
Input
| Field | Type | Default | Notes |
|---|---|---|---|
usernames | string list | — | Required. spez, u/spez, /user/spez and https://www.reddit.com/user/spez/ all work, mixed freely in one list. Duplicates are collapsed case-insensitively, so listing both spez and u/spez pays for the account once. |
maxPostsCount | integer | 50 | Per account. 0 exports the profile and comments without fetching any posts — not "fetch a page and discard it". |
maxCommentsCount | integer | 50 | Per account. 0 likewise skips the comment feed entirely. |
includeNSFW | boolean | false | A post whose own NSFW flag is set is dropped unless this is on. The dropped posts are counted in the Run log. |
Entries that cannot be read as an account (a comment permalink, a blank cell from a pasted spreadsheet column) are reported, not fatal: the run names them in the Run log and processes the rest. Only an input that names no usable account at all is refused.
Output columns
All results are written into a single dataset. Each row carries a dataType column
(user_profile / post / comment) that tells the row's kind apart. One input account
produces one user_profile row, up to maxPostsCount post rows and up to
maxCommentsCount comment rows.
The three Console tabs — Profiles, Posts, Comments — are column presets,
not separate result sets. Apify dataset views only pick which columns to show; they do
not filter rows. So every tab renders all rows, and a row of the "wrong" type shows
undefined in columns that do not apply to it (for example, the Profiles tab shows the
single profile row filled in and the post/comment rows mostly blank). To read a clean
type-specific table:
- switch to All fields and read by the
dataTypecolumn, or - fetch the dataset over the API and filter on
"dataType"(e.g."dataType": "post").
Three conventions worth knowing before you build on the data:
idis a fullname on profile and post rows (t2_…,t3_…) and a bare id on comment rows.parsedIdis always the bare form.images/galleryImages/mediaAssets/previousNamesare[], nevernull— a pinned key set has nowhere to put "absent", andnullwould read as "the source failed to answer" for a post that simply has no images.authorIdon a comment row is the comment's author, which for an account's own comment history is the account itself.
Coverage and known limits
This Actor is explicit about what it cannot fill, because an empty column with no explanation is
worse than a documented one. null means "this data source did not answer", and the fields where
that is true on every row of a type are listed here.
Always null on user_profile: awardeeKarma, awarderKarma, hasVerifiedEmail,
bannerImg, profileVisibility, hideFromRobots.
Always null on post: isOriginalContent, numCrossposts, numDuplicates, removedBy,
bannedBy, removalReason, modReasonTitle, isRobotIndexable, totalAwardsReceived,
gilded, videoUrl, media, secureMedia, mediaMetadata, galleryData.
Always null on comment: stickied, edited, editedAt, distinguished,
scoreHidden, totalAwardsReceived, gilded, authorFlairText, depth, controversiality,
isSubmitter, collapsed, collapsedReason.
Three limits are worth calling out on their own, because they are losses rather than absences:
- Comment parents are only resolved when someone pays for them.
parentId,parsedParentIdandparentKindarenullon every comment row unless the Actor is configured to look them up, because each one costs a request of its own. See Cost for the switch and what it buys. - Comment bodies are the first 300 characters. The comment feed serves a plain-text preview
with the markdown stripped and a hard 300-character cap, so
body,bodyHtml,bodyLengthandwordCounton a comment row are all measured on that preview. Reaching the full text means paging whole threads to find one comment — measured at 40–60 extra requests for the handful of long comments in a typical account, roughly doubling a run's request count. That trade was declined on the record rather than made silently. - URLs and bodies are emitted exactly as sent. A URL that arrives with
&keeps its&rather than being escaped to&, so the value in your dataset is the value you can fetch. Reddit's inline media placeholders in a post body are likewise left as the markdown that arrived, because the URLs they would expand into carry signatures that the payload does not contain.
Cost
The run's scope is logged before any data is fetched, as a bound on the number of upstream requests. Two that matter:
- Comment parents are not looked up by default. Resolving what each comment replied to
(
parentId,parsedParentId,parentKind) costs one request per comment: 50 extra requests on a 50-comment account, which turns an 8-request run into a 58-request one. So it stays off, and those three columns come outnullwhilepostCommentsCountstill fills for comments left on the account's own posts. (Switching it on is an Actor-environment setting on the owner's side — it is not part of a run's input, andCOMMENT_PARENT_ENRICHMENT=1is the value it reads.) Either way the opening plan line in the Run log says which one this run did. maxPostsCount: 0ormaxCommentsCount: 0genuinely removes the requests, which is what makes a profile-only run cost one request per account.
A run that hits the maximum charge for its run stops at an account boundary and is reported as a partial success, not as a failure: the rows it collected are delivered. The same label is used for a run stopped by the platform. Watch for it if you are comparing two runs' row counts.
Free plan limits
A free-plan run is capped so it cannot spend more than a few cents of paid upstream calls. The ceilings are a cost guard, not a quota — a paid plan lifts them:
- Accounts: at most 3 usernames per run.
- Posts: at most 20 posts per account.
- Comments: at most 20 comments per account.
The run logs the effective scope (accounts, posts and comments it will attempt, and how many upstream requests that may take) before any data is fetched, so you always see the cap that applied. If your input asks for more than the free ceiling allows, the run delivers the capped volume and reports a partial success rather than an error.
API
The dataset is plain JSON. Every row carries dataType (user_profile / post / comment) and
the same fields in the same order for its type, so you can split or pivot by dataType without a
column shifting. Download it as JSON or CSV from the run's Storage tab, or pull it
programmatically with the Apify API using the run's default dataset ID. crawledAt is identical
across a run's rows, which makes grouping runs by run trivial.
See Output columns for the full field list of each row type.
Privacy
This Actor reads public Reddit profile, post and comment data through a third-party data provider and writes the result into your run dataset. It does not keep a copy of the scraped data, run a browser, or store cookies. Your provider credential is supplied through the Actor's environment variable on Apify and is never written into the dataset or the rows.
Errors and run status
A run ends in one of three states, reported in the Run log and the dataset summary:
- success — every requested account was served to its limits.
- partial_success — rows were delivered but the run did not finish its input. Causes: a reached charge limit, an account the data source could not serve (suspended, deleted, or not found), or a platform stop. The run still exits 0 and the rows it collected are kept.
- failed — nothing usable was delivered. Causes: a missing data-source credential, an input that names no usable account, or a charge that could not be confirmed. The Run log names the cause in a single sentence.
Accounts the data source cannot serve are named in the Run log with the reason and are skipped individually — they never abort the rest of the run.
Support
Found a bug, a data gap, or a row that looks wrong? Open an issue on this Actor's Apify Store page (the Issues tab), or reach the author through the contact channel shown there. When you report, include the Run ID and the relevant lines from the Run log — they are the fastest way to reproduce what you saw.
Development
npm installnpm run build # tsc -> dist/npm test # unit + contract + parity testsnpm run lint
Offline end-to-end verification (no credential, no requests)
npm run e2e replays the run against captured responses and asserts on what the built Actor
actually wrote, in three scenarios: a run the captures can serve end to end with parent resolution
turned on, a capture gap (which must fail loudly and still send zero real requests), and a bare
run — no environment at all — which is the assertion that the default is the cheap one.
It builds first, needs no credential, and exits non-zero on any failed expectation.
Recording and replaying a real run
cp .env.example .env.local # then fill in the data-source credential# write the run's input to storage/key_value_stores/default/INPUT.jsonnpm run capture # one real run, responses recordedUPSTREAM_REPLAY=1 UPSTREAM_REPLAY_DIR=storage/upstream-responses npm run replay
Replay answers every request from a recorded body and aborts on a request that was never recorded — it never falls through to the network, so it cannot spend money and cannot quietly produce a degraded run that looks successful. It is a local-only switch: it is unreachable on the platform, which always reads live data.
Field parity
npm run check:parity pairs this Actor's output against a saved sample of expected rows, field by
field, and reports the differences in three buckets: fields this data source does not provide,
differences that were chosen deliberately, and values that are live counters and therefore not
comparable. It fails only on undeclared differences. The reasoning for every declared
difference is in src/contract.ts and docs/.
Local run notes: the local storage root is ./storage (override with CRAWLEE_STORAGE_DIR), and
a request count is a cost — an offline replay is the cheap way to iterate on anything downstream
of the fetch.