# Reddit User Profile History Scraper (`apple_yang/reddit-user-profile-history-scraper-api`) Actor

Export a Reddit account's profile, its posts and its comments. Paste one or many usernames, handles or profile URLs and get one clean row per profile, per post and per comment, with karma, community, flair, media, engagement and URL fields ready to export.

- **URL**: https://apify.com/apple\_yang/reddit-user-profile-history-scraper-api.md
- **Developed by:** [APISmith](https://apify.com/apple_yang) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit User Profile History Scraper

Export a Reddit account's **profile**, its **posts** and its **comments** — one clean row per
profile, per post and per comment.

Paste one or many accounts in any format you have lying around (a bare handle, a `u/`-prefixed
one, a full profile URL) and get a dataset that exports to CSV without a column shifting:
every row of a type carries the **same fields in the same order**, with `null` where the data
source has no value.

### What you get

| `dataType`     | Rows                                 | Fields | What it is                                                                         |
| -------------- | ------------------------------------ | ------ | ---------------------------------------------------------------------------------- |
| `user_profile` | one per account                      | 29     | Karma breakdown, icon, bio, verification, account age, profile URL                 |
| `post`         | up to `maxPostsCount` per account    | 75     | Title, body, community, score, ratio, flair, media, crossposts, derived engagement |
| `comment`      | up to `maxCommentsCount` per account | 41     | Body, score, the commented post, derived length/word counts                        |

Rows of all three types land in the **same default dataset** and are told apart by `dataType`.
`crawledAt` is the same value on every row of a run, so a run groups by run rather than
scattering across however long it took.

Every run also produces, for each account, a profile row **whatever else happens** — an account
that has no posts, or that the data source could not serve, still gets its profile row and a
line in the Run log saying so.

### Input

| Field              | Type        | Default | Notes                                                                                                                                                                                                                                |
| ------------------ | ----------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `usernames`        | string list | —       | **Required.** `spez`, `u/spez`, `/user/spez` and `https://www.reddit.com/user/spez/` all work, mixed freely in one list. Duplicates are collapsed case-insensitively, so listing both `spez` and `u/spez` pays for the account once. |
| `maxPostsCount`    | integer     | `50`    | Per account. `0` exports the profile and comments **without fetching any posts** — not "fetch a page and discard it".                                                                                                                |
| `maxCommentsCount` | integer     | `50`    | Per account. `0` likewise skips the comment feed entirely.                                                                                                                                                                           |
| `includeNSFW`      | boolean     | `false` | A post whose own NSFW flag is set is dropped unless this is on. The dropped posts are counted in the Run log.                                                                                                                        |

Entries that cannot be read as an account (a comment permalink, a blank cell from a pasted
spreadsheet column) are **reported, not fatal**: the run names them in the Run log and processes
the rest. Only an input that names no usable account at all is refused.

### Output columns

All results are written into a **single dataset**. Each row carries a `dataType` column
(`user_profile` / `post` / `comment`) that tells the row's kind apart. One input account
produces one `user_profile` row, up to `maxPostsCount` `post` rows and up to
`maxCommentsCount` `comment` rows.

The three Console tabs — **Profiles**, **Posts**, **Comments** — are **column presets,
not separate result sets**. Apify dataset views only pick which columns to show; they do
not filter rows. So every tab renders all rows, and a row of the "wrong" type shows
`undefined` in columns that do not apply to it (for example, the Profiles tab shows the
single profile row filled in and the post/comment rows mostly blank). To read a clean
type-specific table:

- switch to **All fields** and read by the `dataType` column, or
- fetch the dataset over the API and filter on `"dataType"` (e.g. `"dataType": "post"`).

<details>
<summary><code>user_profile</code> — 29 fields</summary>

`dataType`, `createdAt`, `crawledAt`, `id`, `parsedId`, `username`, `totalKarma`, `linkKarma`,
`commentKarma`, `awardeeKarma`, `awarderKarma`, `isGold`, `isMod`, `isEmployee`,
`hasVerifiedEmail`, `verified`, `iconImg`, `snoovatarImg`, `acceptFollowers`, `bio`,
`followersCount`, `bannerImg`, `isNsfw`, `previousNames`, `profileTitle`, `profileDescription`,
`profileVisibility`, `hideFromRobots`, `profileUrl`

</details>

<details>
<summary><code>post</code> — 75 fields</summary>

`dataType`, `title`, `body`, `authorName`, `communityName`, `upVotes`, `commentsCount`,
`postUrl`, `createdAt`, `crawledAt`, `id`, `parsedId`, `postType`, `flair`, `contentUrl`,
`images`, `bodyHtml`, `authorId`, `parsedAuthorId`, `communityId`, `parsedCommunityId`,
`parsedCommunityName`, `score`, `upvoteRatio`, `over18`, `isSelf`, `isVideo`, `isGallery`,
`spoiler`, `locked`, `hidden`, `archived`, `pinned`, `stickied`, `edited`, `editedAt`,
`distinguished`, `scoreHidden`, `isOriginalContent`, `numCrossposts`, `totalAwardsReceived`,
`gilded`, `domain`, `thumbnail`, `urlOverriddenByDest`, `subredditSubscribers`,
`authorFlairText`, `authorPremium`, `numDuplicates`, `removedByCategory`, `removedBy`,
`bannedBy`, `removalReason`, `modReasonTitle`, `isRobotIndexable`, `mediaType`, `hasMedia`,
`galleryCount`, `galleryImages`, `mediaAssets`, `videoUrl`, `media`, `secureMedia`,
`mediaMetadata`, `galleryData`, `ageHours`, `scorePerHour`, `commentsPerHour`,
`engagementTotal`, `commentToScoreRatio`, `isHighEngagement`, `titleLength`, `bodyLength`,
`wordCount`, `outboundUrlHost`

</details>

<details>
<summary><code>comment</code> — 41 fields</summary>

`dataType`, `body`, `authorName`, `subredditName`, `commentUpVotes`, `url`, `commentCreatedAt`,
`crawledAt`, `id`, `postTitle`, `bodyHtml`, `authorId`, `parsedAuthorId`, `subredditId`,
`parsedSubredditId`, `postId`, `parsedPostId`, `parentId`, `parsedParentId`,
`postCommentsCount`, `score`, `stickied`, `edited`, `editedAt`, `distinguished`, `scoreHidden`,
`totalAwardsReceived`, `gilded`, `authorFlairText`, `authorPremium`, `ageHours`, `scorePerHour`,
`bodyLength`, `wordCount`, `authorFullname`, `parentKind`, `depth`, `controversiality`,
`isSubmitter`, `collapsed`, `collapsedReason`

</details>

Three conventions worth knowing before you build on the data:

- **`id` is a fullname on profile and post rows (`t2_…`, `t3_…`) and a bare id on comment
  rows.** `parsedId` is always the bare form.
- **`images` / `galleryImages` / `mediaAssets` / `previousNames` are `[]`, never `null`** — a
  pinned key set has nowhere to put "absent", and `null` would read as "the source failed to
  answer" for a post that simply has no images.
- **`authorId` on a comment row is the comment's author**, which for an account's own comment
  history is the account itself.

### Coverage and known limits

This Actor is explicit about what it cannot fill, because an empty column with no explanation is
worse than a documented one. `null` means "this data source did not answer", and the fields where
that is true on **every** row of a type are listed here.

**Always `null` on `user_profile`:** `awardeeKarma`, `awarderKarma`, `hasVerifiedEmail`,
`bannerImg`, `profileVisibility`, `hideFromRobots`.

**Always `null` on `post`:** `isOriginalContent`, `numCrossposts`, `numDuplicates`, `removedBy`,
`bannedBy`, `removalReason`, `modReasonTitle`, `isRobotIndexable`, `totalAwardsReceived`,
`gilded`, `videoUrl`, `media`, `secureMedia`, `mediaMetadata`, `galleryData`.

**Always `null` on `comment`:** `stickied`, `edited`, `editedAt`, `distinguished`,
`scoreHidden`, `totalAwardsReceived`, `gilded`, `authorFlairText`, `depth`, `controversiality`,
`isSubmitter`, `collapsed`, `collapsedReason`.

Three limits are worth calling out on their own, because they are losses rather than absences:

- **Comment parents are only resolved when someone pays for them.** `parentId`, `parsedParentId` and
  `parentKind` are `null` on every comment row unless the Actor is configured to look them up, because
  each one costs a request of its own. See Cost for the switch and what it buys.
- **Comment bodies are the first 300 characters.** The comment feed serves a plain-text preview
  with the markdown stripped and a hard 300-character cap, so `body`, `bodyHtml`, `bodyLength`
  and `wordCount` on a comment row are all measured on that preview. Reaching the full text means
  paging whole threads to find one comment — measured at **40–60 extra requests** for the handful
  of long comments in a typical account, roughly doubling a run's request count. That trade was
  declined on the record rather than made silently.
- **URLs and bodies are emitted exactly as sent.** A URL that arrives with `&` keeps its `&`
  rather than being escaped to `&amp;`, so the value in your dataset is the value you can fetch.
  Reddit's inline media placeholders in a post body are likewise left as the markdown that
  arrived, because the URLs they would expand into carry signatures that the payload does not
  contain.

### Cost

The run's scope is logged **before** any data is fetched, as a bound on the number of upstream
requests. Two that matter:

- **Comment parents are not looked up by default.** Resolving what each comment replied to
  (`parentId`, `parsedParentId`, `parentKind`) costs one request **per comment**: 50 extra requests on
  a 50-comment account, which turns an 8-request run into a 58-request one. So it stays off, and those
  three columns come out `null` while `postCommentsCount` still fills for comments left on the
  account's own posts. (Switching it on is an Actor-environment setting on the owner's side — it is
  not part of a run's input, and `COMMENT_PARENT_ENRICHMENT=1` is the value it reads.) Either way the
  opening plan line in the Run log says which one this run did.
- **`maxPostsCount: 0` or `maxCommentsCount: 0` genuinely removes the requests**, which is what
  makes a profile-only run cost one request per account.

A run that hits the maximum charge for its run stops at an account boundary and is reported as a
**partial success**, not as a failure: the rows it collected are delivered. The same label is used
for a run stopped by the platform. Watch for it if you are comparing two runs' row counts.

### Free plan limits

A free-plan run is capped so it cannot spend more than a few cents of paid upstream calls. The
ceilings are a **cost guard, not a quota** — a paid plan lifts them:

- **Accounts:** at most **3** usernames per run.
- **Posts:** at most **20** posts per account.
- **Comments:** at most **20** comments per account.

The run logs the effective scope (accounts, posts and comments it will attempt, and how many
upstream requests that may take) **before** any data is fetched, so you always see the cap that
applied. If your input asks for more than the free ceiling allows, the run delivers the capped
volume and reports a **partial success** rather than an error.

### API

The dataset is plain JSON. Every row carries `dataType` (`user_profile` / `post` / `comment`) and
the same fields in the same order for its type, so you can split or pivot by `dataType` without a
column shifting. Download it as JSON or CSV from the run's **Storage** tab, or pull it
programmatically with the Apify API using the run's default dataset ID. `crawledAt` is identical
across a run's rows, which makes grouping runs by run trivial.

See [Output columns](#output-columns) for the full field list of each row type.

### Privacy

This Actor reads **public** Reddit profile, post and comment data through a third-party data
provider and writes the result into **your** run dataset. It does not keep a copy of the
scraped data, run a browser, or store cookies. Your provider credential is supplied through the
Actor's environment variable on Apify and is never written into the dataset or the rows.

### Errors and run status

A run ends in one of three states, reported in the Run log and the dataset summary:

- **success** — every requested account was served to its limits.
- **partial\_success** — rows were delivered but the run did not finish its input. Causes: a reached
  charge limit, an account the data source could not serve (suspended, deleted, or not found), or
  a platform stop. The run still exits 0 and the rows it collected are kept.
- **failed** — nothing usable was delivered. Causes: a missing data-source credential, an input
  that names no usable account, or a charge that could not be confirmed. The Run log names the
  cause in a single sentence.

Accounts the data source cannot serve are named in the Run log with the reason and are skipped
individually — they never abort the rest of the run.

### Support

Found a bug, a data gap, or a row that looks wrong? Open an issue on this Actor's Apify Store
page (the **Issues** tab), or reach the author through the contact channel shown there. When you
report, include the **Run ID** and the relevant lines from the **Run log** — they are the fastest
way to reproduce what you saw.

### Development

```bash
npm install
npm run build          # tsc -> dist/
npm test               # unit + contract + parity tests
npm run lint
```

#### Offline end-to-end verification (no credential, no requests)

`npm run e2e` replays the run against captured responses and asserts on what the built Actor
actually wrote, in three scenarios: a run the captures can serve end to end with parent resolution
turned on, a capture gap (which must fail loudly and still send **zero** real requests), and a bare
run — no environment at all — which is the assertion that the default is the cheap one.
It builds first, needs no credential, and exits non-zero on any failed expectation.

#### Recording and replaying a real run

```bash
cp .env.example .env.local     # then fill in the data-source credential
## write the run's input to storage/key_value_stores/default/INPUT.json
npm run capture                # one real run, responses recorded
UPSTREAM_REPLAY=1 UPSTREAM_REPLAY_DIR=storage/upstream-responses npm run replay
```

Replay answers every request from a recorded body and **aborts** on a request that was never
recorded — it never falls through to the network, so it cannot spend money and cannot quietly
produce a degraded run that looks successful. It is a local-only switch: it is unreachable on the
platform, which always reads live data.

#### Field parity

`npm run check:parity` pairs this Actor's output against a saved sample of expected rows, field by
field, and reports the differences in three buckets: fields this data source does not provide,
differences that were chosen deliberately, and values that are live counters and therefore not
comparable. It fails only on **undeclared** differences. The reasoning for every declared
difference is in `src/contract.ts` and `docs/`.

Local run notes: the local storage root is `./storage` (override with `CRAWLEE_STORAGE_DIR`), and
a request count is a cost — an offline replay is the cheap way to iterate on anything downstream
of the fetch.

# Actor input Schema

## `usernames` (type: `array`):

The accounts to export. A bare handle (spez), a u/-prefixed one (u/spez) and a full profile URL (https://www.reddit.com/user/spez/) all work, and duplicates are collapsed. Each account produces one profile row, up to maxPostsCount post rows and up to maxCommentsCount comment rows. Free-plan runs are limited to 3 accounts per run; a paid plan can export more.

## `maxPostsCount` (type: `integer`):

How many of each account's posts to export. Set 0 to export the profile and its comments without any posts. Free-plan runs are limited to 20 posts per account; a paid plan can export up to the entered value.

## `maxCommentsCount` (type: `integer`):

How many of each account's comments to export. Set 0 to export the profile and its posts without any comments. Free-plan runs are limited to 20 comments per account; a paid plan can export up to the entered value.

## `includeNSFW` (type: `boolean`):

Keep posts whose own NSFW flag is set. Off by default, so a run does not export adult content unless it was asked for.

## Actor input object example

```json
{
  "usernames": [
    "u/spez"
  ],
  "maxPostsCount": 50,
  "maxCommentsCount": 50,
  "includeNSFW": false
}
```

# Actor output Schema

## `results` (type: `string`):

Default dataset: one row per requested account, one row per post and one row per comment. Every row of a type carries the same fields in the same order, with `null` where this data source has no value — so the dataset exports to CSV without a column shifting. `dataType` tells the three apart.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "usernames": [
        "u/spez"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("apple_yang/reddit-user-profile-history-scraper-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "usernames": ["u/spez"] }

# Run the Actor and wait for it to finish
run = client.actor("apple_yang/reddit-user-profile-history-scraper-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "usernames": [
    "u/spez"
  ]
}' |
apify call apple_yang/reddit-user-profile-history-scraper-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,apple_yang/reddit-user-profile-history-scraper-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/hMr6dMugQSOZzd9bg/builds/AOcQhqw2XvwvnWtQ1/openapi.json
