Transcript rows can now be told apart. The transcript endpoint returns caption data only - no title, subreddit, author or date - so a run of fifty videos produced fifty walls of text identified by nothing but a URL. You could not sort them, group them, or find one without reading it.
Two fields are now recovered from the post URL itself, at no extra request and no extra credit:
subreddit - exact, taken straight from the permalink
titleFromUrl - the URL slug made readable, named for where it came from because Reddit lowercases, strips punctuation and truncates it. It is a label to recognise a row by, not the real title.
A bare v.redd.it link or a URL without a slug yields null rather than a guess, and author, score and post date are never invented - they are not in the URL and would cost a second billed request per video to fetch.
videoId, captionUrl and rawVtt are now declared in the dataset schema too; they were being emitted but not documented.
Two output bugs, both found in real run data across all six jobs.
Every result from a site-wide search was typed mediaType: "unknown" - 60 rows out of 60. That endpoint omits both domain and is_self, the two signals the media check relied on, so nothing could be classified. A self post's url is its own permalink, which is the one signal the search response does carry, and that is now used. Image and video are also detected from an i.redd.it or v.redd.it url directly, rather than only from domain.
Subreddit profiles came back with no subscriber count. The field the job is named after, and the one the row is billed for, was absent for all three subreddits tested - the upstream response carried weekly_active_users and weekly_contributions but no subscribers. The transform now also accepts subscriber_count, subscribers_count, num_subscribers and total_subscribers, and when none of them arrives the run log says so for that subreddit instead of quietly returning a row without it.
Verified in the same data: comment depth was honoured, the per-query limit was exact, truncated reply chains were flagged, and removed comments kept their place in the thread.
The run log now names any ranking it had to swap. No single Reddit endpoint accepts all thirteen options in "Which results, and from when?", so a choice the target endpoint cannot take is converted to the nearest one that works - sending the original would be rejected outright. That conversion was silent, so asking for "most argued about" on a subreddit returned top-of-all-time with nothing to explain the difference, which reads as the actor ignoring its own input.
For subreddit listings there is no "most argued about" ranking, so this run uses "top" instead - the closest it supports.
The field description says the same thing up front, with an example.
A pure rename is deliberately not reported: site-wide search calls "most discussed" comment_count where the others call it comments, and the user gets exactly what they picked, so announcing a swap would be false.
Two wording fixes in the form.
"Which posts, and from when?" said posts, but the control is also what orders comments and search results. It is now "Which results, and from when?", and the description names exactly which jobs use it: subreddit posts and both searches use it for ranking and time window, comment threads pass it on as the requested comment order, and subreddit profiles and video transcripts ignore it entirely.
The "Narrow down the results" note still described a section of six filters after four were removed in 4.11. It now matches the two that remain and says plainly that filtering never increases what you pay.
Labels only - no field, type, default or behaviour changed.
Every job-specific option now says which job it belongs to. Apify forms cannot hide a field based on another field, so picking "Comments on a post" still showed the keyword box, both search toggles and the transcript language - four controls that do nothing for that job, with the explanation buried in each description.
Each of those five fields now opens with the same emoji as its job in step 1, so the form can be skimmed:
- 🔍 Search inside subreddits - the words to find
- 💬 Comments on a post - how deep into reply chains?
- 🎬 Video transcript - language
They are also grouped by job rather than interleaved, and the section note now says plainly that an option whose label does not match your choice is ignored.
The longest description in the form - "How many results?" - was cut roughly in half.
Labels and ordering only. Every field key, type, enum, default and prefill is byte-identical, verified against the previous build, so saved Tasks and API callers see no change at all.
Removed the four quality filters from the form: minimum upvotes, minimum comments, keyword matching and its include/exclude mode. The form is down to 14 fields.
They worked, and the upstream API has no equivalent for any of them. The problem was what they cost to satisfy. The paging loop counts rows that survive filtering, so a strict filter makes the actor buy many more upstream pages to deliver the same number of billable rows - a filter passing one row in twelve turns a four-page run into a fifty-page one, for identical revenue. The existing guard only stops after three consecutive pages where nothing survived, so a filter letting through one or two rows per page never triggered it.
Still available: "How far back to look" and "Include adult content".
Existing callers are unaffected. A Saved Task or API integration still sending any of the four keys gets exactly the filtering it configured - suddenly returning rows someone had deliberately excluded would be worse than the cost. compat-test pins both the removal and the legacy path.
The keyword filter said "must appear" but matches any one word. The code keeps a result when it contains at least one of the words, not all of them. The dropdown below it was already correct ("containing one of those words"), but the field title read as though every word had to be present - so adding three words looked like it would narrow results when it actually widened them.
Retitled to "Words to look for (any one is enough)", with the rule stated in the description and in the mode dropdown. Behaviour is unchanged, and the field key is unchanged, so saved Tasks and API callers are unaffected.
Removed the two exact-date inputs from the form. Exact start date and Exact end date are gone; How far back to look remains and covers the same job with one choice instead of three overlapping controls. The form is down to 18 fields.
What is lost: an exact calendar window such as 1-15 March. What is kept: every relative window from the last hour to the last year.
Existing callers are unaffected. A Saved Task or API integration still sending postedAfter or postedBefore gets exactly the window it asked for, including the rule that an explicit start date overrides the dropdown. Only the form entries are gone, and compat-test now pins that behaviour.
Stopped sending a parameter the search API does not have. /v1/reddit/search has no type parameter - the real one is filter (posts or comments). The actor had been sending type: 'link' on every site-wide search, where it was silently ignored. It now sends filter: 'posts', the documented default, stated explicitly so a change of default upstream cannot quietly alter results. The in-subreddit code path does not send it, because that endpoint has no filter parameter either.
No change to what any search returns.
Removed the "Speed and cost" option from the input form. It asked users to decide how stale a cached result could be, in exchange for a saving they could not see, on a question most people have no basis to answer. The form is down to 20 fields.
Its description also overstated its reach: it claimed to cover subreddit profiles, which never read it.
Existing callers are unaffected. A Saved Task or API integration created while the field existed still sends cacheMaxAge, and the actor still honours it - silently ignoring a value someone deliberately set would change their results without telling them. Only the form entry is gone.
Output tab fixed properly. 4.6 removed clean=true but left format=json, which failed the same way for a different reason: format is not a dataset listing option at all, so the Console rejected it outright.
The Console feeds a link's query string straight into the Apify API client's listItems(), which validates it. A query string can only ever supply strings, so the only options usable in a template are the string-typed ones - view, unwind, signature. Booleans (clean, desc, skipEmpty, skipHidden), numbers (limit, offset) and arrays (fields, omit, flatten) all fail on type, and anything outside that list fails as unexpected.
Templates now carry view and nothing else. The CSV link was removed - CSV requires format, which the viewer forbids; Export above the results table still offers CSV, Excel and XML.
The test now checks every query parameter against the real option list and its type, rather than guessing at one symptom at a time.
Fixed "Cannot load data from dataset" on the Output tab. The output schema links added in 4.4 carried clean=true. The Console parses a link's query string into a typed object and expects clean to be a real boolean, so the string "true" was rejected and the Output tab refused to load anything.
clean=true only skips empty items and hidden fields, neither of which this actor produces, so it was removed rather than reworked. The test now fails on any boolean-valued query parameter in a template.
The actor now honors the user's maximum-charge limit. Apify lets a user cap what a single run may charge them, and warns publishers that an actor ignoring that cap earns "no more revenue, but growing costs".
Actor.charge() does not throw at the cap - it quietly charges fewer events than asked and reports the real number in chargedCount, which this actor was discarding. A run that hit its cap therefore kept paging, kept delivering rows, and paid the upstream API for every one of them while earning nothing.
Rows are now trimmed to what is actually chargeable before they are delivered, so a delivered row is always a billed row. The charging manager is consulted before the first batch as well, otherwise a cap smaller than one page would give that whole page away. Reaching the cap stops the fetch loops, the same way the free-plan cap does: a 5,000-post request against a 10-event cap now makes one upstream request instead of two hundred.
Added an output schema. The Apify Console was flagging it as missing, and without one the Output tab has nothing curated to show - just a raw dataset.
Six links now appear on every finished run: all results as JSON, the same rows as CSV for a spreadsheet, posts only, comments only, AI analysis, and the run summary record. The three filtered links use the dataset views that already existed, so a run mixing posts and comments can be read one record type at a time without exporting first.
Validated locally against the constraints in Apify's own schema-definition repo rather than pushed and hoped for: version pinned to 1, every property typed string with a template, no keys outside the spec, every {{variable}} a real run variable, every referenced dataset view present, and the key-value record actually written by the code.
The Apify free plan now returns 5 results per run. Any paid Apify plan is unaffected and uncapped.
Apify excludes free-plan runs from publisher revenue, so those runs earn nothing while every row still spends upstream API credit. A small allowance keeps the actor genuinely free to try without it being usable in bulk for free. Apify permits this and requires it to be disclosed up front, so it is stated in the actor description, the input form, the README and the run log - and a capped run says plainly that it was capped rather than just returning short.
The 5 rows are the real output: same fields, same nesting, same structure. You are never billed for results the cap held back, and AI analysis respects the cap too, so a capped run never spends your Claude credits on rows you will not see.
The cap stops the fetch, not just the write. Capping only where rows are written would have left the paging loop running to whatever the user asked for, so a free-plan run requesting 1,000 posts would still have spent 1,000 posts of upstream API credit to deliver 5 rows - earning nothing and costing full price. The allowance is now applied when the fetch limit is resolved and re-checked before every page and every target, so a capped run makes roughly one upstream request instead of forty. Budget is reserved in a single synchronous step, because targets run concurrently and a read-then-await-then-write sequence let every concurrent target claim the same allowance.
Detection fails open by design. If the plan cannot be determined - no token, a scoped token, an API blip, an unfamiliar plan id - the run is treated as paid and nothing is capped. Wrongly capping a paying customer silently truncates their results and looks like a broken actor; wrongly letting a free run through costs a handful of rows. The tests assert every ambiguous case resolves to "paid".
Ten pricing events became five, and one of them was over-charging.
The old list split posts three ways depending on which endpoint produced them - subreddit-post, subreddit-search-match, reddit-search-result - all at the same $0.0015. Nobody could tell them apart on an invoice, and the distinction never changed what you paid. They are now one event, Post, at the same price. Nothing costs more than it did.
AI analysis had one event per kind of analysis. That over-charged: a row is marked analysed as soon as any selected kind comes back, but the bill was raised for every kind selected - so a post that got a sentiment label and no pain points was still charged the pain-point fee. There is now a single AI analysis event at $0.003 per analysed row, charged once however many kinds you tick, and never on a row the model failed to analyse. Every kind is answered in the same request on your own key, so charging five times for one call was never justified.
The five events now are Post ($0.0015), Comment ($0.001), Subreddit profile ($0.004), Video transcript ($0.008) and AI analysis ($0.003).
A misspelled event can no longer ship. Actor.charge() does not throw on an unknown event name, so a typo delivers rows for free while the run reports success - the failure mode behind the 4.1 billing-summary fix. Event names now live in one file that every mode imports, and a test asserts the code and the pricing file name exactly the same set, with no event priced but never charged and none charged but never priced.
Two bugs surfaced by the first live comment-thread run.
A successful run reported "0 dataset items" and warned that it had produced nothing. The platform's item count can lag a write that has only just happened, and the run read it immediately. A run that delivered 41 rows said it delivered none. The rows actually pushed are now the source of truth for that check.
The billing summary claimed results were charged when they were not. Actor.charge() does not throw when an event is missing from the actor's pricing configuration - it logs a warning and returns - so counting our own charge calls overstated what was billed. The summary now reads the platform's own tally per event and reports that instead, and names any event that is not configured:
These events are not configured in this actor's pricing, so nothing was charged for them: subreddit-post, reddit-comment. Your results are complete and free.
Confirmed working in the same run: nested comment replies at depth 0, 1 and 2 with parent links intact, cursor pagination across 3 pages, isOP on the post author's own reply, isDeleted preserving a removed comment's position, and repliesTruncated where Reddit held more back.
Reddit ad data is gone, and this actor no longer claims to provide it.
Reddit shut down its public Ad Library. The upstream endpoints that read it were permanently retired and now refuse every request, so no tool can return Reddit ad data any more. This is a change at Reddit, not a fault here, and nothing can bring it back.
What that means in practice:
- The two ad modes are removed from the form. A Saved Task or API call that still asks for them gets a plain explanation and finishes cleanly, and is never charged.
- The two ad pricing events have been deleted, rather than left in place to charge for something that cannot be delivered.
- The title, description and README no longer mention ad data.
Everything else is unaffected: subreddit posts, site-wide search, search inside subreddits, complete comment threads, subreddit profiles, video transcripts and AI analysis.
Three output bugs fixed, all found in real run data.
- Text posts were labelled
mediaType: "unknown". The upstream response does not always include Reddit's is_self flag, so self-posts fell through to unknown. A self.* domain now settles it, and they are labelled text.
mediaUrls contained the post's own permalink on text posts. A self-post has no media; the field is now empty for them.
- Galleries were labelled
link. A reddit.com/gallery/ URL is now recognised as gallery, and v.redd.it as video.
Upvote ratios are rounded to three decimals instead of being returned as 17-digit floats.
Dataset views rebuilt. There were six, and four of them showed two irrelevant columns on any single-mode run, because an Apify view can pick columns but cannot filter rows. There are now four - Everything, Posts, Comments, AI analysis - and each leads with a Row type column so a mixed dataset reads clearly.
A run that returns nothing now tells you why.
Filters that remove everything used to look identical to a broken actor: the log repeated the same "0/15" line dozens of times and the run ended clean with an empty dataset. Three fixes:
- Every page now reports what happened -
fetched 25, kept 0 (25 removed by filters) instead of a bare running total.
- Paging stops after three straight pages with nothing kept, and says which filter is the likely cause. Before this, a run could make dozens of requests to collect nothing.
- A run ending at zero results warns explicitly, pointing at filters first and then at spelling and capitalisation, since subreddit names are case-sensitive at the source.
Keyword filtering is now documented honestly. Reddit's subreddit-search endpoint returns titles without post bodies, so in that mode keywords can only match the title. A post titled "Cooling advice needed" is dropped by a keyword filter of "fan" even though the body is all about fans. The field description says so and recommends running without it first.
The exact-date boxes are now calendar pickers. Click a day instead of typing and hoping the format is right.
The date reader was widened at the same time, so it now understands the longer phrasings a picker can produce - "7 days", "1 week", "3 months" - alongside the short forms it already took (7d, 1w, 3mo) and full calendar dates. An unreadable date is skipped with a message naming the formats that work, rather than silently doing nothing.
Three fields that made you type now let you choose.
- How far back to look - a new dropdown (Last 24 hours, Last 7 days, Last 30 days, and so on) replaces guessing at a date format. The two exact-date boxes are still there for a specific window like 1 to 15 March, and an exact date overrides the dropdown.
- Transcript language - a list of 25 languages instead of a two-letter code you had to know.
- Which Claude model - three options instead of a free-text model name.
AI analysis is Claude-only now. The provider picker and service-address boxes are gone, so the section is three fields: what to analyse, your key, and which model. Choose from:
| Model | When to use it |
|---|
| Haiku 4.5 (default) | Sentiment and intent on short Reddit text. Cheapest by a wide margin and good enough for most work. |
| Sonnet 5 | Extraction where quotes need to be sharper. |
| Opus 5 | Nuanced material where accuracy beats cost. |
A rejected key now says so plainly and points at console.anthropic.com, instead of showing a raw status code.
Input rewritten for people who have never used this before.
Every title and description is now plain business English with a concrete example, and no internal field names leak into the interface. "How many results? (0 = every result available)" instead of limit.
Ad filters are now checkboxes. Picking industry, budget, format, placement and campaign goal used to mean typing exact tokens like FINANCIAL_SERVICES from a list buried in the description. All 30 are now tickable, each labelled by what it does ("Industry: Financial services", "Budget: High spend"). Still one field, so the form stays short.
One click now works. The form arrives pre-filled with a real, working example - r/SaaS, top of this week, 50 results - so pressing Start with no editing returns data.
Sections consolidated from seven fragmented groups down to four: the four things that matter at the top, then job-specific options, filters, speed and cost, and AI analysis.
Date fields validate as you type. postedAfter and postedBefore accept a date (2026-03-01) or an age (7d, 24h, 3mo, 1y) and reject anything else with a clear message instead of being quietly ignored.
Over-limit message is now actionable. Asking for more than the ceiling explains that your results are complete up to that point, suggests splitting the work across runs, and states that the ceiling exists to protect the shared data source and prevent surprise bills - not because of any technical limit.
No field keys changed. Existing Saved Tasks, schedules and API integrations keep working. A field that used to accept a single value and is now multi-select also accepts the old bare string (and a comma-separated one), converted automatically.
AI enrichment, using your own LLM key.
Pick any combination under 🤖 AI enrichment and every post and comment row gains AI-derived fields:
- Sentiment - label, score from -1 to 1, and a one-word emotion
- Intent - question, complaint, recommendation, announcement, discussion, comparison, other
- Pain points - verbatim spans describing a problem, with a category and severity
- Feature requests - verbatim spans asking for something
- Buying intent - a post-level score and the signals behind it
You supply the key and pay your provider for tokens directly. This actor charges $0.002 per row for sentiment or intent and $0.004 for the extraction layers, covering the batching, prompting, validation and retries rather than the inference. Works with Anthropic or any OpenAI-compatible endpoint, including Groq, Together, OpenRouter and a local server via the base URL field.
Things worth knowing:
- Nothing runs without an explicit opt-in. No enrichment selected, or no key supplied, means no LLM call and no enrichment charge. The scrape is unaffected.
- Enrichment never fails a run. A missing key, an unparseable response or a provider outage degrades to "no enrichment" and the scrape completes normally.
- You are only billed for rows that came back enriched. A row carries
enrichedAt when the model actually returned data for it; rows without it cost nothing extra.
- Enrichment runs after filtering, so rows your filters drop never consume your tokens.
- Quotes are verified. Pain points and feature requests are checked against the source text and discarded if the model invented them.
- All layers run in one batched call per group of rows, to keep your token spend down.
- Reddit text is treated as untrusted data. Post bodies are passed inside delimiters with an explicit instruction to ignore any directives they contain, so a post cannot steer the model.
- Buying intent describes the post, not the person. Person-level lead scoring is out of scope by design - it conflicts with Reddit's user agreement and with this actor's own terms.
Read of the upstream API specification found nine places where this actor's assumptions did not match the real endpoints. All are fixed. Two of them were losing data silently.
Nested replies were being dropped entirely. The API nests replies as replies.items; the code checked for a plain array and for Reddit's native data.children. Neither matched, so every comment below top level was discarded and commentDepth did nothing. Comment threads now come back complete.
Comment threads now paginate. The endpoint returns one page plus a cursor and the actor only ever made a single call. Deep threads are followed to the cap, and a post row carries commentsReturned and commentsTruncated so a partial thread says so instead of just ending.
Subreddit search was broken three ways. It paginated with after where the endpoint expects cursor, so only the first page ever returned. Its posts use a slimmer shape (votes, created_at_iso, subreddit as an object) that the native transform mangled into zero scores and null dates. And the endpoint returns matching comments and media alongside posts on the same request, which were being thrown away - both are now returned, controlled by two new toggles.
Subreddit info was wrong on nearly every field. The response is flat, not nested, and uses different names. activeUserCount, createdAt, bannerUrl and rules were null or empty on every run. rules is a markdown string, not a structured list, and is now returned as rulesText.
moderators has been removed. The endpoint returns no moderator roster. It was documented as a field no competitor offers and was always empty. Removing a false claim.
Sort and timeframe values now match each endpoint. Each accepts a different enum: subreddit listings have no controversial and no hour; site-wide search uses comment_count, not comments, and has no hot. Presets are translated per mode to the nearest supported value instead of being sent verbatim and rejected.
New: video transcripts. 🎬 Video transcript pulls the plain-text caption track from Reddit video posts. Nothing else in this category returns transcripts. $0.008 per transcript. When Reddit publishes no captions the row still comes back with transcriptAvailable: false, so "no captions" is distinguishable from "failed".
New: caching. Pick a window under ⚡ Caching and a cached response inside it is served instead of a fresh scrape, at no upstream credit cost. Useful for daily monitoring where slightly stale data is fine.
New fields on posts and comments, all previously fetched and discarded: pollData (options and vote counts), controversiality, isEdited / editedAt, isOriginalContent, flairRich, removal, crossPostParent, awardCount, gilded, repliesTruncated, subredditSubscribers, distinguished, isArchived, isPinned.
Ads now return every carousel card. content[] was flattened to card one; multi-card creative was fetched and then discarded. New adCreatives[] keeps all of them, plus creativeType, allowComments and advertiserAvatarUrl.
Not possible, and dropped from the plan: author karma, account age and gold/mod status. There is no user endpoint in this API, so those fields cannot be delivered at all.
Filters. Six optional filters in a new collapsed section. Leave it empty and nothing changes; every filter narrows what reaches the dataset, so it also lowers what you pay - filtered rows are never written and never billed.
postedAfter / postedBefore - absolute (2026-03-01) or relative (7d, 24h, 3mo, 1y). This closes the last gap against the rest of the category. Runs client-side after each page, so it works on every mode and every sort order, unlike Reddit's own coarse relative windows.
includeNSFW - on by default; turn off to drop over-18 content.
minScore - skip posts and comments below a score.
minComments - skip posts with little discussion.
filterKeywords + filterKeywordMode - match on title and body, then keep only matches or drop them.
For ads, a date range matches the ad's flight window: an ad is kept when the period it ran overlaps your range.
Input goes from 7 fields to 13, still around half the category median of 26.
Breaking: comments are now one dataset row each.
Previously, post_comments billed one result per comment but delivered a single dataset row containing the whole tree nested inside it. A 200-comment thread charged for 200 results and showed you 1 row. That is fixed.
- Every billed result is now one row you can see.
post_comments emits one _recordType: "post" row plus one _recordType: "comment" row per comment, and bills exactly once per row.
- Comment rows carry
parentCommentId, depth and the new replyCount, so the thread is fully reconstructible. They also carry postTitle, postUrl and subreddit so a CSV export stands on its own without a join.
- The pre-assembled nested tree is still written to the key-value store under
thread-<postId>, at no extra charge, for anyone who wants it ready-nested.
If you parse output[0].comments[], switch to filtering rows on _recordType === "comment". The legacy post_with_comments shape is no longer emitted.
Repricing. Moved from a flat $4.00 / 1,000 results to per-event pricing, so commodity content is cheap and the Ads Library data is priced for what it is.
| Event | Was (flat) | Now |
|---|
| Reddit post | $0.004 | $0.0015 |
| Subreddit search match | $0.004 | $0.0015 |
| Reddit search result | $0.004 | $0.0015 |
| Reddit comment | $0.004 | $0.001 |
| Subreddit profile | $0.004 | $0.004 |
| Reddit ad (Ads Library) | $0.004 | $0.012 |
| Reddit ad - full detail | $0.004 | $0.025 |
Content scraping is now 62-75% cheaper. Ads Library data, which no other Reddit scraper on Apify returns, is priced accordingly. Still no run-start fee.
Dataset schema. Documented 12 fields that the actor emitted but never declared: commentId, parentCommentId, depth, replyCount, text, isOP, isDeleted, isRemoved, postTitle, postUrl, awards, description. The 💬 view now shows flat comment rows instead of a nested blob.
- Listing rewritten around the Reddit Ads Library, the one capability no other Reddit scraper on Apify offers.
- README restructured into question-shaped sections; Terms of Service collapsed so product content is above the fold. Added a mode-picker, a full input reference, integration examples for n8n / Make / Python / JavaScript, and an honest limitations section.
- Seven modes: subreddit posts, Reddit search, post comments, subreddit search, subreddit info, Ads Library search, ad details.
- Simplified input to four core fields with a universal
targets box and a combined sort + time preset.
0 = all convention on limit and commentDepth, with ceilings enforced in code rather than the schema so they cannot be bypassed via the API. Exceeding a ceiling warns in the log and on the run status instead of truncating silently.
- Five dataset views, key-value store snapshots, and an OpenAPI 3 description.
- Concurrent execution across subreddits, queries, post URLs and ad IDs; dataset writes batched per page.