Search works again. PullPush has closed, and the request layer was asking the
wrong network for each host.
A search for a plain word like shoes failed on the first query with HTTP 429,
and the run's advice — switch to the residential proxy group — was wrong.
Measured from three separate networks (a home connection, Apify's datacenter
pool and Apify's residential pool), PullPush answers every free request with
"this website does not provide free scraping resources". It is not an IP block,
so no proxy group restores it. PullPush is now dropped from a run the moment it
says so, instead of being retried for six minutes.
- Unscoped searches work again. Arctic-Shift only runs a text search that is
scoped to a subreddit or author, and PullPush was what served the unscoped
ones. Those queries now go through Reddit's own search feed, and the post ids
it returns are hydrated back through Arctic-Shift — so a searched post carries
exactly the same columns as a listed one, not a thinner row.
- Each host is now asked over the network that answers it. Measured from
inside the platform, three attempts per combination: Arctic-Shift answers the
platform's own egress every time and refuses the shared datacenter pool, while
Reddit's feed does the reverse — it rate-limits the platform's address within
a request or two and is content with a proxied one. Requests now start on the
route that works for that host and fall through to the others if it stops.
- A refused exit IP is swapped immediately rather than waited out. Apify's
shared pools contain addresses these hosts have already refused; the previous
code backed off 5, 10, 20, 40 and 60 seconds on the same kind of refusal,
which cannot help — the address is no better a minute later. A 403 now draws a
fresh IP within 250ms, up to four times, before changing route. Genuine rate
limiting (429) still backs off, because there the address is fine and the pace
is not.
- A 2xx carrying a block page no longer counts as success. Arctic-Shift
answered some proxied requests with Reddit's own block page; parsing that as
JSON threw an ordinary error, which left the run reporting SUCCEEDED with zero
rows. An unparseable or HTML body is now treated as a refusal, and a run that
saves nothing because every target failed now fails loudly whatever the cause.
- Searches can carry their own scope.
shoes subreddit:malefashionadvice or
shoes author:someuser is read out of the query and sent straight to
Arctic-Shift, which is faster and deeper than going via the feed.
- Unscoped comment search says so up front. No archive still serves a
comment text search without a subreddit or author scope. That target now
reports what to add instead of running for minutes and returning nothing.
- Records are de-duplicated per target, so a page boundary can no longer bill
the same post twice.
- Two IP rotations now happen even at
maxRetries: 0. Drawing a clean address
is not the same as hammering a source that has refused us — and the source
that has, PullPush, is no longer retried at all.
Measured after the change, from the platform: the reported failing input
(shoes, from 2026-09-01, 5 items, maxRetries: 0, default proxy) returns 5
rows in 15-24 seconds across four consecutive runs.
Investigated the seven columns reported as empty — the extraction is correct;
the shipped example was the problem. selftext, authorFlairText, media,
galleryImages, edited, distinguished and suggestedSort were all mapped
to the right source fields. They came back empty because the Actor's prefilled
example ran r/technology, a subreddit that shares nothing but external links.
Verified against the live archive: of 100 recent r/technology posts, 0% are self
posts, 0% carry flair, 0% carry media or galleries, and 0% are distinguished or
have a pinned sort. There is nothing there to extract.
-
Changed the prefilled example from ["technology"] to
["personalfinance", "wallstreetbets", "pics"] at 40 posts each, so a first run
actually exercises the schema: post bodies, flair, media, galleries, pinned
sorts and moderator posts all appear. Measured against the live archive these
subreddits return selftext on 55.8%, authorFlairText on 31.7%,
suggestedSort on 33.3%, galleryImages on 10.0% and media on 2.5% of a
120-row run — every one of them zero before.
-
edited and distinguished stay near zero on recent posts by nature: the
archives snapshot a post within minutes of publication, and moderator posts
are rare in ordinary traffic. Both fill properly once you query a historical
window, which is what this Actor is for — a run over r/announcements and
r/AmItheAsshole before 2024-01-01 returns edited on 26.2% and
distinguished on 48.8% of rows.
-
Documented, in the README and in each field's description, exactly which post
type each of these columns depends on, so an empty column is recognisable as
"this subreddit has no such posts" rather than a broken scrape.
Fixed silent data loss on any subreddit containing a gallery post. The
dataset schema declared galleryImages as a string while the Actor has always
emitted an array of image objects. Pushing a single Reddit gallery post
therefore failed schema validation and aborted that whole subreddit — every
remaining row for it was lost, and the run still reported success. Verified on
a live run before the fix: r/wallstreetbets and r/pics both died with "Schema
validation failed" and contributed 0 of their 40 rows. galleryImages is now
declared as an array and those targets complete. This also explains part of why
galleryImages looked empty in audits: the rows that would have carried it were
the exact rows being rejected.
-
isGallery was emitted as undefined for non-gallery posts, so the column
vanished from exports entirely instead of reading false. It now defaults to
false.
-
No change to the parsers, the backends or the set of output columns. Nothing
was removed: every one of the seven columns is real Reddit data that fills for
the right subreddit.
- Fleet-wide quality audit. Verified end to end against live data and re-checked the input schema, the output columns and the run configuration.
- Output verified on a live run: 40 columns returned, 80.8% of cells populated.
- Run reliability reviewed: 91.4% of public runs succeeded in the last 30 days.
- Input schema, output schema and pricing configuration reviewed.
- Noted that 7 column(s) came back empty in this sample (
selftext, authorFlairText, media, galleryImages, edited, distinguished and others); these are under review.
- Fixed the most common cause of a failed run. The Reddit archive backends block Apify's datacenter egress, and the run used to stop and tell you to switch the Proxy input to RESIDENTIAL by hand — after you had already lost the run. It now escalates to RESIDENTIAL automatically the first time a request chain is exhausted by blocks, and retries. Datacenter is still tried first, so a run that works on it is unchanged.
- A run is still failed loudly, with the reason, if even residential is blocked — a total block never masquerades as an empty result.
- Maintenance release: refreshed the build and dependencies.
- Re-verified live execution, non-empty structured output and dataset field/type integrity.
- Reviewed reliability (retries, pagination) and output quality as part of a full-fleet QA pass.
- Completed the August 2026 full health check: verified empty/programmatic default, Console UI default, and two source-informed alternative inputs on Apify.
- Confirmed successful live execution, non-empty structured output, dataset-field/type integrity, and logical sample quality within the 5-minute quality window.
- Expanded the dataset contract to 54 nullable and exhaustive post/comment fields, corrected media/gallery/upvote types, and removed non-emitted user/subreddit profile views and claims.
- Added a stable archived-thread variation that exercises comment
body, parent/link IDs, vote fields, and reconstructed depth.
- Stopped treating replies with omitted ancestors as top-level comments; depth is now null when the capped archive result cannot prove the ancestor chain.
- Fixed silent empty runs caused by IP blocks. Both archive backends block Apify's shared/datacenter egress IPs (Arctic-Shift → HTTP 522, PullPush → HTTP 403); the old code swallowed those errors and exited SUCCEEDED with 0 items.
- Apify Proxy is now ON by default (
proxyConfiguration default {useApifyProxy:true}). Every request mints a fresh rotating proxy session (new IP) per attempt, with retry/backoff and IP rotation on 403/429/5xx and Cloudflare 520-524 (522). RESIDENTIAL is documented as the escalation if AUTO/datacenter IPs stay blocked.
- A total block now FAILS LOUDLY (
Actor.fail) with a clear "enable Apify Proxy / switch to RESIDENTIAL" message instead of masquerading as an empty success. A valid-but-empty query (reachable backend, no matching data) still exits cleanly.
- No output field keys, input parameter names, or advertised behavior changed.
- Added human-readable dropdown labels (enumTitles) for the Sort option so the finite choices read clearly in the input editor.
- Empty input
{} now returns a sensible default (recent posts from a popular subreddit) instead of erroring — provided subreddits/posts/users/searches behave exactly as before.
- Hardened numeric output fields (score, comment counts, awards, ratios, karma) so numeric-looking values from either archive backend are always stored as real numbers; null/absent values are preserved unchanged. No output keys renamed.
- Confirmed output & dataset schemas are present and linked; verified end-to-end on live data.
- Health check passed — actor verified working end-to-end on Apify platform.
- Changelog refreshed for Store quality compliance.
- Maintenance & reliability pass: re-verified end-to-end against live data and confirmed the Actor completes successfully within the 5-minute quality window on the default input.
- Refreshed the prefilled example input and tuned run defaults for faster, lower-cost runs.