Groups App Store and Google Play reviews by app version and shows how each release changed the rating, with significance flags and the complaints that are new in that version.
The coverage threshold is set per store, each on its own measurement. Under 0.18 one constant of 20 % stood on Google Play data alone and was applied to the App Store too. App Store im:version coverage measured on 2 October 2026 (three feed pages per storefront): 100 % on 2,437 reviews from 18 storefronts (1064216828 us/jp, 284882215 us, 389801252 us, 324684580 gb, 310633997 de, 835599320 us, 363590051 us/br, 327630330 us, 585027354 de, 1459969523 us/in, 333903271 us, 422689480 gb, 1130498044 us, 1094591345 jp, 982107779 us). The App Store threshold is now 90 % (about 10 points under the lowest read, the margin Google Play has under its 29.5 % lowest read); Google Play stays at 20 %. A read of 21 % from the App Store used to give rows carrying only a note; it is now a left-out storefront. Fixture check: im:version kept on 80 % of the entries gives exit 91 under 0.19 and green rows under 0.18.
The status row of an app with no qualifying version no longer lists a storefront that was left out of the statistics. Under 0.18 the fixture 1064216828 with us healthy and gb at 10 % gave, with minReviewsPerVersion 3, rows with countries ["us"], and with minReviewsPerVersion 300 a status row with countries ["us","gb"] and sourceUrl of us next to a note that said gb was left out. Now both branches take countries and sourceUrl after the coverage filter.
The run status message separates the two reasons for leaving a storefront out. Under 0.18 "Storefronts left out because they carry no app version: ... gb (5 of 49 reviews)" named a storefront that does carry the field in 5 reviews. Now "because they carry no app version" lists storefronts with 0 versioned reviews and "because too few of their reviews carry an app version" lists those below the threshold with their counts and the threshold ("gb (5 of 49 reviews, below 90%)").
The failure message names the threshold of the store that was read ("below 90% per storefront" for the App Store, "below 20%" for Google Play).
Tests: 139 (new: App Store at 80 % fails, gb at 80 % next to a healthy us with the two sentences of the status message, status row countries and sourceUrl with gb at 10 % and with gb at 0 %); the tests that used 75 % App Store coverage as the healthy partial case now use 95 %. clockedHttp in http.test.js runs on a fixed clock: under 0.18 the real time of the failed attempt leaked into remainingMs() and the assertion 44999 !== 45000 failed in 3 of 14 loaded runs of node --test.
Verified 02.10: smoke 38/38 offline and 28/28 live, node --test 139/139, prefill live exit 0 with 7 rows, cost measured again $0.0274 per 1k rows (23.7 s, 24 rows from the example input), the price of $4 per 1k is x145.7 the cost.
README: the paragraph "Good to know" no longer describes the 0.17 rule (fewer than half of the reviews of the app).
0.18
The version coverage is judged per storefront, not on the sum of the countries. Under 0.17 the fixture 1064216828 with us and gb and im:version removed from gb gave a green run with two billable rows whose countries were ["us","gb"] although gb contributed no review, with only the aggregate "50 of 99 reviews carry an app version"; with the field removed from us instead, the whole app failed with exit 91 and 0 rows although gb held 49 reviews with a version. Now a storefront with no versioned review, or with less than 20 % of them, is left out of the statistics, named in the note of every row ("gb: none of 49 reviews carries an app version, so this storefront was left out of the version statistics (its feed may have changed)") and in the status message, and removed from countries and sourceUrl; a healthy storefront next to it keeps its rows. The app fails only when no storefront reaches the threshold.
The threshold is 20 %, down from 50 %, and it rests on a measurement. Google Play coverage of raw[10] on 2 October 2026, 400 to 500 reviews per read: Snapchat us 62.0 %, jp 54.0 to 54.8 %, br 53.2 %, ru 56.6 %, id 29.5 to 32.2 %; Telegram us 63.3 to 64.2 %, br 74.6 to 75.5 %, de 83.6 %; Duolingo 92 to 93 %, Spotify 84.8 %; the other 20 apps in us/en 79.8 to 93.0 %. Coverage does not depend on the depth of the read (Telegram 69.3 % at 5000 reviews). Under 0.17 the input {"apps":["com.snapchat.android"],"googlePlayCountry":"id","googlePlayLanguage":"id"} gave exit 91 and 0 rows with the diagnosis "the source may have changed", while us gave 4 billable rows; now both give rows, the Indonesian one with "id: 161 of 500 reviews carry an app version" in the note. The lowest healthy read is about 10 points above the threshold; a field that disappears (0 %) or drops to 10 % still fails.
The counters of the status message are taken before the verdict. Under 0.17 a run that failed on coverage said "Version found on 0 of 0 reviews" next to "Read 500 reviews ... but only 161 carry"; now it says "Version found on 161 of 500 reviews", and the failure message names each storefront ("us: 5 of 50, gb: 5 of 49; below 20 % per storefront"). The fail-on-nothing-delivered check counts reviews that entered the analysis, not reviews read.
Tests: 135 (new: Google Play with 35 % and 10 % coverage, gb without versions next to a healthy us, us without versions next to a healthy gb, gb at 10 % next to a healthy us, two storefronts without versions, two storefronts at 10 % with the counter and the per-storefront numbers); run.test.js is split into three files so that node --test runs them in parallel (suite wall time 38 s instead of 110 s).
0.17
The app version stored with a review is now checked for coverage on every read. Under 0.16 an app whose reviews carried no version failed the run only from 20 reviews up: with 15 reviews the run ended green and the single status row advised lowering minReviewsPerVersion or adding App Store countries, which cannot produce a row when no version was seen. Measured on a fixture (1064216828, us, im:version removed from every entry): 25 reviews gave exit 91, 15 reviews gave exit 0 with the filter advice. Now any read where no review carries a version is a changed-source failure of that app (free status row, run fails when no other app delivered).
A read where fewer than half of the reviews carry a version is the same failure ("only 20 carry an app version"). Under 0.16 the fixture with im:version kept on 20 of 200 entries gave the same three billable rows with reviewCount 3 / 11 / 6 instead of 23 / 108 / 69 and avgRating 1.00 instead of 2.09, with no sign on the rows.
When at least half of the reviews carry a version but not all, the note of every row and of the status row says "N of M reviews carry an app version, so the version statistics rest on fewer reviews." (Google Play leaves the version empty on a share of reviews: 31 of 200 in the recorded Spotify sample). A complete read keeps note null.
Tests: 128 (new: zero coverage below 20 reviews, 10 % coverage, 75 % coverage with the counter on the rows and none on a healthy read, 75 % coverage with no qualifying version).
0.16
A status row for an app in which no version reached minReviewsPerVersion (or onlyRegressions or the other filters left nothing) now carries the partial-read notice that billable rows already carried. Under 0.15 the notice went only to billable rows, so the status row of an app with no qualifying version dropped it. Measured on a fixture (1064216828, us, page 2 failing with HTTP 403, 50 of 200 reviews read): minReviewsPerVersion 3 gave rows with "reading stopped after 50 reviews (HTTP 403 ...)", while 300 gave a single status row that advised lowering the filter or adding App Store countries and quoted the 500 reviews per country limit, although the real cause was the HTTP 403, which passes on a rerun.
The status row now ends with "Partial data: ..." in all three branches that produce it (no version reaches the threshold, onlyRegressions finds nothing, other filters match nothing) and for an empty store answer. When the read was incomplete the threshold advice is replaced with "Reading was incomplete, so run again before changing the filters."; the advice about countries and maxReviewsPerApp stays for a complete read.
The same holds for a 403 in one of several App Store countries, for entries dropped by validation and for a Google Play batch that failed after the first one.
Tests: 124 (new: App Store page 2 with HTTP 403 and no qualifying version compared with the healthy read, one country failing next to a healthy one, half of the entries failing validation, Google Play second batch 403 with no qualifying version and with onlyRegressions).
Verified 02.10: smoke 38/38 offline and 28/28 live, tests 124/124, cost measured again $0.01321 per 1k rows (10.9 s for the example input), the price of $4 per 1k is x302.8 the cost.
actor.json, package.json and this file agree on the version again.
0.15
App Store feed entries without im:rating are no longer filtered out before validation. A feed that changed its format gives a status row with the reason after one try (not a green empty run with a "run again" hint), and a page where only some entries lack the field gives billable rows with a note and the number of dropped entries in the run status.
Cost measured again: $0.01097 per 1k rows (9.1 s for 24 rows), the price of $4 per 1k is x364.5 the cost.
Tests: 119.
0.14
Google Play answering 200 without reviews is no longer reported as "The app has no written reviews in the selected storefront". Under 0.13 an empty payload, an empty list or a changed shape of the first batch gave exit 0, zero rows and that sentence, and next to a healthy App Store app the run was green and silent. Now the Actor asks up to three times (pauses 3 s and 5 s inside the per-app budget) and compares with the rating count read from the already fetched app page (aggregateRating.ratingCount of the JSON-LD; measured 02.10: Spotify 36,412,939, Reddit 4,849,780, Firefox Focus 271,806): with 50 or more ratings it is a status row of kind blocked with the count and the number of tries ("this does not show whether the app has written reviews"), the run fails when no other app delivered reviews; with fewer ratings or none the status row says the app may have no written reviews yet.
A reviews payload whose first element is not a list is a changed format (kind shape), not an empty app.
An empty batch in the middle of the Google Play reviews gets one more try; if it stays empty, note carries "the endpoint answered empty for the batch after N reviews and reading stopped there (reason), so the oldest reviews of the window may be missing". Under 0.13 reading ended with note null (measured: 200 of 400 requested reviews, no trace).
The run status says "The store returned no reviews for ..." for both stores.
Tests: 115 (new: first batch empty payload, empty list and not-a-list, healing on the second try, few or no ratings, budget used up, transport error after an empty answer, empty batch in the middle with and without healing, end-to-end runs for a lone Google Play app and next to App Store; the healthy end of the feed is now a shorter last batch without a token). The fixture app page carries the JSON-LD rating count.
0.13
An error in the middle of reading Google Play reviews no longer discards the reviews already read. Under 0.12 an HTTP 403 or a changed payload on the second batch threw away everything: measured on a fixture (com.spotify.music, 400 reviews, batch 1 = 200 reviews, batch 2 failing), one app went from 4 billable rows to exit 91 and zero rows, and next to a healthy App Store app the run was green and silent with no Google Play rows at all. Now the 200 reviews stay in the version rows (identical rows and review counts as with a healthy end of the feed) and the note says "us: reading stopped after 200 reviews, so the older reviews of the window are missing (HTTP 403 ...)". An error before the first review is still a failure of that app with a status row.
Google Play entries that fail validation are named in the note like the App Store ones: "us: 100 of 200 review entries failed validation and were left out, so the version statistics of this storefront rest on fewer reviews". Under 0.12 half of the entries without id gave fewer rows (one version disappeared) and lower review counts with note null.
README: the paragraph on a feed that stops partway and on invalid entries now covers both stores.
Verified 02.10: smoke 38/38 offline and 28/28 live, example input about 22 rows in 8-10 s, cost about $0.010 per 1k rows (x386 against $4).
Tests: 102 (new: Google Play 403 and changed payload on the second batch for one and for two apps, a healthy end of the feed without a note, failure on the first batch, part of the entries without id; source-level tests of fetchPlayApp).
0.12
A storefront whose entries fail validation (measured on a fixture: gb without the id field or with im:rating = "4 stars", next to a healthy us) is now named in the note of the version rows and in the run status: "gb: N review entries were read but none passed validation, so the review format of this storefront may have changed", or "gb: N of M review entries failed validation and were left out" when only part of them fail. Under 0.11 the rows were billed with note null and the only trace was the global count of dropped reviews.
Verified 02.10: smoke 38/38 offline and 28/28 live, example input 24-25 rows in 9.7-10.9 s, cost $0.011-0.012 per 1k rows (x331-357 against $4).
Tests: 94 (new: all entries of one storefront invalid, entries without id, part of the entries invalid, healthy storefronts without a note, end-to-end run with an invalid gb next to a healthy us). Removed a no-op substitution in the term stemmer.
0.11
A storefront that fails for any reason other than transport no longer discards the storefronts already read. Under 0.10 an HTTP error (403) or a changed feed format in the second country threw the whole app away: measured on a fixture with us healthy and gb page 1 answering 403, the run went from 5 billable rows (249 reviews) to exit 91 and zero rows, and with a healthy Google Play app next to it the run was green and silently returned only the Google Play rows (4 rows instead of 9). Now the failing country is listed in the note ("gb: not read (HTTP 403 ...)") and the other countries are delivered, in either order; an app fails only when no country could be read (all-403 still ends the app with the original error kind and app data).
A transport error or the app's own budget after an empty answer in the middle of a feed is named as such. Under 0.10 page 2 empty and then ENOTFOUND gave "the feed answered empty for page 2 after page 1" and the budget case blamed the store; now the note says "... and reading stopped there (ENOTFOUND ...)" or "(the waiting budget of the app was used up)".
README: the trend counts and the term counts of the example input are given as ranges of two runs on 1 October 2026 (trends 1 to 3 regressed, 0 to 1 improved, 1 to 2 stable, 11 to 15 inconclusive; terms on 7 and 8 of 23 rows), one count per row population instead of the two conflicting numbers of 0.10; the sentence about a failed country describes the behaviour above.
Verified 01.10: smoke 38/38 offline and 28/28 live, example input 23 rows in 9.5-11.3 s, cost $0.0115 per 1k rows (x349 against $4).
Tests: 89 (new: HTTP 403 and changed feed format in the second and in the first country, all countries failing, transport error and exhausted budget after an empty page in the middle of a feed).
0.10
The run budget no longer starves an app that was not touched yet. Under 0.9 every app had up to 75 s for failures but the whole run had a 150 s cap, so two silent App Store apps used it up and a healthy third app (measured: com.reddit.frontpage on Google Play, 500 reviews in 2 s when read alone) got "No request was sent" and the run failed with 3 status rows and zero billable rows (152.3 s, exit 91). Now each app starts with at most an equal share of what is left of the 150 s (apps left counted including itself), capped at 75 s, and never below 9 s: with three apps two silent App Store apps used 50.0 s each, got status rows, and the Google Play app delivered 500 reviews in 2.2 s (102 s for the run). With two apps the first still gets 75 s.
If the App Store answers with entries but every one fails validation (measured on 70 entries whose im:rating was "4 stars"), the Actor reads no further pages of that storefront and the app gets a status row "Read N reviews of App Store ... but none passed validation, so the review format may have changed", with the app name and storefront. Under 0.9 the error was "Could not read: App Store ... ()" with an empty note, because the storefront was counted as blocked or as a failure with an empty list.
A transport error after an empty answer is named in the empty-feed note instead of being reported as a spent budget: for page 1 empty and the second try failing with ENOTFOUND the note says "after 1 attempt per storefront, after which the next try failed (ENOTFOUND ...)" and does not say that the waiting budget was used up. Budget errors carry a flag instead of being recognised by their text.
The message of an app that was refused any request because of the budget no longer exposes "No attempt was made for ; not attempted"; it says which host got no request and which share of the budget was used up.
README: the line about the run cap and the measurement describe the shares; the 5-trend counts are no longer given as fixed numbers. Verified 01.10: smoke 38/38 offline and 28/28 live, example input 23 rows in 13.7 s (2 regressed, 1 stable, 14 inconclusive, 6 baseline, 0 status rows), cost $0.0251 per 1k rows (x159.5 against $4).
Tests: 83 (new: three apps with two silent ones, minimum share after the run cap is used up, all entries invalid, transport error after an empty page, status row with a changed format in the run).
0.9
The budget for failures is per app. Under 0.8 one budget of 75 s covered the whole run: a hanging App Store host used all 75.0 s on the first app and every later app got "No attempt was made ... not attempted" without a single request, although it was healthy (measured: the healthy com.reddit.frontpage got zero requests and then delivered 500 reviews in 1.3 s on its own; with the prefill shape both rows were status rows and the run ended FAILED). Now every app starts a fresh 75 s (beginApp) and a cap of 150 s covers the whole run. Measured on 1 October 2026 with the App Store host accepting connections and never answering and Google Play live: the App Store app 75.0 s and a status row, then 500 reviews from Google Play in 18 s, run total 93 s. With the feed answering 200 without entries for us, gb and de: App Store 81.5 s, then Play delivered 500 reviews (the 0.8 run took 85 s and gave Play nothing).
A transport failure or an exhausted budget in the middle of an App Store app no longer discards the app. A country that fails after some reviews were read keeps them (the rows carry a note such as "gb: reading stopped after N reviews ..."); a country that fails without any review is listed in note while the other countries are delivered; only when no country delivered anything the app gets a status row. An exhausted budget after empty answers keeps the empty-feed explanation instead of the transport message "No attempt was made".
The empty-feed note counts the tries that were actually made: when the budget cuts the ladder it says "after 3 attempts per storefront (the waiting budget of the app was used up before all 6 tries)" instead of a fixed "after 6 attempts" (measured: gb 4 tries, us 6).
Status rows of apps that could not be read (feed blocked, HTTP error, changed format, transport failure) now carry appName, countries (the storefronts looked up) and sourceUrl whenever the store answered the lookup, the same fields as the other status rows. The dataset schema describes it.
Verified 01.10: smoke 38/38 offline and 28/28 live, example input 22 rows in 7.1 s, cost $0.0090 per 1k rows (x446 against $4).
Tests: 78 (new: per-app budget with a cap for the run, tries counted in the note, budget falling in the middle of the empty ladder, a failing second country keeps the first, app data on blocked and failed errors, status row fields in the run).
0.8
The run has a deadline for failures. Under 0.7 the 75 s budget counted only sleep, while time spent inside a request that failed was unlimited: on a local server that accepts the connection and never answers, one request cost 250.6 s of real time (4 attempts of 60 s) and 10.5 s of budget, and 15 requests in the shape of the example input on a 5 s timeout cost 261.1 s (at the default 60 s about 2309 s). Now the budget counts sleep and the time of failed attempts together, every attempt has a hard deadline of 25 s (Promise.race with an AbortController, because the socket timeout does not fire on a connection that hangs) shortened to what is left of the budget, and no attempt starts when less than 3 s is left. Measured on the same silent server: 15 requests in 75.0 s together (the first used the whole budget, the other 14 were refused at once). Requests that succeed do not use the budget.
lowStarTopTerms: olderShare and lift are null when all older versions hold fewer than 12 low-star reviews (6 of 9 rows of the example input had a pool of 4 to 10 and a lift that was only the size of the pool: versionShare times (pool + 1) / 0.5). The lift now divides by the older share taken at least at 1 % and at least at one review of the pool, so an absent term gets the same lift at pools of 200 and 2000 and a smaller one at small pools. newComplaintTerms uses the same function for its lift gate.
lowStarTopTerms is [] when the version has 8 or more low-star reviews but no word or pair occurs in 4 of them (lowStarTopTermsText is null for it) and null below 8 reviews; schema, README and the text field now describe both. On the example input: terms on 8 of 24 rows, [] on 1, null on 15; lift on 2 rows.
README: the Netflix row shows olderShare and lift as null (its older pool has 6 low-star reviews); the text no longer says that a lift near 1 is old and a high one new without the 12-review condition.
package.json version 0.8.0 (it was 0.1.0).
Tests: 72 (new: virtual-clock tests for one silent request, 15 silent requests and slow requests that succeed, a real silent TCP server, a send that never answers, lift against pool size, [] against null).
0.7
The significance test is now a proper Welch t test. Under 0.6 compareRatings used the normal constant 1.96 as the threshold and as the margin multiplier, without the t quantile and the Welch degrees of freedom, and with zero variance it called any difference of means significant with a margin of 0. Now the margin is the t quantile for the Welch-Satterthwaite degrees of freedom times the standard error, and each group's variance is floored at 1/12 (star ratings are whole numbers), so identical ratings never give a margin of zero. Example: [1,1,1] against [2,2,2] was regressed with margin 0; it is still regressed with a margin of 0.65 stars, and [5,5,5] against [4,5,5] is no longer significant.
Measured on 20 000 simulated pairs per cell drawn from the same rating distribution (spread over five stars): a change was flagged in 2.6 % of the pairs at 3 reviews per version, 2.6 % at 5, 4.3 % at 10 and 5.2 % at 50 (0.6 had 9.7, 9.0, 6.5 and 5.2 %); a skewed 5-star population gives 0.4, 0.3, 1.2 and 4.5 %, a bimodal one 3.3, 3.7, 5.0 and 5.4 %. A margin of exactly 0 no longer occurs (was 3.22 % of the pairs at 3 reviews).
Live check with minReviewsPerVersion 3 on the six example apps: 49 comparable rows, 17 rows with 5 reviews or fewer, 4 significant changes and none of them from a version with 7 reviews or fewer (0.6: 5 significant, 3 of them from 4 to 7 reviews, and 2 rows with a margin of 0.05 or less); the narrowest margin is 0.45 stars on a row with 4 reviews.
views.overview (the Version impact table in the Store) no longer has the column New complaints: it was empty on 23 of 23 rows of the Actor's own example input. The field stays in the dataset, in fields and in the API; the table keeps Top low-star terms (8 of 23 rows), the other columns are filled on 17 or 23 of 23 rows.
README: the example row is a real Netflix row (9.85.0, significant drop of 0.51 stars) and the text no longer says that a 500-review run typically gives a regressed row. Measured on the example input on 1 October 2026: at 500 reviews per app 17 comparable rows gave 3 regressed, 1 improved, 2 stable and 11 inconclusive (6 baselines); at 5000 reviews 100 comparable rows gave 7 regressed, 5 improved, 8 stable and 80 inconclusive. newComplaintTerms is filled on 0 of 23 rows at 500 and on 11 of 106 at 5000. The README states the test, the variance floor and the simulated false-alarm rate.
Verified 01.10: example input finished with 23 rows and exit 0 (the 5000-review variant took 148 s for 106 rows); smoke 38/38 offline and 28/28 live, cost $0.0093 per 1k rows (x430 against $4).
Tests: 64 (new: t quantiles against table values, zero variance, false-alarm rate at n of 3, 5, 10 and 50 with a fixed seed, overview columns).
0.6
One waiting budget of 75 seconds per run now lives in the HTTP layer and covers every sleep of the run: the pauses after empty App Store feed pages and the waits before retrying a request that answered 429 or a server error. Under 0.5 the budget covered only the empty-page ladder, so retries of the transport had no limit of their own: with one 429 per page, three countries slept 78 s, and a first page answering 429 Retry-After: 120 slept 362 s and ran 363 s, more than the 5 minutes of the daily Apify test.
A wait that does not fit into what is left of the budget is not taken. Retry-After is honoured only while it fits: 429 Retry-After: 120 on a first page now ends that request at once with an error that names the 120 s wait and the 75 s budget, the app gets a free status row and the other apps are read. Once the budget is spent every request gets a single try. The sleep of a run is at most 75 seconds; the run also spends time on the requests themselves and on the 1.2 second gap between App Store requests.
The earlier wording "worst case sleep of a run is 75 seconds whatever the input" (0.5) held only for empty pages; the README and the input description now describe the single budget that includes retries after 429 and server errors.
views.overview (the Version impact table in the Store) gets the column "Top low-star terms" from the new field lowStarTopTermsText (the five terms of lowStarTopTerms as one string with the lift, for example "log (lift 11.1), working (lift 4.2)"). The column "New complaints" stays, but it is empty at the default depth, so the table no longer rests on it alone.
Verified 01.10 (feed healthy at the time; the budget under 429 is covered by unit tests with a virtual clock): prefill 24 s with 8 rows, smoke 38/38 offline and 28/28 live, cost $0.0209 per 1k rows (x191 against $4).
Tests: 60 (new: 429 on every page of three countries stays within 75 s together with feed pauses; Retry-After: 120 on the first page sleeps nothing and makes one request; Retry-After within the budget is honoured and subtracted from the feed pool; the text field of the top terms).
0.5
All waiting between tries of the App Store feed shares one budget of 75 seconds per run (was six tries with 54 seconds of sleep for every isolated empty page, multiplied by pages, countries and apps). An empty first page keeps six tries with pauses of 3, 5, 8, 12 and 15 seconds, an empty page in the middle of the feed gets three tries and a page after another empty page two. When the budget is spent a page gets one try and is reported. Worst case sleep of a run is 75 seconds whatever the input; the same input that took 125 s (prefill), 181 s (us and gb, 200 reviews) and 205 s (three countries, 500 reviews) under 0.4 with empty pages is bounded by this budget plus the 1.2 s gap between requests.
The note of a row now tells two cases apart. A gap (an empty page followed by a served page) says that up to 50 reviews per page are missing and that page 1 holds the newest reviews; an empty page after the last served page says that the oldest reviews of the window may be missing. Before, every empty page, also the last one of the window, was reported as "missing from the newest part" and counted 50 reviews each.
countries with only unknown storefront codes (for example eu) and no other app delivering anything ends with "No App Store storefront code to read was valid for App Store (eu) ..." instead of "None of the requested apps exists"; the status message names codes that were not looked up.
newComplaintTerms keeps its strict test; the README, the dataset schema, PRICING.differentiator and the use cases no longer present it as filled at the default depth. Measured on 1 October 2026 on the six example apps from Google Play: 0 of 25 rows at 500 reviews per app, 2 of 62 at 2000, about 10 of 73 at 5000 (older versions inside the window hold only a few 1-2 star reviews, so null is structural). lowStarTopTerms is filled on 8 of 25 rows at 500 and is the field to use for triage. The README example is now a real 500-review row with a significant drop and newComplaintTerms: null.
dataset_schema.json states z of at least 3 (the code always used 3, the schema said 3.3).
Complaint terms: movies is read as movie (it was movy, next to movie in the same top five), and suspend and suspended, log and logging are no longer listed as two terms of one top list.
Verified 01.10 (feed healthy at the time, so empty pages could not be reproduced live; the budget is covered by unit tests): prefill 16 s, us and gb at 200 reviews 15 s, one app in three countries at 500 reviews 41 s (36 rows). Smoke 38/38 offline and 28/28 live, cost $0.0209 per 1k rows (x191 against $4).
Tests: 57 (new: budget shared by apps, pages and countries; default budget cap; gap and window-end notes; only invalid storefront codes; stemmer and duplicate forms).
0.4
trend has a fifth value, inconclusive. Before, stable was also given when the sample could not tell: in the 0.3 review 10 of 48 rows with a rating change of 0.4 stars or more were marked stable, next to a Reddit change of -0.95. Now stable needs a change under 0.15 stars and a 95 % margin of at most 0.3 stars, and regressed or improved need a significant change of at least 0.15 stars; everything else is inconclusive. New field ratingDeltaMargin (half-width of the 95 % interval, stars), and deltaSignificant and the margin are columns of the Store table view. The onlyRegressions status row counts the drops that are not significant.
newComplaintTerms no longer needs a 15 % share and a threefold lift together, which left it empty on 46 of 48 rows (large samples: a term in 15 % of the reviews is also common in older versions; small samples: zero hits in older versions gave an automatic lift). A term now has to pass a one-sided two-proportion test against all older versions (z of at least 3), appear in at least 4 reviews and 3 % of the low-star reviews, be at least twice as frequent as before and, when the previous version has 8 or more low-star reviews, at least twice as frequent as in the previous version. The field is null when no comparison is possible (oldest version, fewer than 12 low-star reviews in the version or in the older ones) and [] only when the comparison was made and nothing stood out; before, both were [].
New fields newComplaintTermStats (term, lowStarReviews, versionShare, olderShare, lift) and lowStarTopTerms (the five most frequent terms of the low-star reviews with the same numbers, for every version with at least 8 low-star reviews, so a regression row always shows what the unhappy reviews are about and how much of it is new). Contractions such as "you're" are stop words.
Measured on 1 October 2026 with 5000 Google Play reviews each of Reddit, Spotify, WhatsApp and Duolingo (60 version rows): rows with a rating change of 0.4 stars or more labelled stable fell from 12 to 0; newComplaintTerms is non-empty on 9 of the 23 rows that can be compared (0.3: 3 of 60 rows), null on the 37 rows with too few low-star reviews to compare, and lowStarTopTerms is filled on every row with 8 or more low-star reviews. Large-sample regressions can still return [] by design; the top terms carry the signal there.
One bad input or one failing app no longer ends the run. A storefront code that iTunes answers with HTTP 400 (eu, en) is reported in the status message and the row note, and uk is read as gb; an iTunes lookup that answers with any other error ends the run, because it means the lookup API changed; any other source error of one app (HTTP error from Google Play, changed format) becomes a free status row, the other apps are delivered, and the run fails only when no app delivered any reviews. Before, countries: ["us", "uk"] ended the run as an error after two rows had been billed.
An empty App Store feed page is no longer read as the end of the feed. Measured for gb: pages 1 and 2 empty and page 3 with 50 reviews in the same minute. After an empty page the next two pages are tried (two attempts each), up to three empty pages in a row; a gap is reported as "the feed answered empty for page N" and not as the store stopping to serve pages.
Hint for an app without a qualifying version names the right lever per store: App Store countries and the 500 per country limit, or maxReviewsPerApp up to 5000 for Google Play.
Tests: 52 (new: inconclusive trend, term numbers and null versus empty, previous-version gate, top terms, uk alias, HTTP 400 storefront, empty pages 1 and 2 with data on page 3, one app failing with HTTP 403 next to a working one, hint per store).
0.3
A run fails only when no app delivered any reviews. An App Store app whose feed stays empty gets a free status row and its name in the final status message, and the run succeeds when other apps delivered reviews, so the daily Apify test with the prefill (an App Store app and a Google Play app) is no longer red on days when the feed answers with empty pages.
The status row, README and proxy description no longer claim that the feed refuses some addresses. Measured on 2026-10-01: the same page stays empty for several fresh addresses and for the direct connection, while another app or the next minute returns 50 entries, so the empty answer is random. The text now says that the feed returned no reviews, which does not show whether the app has written reviews, and to run again.
Page 1 is retried six times for every storefront, whatever number of ratings the store lists. Only a storefront with 50 or more listed ratings and no reviews counts as a refused feed that can fail a run; a smaller one gets a neutral status row with the listed ratings and no claim about the app.
Documentation matches what the feed serves: Google Play up to 5000 reviews, App Store 50 reviews per page, up to 500 per country, often only the first page. The README example is a real Google Play row (Spotify 9.1.86.2432, 204 reviews).
newComplaintTerms needs at least 12 low-star reviews in the version and in all older versions together (was 8) and a term in at least 4 reviews (was 3), because terms from small samples were noise.
The final status message names the row limit when maxItems stops the run.
Price $4 per 1k rows (was $20): a measured row holds 12 to 315 reviews (median 37), so the price is about $0.11 per 1k reviews at the median, against a top-3 per-review median of $0.10 per 1k. Cost measured again with a run of 22 rows.
Tests: retries for apps with few or zero ratings, a page 2 that recovers, a dead App Store app next to a working one (success), a dead app alone (failure), a dead app next to a working one with a filter that removes every row (success).
0.2
App Store review feed: an HTTP 200 page without any review is no longer read as "end of reviews". Page 1 empty for a storefront whose lookup lists at least 50 ratings is retried six times with growing pauses, then counted as a refused feed. An app whose feed is refused in every selected country gets a free status row that names the ratings the store lists and says it is not an app without reviews, the other apps are delivered, and the run ends as failed.
A full page followed by an empty one is retried and then reported: version rows keep the newest reviews and carry a note with the number of reviews that arrived; a page shorter than 50 ends the paging without another request.
countries and sourceUrl of a row now name only storefronts that supplied reviews.
Apify Proxy is on by default, every request takes a fresh address; if the proxy cannot be created the run continues without it. Requests to the App Store keep a 1.2 s gap.
newComplaintTerms compares a version only with older versions, so the field means what its name says; a complaint that persists into the next release is no longer credited to it, and the oldest version has no terms. Field description and README corrected (15 % share, three times the frequency).
The "no reviews" note no longer advises lowering filters when the cause is the feed; the filter advice names the 500-reviews-per-country limit.
Verified (smoke offline 38/38, tests 36/36, live Google Play run on six apps: 23 rows in 18.7 s, $0.0226 per 1k rows without proxy, x886 against the price). Live check of the refused feed (01.10, this host): Reddit App Store returned 200 without reviews 6 times, the run delivered Spotify (3 rows), wrote the explaining status row and ended with exit 91. The App Store happy path could not be re-read live because the feed refused this address for hours; it is covered by the recorded fixtures.
Tests: source retry and truncation cases, a run with one dead storefront and one with a dead app. Fixture pages that the store had served empty during recording were reduced to real partial pages.
0.1
Initial release: App Store and Google Play reviews grouped by app version, with the rating change against the previous version, a significance flag, a trend and new complaint terms.
Verified with three inputs (see the smoke report):
Typical: two apps from both stores, 200 reviews each; every version with enough reviews gets one row with a trend.
Edge: only the latest version, only regressions, a single country and a higher minimum of reviews per version.
No results: an app without reviews or an unknown app returns one unbilled status row with found: false; when no requested app exists the run fails with the reason.