App Release Review Impact: Rating by Version avatar

App Release Review Impact: Rating by Version

Pricing

Pay per event

Go to Apify Store
App Release Review Impact: Rating by Version

App Release Review Impact: Rating by Version

Groups App Store and Google Play reviews by app version and shows how each release changed the rating, with significance flags and the complaints that are new in that version.

Pricing

Pay per event

Rating

0.0

(0)

Developer

datagrit

datagrit

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

What does App Release Review Impact: Rating by Version do?

App Release Review Impact reads the newest public reviews of Apple App Store and Google Play apps and groups them by app version. For every version you get the review count, the average rating, the star distribution, how the rating changed against the previous version, whether that change is statistically significant, and the complaint terms that are new in that release. It is built for product managers, QA and support leads, app store optimization teams and analysts who want to know which release made users unhappy without reading thousands of reviews.

Review scrapers return one row per review with a raw version string. This Actor returns one row per version, so a regression shows up as a single line with a trend of regressed instead of a spreadsheet pivot you have to build yourself.

Use cases

  • Release monitoring: run the Actor weekly for your own app and check whether the latest version has a trend of regressed.
  • Competitor tracking: follow the apps of competitors and see which of their releases were received badly and for what reason.
  • Support and QA triage: lowStarTopTerms with its lift against older versions and the sample low-star reviews point to the feature or bug behind a rating drop; on deep Google Play reads newComplaintTerms adds a strict statistical flag for terms that are new in the release.
  • Due diligence and market research: compare release quality across apps of one category, in both stores, in one dataset.
  • Reports: feed per-version ratings into a dashboard or a spreadsheet without cleaning review dumps.

How to use it

  1. Add apps: App Store IDs (digits such as 1064216828), apps.apple.com links, Google Play package names (such as com.spotify.music) or play.google.com links. You can mix both stores in one run.
  2. Choose the App Store countries to read. Reviews of all chosen countries are merged into one row per version.
  3. Set how many newest reviews to read per app and the minimum number of reviews a version needs to get a row.
  4. Optionally keep only the latest versions or only the versions that got worse, then run the Actor.

How the analysis works

  • Reviews are grouped by the app version stored with each review and sorted by version number, not alphabetically.
  • Each version is compared with the nearest older version that has at least the minimum number of reviews, so a version with three reviews does not break the chain.
  • ratingDelta is the change of the average rating and ratingDeltaMargin is the half-width of its 95 % confidence interval, in stars. deltaSignificant is a Welch two-sample t test at 95 % confidence: the margin is the t quantile for the Welch degrees of freedom times the standard error, so a version with 3 to 5 reviews gets a wide margin, and the variance of each group is floored at 1/12 of a star squared (ratings are whole numbers), so identical ratings never give a margin of zero. Simulated on 20 000 pairs of versions drawn from the same rating distribution, the test flagged a change in 0.3 to 5.4 % of the pairs at 3, 5, 10 and 50 reviews per version (95 % confidence promises 5 %).
  • trend has five values. regressed or improved: the change is significant and at least 0.15 stars. stable: the change is under 0.15 stars and the margin is at most 0.3 stars, so the sample is large enough to rule out a bigger change. inconclusive: the sample cannot tell, either a visible change that is not significant or a margin wider than 0.3 stars; read ratingDelta next to ratingDeltaMargin. baseline: the oldest version of an app, which has no predecessor. The onlyRegressions filter keeps regressed rows only; when it returns nothing, the status row says how many versions dropped by at least 0.15 stars without enough reviews to call it significant.
  • lowStarTopTerms lists the five most frequent words and word pairs in the 1 and 2 star reviews of the version, each with lowStarReviews, versionShare (their share of the version's low-star reviews), olderShare (the same share in all older versions) and lift (versionShare divided by olderShare, where olderShare is never taken below 1 % or below one review of the older pool, so a term that is absent from a small older pool does not get a lift that merely reflects the size of the pool). olderShare and lift are null when all older versions together hold fewer than 12 low-star reviews, because a share of a handful of reviews says nothing; they are filled on 3 of the 23 rows of the example input and are mostly null at 500 reviews per app. The list itself is null when the version has fewer than 8 low-star reviews and [] when it has 8 or more but no word or pair occurs in at least 4 of them (their complaints are too scattered). On the 23 rows of the example input (500 reviews on six Google Play apps) two runs on 1 October 2026 gave terms on 7 and 8 rows, [] on 3 and 1 rows and null on 13 and 14 rows; the counts move with the reviews available. Inflected forms of a word are listed once. Where a lift exists, a value near 1 means the term is about as frequent as in older versions and a high value points to a topic that is more frequent in this release.
  • lowStarTopTermsText is the same list as one readable string (log (lift 11.1), working (lift 4.2), login) and is the column of the Version impact table in the Store; null when lowStarTopTerms is null or [].
  • newComplaintTerms is the strict version of that list: terms that are significantly more frequent in the low-star reviews of the version than in all older versions (one-sided two-proportion test with z of at least 3, at least twice the older frequency, in at least 4 reviews and 3 % of the low-star reviews) and that were not already as frequent in the immediately previous version. newComplaintTermStats holds the same numbers per term. An empty array means the comparison was made and no term stands out, which is common on large samples where an old complaint is simply louder; null means no comparison is possible (the oldest version, or fewer than 12 low-star reviews in the version or in the older versions together). This field is strict and mostly empty or null at the default depth: measured on 1 October 2026 on six popular Google Play apps, it was filled with at least one term on 0 of 23 rows at 500 reviews per app and on 11 of 106 rows at 5000, because the older versions inside the window hold only a few 1-2 star reviews each. Use lowStarTopTerms for day-to-day triage and newComplaintTerms when you read Google Play at 5000 reviews and want only terms that stand out statistically. Read the terms next to the sample reviews: they point to a topic, not to a confirmed bug.
  • sampleLowStarReviews holds up to three low-rated reviews that contain those terms. Author names are never collected.

Example output

{
"store": "google-play",
"appId": "com.netflix.mediaclient",
"appName": "Netflix",
"version": "9.85.0 build 4 74576",
"versionRank": 1,
"reviewCount": 234,
"avgRating": 4.21,
"lowStarShare": 0.167,
"previousVersion": "9.84.0 build 3 74548",
"previousAvgRating": 4.73,
"ratingDelta": -0.51,
"deltaSignificant": true,
"ratingDeltaMargin": 0.3,
"trend": "regressed",
"newComplaintTerms": null,
"newComplaintTermStats": null,
"lowStarTopTerms": [
{
"term": "movie",
"lowStarReviews": 8,
"versionShare": 0.205,
"olderShare": null,
"lift": null
},
{
"term": "watch",
"lowStarReviews": 5,
"versionShare": 0.128,
"olderShare": null,
"lift": null
},
{
"term": "work",
"lowStarReviews": 5,
"versionShare": 0.128,
"olderShare": null,
"lift": null
},
{
"term": "playing",
"lowStarReviews": 4,
"versionShare": 0.103,
"olderShare": null,
"lift": null
}
],
"lowStarTopTermsText": "movie, watch, work, playing",
"sampleLowStarReviews": [
"Why are the ads 3× louder than any of the shows?? I damn near blow out my eardrums every single time one comes on and am rushing to turn the volume down.. just to turn it back up in order to hear the shows..."
],
"countries": [
"us"
],
"found": true,
"scrapedAt": "2026-10-01T12:27:15.699Z"
}

This is a real row from a Google Play run on 1 October 2026 (the 500 newest reviews of Netflix; the sample reviews are shortened to one of three and the star counts and source URL are left out). Real numbers depend on the reviews available when you run the Actor. versionRank 1 is the newest version. It is one of the rarer rows: on the same six-app example input (500 Google Play reviews each) two runs on 1 October 2026 gave 23 rows each, 17 of them with a predecessor to compare, and the 17 split into 1 to 3 regressed, 0 to 1 improved, 1 to 2 stable and 11 to 15 inconclusive (6 rows are baseline); at 5000 reviews per app, 100 comparable rows gave 7 regressed, 5 improved, 8 stable and 80 inconclusive. Most versions have too few reviews to call a change, so expect inconclusive on most rows and read ratingDelta next to ratingDeltaMargin. newComplaintTerms is null here because the older versions in a window of 500 reviews have too few 1-2 star reviews to compare against (see the analysis notes above). Fields that cannot be computed, such as the change of the oldest version, are null.

Input

  • Apps: App Store IDs, store links or Google Play package names.
  • App Store countries: two-letter storefronts such as us, gb, de (uk is read as gb). A code the App Store does not know, such as eu, is reported in the status message and the row note and does not stop the run while another code or another app delivers reviews. When none of the codes is valid and no other app delivered anything, the run fails with a message that names the codes and the app gets a free status row. Ignored for Google Play apps.
  • Google Play country and language: which Play storefront and interface language to read.
  • Reviews to read per app: 50 to 5000. Google Play serves up to 5000 newest reviews. The App Store feed pages 50 reviews at a time up to 500 per country, but it often stops answering after the first page, so plan for 50 to 500 reviews per App Store country.
  • Minimum reviews per version: versions below this limit are left out because averages over a handful of reviews are noise.
  • Only the latest N versions and Only versions that got worse: narrow the result.
  • Maximum version rows: stops the run after this many rows.

Pricing

You pay per version row returned. Apps that are not found, apps without enough reviews and duplicates produce a free status row with found: false and a note, so you are never charged for an empty result. The price per row is set on the Actor page and shown before you run it; set a maximum charge per run in the Apify console to cap spending, and the Actor stops cleanly when the limit is reached. Reading the reviews costs nothing extra.

Good to know

  • Rating changes depend on how many reviews a version has. Google Play runs reach back several versions (in our check of 1 October 2026: Spotify 4 versions and Reddit 3 versions from 500 reviews each, 12 to 315 reviews per row). A busy App Store app with only the first feed page (50 reviews) usually gives one or two rows, the oldest of them a baseline without a rating change and with a note saying that the data is partial. Add countries, lower the minimum or use the Google Play listing of the same app for a longer history, and read the reviewCount next to every change.
  • A version appears only in the reviews written while it was installed. A new release collects reviews over days, so the newest row can rest on a few reviews.
  • Google Play shows the version the reviewer used when writing, and a share of reviews has none. Measured on 2 October 2026 on 400 to 500 reviews per storefront: 62 to 93 % in the United States for 22 apps (Snapchat 62 %, Telegram 64 %, Spotify 85 %, Duolingo 93 %) and 53 to 84 % in Germany, Russia, Japan and Brazil; the lowest read was Snapchat in Indonesia with 30 to 32 %. The App Store feed gave a version on every review we read: 100 % on 2,437 reviews from 18 storefronts (measured 2 October 2026). The coverage is therefore judged per storefront (one country for Google Play, each country for the App Store) against a threshold per store, 20 % for Google Play and 90 % for the App Store: at or above it the rows are returned and, when some reviews have no version, the note says "us: 190 of 200 reviews carry an app version, so the version statistics of this storefront rest on fewer reviews"; below it that storefront is left out of the statistics, named in the note and in the status message (separately for storefronts with no version at all and for those under the threshold), and its country disappears from countries and sourceUrl, also on the status row of an app with no qualifying version. The app fails only when no storefront reaches the threshold. The status message of each run reports how many reviews carried a version.
  • If a store changes its format so that no review can be read or no version is found, the app gets a status row that says so, and the run fails instead of returning empty rows when no other app delivered anything.
  • The App Store review feed sometimes answers with an empty page instead of an error, and it does so unpredictably, also for pages that were served a moment earlier. The Actor retries an empty first page up to six times with growing pauses and compares the result with the number of ratings the store lists; an empty page in the middle of the feed gets two more tries. Everything that goes wrong while one app is read shares a budget of 75 seconds for that app: the pauses between tries of empty pages, the waits before retrying a request that the store answered with HTTP 429 or a server error (a Retry-After header is honoured only while it fits into what is left of the budget) and the time spent in attempts that failed. An app gets at most 75 seconds, and at most an equal share of what is left of a 150-second budget for the whole run (with three apps the first one gets 50 seconds), and never less than 9 seconds, so an app that has not been touched yet always keeps enough time to make its requests and a source that hangs for the first apps cannot take the attempts of the next one. Each attempt is limited to 25 seconds by a hard deadline (and to what is left of the budget), so a source that accepts the connection and never answers costs that app 75 seconds in total, not minutes per request. Requests that succeed are not counted, so a large read of a healthy source is not cut short; the run also spends time on a 1.2 second gap between App Store requests. Measured on 1 October 2026 with the App Store host set to accept connections and never answer and Google Play live, on three apps (two App Store apps and com.reddit.frontpage): each App Store app used 50.0 s and got a status row, then com.reddit.frontpage delivered 500 reviews from Google Play in 2.2 s (102 s for the run; with a cap of 150 s for the run shared in turn, the third app used to get no request at all). With the App Store feed answering 200 without entries for the three countries us, gb and de the App Store app used 81.5 s (73 s of pauses) and Google Play delivered its 500 reviews afterwards (4 s). An empty page is not taken as the end of the feed: later pages are tried too (measured: pages 1 and 2 empty, page 3 with 50 reviews in the same minute). The note tells a gap (an empty page followed by a served page, so those reviews are missing from the window) apart from an empty page after the last served one (the oldest reviews of the window may be missing). An app whose feed stays empty gets a free status row that says the feed returned no reviews and how many tries were made (fewer than six when the budget ran out), which does not show whether the app has written reviews; run again to read them. A failed app or country keeps what was already read, whatever the failure was (an empty feed, a transport error, an HTTP error such as 403 or a changed format of the feed in one of several storefronts): the storefronts that were read stay in the version rows, and the note names the storefront that was not read and why. An app is a failure only when none of its storefronts could be read. The status row carries the app name, the storefronts that were looked up and the store page whenever the store answered the lookup. The other apps are delivered and the run succeeds. The same holds for any other error of one app (a store answering with an HTTP error, a changed format): it gets a free status row and the other apps are delivered. The run ends as failed only when no app delivered any reviews. If the feed stops partway (in the App Store, or in Google Play with an HTTP error such as 403 or a changed payload after the first batch of 200 reviews), the reviews read so far stay in the version rows and the note carries the number of reviews that arrived and the reason reading stopped; an error before the first review is a failure of that app. If the store answers with reviews but none of them passes validation (for example the rating field changed its format or was renamed, so the entries carry no rating at all), the app gets a status row that says so and the run fails only when no other app delivered reviews. When that happens in one of several storefronts, the other storefronts stay in the version rows and the note names the storefront and the number of entries that failed validation (all of them, or how many of the entries it returned); the version statistics of that storefront then rest on fewer reviews. Entries of the App Store feed that lack the rating field are counted as invalid in the same way, never skipped silently, and a page whose entries all lack it is not read again as if it were empty. Google Play entries that fail validation are reported the same way ("us: 100 of 200 review entries failed validation and were left out"). Google Play can also answer 200 without a single review. The Actor then asks up to three times (pauses of 3 and 5 seconds, inside the budget above) and compares the result with the rating count in the app page: with 50 or more ratings listed the app gets a free status row that names the count and the number of tries and says that this does not show whether the app has written reviews, and the run fails when no other app delivered reviews; with fewer ratings, or none listed, the status row says that the app may have no written reviews yet. An empty batch in the middle of the Google Play reviews gets one more try, and if it stays empty the note says that reading stopped after that many reviews and that the oldest reviews of the window may be missing. A reviews payload whose shape is not a list is a changed format, not an empty app. The version stored with a review is the field the grouping stands on, so it is checked on every app, however few reviews were read: when no storefront of an app reaches the threshold above (no review carries a version, or fewer than 20 % of the Google Play reviews, or fewer than 90 % of the reviews of each App Store country), the app gets a free status row that says the source may have changed (never advice to lower the filters), and the run fails when no other app delivered reviews. When a storefront reaches it, the note of every row and of the status row says how many of the reviews of that storefront carry a version (for example "us: 209 of 240 reviews carry an app version, so the version statistics of this storefront rest on fewer reviews."), because the review counts and averages of a row rest on those reviews only; Google Play leaves the version empty on a share of reviews (31 of 200 in the Spotify sample we checked).
  • Only public review pages are read. No account, cookies or personal data are involved.

FAQ

Is it legal to scrape app store reviews? The Actor reads public review data from the same public endpoints the stores use for their own pages and the App Store review feed. It collects no data behind a login and does not store reviewer names.

How many reviews can I get? Up to 5000 per Google Play app. The App Store feed allows up to 500 per country, but it often serves only the first 50, so add countries or run again for more.

How often should I run it? Weekly or after each release. A scheduled run shows the newest version as soon as it has enough reviews.

Why is a version missing? It has fewer reviews than the minimum, or it is older than the reviews that the store made available in this run.

Why is a visible drop not marked regressed? Then the trend is inconclusive: with this few reviews the drop cannot be told from chance (see ratingDeltaMargin). stable is used only when the change is under 0.15 stars and the sample is large enough to rule out a bigger one.

Something looks wrong in the data. Open an issue on the Actor page with the input you used and the run link.

Use it together with other datagrit Actors that watch public company and product signals, for example the Hacker News hiring history Actor, to connect release quality with hiring and product activity.