Google Play Scraper: App Details, Reviews, Search, Developers avatar

Google Play Scraper: App Details, Reviews, Search, Developers

Pricing

from $2.00 / 1,000 app returneds

Go to Apify Store
Google Play Scraper: App Details, Reviews, Search, Developers

Google Play Scraper: App Details, Reviews, Search, Developers

Every public fact about an Android app in clean rows: installs, ratings, price, in-app purchases, screenshots, version and update dates, plus paginated reviews with replies, keyword search, a developer's full catalogue and similar apps. Paced and proxied to avoid Google's throttling.

Pricing

from $2.00 / 1,000 app returneds

Rating

0.0

(0)

Developer

Paul Vasquez

Paul Vasquez

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Google Play Suite

Collect Google Play app details, discover apps through search or developer pages, find similar apps, and retrieve public reviews. The actor uses public Google Play responses without a Google account or API key. It normalizes the results into app, review, and discovery-page rows, with separate uncharged error and empty-result rows. Suitable uses include catalogue research, competitor monitoring, and review analysis.

Quick start

Use Python 3.12. Create a virtual environment, install requirements.txt, and run apify run from this directory. The committed INPUT.json requests Spotify and Duolingo details with up to forty reviews each. It explicitly disables Apify Proxy so the daily local test needs no credentials. The measured default validation completed well under two minutes; remote service delays can change that timing.

For hosted runs, the input schema and runtime default proxyConfiguration to {"useApifyProxy":true}. Google can return empty pages to datacenter addresses instead of an explicit HTTP error. Proxy access depends on your Apify account and can have separate costs. Set the configuration explicitly when using a custom proxy or intentionally running directly. Proxy URLs are never included in output rows.

Inputs

Choose mode: details, search, developer, similar, or reviews. Details is the default. Supply packageNames for details, similar, and reviews; entries accept package IDs or HTTPS Google Play details URLs. Search takes queries. Developer mode takes developerIds, using the numeric /store/apps/dev route or named /store/apps/developer route automatically. Arrays accept up to one hundred targets. Package inputs and repeated targets are deduplicated.

language and country default to en and us. maxApps defaults to 100 and caps unique app attempts across the entire run, including multiple targets. It accepts 1–1000. Discovery hydrates each selected package through its details page, producing the same app schema as direct lookup. A failed hydration consumes an attempt and produces a free error row.

Set includeReviews to add reviews after each successful app lookup. Reviews mode retrieves reviews directly without billing for an app-details row. maxReviewsPerApp defaults to 100 and accepts 1–10000. reviewSort supports newest, rating, and helpfulness; rating means Google's rating sort. Optional reviewScoreFilter accepts 1–5 and is applied in both the request and normalization path. timeoutSecs is a per-request timeout, defaults to 20, and accepts 1–120.

Output and prices

Every row includes rowType and source. App rows include package name, title, developer identity and contact fields, descriptions, category, price and currency, purchase and advertising flags, installs, ratings, review count, a five-element histogram ordered one through five stars, version, dates, Android requirement, content rating, images, video, and canonical URL. Missing values remain null; absent screenshots become an empty array. Update timestamps use UTC ISO notation; release dates retain Google's localized display text when no machine timestamp exists.

Review rows contain package name, review ID, author, rating, content, thumbs-up count, app version, creation timestamp, developer reply, and reply timestamp. Discovery-page rows contain mode, target, page number, selected package IDs, count, and source. A page represents successful discovery even if a later app-details request fails. Dataset exports therefore contain several row types; filter rowType before loading a single entity table.

Each successful app row requests one app-returned event at $0.002. Each review requests one review-returned event at $0.0002. Each discovery page containing new selected IDs requests one page-returned event at $0.003. Duplicate-only pages, errors, and zero-result summaries are uncharged. Local SDK runs do not bill. Configure these three events and disable synthetic charges before publication. Charging precedes persistence; storage failure after charging cannot be rolled back by this implementation. A refused charge stops collection before writing that row.

Collection behavior and limits

Requests are sequential and paced. HTTP 429, server errors, and transport failures receive two retries with one- and two-second backoff. Review pages receive an additional half-second delay. Continuation tokens drive review and discovery pagination, with repeated-token detection, review-ID deduplication, and hard bounds of 200 review pages and 50 discovery pages per target. Reaching those bounds emits an uncharged error instead of claiming completeness.

Empty or structurally invalid responses produce warnings or errors identifying possible throttling or schema drift. A valid empty review list cannot prove whether Google has no matches or has throttled the request. Similar mode selects an explicitly labelled similar-apps collection using an English seed page, then requests the collection and app details in your chosen locale.

Development and validation

Run .venv/Scripts/python.exe -m unittest discover -s tests -v, apify validate-schema .actor/input_schema.json, and validation/run_live.ps1. The live harness isolates local storage and excludes inherited actor authentication variables from child processes. It records SDK logs, rows, counts, and elapsed timings under validation/. See VALIDATION.md for measured results and limitations.

The MIT-licensed google-play-scraper supplies detail parsing. This actor owns transport, normalization, pagination, errors, deduplication, and charging. Google's internal protocols can change: the observed reviews endpoint uses oCPfdb; the supplied legacy UsvDTd request returned an RPC error. Developer, similar, and later-page behavior have mocked coverage; hosted proxy operation and real billing still require deployment validation.

Example output

One recorded dataset row, trimmed by omitting fields only. Values are the saved snapshot, not current measurements. Source: validation/results-details.json, first row in the rows array.

{
"rowType": "app",
"packageName": "com.spotify.music",
"title": "Spotify: Music and Podcasts",
"developer": "Spotify AB",
"rating": 4.3471756,
"url": "https://play.google.com/store/apps/details?id=com.spotify.music&hl=en&gl=us"
}

Use cases

  • An Android product team can retrieve reviews for a package list and route recurring complaints, ratings, and developer replies into a review worksheet.
  • A localization agency can collect app details in selected language and country settings, then compare descriptions and displayed metadata outside the actor.
  • A competitive intelligence analyst can search a category phrase and inspect returned packages before adding relevant apps to a recurring watchlist.
  • A mobile portfolio team can refresh public versions, ratings, and purchase flags for known apps and compare saved snapshots in its own database.

Pricing example

Hypothetical batch, calculated from .actor/pay_per_event.json:

EventCountUSD per eventSubtotal
app-returned100$0.002$0.2000
review-returned1,000$0.0002$0.2000
page-returned10$0.003$0.0300

Total declared event charges: $0.43. These counts are a budgeting example, not a promised yield or an actual bill. Any applicable platform or proxy costs are outside this calculation.

Limitations

Discovery-page charges can apply even when subsequent app hydration fails. Empty review responses are ambiguous. Later-page, developer, and similar-mode coverage is not established by the recorded details run.