App Store Scraper: App Details, Reviews, Search & Top Charts avatar

App Store Scraper: App Details, Reviews, Search & Top Charts

Pricing

from $2.00 / 1,000 app returneds

Go to Apify Store
App Store Scraper: App Details, Reviews, Search & Top Charts

App Store Scraper: App Details, Reviews, Search & Top Charts

Everything public about an iOS app in clean rows: ratings, price, version history notes, screenshots, developer and genre data, plus up to 500 recent reviews per country storefront and chart positions. Loop countries for global review coverage. Built for ASO, competitor tracking and app research.

Pricing

from $2.00 / 1,000 app returneds

Rating

0.0

(0)

Developer

Paul Vasquez

Paul Vasquez

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

Apple App Store Suite

Collect public Apple App Store metadata, keyword search results, chart entries, and customer reviews without an API key. This Python 3.12 actor uses Apple's iTunes lookup and search APIs, Apple Marketing Tools charts, and the XML customer review feeds. Optional public product page enrichment looks for additional data inside embedded JSON. It is useful for storefront comparisons, app research, review monitoring, and small competitive datasets. Results describe the public storefront at collection time; this is not an installation or revenue estimator.

Quick start

Install Python 3.12, create a virtual environment, and install requirements:

python -m venv .venv
.venv/Scripts/python.exe -m pip install -r requirements.txt
apify validate-schema .actor/input_schema.json
apify run

The included INPUT.json looks up WhatsApp and Facebook in the US and GB storefronts and requests up to 50 reviews per app per storefront. It needs no keys and is intended for the daily automated smoke test, normally completing within two minutes. The validation harness runs that input plus the requested budget search and top-free chart checks with isolated local storage:

.venv/Scripts/python.exe -m unittest discover -s tests -v
powershell -File validation/run_live.ps1

The harness launches the actual Apify SDK, saves logs and complete JSON results under validation, and checks counts, uniqueness, and eligible event counts. It does not push or publish the actor. Local SDK charges do not bill money.

Inputs and modes

mode selects lookup, search, charts, or reviews; lookup is the default. For lookup and reviews, supply appIds as strings containing numeric Apple IDs, bundle identifiers, or HTTPS apps.apple.com product URLs. Bundle IDs are resolved through lookup before fetching reviews. Unsupported identifiers produce free error rows while other inputs continue. Search accepts queries, an array of phrases, and requests one page of up to 200 candidates per query and country. Apple controls relevance and may return fewer candidates than requested.

Charts accepts chart as top-free or top-paid. The actor fetches the overall top 100 chart and hydrates selected IDs through lookup so chart app rows have the same metadata fields as other modes. Optional genreId filters search and filters chart entries within that overall top 100; it does not request a separate genre chart. No ranking beyond the fetched page is inferred.

countries defaults to ["us"]. Codes are normalized to lowercase and duplicate codes removed. maxApps, default 100, is a global cap on unique app/storefront pairs, including resolved apps in reviews mode. Input order determines which pairs fit. The same app in two storefronts consumes two slots. Search results and lookup aliases are deduplicated within each storefront.

Set includeReviews to true to attach review collection to app discovery. Reviews mode emits review rows without app rows. maxReviewsPerApp defaults to 100 and accepts 1 through 500 per app per review country. reviewCountries defaults to countries; supply additional storefronts to collect more than 500 reviews across countries. Each feed exposes at most ten pages, usually 50 reviews per page. The XML variant is intentional: Apple's JSON variant can return empty feeds. Repeated review IDs are removed per app and storefront; duplicate-only or empty pages stop pagination. These feeds are recent snapshots, not a complete historical archive.

Output and enrichment

Every dataset row has rowType and source. App rows include IDs, name, developer and seller, price and currency, category and genres, overall and current-version ratings and counts, version and dates, release notes, description, size, minimum OS, content rating, languages, screenshots, icon, URL, and country. Unavailable API fields remain null or empty arrays. Size retains Apple's API representation, generally a byte-count string.

Review rows contain appId, country, reviewId, author, rating, title, content, version, updatedAt, voteCount, and voteSum. Page rows identify the search query or chart and report candidate count before the global cap and deduplication. SUMMARY in the default key-value store records row counts, event counts, and processing time. Errors and empty results appear as separate free rows.

includePageDetails defaults to false. When enabled, the actor searches embedded JSON for privacy labels, in-app purchases, rating histograms, and screenshots. These additional fields preserve Apple's nested structures. Missing or changed page structures leave details unavailable; they do not imply no tracking or no purchases. Fetch failures add a warning without discarding the API app row.

Pricing and reliability

The custom PPE prices are $0.002 per app-returned, $0.0002 per review-returned, and $0.005 per page-returned. Each successful emitted row requests one matching event through Actor.charge. Empty search/chart results are free summary rows. Errors and zero-result summaries are uncharged. Configure these events in Console and disable synthetic events before publication; the metadata file alone does not activate Store pricing. A refused charge stops output. Charging precedes persistence, so a storage failure cannot be rolled back transactionally.

Requests retry twice with one- and two-second backoff on 429, server errors, and transport failures. Other HTTP errors fail immediately. timeoutSecs bounds each HTTP operation, not the whole run. Large country lists and optional details increase runtime. Public endpoints can throttle or change. See VALIDATION.md for measured results and the distinction between mocked, live, and hosted checks.

Example output

One recorded dataset row, trimmed by omitting fields only. Values are the saved snapshot, not current measurements. Source: validation/results-lookup.json, first row in the rows array.

{
"rowType": "app",
"appId": "310633997",
"bundleId": "net.whatsapp.WhatsApp",
"name": "WhatsApp Messenger",
"country": "us",
"price": 0.0,
"currency": "USD",
"version": "26.37.76"
}

Use cases

  • A mobile product manager can collect recent competitor reviews and group complaints by app version before prioritizing customer interviews.
  • An app publisher can compare the same app across US and GB storefronts, preserving country alongside prices, currencies, and rating counts.
  • An acquisition analyst can search a category keyword and assemble an app shortlist for manual review of developers and product pages.
  • An app marketing agency can save top-free chart snapshots and compare returned app IDs across runs in its own reporting database.

Pricing example

Hypothetical batch, calculated from .actor/pay_per_event.json:

EventCountUSD per eventSubtotal
app-returned100$0.002$0.2000
review-returned1,000$0.0002$0.2000
page-returned10$0.005$0.0500

Total declared event charges: $0.45. These counts are a budgeting example, not a promised yield or an actual bill. Any applicable platform or proxy costs are outside this calculation.

Limitations

Review feeds expose recent pages, not complete lifetime feedback. Chart genre filtering operates inside the fetched overall top 100. Optional page details may be absent even when an app lookup succeeds.