App Store Scraper: App Details, Reviews, Search & Top Charts
Pricing
from $2.00 / 1,000 app returneds
App Store Scraper: App Details, Reviews, Search & Top Charts
Everything public about an iOS app in clean rows: ratings, price, version history notes, screenshots, developer and genre data, plus up to 500 recent reviews per country storefront and chart positions. Loop countries for global review coverage. Built for ASO, competitor tracking and app research.
Pricing
from $2.00 / 1,000 app returneds
Rating
0.0
(0)
Developer
Paul Vasquez
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Apple App Store Suite
Collect public Apple App Store metadata, keyword search results, chart entries, and customer reviews without an API key. This Python 3.12 actor uses Apple's iTunes lookup and search APIs, Apple Marketing Tools charts, and the XML customer review feeds. Optional public product page enrichment looks for additional data inside embedded JSON. It is useful for storefront comparisons, app research, review monitoring, and small competitive datasets. Results describe the public storefront at collection time; this is not an installation or revenue estimator.
Quick start
Install Python 3.12, create a virtual environment, and install requirements:
python -m venv .venv.venv/Scripts/python.exe -m pip install -r requirements.txtapify validate-schema .actor/input_schema.jsonapify run
The included INPUT.json looks up WhatsApp and Facebook in the US and GB storefronts and requests up to 50 reviews per app per storefront. It needs no keys and is intended for the daily automated smoke test, normally completing within two minutes. The validation harness runs that input plus the requested budget search and top-free chart checks with isolated local storage:
.venv/Scripts/python.exe -m unittest discover -s tests -vpowershell -File validation/run_live.ps1
The harness launches the actual Apify SDK, saves logs and complete JSON results under validation, and checks counts, uniqueness, and eligible event counts. It does not push or publish the actor. Local SDK charges do not bill money.
Inputs and modes
mode selects lookup, search, charts, or reviews; lookup is the default.
For lookup and reviews, supply appIds as strings containing numeric Apple
IDs, bundle identifiers, or HTTPS apps.apple.com product URLs. Bundle IDs are
resolved through lookup before fetching reviews. Unsupported identifiers produce
free error rows while other inputs continue. Search accepts queries, an array
of phrases, and requests one page of up to 200 candidates per query and country.
Apple controls relevance and may return fewer candidates than requested.
Charts accepts chart as top-free or top-paid. The actor fetches the overall
top 100 chart and hydrates selected IDs through lookup so chart app rows have
the same metadata fields as other modes. Optional genreId filters search and
filters chart entries within that overall top 100; it does not request a separate
genre chart. No ranking beyond the fetched page is inferred.
countries defaults to ["us"]. Codes are normalized to lowercase and duplicate
codes removed. maxApps, default 100, is a global cap on unique app/storefront
pairs, including resolved apps in reviews mode. Input order determines which
pairs fit. The same app in two storefronts consumes two slots. Search results
and lookup aliases are deduplicated within each storefront.
Set includeReviews to true to attach review collection to app discovery.
Reviews mode emits review rows without app rows. maxReviewsPerApp defaults
to 100 and accepts 1 through 500 per app per review country. reviewCountries
defaults to countries; supply additional storefronts to collect more than 500
reviews across countries. Each feed exposes at most ten pages, usually 50
reviews per page. The XML variant is intentional: Apple's JSON variant can
return empty feeds. Repeated review IDs are removed per app and storefront;
duplicate-only or empty pages stop pagination. These feeds are recent snapshots,
not a complete historical archive.
Output and enrichment
Every dataset row has rowType and source. App rows include IDs, name,
developer and seller, price and currency, category and genres, overall and
current-version ratings and counts, version and dates, release notes, description,
size, minimum OS, content rating, languages, screenshots, icon, URL, and country.
Unavailable API fields remain null or empty arrays. Size retains Apple's API
representation, generally a byte-count string.
Review rows contain appId, country, reviewId, author, rating, title, content, version, updatedAt, voteCount, and voteSum. Page rows identify the search query or chart and report candidate count before the global cap and deduplication. SUMMARY in the default key-value store records row counts, event counts, and processing time. Errors and empty results appear as separate free rows.
includePageDetails defaults to false. When enabled, the actor searches embedded
JSON for privacy labels, in-app purchases, rating histograms, and screenshots.
These additional fields preserve Apple's nested structures. Missing or changed
page structures leave details unavailable; they do not imply no tracking or no
purchases. Fetch failures add a warning without discarding the API app row.
Pricing and reliability
The custom PPE prices are $0.002 per app-returned, $0.0002 per review-returned, and $0.005 per page-returned. Each successful emitted row requests one matching event through Actor.charge. Empty search/chart results are free summary rows. Errors and zero-result summaries are uncharged. Configure these events in Console and disable synthetic events before publication; the metadata file alone does not activate Store pricing. A refused charge stops output. Charging precedes persistence, so a storage failure cannot be rolled back transactionally.
Requests retry twice with one- and two-second backoff on 429, server errors,
and transport failures. Other HTTP errors fail immediately. timeoutSecs bounds
each HTTP operation, not the whole run. Large country lists and optional details
increase runtime. Public endpoints can throttle or change. See VALIDATION.md
for measured results and the distinction between mocked, live, and hosted checks.
Example output
One recorded dataset row, trimmed by omitting fields only. Values are the saved snapshot, not current measurements. Source: validation/results-lookup.json, first row in the rows array.
{"rowType": "app","appId": "310633997","bundleId": "net.whatsapp.WhatsApp","name": "WhatsApp Messenger","country": "us","price": 0.0,"currency": "USD","version": "26.37.76"}
Use cases
- A mobile product manager can collect recent competitor reviews and group complaints by app version before prioritizing customer interviews.
- An app publisher can compare the same app across US and GB storefronts, preserving country alongside prices, currencies, and rating counts.
- An acquisition analyst can search a category keyword and assemble an app shortlist for manual review of developers and product pages.
- An app marketing agency can save top-free chart snapshots and compare returned app IDs across runs in its own reporting database.
Pricing example
Hypothetical batch, calculated from .actor/pay_per_event.json:
| Event | Count | USD per event | Subtotal |
|---|---|---|---|
app-returned | 100 | $0.002 | $0.2000 |
review-returned | 1,000 | $0.0002 | $0.2000 |
page-returned | 10 | $0.005 | $0.0500 |
Total declared event charges: $0.45. These counts are a budgeting example, not a promised yield or an actual bill. Any applicable platform or proxy costs are outside this calculation.
Limitations
Review feeds expose recent pages, not complete lifetime feedback. Chart genre filtering operates inside the fetched overall top 100. Optional page details may be absent even when an app lookup succeeds.