# Google Play App Check — Android App Domain Enrichment (`brenton8907/google-play-app-check`) Actor

Company domain enrichment: does this company have an Android app on the Google Play Store? Matched on the developer's own listed website, not a name guess. Domains in, one row out: has_app, developer, top app, exact installs, rating, last updated. JSON/CSV/API/MCP.

- **URL**: https://apify.com/brenton8907/google-play-app-check.md
- **Developed by:** [Brenton Keller](https://apify.com/brenton8907) (community)
- **Categories:** Lead generation, SEO tools, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 domain checked, no app founds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google Play App Check — domain to app enrichment

Answers one question per row: **does this company ship an Android app?**

**Company domains in, one row out.** Not an app scraper — you do not need to know the app,
the developer or the package name. You give it `basecamp.com` and it returns `37signals`,
`com.basecamp.bc3`, **1,432,472** installs and the date they last shipped.

Output: `has_app`, the developer and their listed website, and the developer's **top** app
with exact install count, rating, category and last-updated date.

#### Why this is hard, and why the match holds

Every other way of answering this breaks on the same thing: **telling *this* company's
developer account from a same-named one.** Play search never returns zero, so a name match
always finds *something* — `allbirds.com` returns 30 bird-watching apps.

This Actor matches on **the developer's own website as published by Google Play**, so the
answer is a deterministic join rather than a guess.

#### Where this fits

| | Cost | Answers "does this domain have an app?" |
|---|---|---|
| **This Actor** | ~$0.002–0.004 per domain | **Yes, deterministically** |
| Clearbit · Apollo · ZoomInfo | enrichment plans | No — no app-store field at all |
| BuiltWith · Store Leads | $295–995/mo | Only a smart banner on the website; never inside the store |
| Apptopia · Sensor Tower | $12k–150k/year, annual contract | Yes — and priced for market research, not per-lead enrichment |

App-store intelligence has been enterprise-contract software. This is the same question
answered per domain, per credit, with no contract. Drops straight into Clay, n8n, Make or a
`curl` — see [Use it from Clay](#use-it-from-clay) below.

> ### Read this before using `has_app: false`
>
> **`has_app: false` means "no app found", not "this company has no Android app."**
>
> A `true` is strong evidence: it is a deterministic match against the developer's own
> listed website. A `false` is weaker — it means nothing in the search results claimed
> the domain, which is usually right but is still *absence of evidence* rather than
> evidence of absence.
>
> For scoring a prospect list, `true` is safe to act on; a `false` should widen a funnel,
> not disqualify a lead on its own. Check the `hint` column: when a row's search was not
> exhausted, or the per-domain request ceiling was hit, the hint says so and the `false`
> is provisional.

### Why the match is trustworthy

Every Play listing carries the developer's own website, so the match is a **deterministic
join on that website**, not a name guess.

`basecamp.com` → `37signals` is the case that proves it. The developer name shares no word
with the domain, so no amount of name matching finds it, and the website join resolves it
exactly. Name similarity is used only to decide which listings to check first — it never
decides the answer.

### Modes

**`app-check` is the product.** `app-search` is a utility mode kept for the occasional
"what's on Play for this query" lookup; if you want a general Play scraper there are better
and cheaper ones on the Store. What is hard, and what this exists for, is the domain join.

#### `app-check` — enrichment, one row per domain

| Field | Example |
|---|---|
| `has_app` | `true` |
| `developer_name` | `37signals` |
| `developer_website` | `https://basecamp.com` |
| `developer_website_matches_domain` | `true` |
| `matched_on` | `website` or `support_email` |
| `top_app_title` | `Basecamp - Project Management` |
| `top_app_installs` | `1,000,000+` |
| `top_app_installs_exact` | `1432472` |
| `top_app_rating_in_region` | `4.636` |
| `top_app_category` | `Productivity` |
| `last_updated` / `released` | `2026-09-29` / `2015-11-03` |
| `days_since_update` / `app_maintenance_status` | `8` / `active` — see below |
| `developer_app_count`, `top_app_package`, `support_email` | |
| `region` | `us` |
| `hint` | set when an answer is provisional — read it |

**Release recency is the filter worth building on**, because it separates leads that look
identical on install count. `app_maintenance_status` is `active` (updated within 90 days,
so there is a mobile team to sell to), `stale` (90–365), or `abandoned` (over a year — a
poor SDK lead and a strong modernisation-agency lead). `days_since_update` ships alongside
it, so you can re-bucket on your own thresholds without re-running anything.

#### `app-search` — utility mode

Input search queries, get the full Play listing per result: package, title, developer,
bucketed installs, rating, category. Useful for a one-off lookup; it is not what this Actor
is for.

### Input

```json
{
  "domains": ["notion.so", "basecamp.com", "stripe.com", "allbirds.com"],
  "mode": "app-check",
  "country": "us"
}
```

Pass **domains**, not brand names or IP addresses. A brand has no website to join against,
so brand input returns `has_app: null` with a hint telling you to pass the domain — and is
not charged — rather than a guess that would be wrong most of the time. An IP address gets
a free `error` row saying so.

### Cost and runtime

```
1. search Play for the domain                        1 request
2. order candidates by developer-name similarity     free
3. fetch listings until one claims the domain        1 typical
4. fetch the developer's catalogue                   1 request
5. fetch their highest-install app                   1 request
```

Measured over 15 graded domains and a 500-domain run, zero rate limits. A *match* costs
3–5 requests; a *no-match* costs ~11, because there is no early exit; the worst case is
capped by `maxRequestsPerDomain` (25 by default). An exact first-party match is the cheap
case — Play puts it in the hero card, and the hero is always the first candidate checked.

**How many domains per run.** At the no-match pace, the default 3,600 s timeout is reached
at roughly **218 domains**, and about 215 finish because the run stops starting new
domains a minute early. Split larger lists into batches of ~200, or raise the run
timeout. The run's status message carries a measured ETA, and warns up front when the list is
longer than the timeout allows. Rows are saved as each domain finishes, and the run stops
cleanly a minute before its timeout: every domain it did not reach still gets a row, with
`error` starting `not checked:` and no charge, so you can filter those and resubmit them.

**Proxy.** None by default. On Apify's own network Google Play answered 10 of 10 test
domains with 0 rate limits and the same rows as a run from a separate network, so a proxy
adds nothing but latency. If rows come back with a `blocked:` error, enable Apify Proxy
with the RESIDENTIAL group; `useApifyProxy: true` with no group resolves to datacenter,
which Google blocks faster.

### Use it from Clay

Clay has no native app-store column, so add one with **Add enrichment → HTTP API**
(Growth plan or higher):

- **Method** `POST`
- **URL** `https://api.apify.com/v2/acts/brenton8907~google-play-app-check/run-sync-get-dataset-items`
- **Query parameters** `token` = your Apify API token (store it as a Clay secret);
  `maxTotalChargeUsd` = `0.02` as a hard per-row spend ceiling; optionally `fields` =
  `query,has_app,developer_name,developer_website,top_app_title,top_app_installs_exact,top_app_rating_in_region,last_updated,days_since_update,app_maintenance_status,hint`
  to keep Clay's field picker short.
- **Header** `Content-Type: application/json`
- **Body** `{"domains": ["{{domain}}"], "mode": "app-check", "country": "us"}` — replace
  `{{domain}}` with your domain column. A bare domain, `www.` prefix or full URL all work.

The response is a one-element array; map `has_app`, `top_app_installs_exact` and the rest from
item 0. A cold call takes ~10–35 s, so raise Clay's HTTP timeout if it cuts off. For a list of
hundreds, one batch run with every domain in `domains` is faster than a call per row: join the
output back on `query`, which echoes your input verbatim, and resubmit any row whose `error`
starts with `not checked:`.

### Four things that would otherwise give you the wrong app

**The app that matches is usually not the company's main app.** Notion's listing for
*Notion Calendar* (3.6M installs) matches before the 43.9M main app; searching `figma.com`
reaches *Figma for Government* (338 installs) first. Reporting those as "the company's app"
would badly understate a company's footprint, so once the website identifies the
*developer*, the actor reads their catalogue and reports their highest-install app.

**One developer can publish under several websites.** Stripe's Dashboard app lists
`stripe.com` while their Link app lists `link.com`, and only Link appears when searching
`stripe.com`. A failed website match therefore does not clear a developer: if their name
looks right, their other apps are checked too. Without this, `stripe.com` answered "no".

**Some listings name no website at all.** Target's own app (44.7M installs) and Ford's
FordPass list only a support address, `@target.com` and `@ford.com`. When a listing carries
no website, the actor matches on its support-email domain instead and says so with
`matched_on: support_email`; without this, `target.com` answered "no". A listing that *does*
name a website must match on it, unless you set `matchOnSupportEmail`, because an agency can
put its own site on a client's app.

**One company can publish under several developer accounts.** Ford Motor Co. and Ford Motor
Credit Company both point at `ford.com`. After the first match the actor checks up to two
more brand-named developers and reports the one whose top app has the most installs —
FordPass (13.4M), not Ford Credit (285K). A domain with no other brand-named developers
costs nothing extra.

### Known limitations, measured

**A `false` is "not found", not "does not exist".** Play's search ranks by text relevance,
not ownership, so a company's own app is not guaranteed to appear for its domain. The
actor reads Play's hero "top result" card as well as the result grid, which is where an
exact first-party match almost always sits, and it checks candidates in the order Play
returned them. That covers the common case; it cannot prove a company has no app.

Until 2026-10-07 this section claimed `github.com` was an unreachable false negative.
That was wrong, and it was this actor's bug rather than a limitation of Play: GitHub's own
app is the hero result, and the parser was reading only the grid beneath it. Four domains
(`airbnb.com`, `craigslist.org`, `github.com`, `irs.gov`) answered `false` because of it.
All four now answer correctly, in 3 requests rather than 11.

**International domains need their full suffix.** `acme.com.pl` and `toko.acme.co.id`
reduce to `acme.com.pl` and `acme.co.id`, not to `com.pl` and `co.id`. A bare public
suffix (`com.pl` on its own) names no company and comes back as an unbilled error row
telling you so, rather than being silently merged with other inputs.

**Every row joins back to your input.** `query` is the exact string you submitted, so a
500-row input gives a 500-row output in the same order; `query_domain` shows the
normalised domain the match actually used. `news.ycombinator.com` and `ycombinator.com`
stay separate rows.

**Results are scoped to one Play storefront.** Ratings differ by region — Slack is `4.643`
in `us` and `4.333` in `jp` — as do availability and install counts. The rating columns are
named `top_app_rating_in_region` / `rating_in_region` and every row echoes the `region` it
came from, so a cross-region difference reads as a different storefront rather than as
instability. Comparing rows from different regions compares different storefronts.

**Shared domains cannot be joined on website or email alone.** Thousands of independent
developers list their GitHub profile as their homepage, so a `github.com` search surfaces
third-party developers whose listed website is `github.com`. For code hosts, social networks, site builders, free app hosts and
free mail providers, the actor requires the developer *name* to corroborate, and sets
`domain_is_developer_platform`. Enabling `matchOnSupportEmail` will not join on a free-mail
address at all. Both lists are maintained denylists, not complete facts.

**A developer page does not always exist.** The developer name on a listing does not always
resolve to a developer page — `/store/apps/developer?id=Strava%2C%20Inc.` is a live 404.
The match still stands, because it comes from the listing's own website field; the row
reports the matched app as the top app with `developer_app_count` empty, and is billed at
the match rate rather than the premium.

**A deep "no" is reported as provisional.** When fewer candidates were checked than the
search returned, or the per-domain request ceiling was hit, the row carries a `hint` saying
so with the counts, and `budget_exhausted`. Read the `hint` column before treating a
`false` as final.

**EU regions are not a supported configuration.** EU-region proxies may hit a consent
interstitial. The actor handles one — a redirect to a consent or `/sorry` URL is treated as
a block and answered by rotating to a fresh session — but whether rotation escapes consent
when every exit in the pool is in the EU is untested. Use a US exit; if you need EU data,
test a small batch first.

**Fields can go missing if Google reshapes its pages.** Every value is read positionally
out of an undocumented payload embedded in the page. The parsers are written to fail loudly
rather than quietly: if a page does not have the expected shape, the row carries an `error`
instead of a confident answer, and **error rows are never charged**. A single unreadable
listing is tolerated and the check moves on; if no listing for a domain can be read at all,
the row is an `error` with `has_app` empty, not a `false`. A scheduled canary checks the
parsers against known listings daily, but it cannot detect a change in Play's *ranking* —
if Google stopped surfacing companies' own apps for their domains, every check would still
look healthy while answers quietly got worse.

**No change-monitoring mode.** Each run is a point-in-time check; there is no "only new
since last run" option yet.

### Pricing

One event per outcome, because the outcomes cost very different amounts to produce.
FREE-tier prices shown; the platform's Pricing tab lists the volume tiers, which fall to
$0.0007 / $0.0014 / $0.0028 at GOLD and above.

| Event | FREE | When |
|---|---|---|
| `app-no-match` | $0.001 | domain checked, no developer claims it |
| `app-check` | $0.002 | matched, developer's catalogue could not be read |
| `app-detail` | $0.004 | matched, catalogue read, exact installs and dates present |
| `app-listing` | $0.001 | one row of an `app-search` dump |

**The cost asymmetry, stated plainly:** a no-match is the cheapest row to buy and the most
expensive to produce (~11 requests versus 4–5 for a match). That is deliberate — a `false`
is absence of evidence, so charging match price for it would be selling a weaker answer at
the same rate.

Rows that failed, were blocked, or could not be parsed are written to the dataset so the
failure is auditable, and carry **no charge**. Brand-name input, which returns an explicit
non-answer, is also free.

### Output

Results are available as JSON and CSV from the run's dataset, via the API, and over MCP.
The `App checks` dataset view puts `has_app` and `hint` side by side so a provisional
answer is visible in a spreadsheet export, not just in the API payload.

# Actor input Schema

## `mode` (type: `string`):

app-check: one row per domain answering whether that company ships an Android app, matched on the developer's own listed website. app-search: the full Play search listing for each query.

## `domains` (type: `array`):

Company domains to check, one per line (app-check mode). Use the company's real domain: the match is made against the developer's own website on their Play listing, so a domain is answerable and a brand name is not.

## `searchQueries` (type: `array`):

Play Store search queries (app-search mode).

## `maxCandidates` (type: `integer`):

How deep to look before answering 'no app'. Results are ordered by developer-name match first, which is free, so the typical domain still costs one detail request. Lower is cheaper and risks a false 'no': at 3, stripe.com answers 'no' because Stripe's own app ranks 7th behind third-party Stripe tools.

## `maxCatalogueApps` (type: `integer`):

When a developer's name looks like the domain but their matched app lists a different website, how many of their other apps to check. Stripe needs this: their Dashboard app lists stripe.com but only their Link app (link.com) appears when searching stripe.com. Kept separate from maxCandidates so the two limits bound a sum rather than a product.

## `maxRequestsPerDomain` (type: `integer`):

Absolute cap on requests spent on a single domain, whatever the other limits say. A domain whose brand word appears in many third-party developer names (telegram.org, bitcoin.com, vpn.com) would otherwise spend over 100 requests on one row. A domain stopped by this ceiling is reported with budget_exhausted and a hint saying the answer is provisional.

## `matchOnSupportEmail` (type: `boolean`):

A listing that names no website at all is always matched on its support-email domain (Target's and FordPass's listings name only @target.com and @ford.com). This option extends that to listings whose website points somewhere else: it helps companies whose Play website is a marketing domain (Notion's email is @makenotion.com while its site is notion.so), at some risk of matching an agency that lists its own website on a client's app. Free-mail addresses are never matched.

## `includeDeveloperCatalogue` (type: `boolean`):

Adds every app the matched developer publishes to the row. Costs no extra requests — the catalogue is already fetched to pick the top app.

## `maxAppsPerQuery` (type: `integer`):

Row cap per query in app-search mode.

## `requestDelaySeconds` (type: `number`):

Seconds to wait between requests. Google rate-limits per IP; raise this to 3 or more if you see blocks. Fractional values are allowed (1.5 is what the local test suite uses).

## `language` (type: `string`):

Play Store language code (hl).

## `country` (type: `string`):

Play Store country code (gl). This scopes the whole row: ratings differ by storefront (Slack is 4.643 in us and 4.333 in jp), as do availability and install counts. The region is echoed back on every row as `region`, so cross-region differences are not mistaken for instability.

## `proxyConfiguration` (type: `object`):

Off by default: measured on Apify's own network, Google Play served 10 of 10 domains (55 requests, 0 rate limits) with no proxy, and the answers matched a run from a separate network row for row. Turn on Apify Proxy with the RESIDENTIAL group only if rows come back with a 'blocked' error. `useApifyProxy: true` with no group resolves to datacenter, which Google blocks faster.

## Actor input object example

```json
{
  "mode": "app-check",
  "domains": [
    "notion.so",
    "basecamp.com",
    "stripe.com",
    "allbirds.com"
  ],
  "searchQueries": [
    "project management"
  ],
  "maxCandidates": 10,
  "maxCatalogueApps": 5,
  "maxRequestsPerDomain": 25,
  "matchOnSupportEmail": false,
  "includeDeveloperCatalogue": false,
  "maxAppsPerQuery": 30,
  "requestDelaySeconds": 2,
  "language": "en",
  "country": "us",
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per domain checked (or per app in app-search mode), as JSON from the default dataset.

## `csv` (type: `string`):

The same rows as CSV, for spreadsheets.

## `run` (type: `string`):

The run page with the Overview table, status message and log.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "notion.so",
        "basecamp.com",
        "stripe.com",
        "allbirds.com"
    ],
    "searchQueries": [
        "project management"
    ],
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("brenton8907/google-play-app-check").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "notion.so",
        "basecamp.com",
        "stripe.com",
        "allbirds.com",
    ],
    "searchQueries": ["project management"],
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("brenton8907/google-play-app-check").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "notion.so",
    "basecamp.com",
    "stripe.com",
    "allbirds.com"
  ],
  "searchQueries": [
    "project management"
  ],
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call brenton8907/google-play-app-check --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,brenton8907/google-play-app-check"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/e7UAEAMTENYA3177N/builds/oacbtParv6E9xOx4e/openapi.json
