AI Visibility Tracker: GEO/AEO Brand Monitor with Trends avatar

AI Visibility Tracker: GEO/AEO Brand Monitor with Trends

Under maintenance

Pricing

from $60.00 / 1,000 ai visibility checks

Go to Apify Store
AI Visibility Tracker: GEO/AEO Brand Monitor with Trends

AI Visibility Tracker: GEO/AEO Brand Monitor with Trends

Under maintenance

Track how often ChatGPT, Google AI Overviews, Perplexity and Gemini mention your brand — with confidence intervals on every number and a tested month-over-month comparison, so you only act on changes that are real.

Pricing

from $60.00 / 1,000 ai visibility checks

Rating

0.0

(0)

Developer

Paul Bernabeu

Paul Bernabeu

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

AI Visibility Tracker — GEO/AEO monitor with month-over-month trends

Track how often ChatGPT, Google AI Overviews, Perplexity and Gemini mention your brand — with a confidence interval on every number, and a statistically tested comparison against your last run.

Most AI-visibility tools hand you a single percentage from a single pass. That number is noise. Ask an AI assistant the same question twice and you get different answers: SparkToro's study of 2,961 prompt executions found under a 1-in-100 chance that two runs return the same brand list. Report "share of voice: 34%" to a client, get 22% next week, and you have lost the client.

This Actor is built around that problem.


What makes it different

1. Every number carries a confidence interval. Each prompt is asked k times per engine (default 3). Results are reported as a Wilson score interval, which stays correct at small sample sizes and at rates of 0% or 100% — where the usual approximation produces intervals of zero width or bounds outside 0–100%. The interval, not the headline figure, is the result.

2. It remembers. Every other tracker forgets. Runs are stored under a projectId in a named key-value store that survives between runs. Run it again next month on the same projectId and you get a real comparison. This is the difference between a one-off audit and something worth scheduling.

3. It refuses to report noise as a trend. Month-over-month movements go through a two-proportion z-test. Only movements outside the noise band are reported as changes. Everything else is returned in suppressedAsNoise — visible, but not presented as a finding.

A 50% → 60% move at n=20 is noise. The same 10-point move at n=600 is real. Almost every tool in this category reports both as a trend.

4. It separates grounded from parametric visibility. An answer produced after a live web search responds to content and PR work this quarter. An answer produced from model weights does not. They are different metrics with different remedies, and collapsing them into one score hides which lever applies. Both are reported separately.

5. "No answer" and "collection failed" are never conflated. Google decides whether an AI Overview appears; when none appears that is a real observation. A blocked fetch is missing data. Mixing them silently depresses your score. They are counted separately, and rates are computed over successful samples only.

6. It produces something you can hand to a client. A self-contained, printable HTML report is saved under the key REPORT, alongside the raw dataset.


What it costs

What this Actor charges:

EventPriceWhen
actor-start$0.01once per run
visibility-check$0.035per prompt × engine × sample
trend-analysis$1.00only when a previous run exists to compare against
agency-report$2.00only when the HTML report is generated

You are only charged for checks that actually return an answer. Google frequently shows no AI Overview for a given query. Those samples are recorded as real observations, and they are not billed.

Plus downstream cost, which you should read before running at volume. The scraped engines run through Apify's own Actors and bill your Apify account separately:

EngineFree planBronze and above
Google AI Overview$0.003/query$0.002 or less
Google AI Mode$0.20/query$0.005 or less
ChatGPT search$0.20/query$0.005 or less

On the free plan AI Mode and ChatGPT cost roughly 66× more than AI Overview, so only Google AI Overview is enabled by default. Turn the others on deliberately, and preferably on a paid plan.

Worked example, 20 prompts × 3 samples on Google AI Overview: 60 checks = $2.11 from this Actor, plus $1.00 trend and $2.00 report, plus about $0.18 downstream. Roughly $5.30 for a monthly audit.

The full plan, including downstream estimate, is printed in the log before any work starts, and the run stops cleanly at your maxTotalChargeUsd rather than overrunning it.

Compare: Profound starts at $99/month, Scrunch at $300, Otterly at $29 for 15 prompts. Those are subscriptions, and all of them gate API and raw-data access behind an enterprise tier.

Try it free. Set testMode: true to generate deterministic synthetic answers. Nothing is charged, no engine is contacted, and you see the exact output schema and report layout before spending anything.


How the data is collected — disclosed per engine

Vendors in this category are routinely vague about this. Here it is in full.

EngineMethodKey neededBilled by
Google AI Overviewapify/google-ai-overviews-scrapernoyour Apify account (downstream)
Google AI Modeapify/google-ai-mode-scrapernoyour Apify account (downstream)
ChatGPT (search)apify/chatgpt-search-scrapernoyour Apify account (downstream)
Perplexityofficial Sonar APIyesPerplexity
GeminiGemini API + Google Search groundingyesGoogle
ClaudeAnthropic API + web_search (max_uses: 1)yesAnthropic
OpenAIResponses API + web_searchyesOpenAI

The three scraped surfaces run through Apify's own first-party Actors. That is deliberate: proxies, browser fingerprinting, captchas and HTML parsing are the parts that rot fastest, and Apify maintains them daily. Their platform usage is billed to your Apify account separately and is called out in the run log.

The API-based engines use your own keys and are billed to you by the provider. Nothing is resold.

Three details that matter, all found by running the upstream Actors rather than reading their docs:

  • Google's citation links cannot be resolved. AI Overview returns session-scoped stubs like /goto?url=CAESmwEB..., which return HTTP 400 when fetched independently. So for the Google engines this Actor reports the publisher name read from the citation title (reddit.com/r/devops, Chaser, Fluid CRM Blog) rather than pretending to have a URL. Engines you supply a key for return real URLs, and those are reported as resolved domains.
  • The scraped Google surfaces have no locale control. The upstream Actors accept only a list of queries: no country, no language. Results come from wherever the proxy exits, which in testing was not Europe. If you need a specific market, use the API-based engines with your own key.
  • Gemini's grounding URLs are vertexaisearch redirects that expire, so they are resolved at collection time and anything unresolvable is dropped rather than stored. Claude's web_search is pinned to max_uses: 1 so one request cannot silently trigger several billable searches.

On efficiency: each sample round sends every prompt to an engine in a single upstream run, so a 20-prompt, 3-sample job is 3 upstream runs rather than 60. The upstream Actors deduplicate identical queries, which is why repeated samples cannot share a run.


Input

{
"brand": "Linear",
"brandAliases": ["Linear.app"],
"competitors": ["Jira", "Asana", "Notion", "ClickUp"],
"ownedDomains": ["linear.app"],
"prompts": [
"best project management tool for a small design agency",
"what should I use instead of Jira",
"affordable issue tracker for a startup engineering team"
],
"engines": ["google_ai_overview", "chatgpt_search"],
"samplesPerPrompt": 3,
"projectId": "linear",
"compareWithPrevious": true,
"generateHtmlReport": true
}

Write prompts the way a customer would ask them. "best project management tool for a small design agency" measures something. "project management" does not.

Keep projectId stable across runs — it is what makes the comparison possible.


Output

Three record types in the dataset, plus SCORECARD, TREND and REPORT in the key-value store.

Scorecard — overall rates with intervals, per-engine and per-prompt breakdowns, cited source domains, competitor share, and a dataQuality block that says plainly when there is not enough data to make a claim.

Sample — one row per prompt × engine × repetition, so any headline number can be audited back to the raw observations behind it:

{
"type": "sample",
"prompt": "best project management tool for a small design agency",
"engine": "chatgpt_search",
"sampleIndex": 2,
"status": "ok",
"mode": "grounded",
"brandMentioned": true,
"mentionRank": 2,
"recommended": true,
"sentiment": "positive",
"competitorsMentioned": ["Asana", "Notion"],
"citedDomains": ["g2.com", "reddit.com"],
"ownedDomainCited": false
}

Trend — only present from the second run onward:

{
"type": "trend",
"daysBetween": 31,
"changes": [
{
"metric": "mentionRate", "scope": "overall",
"previous": 0.364, "current": 0.625, "delta": 0.261,
"significant": true, "pValue": 0.0002, "direction": "up",
"interpretation": "Increase of +26.1 pts (p=0.0002). This is outside the noise band and worth acting on."
}
],
"suppressedAsNoise": [ "…movements that did not clear the test" ],
"newCitedDomains": ["theverge.com"],
"lostCitedDomains": []
}

The most useful field

The cited sources. They tell you what the engines actually drew on, and where a competitor is cited and you are not, that is your content brief. Mention rate tells you the score; cited sources tell you what to do about it.

Read as topCitedDomains (resolved domains, from engines returning real URLs) or topCitedSources (publisher names, from the Google surfaces, for the reason given above). The report shows whichever is available and labels which it is.


Scheduling

Use Apify's built-in Scheduler and run it weekly or monthly on the same projectId. From the second run onward every report opens with what changed and whether the change is real. That is the whole point of the tool.


Method notes

Brand detection, endorsement and sentiment are deterministic and rule-based, not LLM calls. Running the analysis twice over the same answers always produces the same numbers. In a tool whose entire claim is reproducibility, the analysis layer must not itself be a source of variance. All sampling variation comes from the engines, where it belongs.

Brand matching is accent- and case-insensitive, respects Unicode word boundaries, and handles possessives and names containing dots and hyphens — so Apple matches Apple's but not Applesauce, and Node.js and Coca-Cola match correctly.

On "rank". Average position across samples is reported with a standard deviation, and should be read only in aggregate. Any tool presenting a single stable "your rank in ChatGPT" is overstating what the underlying data supports.


Limitations, stated plainly

  • Google decides whether an AI Overview appears at all. For some queries and countries there simply is not one, and that is recorded as no_answer.
  • Scraped surfaces reflect a logged-out session, which can differ from a personalised one.
  • API-based engines measure that provider's API retrieval, which is not identical to its consumer product.
  • Sentiment is lexicon-based. It is reproducible and auditable, but less nuanced than a human read.
  • Google's answer payload arrives with the site's own share and feedback UI appended, sometimes in an unrelated language. That chrome is stripped before analysis, and in testing it was 37 to 43% of the raw payload.
  • Sample sizes below roughly 9 usable observations produce intervals too wide to support a claim. The Actor tells you when this happens instead of quietly reporting the number anyway.