Chinese AI Visibility Dataset Analyzer avatar

Chinese AI Visibility Dataset Analyzer

Pricing

from $4.25 / 1,000 normalized ai observations

Go to Apify Store
Chinese AI Visibility Dataset Analyzer

Chinese AI Visibility Dataset Analyzer

Normalize existing Kimi, GLM, DeepSeek, Qwen, ChatGPT, Gemini, Perplexity and Google AI Overview datasets; score brand mentions, rank and evidence quality without calling an AI model.

Pricing

from $4.25 / 1,000 normalized ai observations

Rating

0.0

(0)

Developer

Tim Zinin

Tim Zinin

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Turn AI-answer rows you already own into one comparable visibility Dataset. The Actor deterministically normalizes supported ChatGPT, Gemini, Perplexity, Google AI Overviews, Kimi, Qwen, DeepSeek, and GLM-shaped observations; finds your brand and configured competitors; records their order; and classifies citation quality.

It does not browse, call an LLM, or launch another Actor. You bring the observations; this Actor makes them consistent, traceable, and ready for analysis.

Four AI-answer streams becoming normalized visibility records

What you get

Each compatible input observation becomes one aiObservation 1.0 row in the default Dataset, with:

  • a deterministic brand mention and one-based position among the tracked names;
  • detected competitors in first-mention order;
  • a normalized model family and model ID;
  • structured citation URLs, domains, and a grounding classification;
  • source Dataset, Run, row, adapter, timestamp, and bounded evidence provenance;
  • quality flags for conditions such as missing citations, fallback timestamps, or ambiguous aliases.

The default Dataset contains successful results only. Duplicates, incompatible rows, source-read errors, and coverage counters go to free Key-value store records instead of being mixed with analysis-ready data.

This is useful for SEO and brand teams consolidating exports from multiple AI answer collectors, agencies preparing auditable client reporting, and analysts who need a stable schema before loading observations into Sheets, a warehouse, or a dashboard.

Quick start with the public sample Task

Run the public Task analyze-chinese-ai-visibility-sample (Task ID IQB5feJ4eXgPyjSkl) to see the full output without granting access to a private Dataset. Its four inline rows are fictional, clearly labeled examples; they are not live answers or current search evidence. The unchanged sample produces four normalized rows.

For your own small sample, replace the inline observations:

{
"brand": "Acme Running",
"brandAliases": ["Acme"],
"competitors": [
{ "name": "StrideLab", "aliases": [] }
],
"inlineRows": [
{
"sourceActorId": "apify/chatgpt-search-scraper",
"query": "Which running shoe brands should beginners compare?",
"text": "Acme Running and StrideLab are options to compare.",
"sources": [
{ "url": "https://example.com/running-guide" }
],
"checkedAt": "2026-08-01T10:00:00.000Z"
}
],
"market": "United States",
"language": "English",
"maxRows": 100
}

The citation URL in this payload is deliberately illustrative. Supply only observations and URLs you own or are authorized to process.

Analyze your Apify Datasets

Choose up to 10 Datasets in Source Dataset IDs. The Actor receives scoped READ access only to the resources you select. Then provide one adapter name in Dataset adapter hints for each Dataset, in the same order:

{
"brand": "Acme Running",
"competitors": [
{ "name": "StrideLab", "aliases": ["Stride Lab"] }
],
"sourceDatasetIds": ["A1b2C3d4E5f6G7h8I"],
"sourceDatasetHints": ["apify/gemini-search-scraper"],
"market": "China",
"language": "English",
"maxRows": 500
}

The positional hint is mandatory for selected Datasets. Similar ChatGPT and Gemini result shapes cannot otherwise be mapped reliably. To analyze an Actor run, select its default Dataset rather than entering a Run ID. Run provenance is read from that Dataset's metadata when it is available.

Supported adapter hints are:

Source shapeAdapter hint
Bulk LLM Runnerfayoussef/bulk-llm-runner
ChatGPT Search Scraperapify/chatgpt-search-scraper
Google AI Overviews Scraperapify/google-ai-overviews-scraper
Perplexity Search Scraperapify/perplexity-search-scraper
Gemini Search Scraperapify/gemini-search-scraper
LLM Brand Visibilityzinin/llm-brand-visibility
AI Overview Trackerzinin/ai-overview-tracker
Already normalizedai-observation-1.0

For inline JSON or CSV, leave sourceActorHint as auto or set one global adapter. A per-row sourceActorId takes precedence for inline JSON.

Input reference

FieldRequiredPurpose
brandYesCanonical tracked brand, up to 100 characters.
brandAliasesNoAlternative spellings or local-language names.
competitorsNoUp to 20 competitors, each with optional aliases.
inlineRowsOne sourceUp to 1,000 JSON observation rows.
inlineCsvOne sourceHeader plus records; JSON-looking cells are decoded.
sourceDatasetIdsOne sourceUp to 10 buyer-selected Datasets with scoped READ.
sourceDatasetHintsWith DatasetsOne positional adapter per selected Dataset.
sourceActorHintNoAdapter for inline data; defaults to auto.
market, languageNoFallbacks only when a source row omits them.
maxRowsNoHard cap across all sources, from 1 to 1,000.

You may combine inline and Dataset sources. Processing stops at maxRows across the combined source stream.

How scoring works

Matching is deterministic and case-insensitive. The Actor searches the supplied answer for the canonical brand, its aliases, and configured competitors. Brand position is its one-based order among those tracked names, not a general search ranking. If the brand is absent, mentioned is false and position is null.

Aliases shared by the brand and a competitor are flagged as ambiguous and are not double-counted. Citations are accepted only from structured source fields; a URL merely written in answer text is not promoted to evidence. Rows are deduplicated by deterministic observation provenance.

The Actor does not judge whether an answer is factually correct, positive, or negative. groundingStatus describes the supplied evidence structure, not the truth of the underlying answer.

Output Dataset

The default Dataset exposes a table view for brand, query, model, mention, position, competitors, grounding, quality flags, citations, answer excerpt, and timestamp. Every complete record also includes all 25 contract fields:

{
"schemaVersion": "1.0",
"observationId": "c2f2...",
"brand": "Acme Running",
"competitor": "StrideLab",
"competitors": ["StrideLab"],
"query": "Which running shoe brands should beginners compare?",
"modelFamily": "ChatGPT",
"modelId": "chatgpt-search",
"market": "United States",
"language": "English",
"mentioned": true,
"position": 1,
"answerSnippet": "Acme Running and StrideLab are options to compare.",
"citedUrls": ["https://example.com/running-guide"],
"citedDomains": ["example.com"],
"groundingStatus": "grounded",
"qualityFlags": [],
"checkedAt": "2026-08-01T10:00:00.000Z",
"sourceActorId": "apify/chatgpt-search-scraper",
"sourceRunId": null,
"sourceDatasetId": null,
"sourceRowIndex": 0,
"dependencyUsage": null,
"dependencyCostUsd": null,
"evidence": {
"adapter": "apify/chatgpt-search-scraper"
}
}

answerSnippet is capped at 800 characters. Dependency usage and cost are preserved source metadata when present; they are not this Actor's operating cost.

Three free Key-value store records support auditing:

  • OUTPUT — source, delivery, billing, and quarantine counts;
  • COVERAGE — counts by source and normalized model family;
  • QUARANTINE — bounded reasons for withheld rows, without copying raw payloads.

Pricing

This Actor uses Pay Per Event. A run charges one automatic Actor-start event and one result-found event for each normalized row atomically delivered to the default Dataset. Quarantine, duplicates, source errors, and KVS diagnostics do not emit result-found and are free.

Validated rows enter paid output while quarantined rows remain free

TierActor startEach delivered resultDiscount
Free$0.005000$0.0050000%
Bronze$0.004750$0.0047505%
Silver$0.004500$0.00450010%
Gold$0.004250$0.00425015%
Platinum$0.004100$0.00410018%
Diamond$0.004000$0.00400020%

Cost is start price + delivered rows × result price. At the Free tier, the unchanged four-row public sample costs at most $0.025. Set an Apify charge limit appropriate to your chosen maxRows; the theoretical Free-tier event maximum at 1,000 delivered rows is $5.005.

Automation

Start the public Task through the Apify API:

curl -X POST \
"https://api.apify.com/v2/actor-tasks/IQB5feJ4eXgPyjSkl/runs?token=YOUR_APIFY_TOKEN&waitForFinish=600"

Check that the returned Run status is SUCCEEDED, then read its default Dataset URL. Do not automatically retry a timed-out request until you have checked whether the original paid Run was created.

In Make or n8n, use an HTTP Request step for the same Task endpoint, wait for a terminal Run status, and pass the default Dataset items URL to your next step. For recurring production work, duplicate the Task and replace the illustrative inline rows with your authorized Dataset selection or pipeline payload.

Security, privacy, and source rights

  • The Actor runs with limited permissions and requests READ only for Datasets explicitly selected through the Apify resource picker.
  • It does not accept provider API keys, tokens, cookies, or arbitrary Run IDs.
  • It does not call external AI or search services and does not start child Actors.
  • Quarantine and source-error records contain bounded diagnostics rather than raw source payloads, and error text is sanitized before storage.
  • You are responsible for having the right to process the supplied observations, citations, and selected Datasets.

Limitations

  • This is a normalizer and scorer, not a collector. It cannot fetch current AI answers or fill an empty input.
  • Mention matching is deterministic text matching, not entity resolution, sentiment analysis, factual verification, or semantic relevance scoring.
  • Only structured source citations count toward grounding.
  • Up to 1,000 source rows are processed per Run. Answer processing is bounded to 5,000 characters, and output snippets are capped at 800 characters.
  • If the source omits a timestamp, processing time is used and a quality flag is added. Copied Datasets may not expose original Run metadata.
  • The Actor creates one normalized snapshot; trend and historical aggregation belong in your downstream dashboard or warehouse.

Troubleshooting

Why did I get no Dataset rows?

Open QUARANTINE and OUTPUT. Common causes are an unsupported row shape, an incorrect adapter hint, missing usable answer text, or duplicate observations.

Why was Dataset access denied?

Select the Dataset through the input resource picker under the same Apify account or organization that can read it. Pasting a private Run ID is intentionally not supported.

Why does sourceDatasetHints fail validation?

Supply exactly one supported adapter string for each selected Dataset, in the same array order. Do not use auto for selected Datasets.

Why is a URL not listed as a citation?

The source adapter must expose it in a structured citation field. URLs embedded only in answer prose are not inferred as evidence.

How can I verify what was billed?

Compare the default Dataset item count with OUTPUT.counts.delivered and OUTPUT.counts.billed. On-platform they are designed to match; quarantined and failed source rows remain outside the paid Dataset.