Chinese AI Search Citation Analyzer DeepSeek Kimi
Pricing
from $35.00 / 1,000 answers
Chinese AI Search Citation Analyzer DeepSeek Kimi
Analyze citations in Chinese AI search answers with DeepSeek or Kimi. Compare brand mentions, cited domains and changes across serialized scheduled runs.
Pricing
from $35.00 / 1,000 answers
Rating
0.0
(0)
Developer
Tim Zinin
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
12 hours ago
Last modified
Categories
Share
Measure cited domains and literal brand mentions in grounded Chinese AI answers, with explicit coverage and baseline comparisons.
Example provenance. The recorded output examples below describe historical runs from September 6, 2026 on build 0.1.3. They illustrate the response structure and failure boundaries; they are not fresh search results or a promise that another query will produce the same answers.

Start with one query
Paste this input to collect one bounded Chinese-language answer and inspect its citations. The example uses the supported keyless model; availability and provider limits still apply.
{"queries": ["中国企业有哪些开源大语言模型可用于部署?请引用官方资料。"],"models": ["deepseek/deepseek-v4-flash:online"],"brand": "DeepSeek","competitors": ["Qwen"],"maxItems": 1,"comparePrevious": false}
Open the Dataset after the run. Use complete answer rows for citation review, and read free diagnostic or summary rows before interpreting missing coverage. To compare later observations, keep the same query and baseline settings and enable comparePrevious.
What you get
The Chinese AI Search Citation Analyzer runs a bounded query/model comparison through OpenRouter web search. For each complete grounded Answer it extracts provider citation annotations, normalizes cited domains, checks literal brand mentions and compares domains with a prior successful baseline. A free domain_summary reports counts, share and coverage across the successful pairs. Answers without valid citations do not receive a result-found charge.
Grounding here means a completed nonempty answer with at least one valid provider URL annotation. It does not mean the Actor downloaded the source, verified the claim or proved that an annotation supports every sentence. Citation URLs remain inert strings. The product makes the provider's returned evidence easier to inspect and compare; independent source verification remains a buyer review step.
Who uses it
Brand research teams use it to observe which domains appear in a controlled set of Chinese-language AI answers. Content teams use repeatable query sets to identify candidate evidence gaps for editorial review. Agencies can produce a bounded observation table for a client, retaining the query, model and date beside every metric. Product teams can compare model responses without turning free summaries or missing citations into fabricated rankings.
Use it for a defined observation set whose limitations can be explained. It is not a China-wide search market share estimate, a search-engine ranking tracker, a Baidu scraper or a guarantee that content changes caused an answer to change. Repeated model outputs can vary. A small number of query/model pairs should not be presented as a representative population unless a separate sampling design supports that interpretation.
Search adapter and query identity
Version 1 supports OpenRouter web search only. A native Moonshot, DeepSeek or other allowed chat base URL produces unsupported_search before a model request. An OpenRouter model with an optional :online suffix is normalized to the base ID and one Exa web plugin; the Actor does not add a second search operation. Keyless search accepts only the two reviewed flash models. Qwen, Kimi and other compatible IDs require the buyer's OpenRouter key.
There are at most twenty queries and four models, with the product bounded to eighty before expansion. Duplicate queries and normalized model IDs are collapsed. The processing order follows supplied query order and model order. Set maxItems deliberately: the prefill uses one pair. The platform key has a stricter cap of five grounded units and five logical search tasks. A query/model product of twenty-two therefore leaves seventeen pairs unprocessed after five successful Answers.
The web plugin includes an excluded-domain list for the closed platforms named in the access boundary. This is a request to the provider, not an independent audit of provider internals. The Actor itself never fetches a citation target. A URL from a closed platform can be represented as provider-returned evidence without being requested by this code. Review source restrictions before using an external citation in a client deliverable.
Citation and mention interpretation
Only url_citation annotations are considered. URL syntax must be usable HTTP or HTTPS without embedded credentials, fragments or unapproved ports, and known private-address forms are rejected. Invalid annotations are dropped; duplicate identical URLs are collapsed. Titles can be absent. Offsets are retained only if they are valid within the returned answer, otherwise null. Do not invent an exact highlight from a missing or inconsistent offset.
Domains are lowercase hostnames with a leading www. removed. Other subdomains remain distinct. This is not public-suffix or organization-ownership grouping. A domain appearing twice in one answer contributes one successful pair to the domain count. The free summary's share is count divided by successful pairs, while coverage is successful pairs divided by requested pairs. Multiple domains can occur in one answer, so shares are not expected to sum to 100 percent.
mentioned uses literal case-insensitive brand and alias matching after URL text is removed. A brand found only inside a link destination does not count. position is one plus the number of configured competitors whose first appearance precedes the brand; it is null if the brand is absent. competitorsAhead lists those configured names. These measures are textual order, not search rankings, recommendations or sentiment. Broad aliases can match unrelated text and should be reviewed before comparisons.
Domain classification uses a small curated dictionary. Known entries can be official, media, marketplace or qa; otherwise the label is unknown. Each label includes a provenance string. Unknown is a valid result rather than an error, and “official” is not a live proof of current ownership or authority. Keep the raw domain beside the classification so analysts can apply their own reviewed taxonomy.
Baseline behavior
When comparePrevious is true, the Actor opens the named buyer KVS and reads a settings-derived state key before calling the model. The key includes the query/model selection, brand and aliases, competitors, language, endpoint, temperature, output and token limits, working cap and search settings. Credentials are excluded. Changing the selection creates a different comparison context; do not compare two state keys as though the same baseline had been updated.
A malformed stored baseline sets baselineReset:true. A read error stops the run before any LLM call and fails explicitly instead of treating an unavailable record as an empty baseline. Full successful pair delivery and a confirmed free summary are required before writing the next baseline. Partial runs, failed answers, missing citations or uncertain delivery preserve the previous state. Concurrent runs sharing the same key are unsupported because the KVS has no compare-and-swap guarantee here.
newDomains and lostDomains describe changes only for a complete grounded pair. On the first complete observation, domains can appear new because the prior pair set was empty. On an uncited or failed response, both delta fields are null: an outage is not evidence that every prior source disappeared. Complete neighboring rows in a partial run can still carry observation deltas, but baselineCommitted:false tells you the stored comparison point did not advance.
requested and processed count normalized query/model pairs, not individual citations or domains. delivered counts complete Answers. Free domain summaries and notices belong to free. If an answer cites ten pages from two domains it remains one Answer, and each domain contributes at most one count for that pair. A client report should keep both successful and requested denominators visible.
Commercial playbooks
Brand evidence observation report
Define a query list around the client's actual buyer questions and record why each question was selected. Include the exact brand and reviewed aliases; choose a small competitor set. Run a consistent model configuration and inspect each successful answer and its citations. Present cited-domain shares together with coverage and the date. Avoid calling a textual mention position a search rank or treating a few hand-selected queries as market-wide visibility.
Content planning from missing evidence
Review successful grounded answers for domains that recur and buyer questions whose relevant sources are weak or unclear. Send those observations to an editorial review queue. A missing brand mention can suggest a question to investigate, but it does not prove that adding a particular article will change future model behavior. Keep proposed content actions separate from the observation table and reassess them with the same query settings after publication through your own authorized process.
Recurring comparison with baseline discipline
Schedule serialized runs using the same state store and settings. Verify baselineCommitted after each full observation. If a provider outage yields partial work, retain the older baseline and report the coverage gap. The next complete run then compares against the last complete stored observation, not an artificially empty outage snapshot. Changing language, selection or caps creates a new state key and should be labeled as a different series in the report.
Integration detail: pair and domain tables
Use a parent observation table keyed by run ID, query and normalized model. Store citations in a child table retaining the URL and nullable offsets. Store one domain-count row per run/domain from the free summary. For cross-run joins, retain the state key and baseline commit flag. Never multiply the Answer price by the number of citation rows or treat the summary as an additional paid Answer.
中文说明
中国 AI 搜索引用分析:按“查询 × 模型”记录完整回答、引用域名和品牌的字面提及。position 是品牌与已配置竞品在文本中的先后顺序,不是搜索排名。域名份额以成功回答数为分母,并保留总体覆盖率。仅支持 OpenRouter 的联网搜索,平台密钥最多五个搜索任务;Kimi、Qwen 需 BYOK。无引用或失败时不收取结果事件费,域名变化为 null,旧基线保留。本地示例来自模拟服务器,未验证实时来源。

How to run
Begin with the small prefill and inspect both the Dataset and the OUTPUT record. The input form contains examples, but an example is not a credential or a promise that a provider account has access to a model. Keep maxConcurrency at 1. The default runtime allocation is 512 MiB with a 300-second timeout; the Actor stops admitting ordinary work after a 240-second deadline. A smaller workload is the right first check when changing a prompt, field selection, provider or language.
Omit apiKey to use the platform OpenRouter credential within the keyless limits; a buyer key entered in the secret field is authoritative for that run and never silently falls back to the platform account.
Start a run, wait for its final status, and inspect OUTPUT.fatal before downstream processing. Filter Dataset results using found === true, partial === false, resultCount === 1 and error === "". Keep free notices in a separate operational table. A successful process exit alone is insufficient to establish that every requested unit was completed. unprocessed, failed, partial and the Actor-specific counters explain what happened to the rest.
Pricing
The current base price is $0.005 per automatic run start plus $0.05 per complete grounded Answer. The result event is result-found. Bronze receives a 10% result discount ($0.045 per Answer), Silver 20% ($0.04), and Gold, Platinum and Diamond 30% ($0.035). The start price stays $0.005. One complete Answer causes one atomic SDK push with that event. The Actor never manually charges a start event. Check the Store pricing panel and your run receipt for the tariff that applies to your account.
The word “free” on this page refers to the result event. Errors, missing results, partial rows and operational summaries do not carry result-found, but the automatic start can already have been charged. A run with no complete result is therefore not necessarily a zero-dollar run. At the base tariff, the arithmetic for 0, 1 and 2 complete results is $0.005, $0.055 and $0.105. These figures describe Actor event pricing; they do not include an independent invoice from a provider used with BYOK.
BYOK means the buyer supplies the LLM credential and pays that provider under their own account terms, in addition to Actor charges for complete results. Dataset costUsd describes upstream model cost, not the Actor event price. A provider-reported cost is distinguished from a tariff estimate. Unknown model pricing is left null when no usable cost was returned. Never add costUsd to Actor revenue or treat a model usage estimate as proof that an Apify event was charged.
Before expensive work the Actor checks how many results the current buyer budget can cover. Before each paid push it reads the current tariff and remaining money again inside the serialized delivery section. It conservatively checks microdollars rather than copying the SDK's rounded result-count calculation. The final successfully charged row remains a delivered result even if the SDK signals that the spending limit has now been reached. The next unit is not admitted.
On the platform, missing or unreadable pricing, an unreadable balance, a non-PPE tariff, or a paid automatic Dataset-item event causes a failed run. The Actor does not write explanatory Dataset records under unsafe automatic billing; it uses a generic error log and the failed run status. An uncertain write or charge also stops processing. deliveryUncertain requires reconciliation before a retry because a remote write may have succeeded even if its acknowledgement was lost.
The platform credential has separate upstream admission limits: $0.05 per run, at most 5 completed result units, 5 logical LLM jobs and 40 HTTP attempts. These are ceilings; input size, the token reserve, retries, the work deadline or the buyer budget can stop a run earlier. BYOK removes the platform subsidy limits but retains the Actor's bounded input, token and work controls. Raising an input value above a hard cap is rejected instead of silently weakening the protection.
At the pinned DeepSeek Flash output rate of $0.18 per million tokens, the 1200-token default reserves $0.00021600 for output alone (about $0.0002). Claude's observed roughly $0.007 web-plugin component dominates that amount; the existing conservative $0.008/search admission allowance remains in place. The shared costEstimate uses the configured maxOutputTokens for every attempt; reasoning usage is not charged twice.
The cost admission estimate uses the dated tariff evidence recorded on 2026-09-06 for the two allowed flash models. Search additionally reserves an allowance for web work. An admission estimate cannot certify a universal upper bound on an external provider invoice: reasoning, search injection and changed upstream rates can affect actual usage. Reported usage above the reserve stops further work. Provider account limits remain a separate operational control, and local mock costs do not prove production margin.
Input contract
Inputs are JSON objects. Unknown top-level fields, invalid enum values, nonfinite numbers and wrong scalar types are rejected. Do not pass numeric strings where an integer is requested. The inline input is limited to 10 MiB of UTF-8 JSON. Opaque source rows have a 64 KiB limit, depth at most 8, at most 100 elements in each nested array and at most 200 properties in each object. These limits protect memory before prompt or batch expansion.
Where Dataset input is supported, reads use pages of at most 100 rows, with an aggregate limit of 5,000 rows and 10 MiB of retained row JSON. A working Dataset cap produces a partial summary with an omitted count rather than pretending the Dataset ended naturally. A failed Dataset read is an error, not an empty Dataset. The consumer is responsible for stabilizing the supplied Dataset during a run; pagination over a concurrently changing source is not a transactional snapshot.
Field selectors are literal dotted paths, up to 20 paths of 80 characters each. Duplicates are rejected. Property names such as __proto__, constructor and prototype are prohibited. Expressions, JavaScript functions and template evaluation are unsupported. Pass an already prepared field when a transformation needs arithmetic, complex filtering or application-specific access logic. This keeps the boundary between selecting data and executing code explicit.
The complete top-level form follows. prefill values are convenient small examples; defaults apply when a property is omitted. Conditional restrictions described here are also checked at runtime, including tighter platform limits than the maximum shown for BYOK. Secret fields intentionally have no example or prefilled credential.
queries — Chinese search queries
Type: array. Default/sample: ["中国企业有哪些开源大语言模型可用于部署?请引用官方资料。"]. Bounds: maxItems=20.
Up to 20 nonempty queries of at most 500 characters. Duplicate queries are collapsed. Empty queries produce a free not-found control.
models — Models to compare
Type: array. Default/sample: ["deepseek/deepseek-v4-flash:online"]. Bounds: maxItems=4.
Up to four model IDs. Empty uses model. Keyless supports only the two flash IDs. Optional :online suffix becomes one web plugin, never two searches.
brand — Brand to observe
Type: string. Default/sample: "DeepSeek". Bounds: maxLength=100.
Literal brand text sought in the answer, excluding URL text. This is a mention check, not a search-engine rank or a sentiment score.
brandAliases — Brand aliases
Type: array. Default/sample: []. Bounds: maxItems=10.
At most 10 literal aliases. Matching is case-insensitive. Include known spellings and Chinese forms; broad aliases may match unrelated text.
competitors — Competitors and optional aliases
Type: array. Default/sample: ["Qwen"]. Bounds: maxItems=10.
At most 10 names or objects with name and aliases (<=10). competitorsAhead lists competitors whose first mention appears before the brand.
language — Answer language
Type: string. Default/sample: "zh". Bounds: maxLength=20.
Requested answer language passed to the model. Default zh. This setting changes the baseline key and can change citations and mentions.
stateStoreName — Baseline store name
Type: string. Default/sample: "chinese-ai-search-citation-baseline". Bounds: maxLength=63.
Named key-value store in your account. Use serialized schedules; overlapping runs sharing settings are unsupported. Store uses a hash of selection settings.
comparePrevious — Compare with previous full run
Type: boolean. Default/sample: true.
Read baseline before upstream calls and commit only after complete successful delivery of every requested pair. Read errors fail; malformed values are explicitly reset.
maxItems — Maximum processed query/model pairs
Type: integer. Default/sample: 1. Bounds: minimum=1, maximum=80.
Product is bounded to 80 before expansion. Platform key is limited to five grounded Answers and five logical search tasks per run; BYOK allows up to 80.
apiKey — Buyer API key (optional)
Type: string. Secret; no example credential. Bounds: maxLength=512.
Optional buyer credential for the selected endpoint. Never copied to results. Empty uses the platform key only with OpenRouter. Buyer credentials are never replaced after a provider error.
baseUrl — OpenAI-compatible base URL
Type: string. Default/sample: "https://openrouter.ai/api/v1". Bounds: maxLength=300.
HTTPS base path, without /chat/completions. Only documented host/path pairs are accepted. Custom native endpoints require apiKey. Search supports OpenRouter only. No query, credentials, fragments or redirects.
model — Model ID
Type: string. Default/sample: "deepseek/deepseek-v4-flash". Bounds: maxLength=100.
Keyless: deepseek/deepseek-v4-flash or z-ai/glm-5.3-flash only. Qwen, Kimi and other model IDs require BYOK. The provider must support this model and operation.
reasoningMode — Thinking mode (OpenRouter)
Type: string. Default: off. Allowed values: off, low, provider (select editor).
Controls OpenRouter reasoning: off disables thinking where supported, low reduces effort, provider keeps provider defaults. Off uses low effort for z-ai/ models. Direct endpoints retain their own settings. See Thinking models and output tokens below for compatibility repeats and observed limits.
maxOutputTokens — Maximum completion tokens
Type: integer. Default/sample: 1200. Bounds: minimum=1, maximum=4096.
Upper completion allowance per call, including reasoning tokens. Platform hard maximum 2048; BYOK hard maximum 4096. Lower this for small tasks. Length-truncated answers are free incomplete results.
maxTotalTokens — Run token-unit limit
Type: integer. Default/sample: 40000. Bounds: minimum=1, maximum=2000000.
UTF-8 message bytes plus completion reserve before each attempt, reconciled with valid usage afterward. Platform hard maximum 40000; BYOK 2000000. Unknown usage consumes full reserve.
temperature — Temperature
Type: number. Default/sample: 0.2. Bounds: minimum=0, maximum=2.
Sampling temperature passed to the provider. Repeated runs can differ even at a low value. Does not establish factual correctness.
maxConcurrency — Concurrent LLM calls
Type: integer. Default/sample: 1. Bounds: minimum=1, maximum=1.
Version 1 processes one call at a time. Billing and delivery are serialized. Schedule citation comparisons without overlapping runs.
Thinking models and output tokens
maxOutputTokens covers both visible output and provider reasoning tokens. OpenRouter receives reasoningMode: "off" by default as reasoning: {"enabled": false}. Models with the z-ai/ prefix reject that setting, so off uses {"effort": "low"} for them. low always requests {"effort": "low"}; provider omits the parameter and keeps provider defaults. On OpenRouter only, an HTTP 400 in off mode allows one compatibility repeat with effort: low, then one without reasoning if the repeat also returns 400. A third 400 is returned as an error. These repeats count in httpAttempts and reserve tokens and COGS under the existing deadline and run caps; they do not consume the separate allowance of two network retries. At most five requests can result from one task when both allowances are used.
Live checks on 2026-09-06 (three prompts per setting, max_tokens: 512): DeepSeek v4 Flash with provider defaults hit length with null content on 1/3 prompts and spent up to 3 435 reasoning tokens on another; with off it answered 3/3 in 59–78 output tokens and 0 reasoning tokens. GLM 5.3 Flash hit length 3/3 with defaults, rejected enabled: false with HTTP 400 and answered 3/3 with effort: low. These are acceptance observations, not a guarantee for every prompt.
Direct provider endpoints receive no reasoning parameter, regardless of this input, and retain their own thinking settings; maxOutputTokens must cover those settings. A null content with finish_reason: length remains invalid_envelope; truncated string content remains incomplete_completion. Both are free error rows. reasoningTokens reports the provider's nonnegative integer usage.completion_tokens_details.reasoning_tokens, or null when absent or invalid. It is already part of tokensOut (the translator's tokens.output) and is never added again for cost or token admission.
Endpoint and model selection
| Provider host | Accepted base path | Runtime requirement |
|---|---|---|
| openrouter.ai | /api/v1 | Platform key or buyer OpenRouter key |
| api.deepseek.com | empty path or /v1 | Buyer key; provider-native model ID |
| api.moonshot.cn | /v1 | Buyer key; provider-native model ID |
| open.bigmodel.cn | /api/paas/v4 | Buyer key and compatible chat model |
| dashscope.aliyuncs.com | /compatible-mode/v1 | Buyer key and compatible chat model |
| dashscope-intl.aliyuncs.com | /compatible-mode/v1 | Buyer key for the selected region |
| ark.cn-beijing.volces.com | /api/v3 | Buyer key and supported deployment ID |
| api.minimax.io | /v1 | Buyer key and compatible model |
| api.minimaxi.com | /v1 | Buyer key and compatible model |
This table describes the shared transport allowlist, not live certification of every model or provider feature. The citation analyzer supports only the OpenRouter row for grounded search. Other actors append /chat/completions to an accepted base path. Native IDs can differ from OpenRouter slugs; copying a marketplace slug into a native endpoint does not translate it automatically. JSON schema and other model capabilities remain conditional on provider support.
Keyless model IDs are exactly deepseek/deepseek-v4-flash and z-ai/glm-5.3-flash. Qwen, Kimi and other compatible model IDs require BYOK. :free and arbitrary suffixes are not accepted. Only the citation analyzer accepts an optional :online suffix, normalizing it to the base model plus a single web plugin. The runner and translator do not add web search.
HTTPS and approved host/path combinations are required. Userinfo, query strings in the base URL, fragments, unapproved ports, path traversal and redirects are rejected. All resolved IP addresses must pass the public-address check, and the connection uses the verified address set. A public hostname that resolves to even one private address is refused. A redirect never carries the Authorization header to a new destination.
Field dictionary
Every Dataset record has the base fields below. Optional operation-specific fields may be absent from a validation notice or a not-found row when no model call occurred. Do not infer zero tokens, zero cost, an empty answer or high confidence from a missing field. Normalize absent analytical values to null in your warehouse while retaining the original JSON for audit.
| Field | Meaning |
|---|---|
query | One supplied query in the normalized query/model product. Duplicate query strings are collapsed within a run. |
model | Normalized provider model identifier used for this request. Keep it with the prompt or field configuration when comparing results. |
provider | Selected API host. It identifies the transport provider rather than an independent verification service. |
answer | Completed text returned by the model, after credential redaction. It can still contain unsupported factual statements and requires contextual review. |
citations | Provider URL annotations normalized into url, title, startIndex and endIndex. Invalid URLs are dropped and identical URLs deduplicated. |
citedDomains | Distinct lowercase hostnames with a leading www removed. Subdomains otherwise remain distinct; this is not registrable-domain aggregation. |
mentioned | Literal case-insensitive occurrence of the brand or one alias in answer text after URL text is removed. |
position | One plus the number of configured competitors appearing before the brand. Null when the brand is absent; never a search-engine ranking. |
competitorsAhead | Configured competitor names whose first textual mention precedes the brand. Null when the brand is absent. |
domainClasses | Small curated domain dictionary with official/media/marketplace/qa/unknown labels and a provenance string for each decision. |
newDomains | Domains in a complete grounded answer but absent from its prior pair baseline. Null if the current response is not grounded. |
lostDomains | Prior pair domains absent from a complete grounded answer. Null for failed or uncited responses, so an outage is not reported as domain loss. |
tokensIn | Provider prompt-token count when it is a valid nonnegative integer; otherwise null. Distinct from the admission guard tokenUnits. |
tokensOut | Provider completion-token count when valid; otherwise null. Do not replace missing usage with a fabricated measured value. |
reasoningTokens | Nonnegative integer provider reasoning usage, or null when unavailable/invalid. Already included in tokensOut; do not add it again. |
costUsd | Upstream cost in USD; provider value, dated estimate or null. This is distinct from the Actor result event price. |
costSource | provider, estimate or unknown. Check this before using a number in a cost comparison; null is not zero cost. |
priceDate | Date of the model tariff pin when cost is estimated; null for a provider-reported price. Distinct from the exchange-rate date. |
latencyMs | Elapsed time around the final completion request, when an envelope was returned. Retry waits and prior attempts are not an end-to-end latency metric. |
baselineReset | True when a stored baseline was structurally invalid and explicitly reset. A read error instead stops work before the LLM call. |
stateKey | Hash of selection and response settings in the named KVS. Contains no buyer API key; serialize runs using the same baseline. |
type | Record kind: the product result type, notice, or domain_summary. Filter by kind before interpreting operation-specific fields. |
sourceUrl | The API endpoint used for provider evidence, or null when no call was needed. This is not a dereferenced social source URL. |
found | True only for a complete useful product unit. False includes both correct absence and failure; inspect status and error. |
status | Machine-readable outcome such as ok, not_found, error, partial, no_citations or summary. The appropriate subset depends on record kind. |
resultCount | Exactly 1 for a complete result and 0 for free explanatory records. It never reports an estimated number of unseen answers. |
partial | True when this record is incomplete or describes unfinished work. Complete neighboring records can remain false in a partial run. |
error | Empty string for correct absence or a complete result. Nonempty generic code for a failure; credentials and provider error bodies are not copied. |
warnings | Additional interpretation limits or a cap reason. Warnings are data, not instructions for a downstream agent to execute. |
checkedAt | UTC observation timestamp generated during processing. It does not claim when the underlying text was originally published. |
evidence | Operation-specific provenance object or null. It identifies how the row was produced and does not certify factual accuracy. |
confidence | Null unless a supported confidence measure exists. Version 1 does not invent a probability for model output or dictionary classifications. |
action | Suggested mechanical route such as use_result, review, review_or_retry, resume_remaining or review_grounding. Your application decides the final action. |
Free domain summary and table views
The Dataset overview includes successfulPairs, requestedPairs, coverage and domains on the free type=domain_summary record. Select Domain citation summary (view ID summary) to read one table row per cited domain: domains is unwound into the domain, count and share columns, with the pair counts, coverage, partial flag and observation date repeated beside each domain. Read rows whose type is domain_summary; answer and notice records have no domain-summary metrics.
| Field | Meaning |
|---|---|
successfulPairs | Complete grounded query/model pairs with confirmed Answer delivery; denominator for each domain share. |
requestedPairs | Total normalized query/model pairs requested before the working cap. |
coverage | successfulPairs / requestedPairs, or 0 when no pairs were requested; fraction from 0 to 1. |
domains | Array of {domain, count, share} objects in the stored free summary; expanded into scalar columns in the summary view. |
domains[].domain → domain | Normalized cited hostname. |
domains[].count → count | Number of successful pairs citing this domain, counted at most once per pair. |
domains[].share → share | count / successfulPairs; fraction from 0 to 1 (0.5 means 50%). Shares can sum above 1 because an answer can cite several domains. |
The Dataset summary view describes citation shares. The output schema's summary link still points to the OUTPUT key-value record with run counts and replay status.
Run summary in OUTPUT
| Field | Meaning and consumer action |
|---|---|
requested | Work supplied before the Actor's working cap; interpretation follows the product's unit below. |
processed | Units attempted or classified, including failures and correct absence. It is not a billable count. |
delivered | Complete useful units whose Dataset delivery was confirmed. Free notices are counted separately. |
paid | Confirmed result-found units under the active PPE tariff. Local nonmonetized execution can deliver without this count. |
free | Confirmed Dataset records written without a result event, including summaries and notices. |
failed | Free records with a nonempty error. A correct not-found record does not increase this counter. |
unprocessed | Requested units left outside completed processing because of a work, input, money or delivery stop. |
partial | At least part of the requested work is incomplete, erroneous or uncertain. Successful neighbors remain useful. |
fatal | Empty string normally; nonempty means the run must be treated as failed even if some useful rows exist. |
deliveryUncertain | Count of delivery attempts whose outcome could not be confirmed. Stop automatic replay. |
budgetExhausted | Further paid work was refused by the buyer spending gate. It does not invalidate the final paid unit. |
reason / stopReason | Domain-level stop and client-level stop, respectively; preserve both for troubleshooting. |
logicalTasks / httpAttempts | Admitted LLM tasks and actual request attempts, which can differ because of retries. |
tokenUnits | Conservative attempt reserves reconciled with valid returned usage. This is a guard counter, not a universal tokenizer. |
costBoundUsd | Accumulated admission estimates, increased when a provider reports a larger known cost. |
reportedUpstreamCostUsd | Sum of provider cost values actually returned, including unsuccessful completion envelopes. |
unknownCostAttempts | Attempts without an authoritative cost value. It prevents a null invoice from looking free. |
demoInput | Whether the Actor used designated synthetic source data; see the product-specific convention. |
checkedAt / schemaVersion | Observation timestamp and output contract version. Retain both in exports. |
Evidence and boundaries
The historical platform examples in this README refer to build 0.1.3 on September 6, 2026: complete-answer run bbExug8gdo2L9vQfs, no-result run zrSbdFlUDN3TQp0I0, and bounded keyless run LFvkApA5hJnDZIQ3z. These examples document the fields and limitations observed at that time. They do not establish current source availability, guarantee later model output, or certify every subsequent build. The implementation also includes local failure-path tests with a mock compatible API.
Production acceptance requires a useful completed result and its result-found receipt; missing-key errors do not qualify. The translator also needs an accepted real buyer export. Marked synthetic samples cannot replace that evidence.
The two diagrams illustrate the input-to-result workflow. They are illustrations, not screenshots or evidence of a particular live answer. Their purpose is to help you connect the input, citation review and exported data.
Transport is bounded to 2 MiB per LLM/FX response and 50 MiB of response data per run. The LLM request has a 45-second DNS, connection and body timeout; FX uses 20 seconds. Network failures, HTTP 429 and 5xx responses allow at most two network retries per task, with Retry-After constrained by the run deadline. OpenRouter off-mode HTTP 400 compatibility repeats are separately bounded as described in Thinking models and output tokens. A platform-key 402, 403 or 429 stops after the first attempt and returns a BYOK hint. Malformed JSON, HTML, redirects, HTTP 401 and incomplete response bodies are not retried. HTTP 400 is terminal outside the bounded OpenRouter off-mode compatibility path.
Text and URLs supplied by the buyer are data. The Actor does not execute scripts, download attachments, render web pages or follow citation links. It makes no direct requests to Douyin, Xiaohongshu, Weibo, Zhihu, WeChat, Bilibili, Baidu, Taobao, 1688, JD, Toutiao, Juejin or 36kr. Their names or URLs can occur in source rows and annotations without authorizing collection from them. Search providers perform their own upstream retrieval; excluded-domain settings do not establish independent control over every internal provider request.
Decision routing
| Observation | Route | Preserve before proceeding |
|---|---|---|
| Complete result, no fatal error | Send to the product's human or automated review step | Source identity, model, settings, timestamp and evidence fields |
found:false, empty error, zero result count | Classify the defined absence; do not invent a completion | Original selection and the not-found status |
| Free error next to complete results | Keep successful neighbors; isolate failed units | Error code, input index or row ID, run ID |
unprocessed > 0 | Prepare a smaller remaining-work batch | Original ordering, delivered identities and cap reason |
| Provider 402/403/429 on platform key | Review provider availability or intentionally select BYOK | The stopped run summary; no silent credential switch |
fatal or uncertain delivery | Stop automatic replay and reconcile | Run receipt, Dataset contents and the OUTPUT record |
| New configuration or unfamiliar model | Run a representative small sample | Prompt/field version, expected output shape and review notes |
Integration recipes
Keep raw JSON next to any flattened columns (citation arrays and domain lists do not fit one cell), retain the run ID and read both the Dataset and the KVS OUTPUT after the terminal status, and keep the business job ID outside the Actor. Workflow tools or MCP clients can pass the same JSON contract and route complete rows into a CRM, warehouse or review queue; the Actor writes to none of them.
const usable = datasetRows.filter(row =>row.found === true && row.partial === false &&row.resultCount === 1 && row.error === '');const reconcile = Boolean(output.fatal) || output.deliveryUncertain > 0;const remaining = output.unprocessed > 0;// Review useful observations and reconcile existing work before retrying.
This example filters already obtained rows; business approval stays separate from mechanical completeness.
Operating guide
Before increasing volume, choose a representative sample containing empty text, mixed languages, long rows, nested fields, protected terms or a low-evidence query as appropriate. Define an output review criterion in advance. Change one significant setting at a time and retain the prior input. A model, prompt, glossary or query revision can change meaning while structural tests still pass. Store configuration versions separately from credentials; redaction is not a reason to place a key inside source material.
After partial work, distinguish a deterministic correction from a transient failure. Invalid input requires editing the input. A schema incompatibility requires fixing the schema or choosing a supporting model. A budget or token stop requires a smaller remaining workload or an intentionally authorized limit change. Authentication errors require checking the chosen account. Repeating the same invalid configuration adds no useful evidence and may incur another automatic start.
If delivery is uncertain, inspect the existing Dataset and receipt before retrying. Keep successful neighboring results and reconstruct the remaining set using source IDs or the query/model pair, not free-notice row numbers. The Actor reports replaySafe:false because it cannot provide exactly-once processing across separate runs, process crashes or ambiguous acknowledgements. Serialize recurring schedules that share citation baseline settings.
FAQ
Does complete mean accurate? Completion means the provider finished in the required form and mechanical checks passed. Schema validation constrains structure, while factual and linguistic review remain separate decisions.
Can a run finish without a paid result? Yes. Correct absence, validation failure, a missing key, a provider error or an admission limit can yield no result-found event. Free records can still make the Dataset nonempty, and the automatic start may still apply.
Can I use any model or endpoint? Keyless runs accept only the two reviewed flash IDs. BYOK permits compatible model IDs on the listed host/path pairs. A valid identifier is not proof of account access or feature support. Arbitrary proxies and private endpoints are not accepted.
Can I safely retry everything? Not automatically. Inspect the existing run and its receipt first. The Actor has no cross-run delivery or billing ledger and a repeated complete result can incur another start, upstream request and result event.
Are the links and demo evidence verified live? No. URLs are retained as strings. Local fixtures prove software behavior, not source truth, language accuracy, commercial margin or a public Store quality score. Production acceptance remains a separate release step.
Sources and rights
Provider interface references are OpenRouter web search documentation and usage accounting documentation. They are interface references, not a claim that every allowed provider was called in this build. SDK behavior was checked against the installed, exactly pinned Apify 3.7.2 package. Applicable FX rows retain their own route and observation date.
The Actor grants no ownership of source texts, trademarks, linked publications or provider outputs. Obtain data through sources you are authorized to use. Access to an export does not automatically establish a right to republish it. Keep provenance and any source restrictions with shared reports, particularly when results cross customer or organization boundaries.
Recorded happy, not-found and partial output (accepted platform runs)
The three examples below are exact records from the accepted acceptance runs listed in Evidence and boundaries; strings longer than the page limit are cut with an explicit truncation note, nothing else is edited. Only checkedAt and provider latency differ between repeated runs.
happy — paid result-found row
Run bbExug8gdo2L9vQfs on build 0.1.3, 2026-09-06, 16 s, Dataset records: 2, charged events: {"apify-actor-start": 1, "result-found": 1}.
Input:
{"queries": ["中国企业有哪些开源大语言模型?请引用官方资料。"],"maxItems": 1,"comparePrevious": false,"brand": "DeepSeek","language": "zh","stateStoreName": "chinese-ai-search-citation-baseline","baseUrl": "https://openrouter.ai/api/v1","model": "deepseek/deepseek-v4-flash","reasoningMode": "off","maxOutputTokens": 1200,"maxTotalTokens": 40000,"temperature": 0.2,"maxConcurrency": 1}
First Dataset record (exact):
{"schemaVersion": "1.0","type": "answer","sourceUrl": "https://openrouter.ai/api/v1/chat/completions","found": true,"status": "ok","resultCount": 1,"partial": false,"error": "","warnings": ["Citation URLs are not fetched or independently verified."],"checkedAt": "2026-09-06T00:52:11.296Z","evidence": {"kind": "provider_url_citation","count": 5},"confidence": null,"action": "use_result","reasoningTokens": 0,"query": "中国企业有哪些开源大语言模型?请引用官方资料。","model": "deepseek/deepseek-v4-flash","provider": "openrouter.ai","answer": "根据公开资料,多家中国企业已发布并开源了各自的大语言模型。以下是部分代表性模型及其官方信息:\n\n### 1. 腾讯混元 Hy4 preview\n- **发布方**:腾讯\n- **模型特点**:总参数 770B,激活参数 49B,上下文长度突破 1M,面向编程、办公与科研等真实生产力场景。\n- **开… [truncated here for length; the record holds 1504 characters]","citations": [{"url": "https://www.tencent.com/zh-hk/tencent-releases-and-open-sources-tencent-hy4-preview/","title": "騰訊發布並開源Hy4 preview - Tencent","startIndex": 0,"endIndex": 0},{"url": "https://tech.meituan.com/2026/07/12/LongCat-2.0-Open-source.html","title": "正式开源!美团 LongCat-2.0 同步开放国产卡推理代码 | 美团 · 技术团队","startIndex": 0,"endIndex": 0},{"url": "https://qwen.readthedocs.io/zh-cn/stable/","title": "欢迎来到Qwen ¶","startIndex": 0,"endIndex": 0},{"url": "https://github.com/zai-org/GLM-4/blob/main/README_zh.md","title": "README_zh.md at main · zai-org/GLM-4","startIndex": 0,"endIndex": 0},{"url": "https://gitee.com/Tele-AI/TeleChat2","title": "TeleChat2: 星辰语义大模型TeleChat2是由中国电信人工智能研究院研发训练的大语言模型,是首个完全国产算力训练并开源的千亿参数模型","startIndex": 0,"endIndex": 0}],"citedDomains": ["gitee.com","github.com","qwen.readthedocs.io","tech.meituan.com","tencent.com"],"mentioned": true,"position": 2,"competitorsAhead": ["Qwen"],"domainClasses": [{"domain": "gitee.com","class": "unknown","provenance": "no_dictionary_match"},{"domain": "github.com","class": "unknown","provenance": "no_dictionary_match"},{"domain": "qwen.readthedocs.io","class": "unknown","provenance": "no_dictionary_match"},{"domain": "tech.meituan.com","class": "unknown","provenance": "no_dictionary_match"},{"domain": "tencent.com","class": "unknown","provenance": "no_dictionary_match"}],"newDomains": ["gitee.com","github.com","qwen.readthedocs.io","tech.meituan.com","tencent.com"],"lostDomains": [],"tokensIn": 4838,"tokensOut": 771,"costUsd": 0.0075278812,"costSource": "provider","priceDate": null,"latencyMs": 12792,"baselineReset": false,"stateKey": "cn-f4f2ad55e62a27eb064b61ac96d7a35718795e2c4690ae6a65d42742959c9669"}
The run's OUTPUT record carries the counters described in Run summary in OUTPUT (paid: 1, free: 1, partial: false).
not-found — free row, no result charge
Run zrSbdFlUDN3TQp0I0 on build 0.1.3, 2026-09-06, 2 s, Dataset records: 1, charged events: {"apify-actor-start": 1, "result-found": 0}.
Input: the same prefill with a deliberately non-existent brand and maxItems: 1 (see acceptance.json, golden «not-found»).
Dataset record (exact):
{"schemaVersion": "1.0","type": "answer","sourceUrl": null,"found": false,"status": "not_found","resultCount": 0,"partial": false,"error": "","warnings": [],"checkedAt": "2026-09-06T00:52:14.274Z","evidence": null,"confidence": null,"action": "review","reasoningTokens": null}
partial — keyless cap reached, the extra row is never charged
Run LFvkApA5hJnDZIQ3z on build 0.1.3, 2026-09-06, 55 s, Dataset records: 7, charged events: {"apify-actor-start": 1, "result-found": 5}.
Input: the runner prefill with maxItems: 21 (see acceptance.json, golden «keyless cap»).
Free notice record (exact):
{"schemaVersion": "1.0","type": "notice","sourceUrl": null,"found": false,"status": "partial","resultCount": 0,"partial": true,"error": "","warnings": ["working_or_buyer_cap"],"checkedAt": "2026-09-06T00:53:13.320Z","evidence": null,"confidence": null,"action": "review","reasoningTokens": null,"unprocessed": 17}
Related tools
Related tools for adjacent workflows in AI and search visibility.
| Actor | What it does |
|---|---|
| AI Answer & Citation Change Monitor | Pair it in the AI and search visibility workflow: Monitor grounded AI answers by query, model, and language; detect rewrites and cited-domain additions or... |
| AI Crawler Access Checker | Pair it in the AI and search visibility workflow: Audit up to 100 sites for 16 AI crawler policies |
| AI Overview Citation Tracker | Pair it in the AI and search visibility workflow: Track which public URLs and domains selected grounded AI models cite for buyer-supplied queries |
| LLM Brand Visibility Tracker | Pair it in the AI and search visibility workflow: For each query that matters, check whether AI assistants recommend YOUR brand — and which competitors they... |
| llms.txt Auditor & AI Crawler Policy Checker | Pair it in the AI and search visibility workflow: Audit public llms.txt, llms-full.txt, and root robots.txt rules for nine named AI crawlers |