# Schema.org for LLMs Audit (`zinin/schema-org-for-llms-audit`) Actor

Audit public pages for AI-readable Organization, Product, and Article facts. Get a deterministic GEO readiness score, missing high-impact fields, cross-surface contradictions, evidence paths, and prioritized fixes without an LLM or browser.

- **URL**: https://apify.com/zinin/schema-org-for-llms-audit.md
- **Developed by:** [Tim Zinin](https://apify.com/zinin) (community)
- **Categories:** SEO tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.64 / 1,000 llm fact-readiness url audits

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Schema.org for LLMs Audit — GEO entity and product fact checker

**Find the machine-readable facts that AI answer engines can verify about your brand, products, and
articles — before a competitor becomes the cleaner answer.** Submit 1–50 public URLs and receive one
evidence-backed readiness audit per successfully fetched page: grade, score, entity inventory,
high-impact gaps, contradictions, and prioritized fixes. No LLM key, browser, proxy, or search account
is required.

![Schema.org for LLMs Audit turns public pages into a prioritized GEO brief](https://raw.githubusercontent.com/TimmyZinin/apify-actor-assets/0736d0a63770bfda34d13ada9ca15b308d98baef/marketing14/schema-org-for-llms-audit/readme-hero.webp)

The Actor reads only public static HTML that the target host permits through robots.txt. It compares
JSON-LD, Schema.org Microdata and RDFa, OpenGraph, canonical metadata, and the brand name you expect to
see. A successful audit is delivered to the default Dataset and charged as one <code>result-found</code>
event. A denied, unsafe, timed-out, oversized, or over-budget URL is written as a bounded free diagnostic
to the run's key-value store instead of becoming a paid pseudo-result.

Try the published example, **Audit Apify's homepage for AI-readable brand facts**, from the
[Examples tab](https://apify.com/zinin/schema-org-for-llms-audit/examples/audit-apify-brand-facts).
It uses a real public Organization page, requires no secret input, and demonstrates a Grade A result.

### What you get

Each successful URL produces a decision record rather than a loose scrape. The record answers five
practical questions:

1. **What machine-readable entity is present?** The Actor recognizes Organization and LocalBusiness
   subtypes, Product variants, Article variants, JSON-LD graphs, Microdata scopes, and RDFa scopes.
2. **Can an answer engine identify the brand?** It checks stable entity IDs, names, URLs, logo URLs,
   authoritative <code>sameAs</code> domains, publisher references, and an optional buyer-supplied brand.
3. **Can an answer engine understand the offer or content?** Product records expose brand, SKU/GTIN,
   offer price, currency, and availability; Article records expose headline, publication dates,
   publisher, author presence, and image.
4. **Do the surfaces agree?** Canonical URL, OpenGraph URL/title/site name/product price, and structured
   entities are reconciled. Conflicts are named, severity-ranked, and tied to both evidence surfaces.
5. **What should the marketer or developer fix next?** Missing high-impact facts and contradictions are
   translated into bounded actions such as adding a stable Organization <code>@id</code>, completing
   Product Offer fields, or aligning OpenGraph and canonical metadata.

The Dataset includes the original normalized URL, final URL after guarded redirects, HTTP status,
document title and language, detected document types, normalized target entities, OpenGraph summary,
fact inventory, 0–100 readiness score, A–F grade, missing-fact codes, contradictions, fixes, quality
flags, evidence paths, expected-brand match, static-HTML disclosure, and one run-wide observation time.

This is deliberately a **deterministic audit**, not a synthetic AI opinion. The same HTML and settings
produce the same findings. The Actor does not ask a model whether a page “looks good,” invent a citation
probability, pretend that structured data guarantees ranking, or claim access to proprietary answer
engine indexes. It measures whether important public facts are present, attributable, internally
consistent, and straightforward for automated systems to parse.

The run also provides three operational records:

- <code>OUTPUT</code> — requested, normalized, attempted, audited, delivered, billed, quarantined, and
  budget-withheld counts plus pricing seen by the running build.
- <code>COVERAGE</code> — grade distribution, document-type distribution, and diagnostic reason counts.
- <code>QUARANTINE</code> — short free diagnostics for URLs that did not become Dataset rows. Response
  bodies, private addresses, credentials, and Person entity details are never copied there.

Useful platform behavior comes with the product: save the input as a Task, schedule recurring audits,
trigger runs through the API, export the Dataset as JSON/CSV/Excel, connect webhooks, or expose the Actor
to an AI client through Apify's hosted MCP server. The output schema makes the most useful fields visible
in the Store table instead of forcing a buyer to inspect raw logs.

### Who uses it

**GEO and AEO marketers** use the grade and gap list to turn a vague “improve AI visibility” request into
a page-level backlog. A homepage may need a stable Organization identity and authoritative <code>sameAs</code> links; a product page may need complete Offer facts; a thought-leadership article may
need publisher and date provenance.

**Brand and communications teams** use <code>expectedBrandName</code> and <code>brandAliases</code> to
detect pages whose structured entity name or <code>og:site\_name</code> no longer matches the accepted
brand. This is especially useful after a rename, merger, localization project, or multi-brand site
migration. Matching is token-aware and does not treat a short brand such as “AI” as an arbitrary
substring.

**Technical SEO teams** use the normalized fact inventory and evidence paths to reproduce findings
without hunting through a large HTML document. The result distinguishes valid and invalid JSON-LD
blocks, JSON-LD versus Microdata/RDFa entities, canonical presence, OpenGraph completeness, and the
number of consistency checks actually possible on that page.

**E-commerce marketers** use Product findings to identify pages where a crawler sees a product name but
cannot reliably connect it to a brand, offer price, currency, availability, SKU, or GTIN. Prices are
normalized before comparison, so formatting such as <code>1,299.00</code> and <code>1299</code> does not
create a false mismatch.

**Content operations teams** use Article findings to check headline, publication or modification date,
publisher, author presence, and image. The Actor reports that a Person reference exists but deliberately
does not output a person's name, profile, or other details.

**Agencies** use batch inputs to audit the same small, high-value page set for every client and export a
standardized evidence file. The output is suited to a client brief because every score is accompanied by
the facts, gaps, contradictions, and source paths that produced it.

**Product teams and automation builders** use the stable codes for routing. A workflow can open a ticket
only for high-severity contradictions, notify a marketer when the expected brand does not match, or
compare score distributions over scheduled runs without parsing prose.

This Actor is not a general crawler, a JavaScript renderer, a search-rank tracker, an AI citation
monitor, a legal compliance scanner, or a substitute for Schema.org validation. It complements those
tools by answering a narrower commercial question: **does this public page expose a coherent, useful set
of brand, product, or article facts to automated readers right now?**

![Fail-closed workflow from URL normalization to atomic delivery and billing](https://raw.githubusercontent.com/TimmyZinin/apify-actor-assets/0736d0a63770bfda34d13ada9ca15b308d98baef/marketing14/schema-org-for-llms-audit/readme-workflow.webp)

### How to run

#### Fastest Store run

1. Open the Actor and choose **Try for free**.
2. Add one to fifty public HTTP(S) URLs. Start with pages that represent the brand: homepage, flagship
   product pages, evergreen guides, and important editorial pages.
3. Optionally enter the expected public brand name and accepted aliases.
4. Keep maximum concurrency at 4 unless the target site explicitly supports a higher request rate.
5. Set a maximum run charge that covers the start event and the number of results you want.
6. Run the Actor, then open **Output → Dataset** for audits and **Output → Key-value store** for the run
   receipt, coverage summary, or free diagnostics.

Minimal input:

```json
{
  "urls": [
    "https://apify.com/"
  ],
  "expectedBrandName": "Apify",
  "maxConcurrency": 2
}
```

Typical three-page brand audit:

```json
{
  "urls": [
    "https://example.org/",
    "https://example.org/products/flagship",
    "https://example.org/blog/research-report"
  ],
  "expectedBrandName": "Example",
  "brandAliases": [
    "Example Inc.",
    "Example Labs"
  ],
  "maxConcurrency": 3
}
```

The Actor removes URL fragments and exact normalized duplicates before fetching. A URL without a scheme
is normalized to HTTPS. Credentials embedded in a URL are rejected. Only ordinary public HTTP(S)
destinations are eligible; private, loopback, link-local, metadata, documentation, benchmark, transition,
and other special-use address ranges fail closed.

#### Choose pages intentionally

A 50-page random crawl is rarely the best first audit. Start with one representative page for each
decision type:

| Business question | Recommended page | Primary facts |
| --- | --- | --- |
| Can an AI system identify us? | Homepage or About page | Organization name, URL, logo, stable ID, sameAs |
| Can it describe the offer? | Product detail page | Product, brand, SKU/GTIN, Offer price/currency/availability |
| Can it attribute our content? | Article or report | headline, dates, publisher, author presence, image |
| Did a migration create conflicts? | Old and new canonical pages | canonical, OpenGraph URL, entity URL, brand name |
| Are localized pages coherent? | One page per locale | page language, identity, URL and metadata alignment |

For monitoring, save the input as an Apify Task and schedule it at a cadence appropriate to the site.
Daily runs make sense during a migration; weekly or monthly runs are usually enough for stable pages.
The Actor does not maintain history itself. Store Dataset exports externally or compare scheduled run
outputs by <code>url</code>, <code>readinessScore</code>, and stable finding codes.

#### Read the run in the right order

First inspect <code>OUTPUT.counts</code>. For an accepted run, <code>delivered</code> equals <code>billed</code> on platform and equals the default Dataset item count.
Then inspect <code>COVERAGE</code> to see whether the batch is mostly strong, weak, or diagnostic. Finally
use each row's <code>missingHighImpactFacts</code>, <code>contradictions</code>, and <code>evidencePaths</code> to assign work.

An empty Dataset is not automatically a failed run. It may mean every target was denied by robots.txt,
unsafe, unavailable, oversized, non-HTML, or withheld by the buyer's charge limit. The reason will be in <code>QUARANTINE</code> and the aggregate count will be in <code>OUTPUT</code>.

### Pricing

The Actor uses Apify **PAY\_PER\_EVENT** pricing. There are only two billable nouns:

- <code>apify-actor-start</code> — the platform start event for the run.
- <code>result-found</code> — one successfully fetched, audited, and atomically delivered URL result.

The default Dataset event is intentionally not priced. Errors and diagnostic records are not <code>result-found</code> events. The running Actor reads its platform pricing and buyer charge budget
before starting network work; if pricing is unreadable or misconfigured, it fails closed.

| Tier | Start event | Per delivered <code>result-found</code> | Discount from FREE |
| --- | ---: | ---: | ---: |
| FREE | $0.0050000 | $0.0007500 | 0% |
| BRONZE | $0.0047500 | $0.0007125 | 5% |
| SILVER | $0.0045000 | $0.0006750 | 10% |
| GOLD | $0.0042500 | $0.0006375 | 15% |
| PLATINUM | $0.0041000 | $0.0006150 | 18% |
| DIAMOND | $0.0040000 | $0.0006000 | 20% |

#### Pricing calculator

Use:

**estimated Actor charge = start price + delivered audit count × result price**

Examples, before any taxes or account-specific platform terms:

| Delivered audits | FREE | BRONZE | SILVER | GOLD | PLATINUM | DIAMOND |
| ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| 1 | $0.005750 | $0.005463 | $0.005175 | $0.004888 | $0.004715 | $0.004600 |
| 10 | $0.012500 | $0.011875 | $0.011250 | $0.010625 | $0.010250 | $0.010000 |
| 20 | $0.020000 | $0.019000 | $0.018000 | $0.017000 | $0.016400 | $0.016000 |
| 50 | $0.042500 | $0.040375 | $0.038250 | $0.036125 | $0.034850 | $0.034000 |

The input limit is 50 URLs, but fewer than 50 can become billable results. Invalid duplicates, robots
denials, failed DNS, unsafe destinations, non-HTML responses, timeouts, oversized HTML, and pages
withheld by the charge ceiling stay free diagnostics. A successfully fetched page with a low score is
still a real delivered audit and is billable: “this page has no useful machine-readable facts” is often
the most valuable result in the batch.

At the beginning of a run, the Actor estimates how many result events fit after the amount already
charged. URLs beyond that affordable count are not fetched and receive a <code>buyer\_budget\_not\_attempted</code> diagnostic. Immediately before each successful delivery, it
re-checks the budget. This avoids doing paid work that cannot be delivered and avoids silently pushing
unbilled results.

The accepted 20-URL canary used BRONZE pricing and a $0.05 maximum charge. It delivered and billed 20
audits, with a <code>result-found</code> price of $0.0007125 and no Dataset-item price. The source and
network cost measured after the run left the conservative result-only revenue/COGS ratio above 17×.

### Input contract

The input is a strict JSON object. Unknown fields are rejected instead of being silently ignored.

#### <code>urls</code> — required

- Type: array of strings.
- Minimum: 1 item.
- Maximum: 50 items.
- Per-item limit: 2,048 characters.
- Accepted schemes: HTTP and HTTPS only.
- Behavior: surrounding whitespace is removed; a missing scheme becomes HTTPS; fragments are removed;
  exact normalized duplicates are quarantined without a fetch.
- Rejected during input normalization: embedded username/password, malformed URL, empty string, and
  control characters. A syntactically valid URL that resolves to a non-public destination is instead
  stopped before target-content fetch and recorded as a free <code>unsafe\_private\_target</code>
  diagnostic.

#### <code>expectedBrandName</code> — optional

- Type: string, 1–100 characters.
- Purpose: deterministic comparison against Organization names, Product brands, publishers, and <code>og:site\_name</code>.
- Privacy: used inside the run only; it is never sent to an AI service because the Actor calls no AI
  service.
- Interpretation: absence returns <code>expectedBrandMatch: null</code>; a supplied brand returns true
  or false and may add <code>EXPECTED\_BRAND\_NOT\_MATCHED</code>.

#### <code>brandAliases</code> — optional

- Type: unique array of strings.
- Maximum: 20 aliases, each 1–100 characters.
- Requires: <code>expectedBrandName</code>.
- Purpose: accept public variants such as a legal company suffix, former brand, abbreviation, or
  localized spelling without weakening matching to arbitrary substrings.

#### <code>maxConcurrency</code> — optional

- Type: integer from 1 through 10.
- Default and prefill: 4.
- Purpose: number of page inspections that may proceed in parallel.
- Recommendation: keep 2–4 for a single host. Higher concurrency does not bypass target policy and may
  increase throttling risk.

The Actor deliberately has no proxy, browser, cookie, login, user-agent override, JavaScript-rendering,
LLM key, model, prompt, recursive crawl, or “ignore robots” setting. Those omissions are product
boundaries, not hidden advanced options.

### Real happy, partial and failure output

The following examples separate observed live behavior from contract-shaped boundary examples. They are
not invented performance claims.

#### Happy output — public Organization page

The published Example Task audits <code>https://apify.com/</code> with expected brand “Apify.” The page
currently exposes a valid Organization entity, canonical URL, complete OpenGraph core, stable <code>@id</code>, logo, and multiple authoritative <code>sameAs</code> domains. Dynamic web content can
change, so timestamps and peripheral document types will vary. The JSON below is explicitly abridged:
its three <code>documentTypes</code> and three <code>sameAsDomains</code> are representative subsets of
the larger arrays returned by the observed page.

```json
{
  "schemaVersion": "1.0",
  "url": "https://apify.com/",
  "finalUrl": "https://apify.com/",
  "httpStatus": 200,
  "title": "Apify: The largest marketplace of trusted tools for AI",
  "canonicalUrl": "https://apify.com/",
  "pageLanguage": "en",
  "documentTypes": [
    "Organization",
    "SoftwareApplication",
    "WebSite"
  ],
  "entities": [
    {
      "entityType": "Organization",
      "schemaTypes": ["Organization"],
      "entityId": "https://apify.com/#organization",
      "name": "Apify",
      "url": "https://apify.com/",
      "logoUrl": "https://apify.com/img/apify-logo/apify-symbol-200x200.svg",
      "sameAsDomains": ["github.com", "linkedin.com", "youtube.com"],
      "sourcePath": "/jsonLd/0"
    }
  ],
  "readinessScore": 100,
  "grade": "A",
  "missingHighImpactFacts": [],
  "contradictions": [],
  "fixes": [],
  "expectedBrandMatch": true,
  "staticHtmlOnly": true
}
```

#### Weak but successful output — observed paid canary

The accepted private canary fetched 20 distinct public <code>example.com</code> query URLs from the
exact deployed runtime. Every fetch succeeded and therefore produced a real paid audit. All 20 were
Grade F because that intentionally minimal page had no JSON-LD target entity, canonical URL, or
OpenGraph core. The JSON below is an abridged view of that real row: it omits <code>auditId</code>, <code>checkedAt</code>, <code>openGraph</code>, <code>evidencePaths</code>, some inventory fields, and all but one observed fix:

```json
{
  "schemaVersion": "1.0",
  "url": "https://example.com/?schema-canary=01",
  "finalUrl": "https://example.com/?schema-canary=01",
  "httpStatus": 200,
  "title": "Example Domain",
  "pageLanguage": "en",
  "documentTypes": [],
  "entities": [],
  "factInventory": {
    "jsonLdBlocks": 0,
    "validJsonLdBlocks": 0,
    "invalidJsonLdBlocks": 0,
    "targetEntityCount": 0,
    "organizationCount": 0,
    "productCount": 0,
    "articleCount": 0,
    "hasCanonical": false,
    "hasOpenGraphCore": false,
    "consistencyChecks": 0
  },
  "readinessScore": 0,
  "grade": "F",
  "missingHighImpactFacts": [
    "NO_VALID_JSON_LD",
    "NO_TARGET_ENTITY",
    "NO_ORGANIZATION_IDENTITY",
    "NO_CANONICAL_URL",
    "OPEN_GRAPH_CORE_INCOMPLETE",
    "EXPECTED_BRAND_NOT_MATCHED"
  ],
  "contradictions": [],
  "qualityFlags": [
    "no_target_machine_readable_entity",
    "open_graph_absent",
    "canonical_absent"
  ],
  "fixes": [
    {
      "code": "EXPECTED_BRAND_NOT_MATCHED",
      "action": "Align Organization/Product brand names and og:site_name with an accepted buyer-supplied brand name."
    }
  ],
  "expectedBrandMatch": false,
  "staticHtmlOnly": true
}
```

The live run receipt proved the delivery and pricing invariant:

```json
{
  "status": "complete",
  "counts": {
    "requested": 20,
    "normalizedUnique": 20,
    "affordable": 20,
    "attempted": 20,
    "audited": 20,
    "delivered": 20,
    "billed": 20,
    "quarantined": 0,
    "budgetWithheld": 0
  },
  "billing": {
    "resultEvent": "result-found",
    "resultPriceUsd": 0.0007125,
    "datasetPriceUsd": 0,
    "maxTotalChargeUsd": 0.05,
    "spentBeforeUsd": 0.00475
  }
}
```

#### Partial content — supported and explicitly flagged

A page may be fetched successfully and expose a Product while omitting important commerce facts. That
is a billable audit, not a network failure. The precise values depend on the page, but the supported
contract looks like this:

```json
{
  "documentTypes": ["Product"],
  "entities": [
    {
      "entityType": "Product",
      "name": "Example Widget",
      "brand": "Example",
      "offerPrice": null,
      "offerCurrency": null,
      "availability": null,
      "sourcePath": "/jsonLd/0"
    }
  ],
  "readinessScore": 71,
  "grade": "B",
  "missingHighImpactFacts": [
    "PRODUCT_COMMERCE_FACTS_INCOMPLETE"
  ],
  "fixes": [
    {
      "code": "PRODUCT_COMMERCE_FACTS_INCOMPLETE",
      "action": "Add Product brand plus Offer price, priceCurrency and availability."
    }
  ],
  "staticHtmlOnly": true
}
```

This exact missing-field path is covered by the deployed parser's Product fixtures. It is shown as a
contract example, not presented as an observation from the public Apify homepage.

#### Free failure diagnostic

Failures do not enter the default Dataset. A private address, robots denial, timeout, oversized document,
unsupported content type, or exhausted budget produces a bounded record in <code>QUARANTINE</code>:

```json
{
  "url": "http://127.0.0.1/admin",
  "reason": "unsafe_private_target",
  "httpStatus": null
}
```

or:

```json
{
  "url": "https://example.org/private-page",
  "reason": "robots_disallowed"
}
```

No fetched response body is copied into a diagnostic. A completed run with only diagnostics can have
zero Dataset items, zero <code>result-found</code> events, and a non-empty <code>QUARANTINE</code>.

### Field dictionary

#### Top-level audit identity

| Field | Type | Meaning |
| --- | --- | --- |
| <code>schemaVersion</code> | string | Output contract version, currently <code>1.0</code> |
| <code>auditId</code> | string | SHA-256 of requested URL, final URL, and run-wide observation time |
| <code>url</code> | string | Normalized buyer-supplied URL |
| <code>finalUrl</code> | string | Final URL after up to five guarded redirects |
| <code>httpStatus</code> | integer | Successful HTML response status |
| <code>title</code> | string or null | Normalized HTML title |
| <code>canonicalUrl</code> | string or null | First safe absolute canonical URL |
| <code>pageLanguage</code> | string or null | Normalized HTML <code>lang</code> value |
| <code>checkedAt</code> | ISO timestamp | One shared observation time for every row in a run |
| <code>staticHtmlOnly</code> | boolean | Always true; reminds downstream users of the rendering boundary |

#### Entity and document fields

| Field | Type | Meaning |
| --- | --- | --- |
| <code>documentTypes</code> | string array | Up to 40 normalized Schema.org types seen in supported markup |
| <code>entities</code> | object array | Up to 20 normalized Organization, Product, or Article entities |
| <code>entityType</code> | enum | <code>Organization</code>, <code>Product</code>, or <code>Article</code> |
| <code>schemaTypes</code> | string array | Original normalized type family for the entity |
| <code>entityId</code> | string or null | Safe normalized Schema.org <code>@id</code> or embedded identifier |
| <code>name</code> | string or null | Organization/Product name or Article headline |
| <code>url</code> | string or null | Entity URL resolved against the page |
| <code>logoUrl</code> | string or null | Organization logo URL |
| <code>imageUrl</code> | string or null | Product/Article image URL |
| <code>sameAsDomains</code> | string array | Hostnames from Organization <code>sameAs</code>; URLs are reduced to domains |
| <code>brand</code> | string or null | Product brand name |
| <code>sku</code> / <code>gtin</code> | string or null | Product identifiers when published |
| <code>offerPrice</code> | string or null | Product offer price as normalized source text |
| <code>offerCurrency</code> | string or null | Product offer currency |
| <code>availability</code> | string or null | Terminal Schema.org availability term |
| <code>headline</code> | string or null | Article headline |
| <code>datePublished</code> / <code>dateModified</code> | string or null | Source date text |
| <code>publisher</code> | string or null | Article publisher reference |
| <code>authorPresent</code> | boolean or null | Whether an author reference exists; Person details are excluded |
| <code>sourcePath</code> | string | Bounded JSON-LD, Microdata, or RDFa path |

#### OpenGraph and inventory

<code>openGraph</code> contains title, type, URL, site name, description/image presence, and optional
product price/currency. <code>factInventory</code> records:

- total, valid, and invalid JSON-LD blocks;
- total target entities and counts by Organization, Product, and Article;
- ignored Person entities;
- target entities by JSON-LD, Microdata, and RDFa;
- normalized Microdata and RDFa type lists;
- canonical and OpenGraph-core presence;
- number of cross-surface consistency checks that could actually be made.

#### Decision fields

| Field | Type | Meaning |
| --- | --- | --- |
| <code>readinessScore</code> | integer 0–100 | Weighted fact coverage minus contradiction and brand-mismatch penalties |
| <code>grade</code> | A–F | A ≥ 85, B ≥ 70, C ≥ 55, D ≥ 40, otherwise F |
| <code>missingHighImpactFacts</code> | code array | Important absent or incomplete facts, maximum 20 |
| <code>contradictions</code> | object array | Code, severity, and compared surfaces, maximum 20 |
| <code>fixes</code> | object array | One actionable recommendation per gap or contradiction, maximum 20 |
| <code>qualityFlags</code> | code array | Parsing/context flags such as invalid JSON-LD or absent canonical |
| <code>evidencePaths</code> | object | Source paths supporting important normalized facts |
| <code>expectedBrandMatch</code> | boolean or null | Match result when an expected brand was supplied |

An A is reserved for a result with no declared high-impact gap and no contradiction. Even if raw field
coverage would otherwise qualify for A, any declared gap or contradiction caps the score at 84. This
prevents a page with a critical missing fact from receiving a misleading top grade.

#### Run records

<code>OUTPUT</code> is the audit receipt. <code>COVERAGE</code> groups grades, document types, and
diagnostics. <code>QUARANTINE</code> holds free, bounded per-URL diagnostics. These are key-value-store
records, not Dataset rows and not separately priced result events.

### Evidence and boundaries

The Actor collects evidence from five public surfaces:

1. <code>application/ld+json</code>, including arrays, nested objects, and <code>@graph</code> entries up
   to bounded recursion and value limits.
2. Schema.org Microdata scopes using <code>itemscope</code>, <code>itemtype</code>, <code>itemprop</code>, and supported scalar attributes.
3. RDFa scopes using <code>typeof</code>, <code>property</code>, <code>about</code>, <code>resource</code>, and supported scalar attributes.
4. OpenGraph metadata, including common product price and currency properties.
5. HTML canonical, title, and language metadata.

Evidence paths point to normalized source locations such as <code>/jsonLd/0/name</code> or <code>/microdata/2</code>. They make a finding reproducible, but they are not full JSONPath expressions
against a preserved source document. The Actor does not store the HTML body.

Network authorization is fail-closed. Before the first target request and before every redirect hop, the
Actor:

- normalizes and validates the URL;
- loads the current origin's robots.txt;
- applies the most specific matching rule for its named user agent;
- resolves DNS and rejects the host if any returned address is special-use or non-global;
- pins the connection to one already verified address, preferring IPv4 when available;
- follows at most five redirects;
- caps robots.txt at 100 KB, HTML at 750 KB, and each request at 20 seconds.

This design narrows DNS rebinding and server-side request forgery risk. It does not claim to be a
general-purpose security scanner, firewall, or legal authorization service.

The extraction boundary is **initial static HTML only**. Metadata inserted after page load by JavaScript
is invisible. There is no headless browser and no attempt to execute scripts. This keeps cost and
behavior predictable but means a page may score lower than a browser-based validator shows.

Person details are excluded even when nested inside an Article or Organization graph. The Actor may
return <code>authorPresent: true</code> and increment <code>ignoredPersonEntities</code>, but it does not
return a person's name, profile URL, image, email, job title, or social accounts.

The score is an operational prioritization aid. It is not:

- proof that an AI answer engine crawled, indexed, cited, ranked, or recommended the page;
- a Google rich-results eligibility decision;
- a full Schema.org syntax validator;
- a replacement for content quality, authority, accessibility, or conventional technical SEO review;
- a legal opinion about copyright, database rights, privacy, terms of service, or automated access.

### Decision routing

The Actor is easiest to automate when routing on stable fields rather than prose.

| Condition | Suggested route | Why |
| --- | --- | --- |
| Grade A and expected brand true | Monitor | Core machine-readable facts are present and coherent |
| Grade B/C with missing facts | Create structured-data ticket | Coverage is useful but an explicit high-impact gap remains |
| High-severity contradiction | Escalate before publishing campaign | Two public machine surfaces disagree |
| Expected brand false | Brand/governance review | Page identity does not match accepted names |
| No target entity | Technical SEO backlog | Automated readers lack a supported primary entity |
| Product commerce incomplete | E-commerce schema owner | Offer facts cannot be reliably used |
| Article provenance incomplete | Editorial platform owner | Attribution and freshness are weak |
| Diagnostic reason | Access/operations queue, no SEO ticket yet | The page was not audited successfully |

Example pseudo-routing:

```text
if row.contradictions contains severity high:
    open urgent metadata-consistency ticket
else if row.expectedBrandMatch is false:
    notify brand governance
else if row.missingHighImpactFacts is not empty:
    add fixes to technical SEO backlog
else:
    record healthy observation
```

Do not compare <code>auditId</code> across runs as a stable page identifier because its timestamp is
intentionally part of the hash. Join history on normalized <code>url</code>. Compare <code>readinessScore</code>, <code>grade</code>, finding codes, and selected entity facts. Keep the
source <code>checkedAt</code> so a later page change is not mistaken for inconsistent parsing.

For a batch, route operational failures first by <code>QUARANTINE.reason</code>. A robots denial should
not create the same ticket as an expected-brand mismatch because no page audit occurred. A <code>buyer\_budget\_not\_attempted</code> record is solved by changing run budget or splitting the task,
not by changing the target site.

### Commercial playbooks

#### GEO readiness baseline for a new client

Choose the homepage, About page, two revenue-driving Product pages, and two authoritative Articles. Run
with the client's accepted brand name and aliases. Deliver:

- score and grade by page;
- one consolidated list of high-impact gap codes;
- contradictions with both compared surfaces;
- top three fixes per page;
- screenshots or source snippets collected manually only where the client needs implementation detail.

The Actor supplies the evidence map and prioritization. The agency supplies context, ownership, and
implementation. This is a credible paid audit because the deliverable is reproducible and bounded.

#### Rebrand or domain-migration quality control

Audit the old homepage, new homepage, key redirected Product URLs, and major Articles. Compare canonical,
entity URL, OpenGraph URL, entity name, and expected-brand result. High-severity mismatches can reveal a
page whose visible content migrated while its JSON-LD still names the old domain or brand.

Run daily during launch week, then weekly until results stabilize. Keep URLs in the same saved Task so
Dataset diffs remain simple.

#### E-commerce structured-offer backlog

Submit the highest-traffic or highest-margin Product pages, not the entire catalog. Group findings:

- Product entity absent;
- brand absent or inconsistent;
- Offer price/currency/availability incomplete;
- identifier absent;
- OpenGraph product price conflicts with structured Offer.

Route the first three groups to the template/schema owner because one template fix may improve thousands
of pages. Use the result count to estimate rollout impact; do not imply that each audit equals an AI
citation.

#### Editorial authority and provenance review

Audit evergreen guides, original research, and executive thought leadership. Filter Article rows for
missing publisher, author presence, publication dates, or image. Verify that the Organization identity
on the site is also complete. This creates a compact provenance checklist for editorial operations.

#### Multi-brand governance

Create one Task per brand with its own expected name and aliases. Schedule the same page families. Send
only false brand matches and high-severity contradictions to the shared governance queue. This avoids
asking a reviewer to read every healthy result.

#### Lead magnet or freemium audit

An agency can offer a five-page baseline as a low-cost diagnostic, then sell implementation and
monitoring. The Actor's price is measured per delivered page; at FREE tier, a five-result run is
approximately $0.00875 including one start event. Present the report honestly as structured-fact
readiness, not guaranteed AI ranking.

### Integration recipes

#### REST API with cURL

Start a run and let the API wait up to its supported 60-second maximum:

```bash
curl -X POST \
  "https://api.apify.com/v2/acts/zinin~schema-org-for-llms-audit/runs?token=<APIFY_TOKEN>&waitForFinish=60&maxTotalChargeUsd=0.02" \
  -H "Content-Type: application/json" \
  -d '{
    "urls": ["https://apify.com/"],
    "expectedBrandName": "Apify",
    "maxConcurrency": 2
  }'
```

If the returned run status is still transitional, read <code>data.id</code> from that response and poll <code>GET https://api.apify.com/v2/actor-runs/\<RUN\_ID></code> until it is terminal. A terminal <code>SUCCEEDED</code> response carries <code>defaultDatasetId</code>; use that ID to retrieve rows:

```bash
curl "https://api.apify.com/v2/datasets/<DATASET_ID>/items?clean=true&format=json&token=<APIFY_TOKEN>"
```

Keep tokens in a secret manager or request header in production; do not commit them to source control.

#### JavaScript client

```javascript
import { ApifyClient } from "apify-client";

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor("zinin/schema-org-for-llms-audit").call({
  urls: ["https://apify.com/"],
  expectedBrandName: "Apify",
  maxConcurrency: 2
}, {
  maxTotalChargeUsd: 0.02
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map(({ url, grade, readinessScore, fixes }) => ({
  url, grade, readinessScore, fixes
})));
```

#### Python client

```python
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("zinin/schema-org-for-llms-audit").call(
    run_input={
        "urls": ["https://apify.com/"],
        "expectedBrandName": "Apify",
        "maxConcurrency": 2,
    },
    max_total_charge_usd=0.02,
)
rows = client.dataset(run["defaultDatasetId"]).list_items().items
for row in rows:
    print(row["url"], row["grade"], row["readinessScore"])
```

#### Apify MCP server

Connect an MCP-compatible client to <code>https://mcp.apify.com</code> with OAuth, or provide an Apify
token as a Bearer header. Enable the Actor directly:

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com?tools=zinin/schema-org-for-llms-audit"
    }
  }
}
```

The client exposes the selected Actor as a tool and includes <code>get-actor-output</code> for structured results. Alternatively enable the general <code>call-actor</code> tool and pass Actor ID <code>zinin/schema-org-for-llms-audit</code>. Running
Actors and reading their storage requires authenticated MCP access. The Actor uses <code>LIMITED\_PERMISSIONS</code> and is compatible with Apify's MCP execution surface.

#### Saved Task and schedule

Save the desired URL set as a Task, run it once manually, then attach an Apify Schedule. Use a webhook
for terminal run status and read the default Dataset plus <code>OUTPUT</code> after success. If a
downstream system needs only changes, compare the newest rows to the previous run by URL and finding
codes before sending notifications.

#### Make, Zapier, Sheets, BI, or warehouse

Use the run-finished trigger, then export the default Dataset. Flatten:

- <code>url</code>, <code>grade</code>, <code>readinessScore</code>, and <code>expectedBrandMatch</code> as dimensions;
- lengths of gaps and contradictions as triage metrics;
- finding codes as exploded child rows when your warehouse supports them;
- <code>checkedAt</code> as observation time, not ingestion time.

Keep <code>entities</code>, <code>fixes</code>, and <code>evidencePaths</code> as JSON if the destination
supports semi-structured columns. Do not drop <code>staticHtmlOnly</code>; it is an important
interpretation boundary.

### Operating guide

#### Before the first production run

1. Confirm each URL is public and intended for automated access.
2. Choose a charge ceiling using the pricing calculator.
3. Use an expected brand only when the accepted public spelling is known.
4. Start with low concurrency and a small representative sample.
5. Decide where historical outputs will live; the Actor does not create a long-term trend database.

#### After every run

1. Confirm terminal status is <code>SUCCEEDED</code>.
2. Compare Dataset item count with <code>OUTPUT.counts.delivered</code> and, on platform, <code>OUTPUT.counts.billed</code>.
3. Review <code>QUARANTINE</code> before concluding that missing Dataset rows mean missing metadata.
4. Sort Dataset rows by contradiction severity, expected-brand mismatch, missing-fact count, then score.
5. Preserve run ID and checked time with any exported report.

#### Troubleshooting

**Dataset is empty, but the run succeeded.** Open <code>QUARANTINE</code>. The most common reasons are
robots denial, HTML over 750 KB, non-HTML response, timeout, or charge-budget withholding.

**The page looks rich in a browser but scores poorly.** Its structured metadata may be inserted by
client JavaScript. This Actor reads only initial static HTML. Use a browser-based validator to confirm
the difference, then decide whether server-rendering critical facts is appropriate.

**The expected brand is false even though the brand appears in body copy.** Body copy is not a supported
identity surface. Publish the accepted name in Organization/Product/publisher metadata or <code>og:site\_name</code>, or add a legitimate alias to input.

**A Product price conflict looks surprising.** Compare <code>contradictions</code> and <code>evidencePaths</code>. The Actor normalizes numeric formatting but does not perform currency
conversion, tax interpretation, regional price selection, or offer aggregation beyond its bounded
primary-value selection.

**A target is quarantined as unsafe.** The hostname resolved to at least one non-global address or was a
special-use literal. This is not overrideable. Use the Actor only for public web destinations.

**The run stopped near its charge ceiling.** The platform may have charged the start event before result
delivery. Increase <code>maxTotalChargeUsd</code> or reduce/split the URL list. Withheld pages remain free.

#### Monitoring drift

For scheduled runs, alert on:

- new high-severity contradiction codes;
- transition from expected-brand true to false;
- disappearance of the target entity;
- grade decline of two or more levels;
- sudden <code>robots\_disallowed</code> or <code>html\_too\_large</code> diagnostics;
- divergence among requested, delivered, billed, and Dataset counts.

Do not alert only on a one-point score change. Small field changes can alter weighted coverage without
creating a commercially meaningful regression.

### FAQ

#### Does structured data make my brand appear in ChatGPT, Gemini, Claude, or another answer engine?

No guarantee is possible. Structured facts can make a page easier for automated systems to interpret,
but crawl access, indexing, authority, retrieval, model behavior, and answer context remain outside this
Actor. The output measures public fact readiness, not citations or rank.

#### Does the Actor call an LLM?

No. It calls no OpenAI, Anthropic, Google, OpenRouter, Chinese-model, search, or embedding API. Scoring
and recommendations are deterministic code.

#### Does it render JavaScript?

No. It audits initial static HTML. Client-injected JSON-LD is outside scope and explicitly disclosed by <code>staticHtmlOnly: true</code>.

#### Why is a Grade F page billable?

Because the page was safely fetched, audited, and delivered, and “no usable target facts were present”
is a valid decision result. Network failures, denials, unsafe targets, and budget withholding are free
diagnostics instead.

#### Can it crawl an entire domain?

No. You provide exact URLs, up to 50 per run. There is no sitemap discovery, link traversal, or recursive
crawl. This protects scope, cost, and target load.

#### Can I bypass robots.txt or use a proxy?

No. Robots checks are mandatory per origin and redirect hop. The Actor exposes no proxy or bypass
setting.

#### What entity types are audited?

Organization and many LocalBusiness/organization subtypes, Product variants, and Article variants.
Other Schema.org types can appear in <code>documentTypes</code> but are not normalized as primary target
entities.

#### Are Person entities returned?

No. Their details are excluded. For an Article, the Actor reports only whether an author reference is
present.

#### Does it validate every Schema.org rule?

No. It parses bounded facts needed for this decision product. Use the Schema.org validator and relevant
search-engine rich-result tools for full vocabulary and platform-specific validation.

#### Can I compare two runs?

Yes. Join rows by normalized URL and compare score, grade, finding codes, contradictions, and selected
entity facts. Keep <code>checkedAt</code>. Do not use <code>auditId</code> as a time-invariant page key.

#### Why can the output differ tomorrow?

The target page, robots.txt, DNS, redirect chain, or static markup may change. The Actor reports live
public evidence at run time and does not cache page bodies.

#### What if robots.txt is missing?

HTTP 404 or 410 for robots.txt is treated as no declared rule. Unreadable, denied, oversized, or
HTML-shaped robots responses fail closed because policy cannot be established reliably.

#### What counts as a contradiction?

The five implemented contradiction codes are <code>CANONICAL\_OG\_URL\_MISMATCH</code> (canonical URL
versus OpenGraph URL), <code>TITLE\_OG\_TITLE\_MISMATCH</code> (HTML title versus OpenGraph title), <code>ORG\_OG\_SITE\_NAME\_MISMATCH</code> (Organization name versus OpenGraph site name), <code>PRODUCT\_PRICE\_MISMATCH</code> (normalized Product Offer price versus OpenGraph product price), and <code>PRODUCT\_CURRENCY\_MISMATCH</code> (Offer currency versus OpenGraph product currency). Each returned
object names its code, severity, and compared surfaces. An expected-brand mismatch is a missing-fact code, <code>EXPECTED\_BRAND\_NOT\_MATCHED</code>, not a contradiction object.

#### Is my input sent anywhere else?

It is used inside the Apify run to fetch the public targets. There is no child Actor or external AI
service. Normal Apify platform storage and networking still apply to the run.

#### Can I get CSV or Excel?

Yes. Use Apify Dataset export formats. Nested entities, gaps, contradictions, and evidence are most
complete in JSON; flattened formats may serialize them as nested values.

#### How should I report a reproducible issue?

Provide the run ID, exact normalized URL, row <code>checkedAt</code>, relevant finding code, and whether
the discrepancy concerns static HTML or browser-rendered DOM. Never send private credentials.

#### Complete the AI visibility workflow

| Related Actor | Use it for |
| --- | --- |
| [AI Crawler Access Checker](https://apify.com/zinin/ai-crawler-access-checker) | Verify robots.txt access for major AI crawler user agents |
| [llms.txt Auditor](https://apify.com/zinin/llms-txt-auditor) | Audit an llms.txt file and crawler-policy alignment |
| [AI Answer Change Alert](https://apify.com/zinin/ai-answer-change-alert) | Detect changes in collected answer and citation datasets |
| [AI Overview Citation Tracker](https://apify.com/zinin/ai-overview-tracker) | Track cited domains in Google AI Overview observations |

Use this Actor for page-level facts, the access Actors for crawler policy, and citation Actors for
observed answer surfaces. None of them should be interpreted as a guaranteed ranking forecast.

### Sources and rights

The Actor reads buyer-supplied public HTTP(S) pages and each target origin's public robots.txt. It does
not search for URLs, log in, accept cookies, access a private API, bypass a paywall, solve a challenge,
or reproduce a site-specific database.

The target page remains the source of every extracted fact. Output contains normalized facts and bounded
evidence paths, not the HTML body, article text, image binaries, or Person profiles. Logo and image
values are source URLs only; the Actor does not download or redistribute the assets.

Robots.txt is checked as an automated-access signal before each initial and redirected request. It is not
a legal license or a complete statement of site terms. Buyers are responsible for confirming that their
chosen targets and intended downstream use comply with applicable terms, contracts, copyright, database
rights, privacy rules, and law. This page does not provide legal advice.

The scoring framework and fix codes are this Actor's deterministic analysis. Schema.org is a public
vocabulary; OpenGraph and HTML metadata are public web conventions. The Actor does not claim affiliation
with Schema.org, Google, OpenAI, Anthropic, Microsoft, Meta, or any target brand.

Operational behavior follows Apify's public platform contracts:

- [Actor README guidance](https://docs.apify.com/actors/publishing/actor-readme)
- [Pay-per-event SDK behavior](https://docs.apify.com/sdk/js/docs/concepts/pay-per-event)
- [Apify API client](https://docs.apify.com/api/client/js/reference/class/ApifyClient)
- [Apify MCP server](https://docs.apify.com/integrations/mcp)

Target sites can change markup, robots policy, DNS, redirects, or availability without notice. A
successful observation is evidence for its timestamp, not a warranty of future access or format.

### Limits

- 1–50 exact URLs per run; no discovery or recursive crawl.
- Static HTML only; no JavaScript rendering.
- 750 KB maximum HTML and 100 KB maximum robots document.
- 20-second per-request timeout and five redirect hops.
- Maximum 20 normalized target entities, 40 document types, 20 gaps, 20 contradictions, and 20 fixes per
  row; bounded parsing prevents untrusted pages from creating unbounded output.
- Organization, Product, and Article families only as primary entities.
- No Person details.
- No review of visible body copy, writing quality, topical authority, backlinks, accessibility, or
  conventional search rank.
- No full Schema.org syntax validation or search-engine rich-result eligibility.
- No AI citation, sentiment, share-of-voice, prompt, or model-position measurement.
- No currency conversion, regional offer selection, tax/shipping interpretation, or catalog matching.
- No history database, change alert, PDF report, or ticket-system write inside this Actor.
- No proxy, login, session, cookie banner interaction, CAPTCHA handling, or robots override.
- Pages served differently by geography, consent state, or user agent may expose facts different from
  those this named crawler receives.

### Support boundary

Support covers the documented input validation, public network-safety checks, supported static-HTML
fact extraction, score/finding contract, schemas, Dataset/KVS routing, and <code>result-found</code> billing invariant for the current build.

Support cannot guarantee a target site's uptime, permission, HTML shape, server response, continued
publication of any field, or placement in an external AI answer. It cannot provide legal advice, bypass
target controls, recover facts available only after client JavaScript, or reinterpret unsupported entity
types as though they were audited.

For a useful support request, include Actor run ID, exact URL, observation timestamp, status and reason
or finding code, and a concise statement of expected versus actual static-HTML behavior. If the problem
appears only in browser-rendered DOM, say so explicitly.

***

Built by [zinin](https://apify.com/zinin). The product is intentionally narrow: public evidence in,
defensible GEO action out.

# Actor input Schema

## `urls` (type: `array`):

One to 50 public HTTP(S) pages. Fragments and exact duplicates are removed before fetching.

## `expectedBrandName` (type: `string`):

Used only for a deterministic identity-consistency flag. It is not sent to any AI service.

## `brandAliases` (type: `array`):

Optional accepted spellings used with Expected brand name.

## `maxConcurrency` (type: `integer`):

Parallel public-page fetches. A low default protects target sites and run budgets.

## Actor input object example

```json
{
  "urls": [
    "https://example.com/"
  ],
  "maxConcurrency": 4
}
```

# Actor output Schema

## `results` (type: `string`):

One normalized audit per successfully fetched URL.

## `OUTPUT` (type: `string`):

Delivery, billing and timing counts.

## `QUARANTINE` (type: `string`):

Bounded failure reasons; response bodies are never copied.

## `COVERAGE` (type: `string`):

Grades, entity types and diagnostic counts.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com/"
    ],
    "maxConcurrency": 4
};

// Run the Actor and wait for it to finish
const run = await client.actor("zinin/schema-org-for-llms-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://example.com/"],
    "maxConcurrency": 4,
}

# Run the Actor and wait for it to finish
run = client.actor("zinin/schema-org-for-llms-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com/"
  ],
  "maxConcurrency": 4
}' |
apify call zinin/schema-org-for-llms-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zinin/schema-org-for-llms-audit"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5sTccSHfqvFyAvA3s/builds/NR6EjYoUJ8gwUqvUF/openapi.json
