# Internal Link Architecture Verifier (`nexgenwatch/internal-link-architecture-verifier`) Actor

Runs a bounded same-origin crawl unioned with your sitemap, builds the internal link graph, and returns one evidenced verdict per page: orphan-from-sitemap status, click depth, nofollow conflicts, canonical inlinks, inbound/outbound counts and anchor-text concentration.

- **URL**: https://apify.com/nexgenwatch/internal-link-architecture-verifier.md
- **Developed by:** [NexGen Watch](https://apify.com/nexgenwatch) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $33.50 / 1,000 page architecture checks

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Internal Link Architecture Verifier

Runs a bounded same-origin crawl (unioned with your sitemap), builds the internal link graph, and returns one evidenced verdict per reachable page: orphan-from-sitemap status, click depth from root, navigation-only-inlinks, nofollow conflicts, whether the canonical target receives links, inbound/outbound counts and anchor-text concentration. Blocked or unreachable nodes are reported UNKNOWN and never poison the graph verdicts of reachable pages.

### What you submit

`urls` — site root or sitemap urls. A site root URL and/or sitemap URL(s). The actor runs a BOUNDED same-origin crawl (page + depth caps) unioned with the sitemap URLs, builds the internal link graph, and returns one evidenced verdict per reachable page: orphan-from-sitemap status, click depth from root, navigation-only-inlinks flag, nofollow conflicts, whether the canonical target receives internal links, inbound/outbound internal-link counts, and inbound anchor-text concentration. A blocked or unreachable node is reported UNKNOWN, never bills, and never poisons the graph verdicts of reachable pages (the graph is labeled partial and metrics are lower bounds over the reachable subgraph).

### The runtime source gate (per submitted target)

Because you choose the targets, the source contract is enforced at run time, per origin, before any page is read:

- **robots.txt** is fetched once per origin and honored for our crawler. No robots / 404 = permitted; a disallowed path or an unavailable robots file = **BLOCKED**.
- **SSRF guard** — every host is resolved and must be a public address. Private, loopback, link-local, reserved and cloud-metadata addresses are refused before a socket opens.
- **HTTP** — bounded body reads; `403` / `429` / `5xx` = **BLOCKED**; DNS / timeout / connection faults = **UNREACHABLE**.

A **BLOCKED** or **UNREACHABLE** target is delivered as an unbilled status row — never a broken-site verdict, and never charged.

### Output

One row per target. A complete evidenced verdict carries `outcome: answer`; blocked / unreachable / unparseable targets carry that status and are not billed.

### Pricing

Pay per event. The reserved start event is platform-automatic and is never charged from code. The only event charged in code is one complete verdict per target — blocked, unreachable and locally-malformed targets never bill (push-then-charge; failures can only undercharge).

| Event | Price (USD) | When it fires |
| --- | --- | --- |
| `apify-actor-start` | $0.02 | once when the run starts (reserved; never charged in code) |
| `page_architecture_check` | $0.05 | once per **complete evidenced verdict** delivered (blocked / unreachable / unparseable never bill) |

Volume tiers reduce the per-verdict price as usage grows: $0.05 → $0.045 → $0.04 → $0.0335.

### Why this and not a generic crawler

site-level internal-link-GRAPH verdicts (orphans, click depth, nav-only inlinks, nofollow conflicts, anchor concentration) — not a one-page link array (the 117-user store comparator), and a different job from our own sitemap-indexability-auditor (per-URL indexability) and website-migration-link-auditor (old->new redirect acceptance).

# Actor input Schema

## `urls` (type: `array`):

A site root URL and/or sitemap URL(s). The actor runs a BOUNDED same-origin crawl (page + depth caps) unioned with the sitemap URLs, builds the internal link graph, and returns one evidenced verdict per reachable page: orphan-from-sitemap status, click depth from root, navigation-only-inlinks flag, nofollow conflicts, whether the canonical target receives internal links, inbound/outbound internal-link counts, and inbound anchor-text concentration. A blocked or unreachable node is reported UNKNOWN, never bills, and never poisons the graph verdicts of reachable pages (the graph is labeled partial and metrics are lower bounds over the reachable subgraph).

## `maxPages` (type: `integer`):

Hard cap on pages fetched across all inputs. Bounds the crawl and billing (one verdict per reachable page).

## `maxDepth` (type: `integer`):

How many link-hops from each root to crawl (same-origin only).

## `userAgent` (type: `string`):

Override the transparent crawler UA. Default is the fleet's contactable bot UA.

## Actor input object example

```json
{
  "urls": [
    "https://www.gov.uk/foreign-travel-advice/france",
    "http://169.254.169.254/latest/meta-data"
  ],
  "maxPages": 25,
  "maxDepth": 2,
  "userAgent": "Mozilla/5.0 (compatible; NexGenWatchBot/1.0; +https://apify.com/nexgenwatch)"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.gov.uk/foreign-travel-advice/france",
        "http://169.254.169.254/latest/meta-data"
    ],
    "maxPages": 25,
    "maxDepth": 2,
    "userAgent": "Mozilla/5.0 (compatible; NexGenWatchBot/1.0; +https://apify.com/nexgenwatch)"
};

// Run the Actor and wait for it to finish
const run = await client.actor("nexgenwatch/internal-link-architecture-verifier").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://www.gov.uk/foreign-travel-advice/france",
        "http://169.254.169.254/latest/meta-data",
    ],
    "maxPages": 25,
    "maxDepth": 2,
    "userAgent": "Mozilla/5.0 (compatible; NexGenWatchBot/1.0; +https://apify.com/nexgenwatch)",
}

# Run the Actor and wait for it to finish
run = client.actor("nexgenwatch/internal-link-architecture-verifier").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.gov.uk/foreign-travel-advice/france",
    "http://169.254.169.254/latest/meta-data"
  ],
  "maxPages": 25,
  "maxDepth": 2,
  "userAgent": "Mozilla/5.0 (compatible; NexGenWatchBot/1.0; +https://apify.com/nexgenwatch)"
}' |
apify call nexgenwatch/internal-link-architecture-verifier --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nexgenwatch/internal-link-architecture-verifier"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/PaMlUDGGLSmIWjWJQ/builds/EAK0skhhRLNvgcXte/openapi.json
