# HTML Sitemap Link Extractor (`junipr/html-sitemap-link-extractor`) Actor

Audit html sitemap link extractor inputs and return structured findings, evidence rows, and a buyer-ready report for SEO, developer, and operations teams.

- **URL**: https://apify.com/junipr/html-sitemap-link-extractor.md
- **Developed by:** [junipr](https://apify.com/junipr) (community)
- **Categories:** SEO tools, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.90 / 1,000 html sitemap link extractor sitemap file checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## HTML Sitemap Link Extractor

### Store Positioning

**Store title:** HTML Sitemap Link Extractor

**Short description:** Audit html sitemap link extractor inputs and return structured findings, evidence rows, and a buyer-ready report for SEO, developer, and operations teams.

**SEO title:** HTML Sitemap Link Extractor — technical SEO, web, and domain audit

**SEO description:** Audit html sitemap link extractor inputs and return structured findings, evidence rows, and a buyer-ready report for SEO, developer, and operations teams. Use it to find crawlability, indexability, security, metadata, and page-quality issues with evidence-backed rows and audit reports.

**Categories:** SEO\_TOOLS, DEVELOPER\_TOOLS

**Keywords:** html, sitemap, link, extractor, structured data, public data, web/domain audit

### Fixed-Inclusive PPE Pricing

This actor uses pay-per-event pricing. Event prices include Apify platform usage; users are not expected to pay a separate platform-usage pass-through charge for the configured pricing model.

- Tier: W1 — Web/domain audit
- Primary event: `sitemap-file-checked` at $0.00490 base
- Default max charge: $10.00
- Store discounts: FREE/BRONZE base, SILVER discounted, GOLD deepest approved discount

Event set:

- `actor-start`: base $0.00500, GOLD $0.00400. HTML Sitemap Link Extractor: charged when actor start is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `sitemap-file-checked`: base $0.00490, GOLD $0.00392. HTML Sitemap Link Extractor: charged when sitemap file checked is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `record-extracted`: base $0.00372, GOLD $0.00298. HTML Sitemap Link Extractor: charged when record extracted is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `finding-emitted`: base $0.00372, GOLD $0.00298. HTML Sitemap Link Extractor: charged when finding emitted is completed. The price includes Apify platform usage; no separate usage pass-through is intended.
- `audit-report-generated`: base $0.05000, GOLD $0.04000. HTML Sitemap Link Extractor: charged when audit report generated is completed. The price includes Apify platform usage; no separate usage pass-through is intended.

### Public Task Concepts

- Audit HTML Sitemap Link controls on a capped public sample
- Find high-priority HTML Sitemap Link issues before release
- Validate HTML Sitemap Link evidence from supplied pages
- Prioritize HTML Sitemap Link fixes with severity and proof
- Export HTML Sitemap Link QA rows for client review

HTML Sitemap Link Extractor turns visible sitemap pages into structured link inventories while retaining section headings, nested list depth, hierarchy paths, anchor labels, normalized targets, and link health.

### What it does

- Fetches public candidate pages or parses supplied HTML through the same hierarchy-aware extractor.
- Detects likely HTML sitemap pages from URL, title, and visible content signals.
- Preserves heading sections, nested list depth, link order, hierarchy path, internal/external classification, and evidence excerpts.
- Detects duplicate targets, empty anchors, broken links when enabled, external links, and low-structure sitemap pages.

### Reports

The key-value store contains the extraction report, full link CSV, section hierarchy JSON, duplicate-link CSV, and link-status CSV.

### Example input

```json
{
  "startUrls": [
    "https://www.wikipedia.org/"
  ],
  "htmlInputs": [
    {
      "sourceUrl": "https://www.wikipedia.org/wiki/Main_Page",
      "html": "<!doctype html><html lang=\"en\"><head><title>HTML Sitemap Link Extractor Demo</title><meta name=\"description\" content=\"A public page for technical SEO analysis.\"><meta name=\"viewport\" content=\"width=device-width, initial-scale=1\"><link rel=\"canonical\" href=\"https://www.wikipedia.org/wiki/Main_Page\"><link rel=\"stylesheet\" href=\"https://cdn.jsdelivr.net/npm/site.css\"><link rel=\"stylesheet\" href=\"/local.css\" integrity=\"sha384-bad\"><script src=\"https://cdn.jsdelivr.net/npm/app.js\"></script><script src=\"/ok.js\" integrity=\"sha256-q3bAT6PqU/NkM+nmCzOxM1iv7FPl/JTMqy0YgkNF7+k=\" crossorigin=\"anonymous\"></script></head><body><h1>HTML Sitemap Link Extractor</h1><h3>Skipped level</h3><p>Updated on 2026-06-20. This page has enough text to support readability, content quality, architecture, link, image, and technical SEO checks with deterministic computed evidence generated from markup.</p><nav><a href=\"/wiki/Main_Page\">Main Page</a><a href=\"https://www.mediawiki.org/\">External docs</a></nav><img src=\"/logo.png\" alt=\"\" width=\"120\"><form action=\"https://forms.example.org/submit\"><label>Email<input type=\"email\" name=\"email\" required></label><input type=\"hidden\" name=\"token\" value=\"abc\"></form><script type=\"application/ld+json\">{\"@context\":\"https://schema.org\",\"@type\":\"Product\",\"name\":\"Demo Product\",\"sku\":\"sku-1\",\"offers\":{\"@type\":\"Offer\",\"price\":\"19.00\",\"priceCurrency\":\"USD\",\"availability\":\"https://schema.org/InStock\"}}</script></body></html>",
      "headers": {
        "cache-control": "max-age=3600",
        "content-encoding": "br",
        "last-modified": "Wed, 01 Jul 2026 00:00:00 GMT"
      }
    }
  ],
  "allowedDomains": "",
  "maxPages": 3,
  "maxLinks": 3,
  "detectSitemapPages": true,
  "includeSectionHierarchy": true,
  "checkLinkStatus": true,
  "includeExternalLinks": true,
  "requestDelayMs": 3,
  "timeoutMs": 3
}
```

### Output fields

sourceUrl, sourceTitle, sitemapPageDetected, sectionHeading, sectionDepth, linkOrder, anchorText, targetUrl, normalizedTargetUrl, internal, external, statusCode, duplicateLinkGroupId, emptyAnchor, hierarchyPath, issueCode, recommendation, evidenceSnippet

### Cost controls

Use `maxPages` and `maxLinks` to bound extraction; link-status requests can be disabled for inventory-only runs.

### Limitations

This actor does not parse XML sitemaps, render client-side navigation, generate sitemap pages, or change site links.

# Actor input Schema

## `startUrls` (type: `array`):

start urls used by Html Sitemap Link Extractor.

## `htmlInputs` (type: `array`):

html inputs used by Html Sitemap Link Extractor.

## `allowedDomains` (type: `string`):

allowed domains used by Html Sitemap Link Extractor.

## `maxPages` (type: `integer`):

max pages used by Html Sitemap Link Extractor.

## `maxLinks` (type: `integer`):

max links used by Html Sitemap Link Extractor.

## `detectSitemapPages` (type: `boolean`):

detect sitemap pages used by Html Sitemap Link Extractor.

## `includeSectionHierarchy` (type: `boolean`):

include section hierarchy used by Html Sitemap Link Extractor.

## `checkLinkStatus` (type: `boolean`):

check link status used by Html Sitemap Link Extractor.

## `includeExternalLinks` (type: `boolean`):

include external links used by Html Sitemap Link Extractor.

## `requestDelayMs` (type: `integer`):

request delay milliseconds used by Html Sitemap Link Extractor.

## `timeoutMs` (type: `integer`):

timeout milliseconds used by Html Sitemap Link Extractor.

## Actor input object example

```json
{
  "startUrls": [],
  "htmlInputs": [
    {
      "sourceUrl": "https://shop.example.org/sitemap",
      "html": "<html><head><title>HTML Sitemap</title></head><body><h1>Site map</h1><h2>Products</h2><ul><li><a href='/products'>All products</a><ul><li><a href='/products/shoes'>Shoes</a></li></ul></li></ul><h2>Company</h2><ul><li><a href='/about'>About</a></li></ul></body></html>"
    }
  ],
  "allowedDomains": "",
  "maxPages": 3,
  "maxLinks": 3,
  "detectSitemapPages": true,
  "includeSectionHierarchy": true,
  "checkLinkStatus": false,
  "includeExternalLinks": true,
  "requestDelayMs": 0,
  "timeoutMs": 10000
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `reports` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("junipr/html-sitemap-link-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("junipr/html-sitemap-link-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call junipr/html-sitemap-link-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,junipr/html-sitemap-link-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wdp4uGU6mhHoUqtED/builds/1ipyuSUnEax38hUwY/openapi.json
