# GOV.UK Content Sitemap Scraper (`neuton/govuk-content-sitemap-scraper`) Actor

Discover GOV.UK URLs from sitemaps and enrich them with GOV.UK Content API metadata for policy monitoring, compliance research, and public-content inventories.

- **URL**: https://apify.com/neuton/govuk-content-sitemap-scraper.md
- **Developed by:** [Ashwin Prasad](https://apify.com/neuton) (community)
- **Categories:** Business, AI, Automation
- **Stats:** 2 total users, 1 monthly users, 71.4% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## GOV.UK Content Sitemap Scraper

Turn GOV.UK sitemaps into a clean, searchable dataset of official pages, titles, document types, publishing organisations, and update dates. Use it for UK policy monitoring, compliance research, public-content inventories, and AI/RAG ingestion.

### Start in 30 seconds

1. Keep the default GOV.UK sitemap.
2. Add a topic such as `employment` or `tax`.
3. Run with 25-50 URLs, then inspect the Dataset tab.

```json
{
  "startUrls": [{ "url": "https://www.gov.uk/sitemap.xml" }],
  "keywords": ["employment"],
  "fetchContentApi": true,
  "maxUrls": 25
}
```

Expected result: up to 25 official GOV.UK page rows with source URLs and public Content API metadata where available. Empty topic matches return zero rows rather than diagnostic records.

### Why this actor

GOV.UK is a major source for UK policy, compliance, tax, employment, immigration, procurement, business-support, and public-service information. Organizations often need to discover and monitor official pages by topic, but raw sitemaps and content metadata are awkward to join manually. This actor creates structured rows that work for UK compliance dashboards, public-sector research, site inventories, and RAG pipelines.

SEO keywords covered by this Store page include GOV.UK scraper, GOV.UK content API scraper, UK government website scraper, government guidance monitor UK, GOV.UK sitemap scraper, UK compliance content scraper, UK policy monitoring, and GOV.UK RAG dataset.

### Output

- GOV.UK page URL and base path
- Title and description
- Document type and schema name
- Public update date and first published date
- Publishing app, organisations, and topical links
- Raw GOV.UK Content API payload when available

### Common Use Cases

- Monitor regulatory, tax, immigration, employment, or policy guidance
- Build GOV.UK content inventories for search and AI retrieval
- Track public-sector information updates by topic
- Feed official UK government pages into compliance dashboards
- Discover pages for downstream sitemap or content-change monitors
- Trigger alerts when pages matching a policy, tax, employment, or immigration topic are discovered or updated

### Agent and automation workflows

Schedule the actor for topic lists, send rows into Airtable, Slack, Make, n8n, a warehouse, or a vector database, and let agents summarize new guidance, route pages to policy owners, or prepare compliance briefs from official source URLs.

### Responsible use

This actor extracts public GOV.UK sitemap and Content API metadata. Use it for monitoring, compliance support, public research, content inventory, and AI retrieval workflows. Verify legal, policy, or compliance decisions against the official page and qualified advice.

### Pricing

Pay **$2 per 1,000 returned content-page rows**. Content API enrichment is included when enabled. Filtered URLs, duplicate URLs, failed enrichment requests, and diagnostics are not returned as billable rows.

# Actor input Schema

## `startUrls` (type: `array`):

GOV.UK sitemap or page URLs. The default uses the main GOV.UK sitemap index.

## `keywords` (type: `array`):

Optional URL/title/description keywords to keep.

## `fetchContentApi` (type: `boolean`):

Fetch GOV.UK Content API metadata for discovered pages.

## `maxUrls` (type: `integer`):

Maximum GOV.UK page rows to return.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.gov.uk/sitemap.xml"
    }
  ],
  "fetchContentApi": true,
  "maxUrls": 50
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.gov.uk/sitemap.xml"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("neuton/govuk-content-sitemap-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.gov.uk/sitemap.xml" }] }

# Run the Actor and wait for it to finish
run = client.actor("neuton/govuk-content-sitemap-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.gov.uk/sitemap.xml"
    }
  ]
}' |
apify call neuton/govuk-content-sitemap-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,neuton/govuk-content-sitemap-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/LJFyFJrbbwDQgY9lo/builds/rVqUVLi6lPpde3Wbw/openapi.json
