# Ai Web Intelligence (`hakashi_katake/ai-web-intelligence`) Actor

AI web crawler and scraper for clean Markdown, text, metadata, JSON-LD, PDFs, RAG chunks, screenshots, and structured web data. HTTP-first crawling with automatic browser fallback, built for AI agents, RAG pipelines, research, and knowledge bases.

- **URL**: https://apify.com/hakashi\_katake/ai-web-intelligence.md
- **Developed by:** [Hakashi Katake](https://apify.com/hakashi_katake) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 http pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## AI Web Intelligence

Crawl public websites into clean, structured, AI-ready content for LLMs, RAG pipelines, vector databases, AI agents, search systems, monitoring and content analysis.

### Why it exists

Website crawlers often force a choice between expensive browser rendering and incomplete raw HTML. AI Web Intelligence starts with fast HTTP extraction, uses bounded Playwright fallback only when a page looks client-rendered, and emits a consistent record shape with content hashes and optional RAG chunks.

### Current features

- **Seed Discovery**: Automated XML sitemap parsing (`sitemap.xml` & `robots.txt` Sitemap directives) and emerging standard `llms.txt`/`llms-full.txt` discovery.
- **Hybrid Crawling**: High-speed HTTP-first Cheerio crawler with bounded Playwright browser fallback for client-rendered Single-Page Applications (SPAs).
- **Clean Content Extraction**: Mozilla Readability article parsing, boilerplate stripping, and clean ATX Markdown and plain text generation.
- **Deep Content Intelligence**: Author heading hierarchy (`H1-H6`), reading time estimation, and content classification (`article`, `documentation`, `product`, `faq`, etc.).
- **Rich Structured Data**: OpenGraph (`og:*`), Twitter Cards, JSON-LD schema.org entities, breadcrumb lists, and FAQ question-answer pairs.
- **Native Document Processing**: High-performance PDF text, author, title, and page count extraction directly in the HTTP stream without browser bloat.
- **Screenshots & Media Optimization**: Optional Playwright full-page screenshots stored in Key-Value Store (`SCREENSHOT_{hash}.png`) with URLs in dataset records; network route aborts for images/media and ad trackers.
- **Privacy & Safety**: Zero-cost regex engine for PII redaction (email, phone, SSN) and hardened SSRF protection (RFC 1918, CGNAT, AWS/GCP instance metadata, decimal IPs, and IPv6).
- **RAG-Ready Pre-Chunking**: Exact `o200k_base` token counting, deterministic chunk IDs (`{hash}-c{index}`), character offsets, and hierarchical heading paths.
- **Section-Level Change Detection**: Cross-run snapshot comparison tracking `NEW`, `MODIFIED`, `REMOVED`, `UNCHANGED` states, character diff percentages, added/removed text snippets, and modified sections.
- **Run Accounting**: Generates comprehensive compute unit and USD cost estimates in the default Key-Value Store record `OUTPUT`.

### Quick start input

```json
{
  "startUrls": [{ "url": "https://docs.apify.com/" }],
  "maxPages": 50,
  "maxDepth": 2,
  "discoverSitemaps": true,
  "useLlmsTxt": true,
  "renderingMode": "auto",
  "respectRobotsTxt": true,
  "extractPdfText": true,
  "includeMarkdown": true,
  "includeStructuredData": true,
  "enableChunking": true,
  "chunkSize": 512,
  "chunkOverlap": 50,
  "enableChangeDetection": false
}
```

For scheduled snapshots and change diffing, pass a persistent named `snapshotStoreName` and set `enableChangeDetection` to `true`.

### Output example

```json
{
  "url": "https://example.com/docs/guide",
  "finalUrl": "https://example.com/docs/guide",
  "statusCode": 200,
  "title": "Getting Started Guide",
  "metadata": {
    "description": "Comprehensive developer guide",
    "author": "Apify Team",
    "language": "en",
    "canonicalUrl": "https://example.com/docs/guide",
    "publishedAt": "2026-01-15T00:00:00.000Z",
    "modifiedAt": "2026-03-20T00:00:00.000Z"
  },
  "structuredData": {
    "openGraph": { "title": "Guide", "type": "article", "image": "https://example.com/og.jpg" },
    "twitter": { "card": "summary_large_image", "title": "Guide" },
    "breadcrumbs": [{ "position": 1, "name": "Docs", "url": "https://example.com/docs" }],
    "faqs": [{ "question": "How do I install?", "answer": "Run npm install" }],
    "schemaEntities": [{ "type": "TechArticle", "name": "Getting Started" }]
  },
  "contentIntelligence": {
    "headings": [{ "level": 1, "text": "Introduction", "id": "intro" }],
    "readingTimeMinutes": 3,
    "contentType": "documentation"
  },
  "markdown": "# Introduction\n\nWelcome to the guide...",
  "text": "Introduction Welcome to the guide...",
  "tokens": { "count": 245, "encoding": "o200k_base" },
  "chunks": [
    {
      "id": "e3b0c44298fc1c14-c0",
      "index": 0,
      "text": "Introduction\n\nWelcome...",
      "tokens": 245,
      "sourceUrl": "https://example.com/docs/guide",
      "headingPath": ["Introduction"],
      "characterStart": 0,
      "characterEnd": 1200
    }
  ],
  "contentHash": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
  "crawl": { "depth": 1, "durationMs": 420, "crawledAt": "2026-09-28T12:00:00.000Z", "rendering": "http" },
  "changeDetection": {
    "status": "MODIFIED",
    "changed": true,
    "changePercentage": 12.5,
    "addedText": "newly added section...",
    "removedText": null,
    "modifiedSections": ["Installation"],
    "contentVersion": 2
  },
  "errors": []
}
```

### API and scheduling

The Actor uses the standard Apify Run API and Dataset API. Create a schedule in Apify Console for recurring crawls; named snapshots allow the Actor to compare runs. Webhooks are supported by Apify’s platform, but this project does not create webhooks automatically.

### Cost and performance

The Actor uses no custom paid events and does not enable proxies by default, so the default cost model is Apify platform usage. Raw HTTP is materially cheaper/faster than browser rendering; Playwright is used only when selected or when auto fallback detects a likely client-rendered shell. See [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) for measured results and limitations.

### Local development

```bash
npm install
npm run typecheck
npm test
npm run build
```

Run the Actor with the local Apify CLI:

```bash
./node_modules/.bin/apify run --input-file storage/key_value_stores/default/INPUT.json
```

Run the five-site benchmark:

```bash
npm run benchmark
```

The opt-in live integration tests are intentionally small and should be added as the external test suite grows:

```bash
npm run test:integration
```

### Research and architecture

- [`docs/COMPETITOR_RESEARCH.md`](docs/COMPETITOR_RESEARCH.md)
- [`docs/APIFY_API_RESEARCH.md`](docs/APIFY_API_RESEARCH.md)
- [`docs/CRAWLER_ARCHITECTURE.md`](docs/CRAWLER_ARCHITECTURE.md)
- [`docs/PRODUCT_GAP.md`](docs/PRODUCT_GAP.md)
- [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md)

### Legal and safety boundaries

This Actor is intended for publicly accessible web content. It does not bypass authentication, paywalls, CAPTCHAs or access controls. It validates URLs, blocks common private/local targets, respects robots.txt by default, bounds response sizes and continues with structured errors when individual pages fail.

# Actor input Schema

## `startUrls` (type: `array`):

One or more public HTTP(S) URLs where crawling starts.

## `maxPages` (type: `integer`):

Hard cap on stored page records per run.

## `maxDepth` (type: `integer`):

Link distance from each start URL. Zero crawls only start URLs.

## `includeSubdomains` (type: `boolean`):

Allow links on subdomains of the start URL’s registrable domain.

## `includeUrlGlobs` (type: `array`):

Optional glob patterns that links must match.

## `excludeUrlGlobs` (type: `array`):

Optional glob patterns for URLs that must not be crawled.

## `discoverSitemaps` (type: `boolean`):

Read sitemap locations from robots.txt and common sitemap paths, then seed the crawl.

## `useLlmsTxt` (type: `boolean`):

Read root llms.txt files and enqueue linked public documentation URLs.

## `renderingMode` (type: `string`):

Auto uses HTTP first and bounded browser fallback; never disables browser fallback; always renders every page in Playwright.

## `maxBrowserPages` (type: `integer`):

Maximum pages that may be retried through Playwright in auto mode.

## `includeHtml` (type: `boolean`):

Include cleaned readable HTML. Disabled by default to reduce Dataset size.

## `includeText` (type: `boolean`):

Include normalized plain text extracted from the readable page content.

## `includeMarkdown` (type: `boolean`):

Include structure-preserving Markdown extracted from the readable page content.

## `includeMetadata` (type: `boolean`):

Include title, description, author, language, canonical and date metadata.

## `includeJsonLd` (type: `boolean`):

Parse valid JSON-LD script blocks when present.

## `includeLinks` (type: `boolean`):

Include normalized internal and external links found in readable content.

## `includeImages` (type: `boolean`):

Include image URLs, alt text, title, width and height without downloading binaries.

## `includeScreenshots` (type: `boolean`):

Save optional Playwright screenshots to Key-Value Store and return public URLs.

## `screenshotFullPage` (type: `boolean`):

Capture full-page screenshots instead of the viewport.

## `blockMedia` (type: `boolean`):

Abort image, font and media subrequests during browser rendering.

## `blockAds` (type: `boolean`):

Abort common advertising and analytics resource domains in browser mode.

## `redactPii` (type: `boolean`):

Best-effort redaction of email, phone-like and SSN-shaped values in content and chunks.

## `extractPdfText` (type: `boolean`):

Extract text and basic metadata from PDF responses within the size limit.

## `maxPdfBytes` (type: `integer`):

Hard limit for PDFs processed by the PDF parser.

## `enableChunking` (type: `boolean`):

Create token-based chunks from extracted text/Markdown.

## `chunkSize` (type: `integer`):

Target token count for optional RAG chunks.

## `chunkOverlap` (type: `integer`):

Token overlap between adjacent optional RAG chunks.

## `respectRobotsTxt` (type: `boolean`):

Skip URLs disallowed by robots.txt. Enabled by default.

## `sameDomainDelaySecs` (type: `number`):

Minimum delay between requests to the same domain.

## `requestTimeoutSecs` (type: `integer`):

Maximum time allowed for one page request and extraction.

## `maxRequestRetries` (type: `integer`):

Bounded retries for transient HTTP or browser failures.

## `maxResponseBytes` (type: `integer`):

Reject responses larger than this limit to prevent runaway memory/storage use.

## `maxConcurrency` (type: `integer`):

Maximum number of HTTP pages processed in parallel.

## `snapshotStoreName` (type: `string`):

Optional named Key-Value Store used for cross-run change detection.

## `enableChangeDetection` (type: `boolean`):

Compare content hashes with a prior snapshot and emit NEW/MODIFIED/REMOVED/UNCHANGED states.

## Actor input object example

```json
{
  "startUrls": [],
  "maxPages": 100,
  "maxDepth": 3,
  "includeSubdomains": false,
  "includeUrlGlobs": [],
  "excludeUrlGlobs": [],
  "discoverSitemaps": false,
  "useLlmsTxt": false,
  "renderingMode": "auto",
  "maxBrowserPages": 10,
  "includeHtml": false,
  "includeText": true,
  "includeMarkdown": true,
  "includeMetadata": true,
  "includeJsonLd": true,
  "includeLinks": true,
  "includeImages": true,
  "includeScreenshots": false,
  "screenshotFullPage": false,
  "blockMedia": true,
  "blockAds": true,
  "redactPii": false,
  "extractPdfText": true,
  "maxPdfBytes": 10000000,
  "enableChunking": false,
  "chunkSize": 512,
  "chunkOverlap": 50,
  "respectRobotsTxt": true,
  "sameDomainDelaySecs": 0.25,
  "requestTimeoutSecs": 30,
  "maxRequestRetries": 2,
  "maxResponseBytes": 5000000,
  "maxConcurrency": 10,
  "snapshotStoreName": "",
  "enableChangeDetection": false
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("hakashi_katake/ai-web-intelligence").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("hakashi_katake/ai-web-intelligence").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call hakashi_katake/ai-web-intelligence --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,hakashi_katake/ai-web-intelligence"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Iorh2swoJF0btkzFf/builds/RgcF7n4H7fCVsCezq/openapi.json
