Ai Web Intelligence avatar

Ai Web Intelligence

Pricing

from $2.00 / 1,000 http pages

Go to Apify Store
Ai Web Intelligence

Ai Web Intelligence

AI web crawler and scraper for clean Markdown, text, metadata, JSON-LD, PDFs, RAG chunks, screenshots, and structured web data. HTTP-first crawling with automatic browser fallback, built for AI agents, RAG pipelines, research, and knowledge bases.

Pricing

from $2.00 / 1,000 http pages

Rating

0.0

(0)

Developer

Hakashi Katake

Hakashi Katake

Maintained by Community

Actor stats

1

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

Crawl public websites into clean, structured, AI-ready content for LLMs, RAG pipelines, vector databases, AI agents, search systems, monitoring and content analysis.

Why it exists

Website crawlers often force a choice between expensive browser rendering and incomplete raw HTML. AI Web Intelligence starts with fast HTTP extraction, uses bounded Playwright fallback only when a page looks client-rendered, and emits a consistent record shape with content hashes and optional RAG chunks.

Current features

  • Seed Discovery: Automated XML sitemap parsing (sitemap.xml & robots.txt Sitemap directives) and emerging standard llms.txt/llms-full.txt discovery.
  • Hybrid Crawling: High-speed HTTP-first Cheerio crawler with bounded Playwright browser fallback for client-rendered Single-Page Applications (SPAs).
  • Clean Content Extraction: Mozilla Readability article parsing, boilerplate stripping, and clean ATX Markdown and plain text generation.
  • Deep Content Intelligence: Author heading hierarchy (H1-H6), reading time estimation, and content classification (article, documentation, product, faq, etc.).
  • Rich Structured Data: OpenGraph (og:*), Twitter Cards, JSON-LD schema.org entities, breadcrumb lists, and FAQ question-answer pairs.
  • Native Document Processing: High-performance PDF text, author, title, and page count extraction directly in the HTTP stream without browser bloat.
  • Screenshots & Media Optimization: Optional Playwright full-page screenshots stored in Key-Value Store (SCREENSHOT_{hash}.png) with URLs in dataset records; network route aborts for images/media and ad trackers.
  • Privacy & Safety: Zero-cost regex engine for PII redaction (email, phone, SSN) and hardened SSRF protection (RFC 1918, CGNAT, AWS/GCP instance metadata, decimal IPs, and IPv6).
  • RAG-Ready Pre-Chunking: Exact o200k_base token counting, deterministic chunk IDs ({hash}-c{index}), character offsets, and hierarchical heading paths.
  • Section-Level Change Detection: Cross-run snapshot comparison tracking NEW, MODIFIED, REMOVED, UNCHANGED states, character diff percentages, added/removed text snippets, and modified sections.
  • Run Accounting: Generates comprehensive compute unit and USD cost estimates in the default Key-Value Store record OUTPUT.

Quick start input

{
"startUrls": [{ "url": "https://docs.apify.com/" }],
"maxPages": 50,
"maxDepth": 2,
"discoverSitemaps": true,
"useLlmsTxt": true,
"renderingMode": "auto",
"respectRobotsTxt": true,
"extractPdfText": true,
"includeMarkdown": true,
"includeStructuredData": true,
"enableChunking": true,
"chunkSize": 512,
"chunkOverlap": 50,
"enableChangeDetection": false
}

For scheduled snapshots and change diffing, pass a persistent named snapshotStoreName and set enableChangeDetection to true.

Output example

{
"url": "https://example.com/docs/guide",
"finalUrl": "https://example.com/docs/guide",
"statusCode": 200,
"title": "Getting Started Guide",
"metadata": {
"description": "Comprehensive developer guide",
"author": "Apify Team",
"language": "en",
"canonicalUrl": "https://example.com/docs/guide",
"publishedAt": "2026-01-15T00:00:00.000Z",
"modifiedAt": "2026-03-20T00:00:00.000Z"
},
"structuredData": {
"openGraph": { "title": "Guide", "type": "article", "image": "https://example.com/og.jpg" },
"twitter": { "card": "summary_large_image", "title": "Guide" },
"breadcrumbs": [{ "position": 1, "name": "Docs", "url": "https://example.com/docs" }],
"faqs": [{ "question": "How do I install?", "answer": "Run npm install" }],
"schemaEntities": [{ "type": "TechArticle", "name": "Getting Started" }]
},
"contentIntelligence": {
"headings": [{ "level": 1, "text": "Introduction", "id": "intro" }],
"readingTimeMinutes": 3,
"contentType": "documentation"
},
"markdown": "# Introduction\n\nWelcome to the guide...",
"text": "Introduction Welcome to the guide...",
"tokens": { "count": 245, "encoding": "o200k_base" },
"chunks": [
{
"id": "e3b0c44298fc1c14-c0",
"index": 0,
"text": "Introduction\n\nWelcome...",
"tokens": 245,
"sourceUrl": "https://example.com/docs/guide",
"headingPath": ["Introduction"],
"characterStart": 0,
"characterEnd": 1200
}
],
"contentHash": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"crawl": { "depth": 1, "durationMs": 420, "crawledAt": "2026-09-28T12:00:00.000Z", "rendering": "http" },
"changeDetection": {
"status": "MODIFIED",
"changed": true,
"changePercentage": 12.5,
"addedText": "newly added section...",
"removedText": null,
"modifiedSections": ["Installation"],
"contentVersion": 2
},
"errors": []
}

API and scheduling

The Actor uses the standard Apify Run API and Dataset API. Create a schedule in Apify Console for recurring crawls; named snapshots allow the Actor to compare runs. Webhooks are supported by Apify’s platform, but this project does not create webhooks automatically.

Cost and performance

The Actor uses no custom paid events and does not enable proxies by default, so the default cost model is Apify platform usage. Raw HTTP is materially cheaper/faster than browser rendering; Playwright is used only when selected or when auto fallback detects a likely client-rendered shell. See docs/BENCHMARKS.md for measured results and limitations.

Local development

npm install
npm run typecheck
npm test
npm run build

Run the Actor with the local Apify CLI:

$./node_modules/.bin/apify run --input-file storage/key_value_stores/default/INPUT.json

Run the five-site benchmark:

$npm run benchmark

The opt-in live integration tests are intentionally small and should be added as the external test suite grows:

$npm run test:integration

Research and architecture

  • docs/COMPETITOR_RESEARCH.md
  • docs/APIFY_API_RESEARCH.md
  • docs/CRAWLER_ARCHITECTURE.md
  • docs/PRODUCT_GAP.md
  • docs/BENCHMARKS.md

This Actor is intended for publicly accessible web content. It does not bypass authentication, paywalls, CAPTCHAs or access controls. It validates URLs, blocks common private/local targets, respects robots.txt by default, bounds response sizes and continues with structured errors when individual pages fail.