Website RAG & Lead Intelligence Crawler avatar

Website RAG & Lead Intelligence Crawler

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Website RAG & Lead Intelligence Crawler

Website RAG & Lead Intelligence Crawler

Convert public websites to clean Markdown for RAG and AI. Crawl pages and extract emails, phones, social profiles, metadata, technologies, headings and links.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Samuel Huirau Atutahi

Samuel Huirau Atutahi

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Turn public websites into structured, AI-ready data.

This Actor crawls public website pages and produces clean Markdown together with useful business and technical intelligence.

What it extracts

For every successfully crawled page:

  • Clean Markdown for RAG, LLM, and knowledge-base workflows
  • Page title and meta description
  • Canonical URL
  • H1/H2/H3 headings
  • Publicly displayed email addresses
  • Publicly displayed phone numbers
  • Linked social profiles
  • Basic website technology detection
  • Internal and external link counts
  • HTTP status
  • Crawl depth
  • Word and character counts
  • Fetch timestamp

Useful for

RAG and AI agents

Convert website content into compact Markdown suitable for embeddings, retrieval systems, AI agents, summarization, and knowledge bases.

Business research

Collect publicly displayed business contact information, website metadata, social links, and technology signals alongside page content.

Website intelligence

Discover page structure, technologies, canonical URLs, headings, and internal-link relationships.

Data pipelines

Export results as JSON, CSV, Excel, XML, JSONL, or retrieve them through the Apify API.

Crawl controls

Configure:

  • Multiple start URLs
  • Maximum page count
  • Crawl depth
  • Same-domain restriction
  • robots.txt compliance
  • Include URL patterns
  • Exclude URL patterns
  • Concurrency
  • Request timeout
  • Maximum Markdown size per page

Output

Each dataset item represents one page.

Typical fields include:

url, final_url, title, description, markdown, emails, phones, social_profiles, technologies, headings, word_count, internal_link_count, and external_link_count.

A run-level SUMMARY record also contains aggregate crawl statistics and unique discovered contacts and technologies.

Responsible use

This Actor is intended for publicly accessible website content.

respect_robots is enabled by default. Users are responsible for ensuring their use complies with applicable website terms, permissions, privacy rules, and laws.

The Actor does not log into websites, bypass access controls, solve CAPTCHAs, or access private content.

Cost-efficient architecture

The Actor uses lightweight HTTP requests rather than a browser by default, keeping runs fast and inexpensive.

DataVault Labs

Built by DataVault Labs for practical automation, structured data, RAG, and AI-agent workflows.