Website RAG & Lead Intelligence Crawler
Pricing
from $2.00 / 1,000 results
Website RAG & Lead Intelligence Crawler
Convert public websites to clean Markdown for RAG and AI. Crawl pages and extract emails, phones, social profiles, metadata, technologies, headings and links.
Pricing
from $2.00 / 1,000 results
Rating
0.0
(0)
Developer
Samuel Huirau Atutahi
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Turn public websites into structured, AI-ready data.
This Actor crawls public website pages and produces clean Markdown together with useful business and technical intelligence.
What it extracts
For every successfully crawled page:
- Clean Markdown for RAG, LLM, and knowledge-base workflows
- Page title and meta description
- Canonical URL
- H1/H2/H3 headings
- Publicly displayed email addresses
- Publicly displayed phone numbers
- Linked social profiles
- Basic website technology detection
- Internal and external link counts
- HTTP status
- Crawl depth
- Word and character counts
- Fetch timestamp
Useful for
RAG and AI agents
Convert website content into compact Markdown suitable for embeddings, retrieval systems, AI agents, summarization, and knowledge bases.
Business research
Collect publicly displayed business contact information, website metadata, social links, and technology signals alongside page content.
Website intelligence
Discover page structure, technologies, canonical URLs, headings, and internal-link relationships.
Data pipelines
Export results as JSON, CSV, Excel, XML, JSONL, or retrieve them through the Apify API.
Crawl controls
Configure:
- Multiple start URLs
- Maximum page count
- Crawl depth
- Same-domain restriction
- robots.txt compliance
- Include URL patterns
- Exclude URL patterns
- Concurrency
- Request timeout
- Maximum Markdown size per page
Output
Each dataset item represents one page.
Typical fields include:
url, final_url, title, description, markdown, emails, phones, social_profiles, technologies, headings, word_count, internal_link_count, and external_link_count.
A run-level SUMMARY record also contains aggregate crawl statistics and unique discovered contacts and technologies.
Responsible use
This Actor is intended for publicly accessible website content.
respect_robots is enabled by default. Users are responsible for ensuring their use complies with applicable website terms, permissions, privacy rules, and laws.
The Actor does not log into websites, bypass access controls, solve CAPTCHAs, or access private content.
Cost-efficient architecture
The Actor uses lightweight HTTP requests rather than a browser by default, keeping runs fast and inexpensive.
DataVault Labs
Built by DataVault Labs for practical automation, structured data, RAG, and AI-agent workflows.