Website Content Crawler – Markdown for AI, Emails & Contacts
Pricing
from $1.20 / 1,000 results
Website Content Crawler – Markdown for AI, Emails & Contacts
Website content crawler that turns a whole site into clean Markdown for LLM and RAG pipelines, with the role emails, social profiles and WhatsApp numbers found on each page. Seeded from the sitemap, honours robots.txt, plain HTTP with no browser.
Pricing
from $1.20 / 1,000 results
Rating
0.0
(0)
Developer
Locomint
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
This website content crawler starts from a site's homepage, follows its internal links and its sitemap, and returns every page as clean Markdown for LLM and RAG pipelines, with the email addresses, social profiles and WhatsApp numbers found on each page. It uses plain HTTP rather than a browser and honours robots.txt.
What it does
For each start URL the crawler reads the site's robots.txt, collects page URLs from the sitemap, and then works through the site breadth-first:
- Where it goes. Only the host you gave.
www.and the bare domain count as one site, and subdomains are followed only when Include subdomains is on. Links to images, CSS, JavaScript, PDFs, Office documents, archives, audio, video, fonts, CSV, TXT, XML and JSON files are skipped. - The sitemap.
sitemap.xml,sitemap_index.xmland everySitemap:line in robots.txt are read, index files included, up to 20 sitemap files and 5,000 URLs. This finds pages the homepage never links to. - The order. The start page first, then the sitemap URLs in the order the sitemap lists them, then links found on crawled pages, level by level down to Max link depth. The crawl of a site stops at Max pages per site.
- Duplicates. Fragments, trailing slashes and tracking parameters (
utm_*,fbclid,gclid) are removed before URLs are compared, so a page is not fetched twice. Other query strings are kept:?page=2is a page of its own. - robots.txt. URLs the file disallows for all crawlers are dropped before they are fetched.
Every page goes through the same extractor as the Contact Details Scraper, and the row adds where the page sits in the site:
| Field | What goes in it |
|---|---|
content | The page body as Markdown (headings, lists, links, quotes and code blocks kept) or plain text, cut at maxChars. Scripts, styles, forms and the nav, header, footer and aside blocks are removed. |
title, description, language, canonical_url, headings, word_count | Page metadata. language is what <html lang> declares; headings holds up to 50. |
emails | Every address the page publishes: up to five company mailboxes such as info@ and sales@ first, then up to five that name a member of staff. |
named_emails | The addresses in emails that name a person, on the site's own domain or on a free mail provider. Remove them from emails to keep company mailboxes only. |
socials | The first profile link per network: Facebook, Instagram, LinkedIn, X, YouTube and TikTok. linkedin is always the company's own page. |
named_profiles | Up to five personal profile links the page publishes for its staff, such as linkedin.com/in/maria-silva. Only the link is returned; the profile is never opened. |
whatsapp | The number in the first WhatsApp click-to-chat link, as + and digits. |
contact_form_url, tech_stack | The first same-site link whose path contains "contact", and technologies matched by 36 page signatures plus four response headers. |
site, depth, found_on | The site's host, how many links from the start page, and the page (or sitemap) that led to this one. |
Who it is for
- AI and RAG builders. Point it at a documentation site, help centre or knowledge base and
load the Markdown straight into a chunker.
canonical_urlandheadingsgive each chunk a source and a section title, andword_countshows which pages are worth embedding. - SEO audits. One crawl gives every page's
http_status,title,description,canonical_url,headingsandword_count. Filter onhttp_errorto find broken internal links (found_onnames the page that links to each one), onthinfor pages with almost no server-rendered text, and on empty titles or descriptions. - Content migration. Move a site to a new CMS from Markdown that already keeps headings,
lists, links and code blocks, with
found_onanddepthrecording the old structure. - Contact discovery. Contacts are read on every page, not only the homepage, so an address
that appears only on
/aboutor on one branch's page is still found. Crawl a list of company sites with a small page cap to collect each company's published mailboxes, social profiles and WhatsApp number.
How to use it
In the Apify Console:
- Paste the homepage of each site into Websites to crawl, as a full URL with
https://, up to 50 sites. - Set Max pages per site (default 50, up to 2,000) and Max link depth (default 3). The page cap is what controls the cost.
- To crawl one section, put a path such as
/docs/in Only crawl URLs containing. To skip archives, put/tag/or?page=in Skip URLs containing. - Leave Use the sitemap and Respect robots.txt on, choose Markdown or Plain text, and start the run. Export JSON for a RAG pipeline, or CSV and Excel for an audit.
From the API, this call starts a run, waits for it and returns the rows:
curl -X POST \"https://api.apify.com/v2/acts/locomint~website-crawler-content-contacts/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"startUrls": [{"url": "https://docs.python.org/3/"}], "maxPagesPerSite": 20}'
The synchronous endpoint waits up to 300 seconds. For a large crawl, start the run with
POST https://api.apify.com/v2/acts/locomint~website-crawler-content-contacts/runs?token=YOUR_APIFY_TOKEN
and read its dataset when the run finishes. The official apify-client packages for Python
and JavaScript do the same in a few lines.
Input example
{"startUrls": [{ "url": "https://docs.python.org/3/" }],"maxPagesPerSite": 500,"maxDepth": 3,"includeSubdomains": false,"includeUrlPatterns": ["/3/library/"],"excludeUrlPatterns": ["/genindex"],"useSitemap": true,"respectRobots": true,"outputFormat": "markdown","maxChars": 50000,"includeContacts": true,"includeLinks": false,"concurrency": 5}
| Field | Default | Allowed | What it does |
|---|---|---|---|
startUrls | required | 1 to 50 sites | Where each crawl starts. Each site is crawled separately. |
maxPagesPerSite | 50 | 1 to 2,000 | Pages delivered per site, and so the most one site can cost. |
maxDepth | 3 | 0 to 10 | How many links away from the start page to follow. Sitemap pages are crawled at any setting. |
includeSubdomains | false | true / false | Also crawl blog.example.com, shop.example.com and the like. |
includeUrlPatterns | none | list of text | Keep only URLs containing one of these (not case-sensitive). The start page is always crawled. |
excludeUrlPatterns | none | list of text | Skip URLs containing any of these. |
useSitemap | true | true / false | Add the sitemap's URLs to the queue after the start page. |
respectRobots | true | true / false | Drop disallowed URLs before they are fetched. Leave on. |
outputFormat | markdown | markdown, text | Markdown keeps structure and links; text is the words only. |
maxChars | 50,000 | 500 to 500,000 | Longer content is cut, ends with [truncated] and sets truncated: true. |
includeContacts | true | true / false | Emails, socials, WhatsApp, contact page and tech stack from every page. |
includeLinks | false | true / false | Adds links: up to 500 absolute URLs per page. |
concurrency | 5 | 1 to 10 | Pages fetched at the same time within one site. Sites are crawled one after another. |
Output example
A real row from a two-page test crawl of https://docs.python.org/3/ on 11 September 2026
(maxPagesPerSite: 2, maxDepth: 1, sitemap off). This is the second page, reached by a link
on the start page. content is trimmed here; the row held 2,722 characters.
{"url": "https://docs.python.org/3/genindex.html","final_url": "https://docs.python.org/3/genindex.html","status": "ok","http_status": 200,"title": "Index — Python 3.14.7 documentation","description": null,"language": "en","canonical_url": "https://docs.python.org/3/genindex.html","headings": ["Navigation", "Index", "Navigation"],"content": "### Navigation\n\n- [index](https://docs.python.org/3/genindex.html)\n\n- [modules](https://docs.python.org/3/py-modindex.html) |\n\n[...]\n\n# Index\n\nIndex pages by letter:\n\n[Symbols](https://docs.python.org/3/genindex-Symbols.html)\n| [_](https://docs.python.org/3/genindex-_.html)\n| [A](https://docs.python.org/3/genindex-A.html)\n[...]","word_count": 162,"truncated": false,"fetched_at": "2026-09-11T18:16:07+00:00","emails": [],"socials": {},"whatsapp": null,"contact_form_url": null,"tech_stack": [],"site": "docs.python.org","depth": 1,"found_on": "https://docs.python.org/3"}
- The
### Navigationblock is Sphinx's navigation bar, built from plaindivelements rather than anavtag, so the extractor keeps it. It repeats on every page of such a site, so a chunker can strip it by that heading. found_onand each row'surlare in the crawler's normalised form, without a trailing slash or tracking parameters;final_urlis the address the server answered from.noteappears when there is something to say about a fetch: robots.txt disallowed it, robots.txt could not be read, or the page gave no answer within 25 seconds.linksappears with Include outgoing links on.
status is one of ok, thin (under 40 words of server-rendered text), parked,
http_error (4xx or 5xx, with the code in http_status), not_html, unreachable (no
connection or no answer in 25 seconds), blocked (403, 429 or a challenge page),
refused (a private network address) or robots_disallowed. Rows other than ok, thin
and parked carry no content.
Pricing
| Event | Price |
|---|---|
| Page delivered (one dataset row) | $0.0008, which is $0.80 per 1,000 until 26 September 2026, then $0.0015, which is $1.50 per 1,000 |
| Actor start | $0.00005 per GB of run memory, charged once per run |
Worked example: crawl 2,000 pages of a docs site: 2,000 x $0.0008 = $1.60, plus $0.00005 for starting a 1 GB run. Contact discovery across 200 company sites at 10 pages each is also 2,000 pages. From 26 September 2026 the price per page becomes $0.0015.
You pay only these event prices; Apify compute is not billed to you separately, and contacts cost nothing on top of the page. Every page that comes back as a row is one charge, including pages that answer 404, time out or are blocked. URLs dropped before fetching, by robots.txt, your patterns or the file-type filter, produce no row and cost nothing. If you set a maximum cost per run, the crawler stops before the row that would pass it.
FAQ
Does it render JavaScript?
No. It reads the HTML each server sends, which is why a whole site costs one small price per page. Pages that build their content in the browser come back thin, and links that exist
only after JavaScript runs are not followed. Server-rendered sites, which covers most
documentation generators, blogs and CMS sites, crawl fully.
Does it honour robots.txt?
Yes, and turning Respect robots.txt off does not stop that. With the setting on, disallowed
URLs are dropped from the queue before they are fetched and cost nothing. Every page is also
checked again just before its fetch, including rules with * and $ wildcards; a page caught
there is not fetched and comes back as a robots_disallowed row, which counts as a page. With
the setting off, every disallowed page takes that second route, so leave it on.
Can I crawl only one section of a site?
Yes: put the section's path, for example /docs/ or /3/library/, in Only crawl URLs
containing. The start page is always crawled, but other pages outside the pattern are not,
so their links are not followed either. Keep Use the sitemap on, because it reaches deep
pages in the section that the start page does not link to.
How do the sitemap and the depth limit work together?
Sitemap URLs enter the queue at depth 1, straight after the start page, and are crawled whatever Max link depth is set to. With a page cap smaller than the sitemap, the sitemap's own order decides which pages you get. For a crawl of the start page and its immediate neighbourhood only, turn Use the sitemap off.
How do I get one list of contacts per site?
Each page carries the contacts found on that page, so an address in a site-wide footer appears
on every row. Export the dataset and group by site, collecting the unique values of
emails, socials and whatsapp.
Limits
- No JavaScript rendering, and no links followed that only JavaScript creates.
- 50 sites per run, 2,000 pages per site, depth 10, 5,000 sitemap URLs from at most 20 sitemap files, 25 seconds and 3 MB of HTML per page, at most 5 redirects.
- The host you gave only, plus its subdomains if you ask. The crawler never follows links to other sites.
- If a site starts answering with 403 or 429 partway through, the crawler does not retry and
does not switch IP. For the next 15 minutes the pages already queued for that site come back
blockedwithout a request, and each of those rows counts as a page. A modest Max pages per site limits what a site that blocks crawlers can cost you. - Query strings other than tracking parameters make distinct pages, so a filtered catalogue can fill the page cap with near-duplicates. Exclude them with Skip URLs containing.
- Up to five company mailboxes and five named addresses per page. A named address on another company's domain is not returned, and Cloudflare-protected, image and "[at]" addresses are not decoded. Phone numbers come only from WhatsApp click-to-chat links.
- Boilerplate is removed by tag, so menus built from plain
divelements stay incontent.
Compliance
Contact data comes only from what the site publishes. That includes addresses that name a
member of staff and links to staff profiles, which are kept in named_emails and
named_profiles so you can use them or leave them out; a profile page is never opened. A
named work address is personal data in the EU and the UK. You are responsible for using the
data lawfully, including whether you may write to a named person (for example under GDPR,
CAN-SPAM, PECR and local law).
Questions, bug reports and feature requests go on this actor's Issues tab. Anyone can ask for their own address or profile link to be left out, and a business can ask for its whole site to be left out, by writing to info@locomint.io; that address is for removal requests only, and removals are honoured in every Locomint listing. This actor keeps no copy of what it reads between runs, so each run returns what the public pages show at that moment.
Other Locomint actors
- Google Maps Scraper & Email Extractor – Business Leads: Search terms and a city in, business records with website contacts out.
- Google Maps Scraper – Multi-City Lead Lists with Emails: Many categories across many cities in one deduplicated run.
- Google Maps Place Details Scraper – Bulk Place ID Lookup: Place IDs or place-page links in, full records out.
- Website Email Scraper – Contact Details, Socials & WhatsApp: Contact points from website URLs you supply.
- Bulk Email Verifier & Validator: Checks whether addresses can receive mail.
- Company Enrichment API – Domain to Emails, Socials & Tech: A domain in, its contacts and technologies out.
- AI Crawler Checker – robots.txt Rules for GPTBot & ClaudeBot: Which AI crawlers a site's robots.txt allows.
- Schema Markup Validator & Generator – JSON-LD Checker: Checks and generates schema.org markup.