Website Content Crawler – Markdown for AI, Emails & Contacts avatar

Website Content Crawler – Markdown for AI, Emails & Contacts

Pricing

from $1.20 / 1,000 results

Go to Apify Store
Website Content Crawler – Markdown for AI, Emails & Contacts

Website Content Crawler – Markdown for AI, Emails & Contacts

Website content crawler that turns a whole site into clean Markdown for LLM and RAG pipelines, with the role emails, social profiles and WhatsApp numbers found on each page. Seeded from the sitemap, honours robots.txt, plain HTTP with no browser.

Pricing

from $1.20 / 1,000 results

Rating

0.0

(0)

Developer

Locomint

Locomint

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Share

This website content crawler starts from a site's homepage, follows its internal links and its sitemap, and returns every page as clean Markdown for LLM and RAG pipelines, with the email addresses, social profiles and WhatsApp numbers found on each page. It uses plain HTTP rather than a browser and honours robots.txt.

What it does

For each start URL the crawler reads the site's robots.txt, collects page URLs from the sitemap, and then works through the site breadth-first:

  • Where it goes. Only the host you gave. www. and the bare domain count as one site, and subdomains are followed only when Include subdomains is on. Links to images, CSS, JavaScript, PDFs, Office documents, archives, audio, video, fonts, CSV, TXT, XML and JSON files are skipped.
  • The sitemap. sitemap.xml, sitemap_index.xml and every Sitemap: line in robots.txt are read, index files included, up to 20 sitemap files and 5,000 URLs. This finds pages the homepage never links to.
  • The order. The start page first, then the sitemap URLs in the order the sitemap lists them, then links found on crawled pages, level by level down to Max link depth. The crawl of a site stops at Max pages per site.
  • Duplicates. Fragments, trailing slashes and tracking parameters (utm_*, fbclid, gclid) are removed before URLs are compared, so a page is not fetched twice. Other query strings are kept: ?page=2 is a page of its own.
  • robots.txt. URLs the file disallows for all crawlers are dropped before they are fetched.

Every page goes through the same extractor as the Contact Details Scraper, and the row adds where the page sits in the site:

FieldWhat goes in it
contentThe page body as Markdown (headings, lists, links, quotes and code blocks kept) or plain text, cut at maxChars. Scripts, styles, forms and the nav, header, footer and aside blocks are removed.
title, description, language, canonical_url, headings, word_countPage metadata. language is what <html lang> declares; headings holds up to 50.
emailsEvery address the page publishes: up to five company mailboxes such as info@ and sales@ first, then up to five that name a member of staff.
named_emailsThe addresses in emails that name a person, on the site's own domain or on a free mail provider. Remove them from emails to keep company mailboxes only.
socialsThe first profile link per network: Facebook, Instagram, LinkedIn, X, YouTube and TikTok. linkedin is always the company's own page.
named_profilesUp to five personal profile links the page publishes for its staff, such as linkedin.com/in/maria-silva. Only the link is returned; the profile is never opened.
whatsappThe number in the first WhatsApp click-to-chat link, as + and digits.
contact_form_url, tech_stackThe first same-site link whose path contains "contact", and technologies matched by 36 page signatures plus four response headers.
site, depth, found_onThe site's host, how many links from the start page, and the page (or sitemap) that led to this one.

Who it is for

  • AI and RAG builders. Point it at a documentation site, help centre or knowledge base and load the Markdown straight into a chunker. canonical_url and headings give each chunk a source and a section title, and word_count shows which pages are worth embedding.
  • SEO audits. One crawl gives every page's http_status, title, description, canonical_url, headings and word_count. Filter on http_error to find broken internal links (found_on names the page that links to each one), on thin for pages with almost no server-rendered text, and on empty titles or descriptions.
  • Content migration. Move a site to a new CMS from Markdown that already keeps headings, lists, links and code blocks, with found_on and depth recording the old structure.
  • Contact discovery. Contacts are read on every page, not only the homepage, so an address that appears only on /about or on one branch's page is still found. Crawl a list of company sites with a small page cap to collect each company's published mailboxes, social profiles and WhatsApp number.

How to use it

In the Apify Console:

  1. Paste the homepage of each site into Websites to crawl, as a full URL with https://, up to 50 sites.
  2. Set Max pages per site (default 50, up to 2,000) and Max link depth (default 3). The page cap is what controls the cost.
  3. To crawl one section, put a path such as /docs/ in Only crawl URLs containing. To skip archives, put /tag/ or ?page= in Skip URLs containing.
  4. Leave Use the sitemap and Respect robots.txt on, choose Markdown or Plain text, and start the run. Export JSON for a RAG pipeline, or CSV and Excel for an audit.

From the API, this call starts a run, waits for it and returns the rows:

curl -X POST \
"https://api.apify.com/v2/acts/locomint~website-crawler-content-contacts/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"startUrls": [{"url": "https://docs.python.org/3/"}], "maxPagesPerSite": 20}'

The synchronous endpoint waits up to 300 seconds. For a large crawl, start the run with POST https://api.apify.com/v2/acts/locomint~website-crawler-content-contacts/runs?token=YOUR_APIFY_TOKEN and read its dataset when the run finishes. The official apify-client packages for Python and JavaScript do the same in a few lines.

Input example

{
"startUrls": [{ "url": "https://docs.python.org/3/" }],
"maxPagesPerSite": 500,
"maxDepth": 3,
"includeSubdomains": false,
"includeUrlPatterns": ["/3/library/"],
"excludeUrlPatterns": ["/genindex"],
"useSitemap": true,
"respectRobots": true,
"outputFormat": "markdown",
"maxChars": 50000,
"includeContacts": true,
"includeLinks": false,
"concurrency": 5
}
FieldDefaultAllowedWhat it does
startUrlsrequired1 to 50 sitesWhere each crawl starts. Each site is crawled separately.
maxPagesPerSite501 to 2,000Pages delivered per site, and so the most one site can cost.
maxDepth30 to 10How many links away from the start page to follow. Sitemap pages are crawled at any setting.
includeSubdomainsfalsetrue / falseAlso crawl blog.example.com, shop.example.com and the like.
includeUrlPatternsnonelist of textKeep only URLs containing one of these (not case-sensitive). The start page is always crawled.
excludeUrlPatternsnonelist of textSkip URLs containing any of these.
useSitemaptruetrue / falseAdd the sitemap's URLs to the queue after the start page.
respectRobotstruetrue / falseDrop disallowed URLs before they are fetched. Leave on.
outputFormatmarkdownmarkdown, textMarkdown keeps structure and links; text is the words only.
maxChars50,000500 to 500,000Longer content is cut, ends with [truncated] and sets truncated: true.
includeContactstruetrue / falseEmails, socials, WhatsApp, contact page and tech stack from every page.
includeLinksfalsetrue / falseAdds links: up to 500 absolute URLs per page.
concurrency51 to 10Pages fetched at the same time within one site. Sites are crawled one after another.

Output example

A real row from a two-page test crawl of https://docs.python.org/3/ on 11 September 2026 (maxPagesPerSite: 2, maxDepth: 1, sitemap off). This is the second page, reached by a link on the start page. content is trimmed here; the row held 2,722 characters.

{
"url": "https://docs.python.org/3/genindex.html",
"final_url": "https://docs.python.org/3/genindex.html",
"status": "ok",
"http_status": 200,
"title": "Index — Python 3.14.7 documentation",
"description": null,
"language": "en",
"canonical_url": "https://docs.python.org/3/genindex.html",
"headings": ["Navigation", "Index", "Navigation"],
"content": "### Navigation\n\n- [index](https://docs.python.org/3/genindex.html)\n\n- [modules](https://docs.python.org/3/py-modindex.html) |\n\n[...]\n\n# Index\n\nIndex pages by letter:\n\n[Symbols](https://docs.python.org/3/genindex-Symbols.html)\n| [_](https://docs.python.org/3/genindex-_.html)\n| [A](https://docs.python.org/3/genindex-A.html)\n[...]",
"word_count": 162,
"truncated": false,
"fetched_at": "2026-09-11T18:16:07+00:00",
"emails": [],
"socials": {},
"whatsapp": null,
"contact_form_url": null,
"tech_stack": [],
"site": "docs.python.org",
"depth": 1,
"found_on": "https://docs.python.org/3"
}
  • The ### Navigation block is Sphinx's navigation bar, built from plain div elements rather than a nav tag, so the extractor keeps it. It repeats on every page of such a site, so a chunker can strip it by that heading.
  • found_on and each row's url are in the crawler's normalised form, without a trailing slash or tracking parameters; final_url is the address the server answered from.
  • note appears when there is something to say about a fetch: robots.txt disallowed it, robots.txt could not be read, or the page gave no answer within 25 seconds. links appears with Include outgoing links on.

status is one of ok, thin (under 40 words of server-rendered text), parked, http_error (4xx or 5xx, with the code in http_status), not_html, unreachable (no connection or no answer in 25 seconds), blocked (403, 429 or a challenge page), refused (a private network address) or robots_disallowed. Rows other than ok, thin and parked carry no content.

Pricing

EventPrice
Page delivered (one dataset row)$0.0008, which is $0.80 per 1,000 until 26 September 2026, then $0.0015, which is $1.50 per 1,000
Actor start$0.00005 per GB of run memory, charged once per run

Worked example: crawl 2,000 pages of a docs site: 2,000 x $0.0008 = $1.60, plus $0.00005 for starting a 1 GB run. Contact discovery across 200 company sites at 10 pages each is also 2,000 pages. From 26 September 2026 the price per page becomes $0.0015.

You pay only these event prices; Apify compute is not billed to you separately, and contacts cost nothing on top of the page. Every page that comes back as a row is one charge, including pages that answer 404, time out or are blocked. URLs dropped before fetching, by robots.txt, your patterns or the file-type filter, produce no row and cost nothing. If you set a maximum cost per run, the crawler stops before the row that would pass it.

FAQ

Does it render JavaScript?

No. It reads the HTML each server sends, which is why a whole site costs one small price per page. Pages that build their content in the browser come back thin, and links that exist only after JavaScript runs are not followed. Server-rendered sites, which covers most documentation generators, blogs and CMS sites, crawl fully.

Does it honour robots.txt?

Yes, and turning Respect robots.txt off does not stop that. With the setting on, disallowed URLs are dropped from the queue before they are fetched and cost nothing. Every page is also checked again just before its fetch, including rules with * and $ wildcards; a page caught there is not fetched and comes back as a robots_disallowed row, which counts as a page. With the setting off, every disallowed page takes that second route, so leave it on.

Can I crawl only one section of a site?

Yes: put the section's path, for example /docs/ or /3/library/, in Only crawl URLs containing. The start page is always crawled, but other pages outside the pattern are not, so their links are not followed either. Keep Use the sitemap on, because it reaches deep pages in the section that the start page does not link to.

How do the sitemap and the depth limit work together?

Sitemap URLs enter the queue at depth 1, straight after the start page, and are crawled whatever Max link depth is set to. With a page cap smaller than the sitemap, the sitemap's own order decides which pages you get. For a crawl of the start page and its immediate neighbourhood only, turn Use the sitemap off.

How do I get one list of contacts per site?

Each page carries the contacts found on that page, so an address in a site-wide footer appears on every row. Export the dataset and group by site, collecting the unique values of emails, socials and whatsapp.

Limits

  • No JavaScript rendering, and no links followed that only JavaScript creates.
  • 50 sites per run, 2,000 pages per site, depth 10, 5,000 sitemap URLs from at most 20 sitemap files, 25 seconds and 3 MB of HTML per page, at most 5 redirects.
  • The host you gave only, plus its subdomains if you ask. The crawler never follows links to other sites.
  • If a site starts answering with 403 or 429 partway through, the crawler does not retry and does not switch IP. For the next 15 minutes the pages already queued for that site come back blocked without a request, and each of those rows counts as a page. A modest Max pages per site limits what a site that blocks crawlers can cost you.
  • Query strings other than tracking parameters make distinct pages, so a filtered catalogue can fill the page cap with near-duplicates. Exclude them with Skip URLs containing.
  • Up to five company mailboxes and five named addresses per page. A named address on another company's domain is not returned, and Cloudflare-protected, image and "[at]" addresses are not decoded. Phone numbers come only from WhatsApp click-to-chat links.
  • Boilerplate is removed by tag, so menus built from plain div elements stay in content.

Compliance

Contact data comes only from what the site publishes. That includes addresses that name a member of staff and links to staff profiles, which are kept in named_emails and named_profiles so you can use them or leave them out; a profile page is never opened. A named work address is personal data in the EU and the UK. You are responsible for using the data lawfully, including whether you may write to a named person (for example under GDPR, CAN-SPAM, PECR and local law).

Questions, bug reports and feature requests go on this actor's Issues tab. Anyone can ask for their own address or profile link to be left out, and a business can ask for its whole site to be left out, by writing to info@locomint.io; that address is for removal requests only, and removals are honoured in every Locomint listing. This actor keeps no copy of what it reads between runs, so each run returns what the public pages show at that moment.

Other Locomint actors