No credit card required

Website Content Crawler

apify/website-content-crawler

No credit card required

Crawl websites and extract text content to feed AI models, LLM applications, vector databases, or RAG pipelines. The Actor supports rich formatting using Markdown, cleans the HTML, downloads files, and integrates well with 🦜🔗 LangChain, LlamaIndex, and the wider LLM ecosystem.

Back to issues Create new issue

crawling takes longer

Closed

yener.yasin030 opened this issue

Crawling crawles unimportant things. and takes more then a hour. its eat my money iwasnt on the pc.

Jiří Spilka (jiri.spilka)

Hi, thank you for using the Website Content Crawler.

However, I couldn’t find the runId you attached.
I reviewed your runs, and the longest one took ~20 minutes. While it was a bit slow, scraping pages like news.google can be challenging.

Could you please provide the runId where you encountered the issue?

Thank you, Jiri

yener.yasin030

I have deleted it falsewise the last run cost me 12 dolar something like that...

22 Oca 2025 Çar 12:56 tarihinde Jiří Spilka notifications@apify.com şunu yazdı:

Jiří Spilka (jiri.spilka)

I found the runId: rBiWKBvKfnA6T6bgh in the database, and it was indeed deleted along with the corresponding dataset.

The last status message I can see is:
Crawled 892/4700 pages, 0 failed requests, desired concurrency 2.

However, I'm unable to debug further as I don’t have access to all the necessary information.

Please confirm, if I can restore (undelete) runId: rBiWKBvKfnA6T6bgh so I can investigate further?

yener.yasin030

Please restore it i confirm it.

22 Oca 2025 Çar 14:00 tarihinde Jiří Spilka notifications@apify.com şunu yazdı:

Jiří Spilka (jiri.spilka)

Thank you. It looks like you started the Website Content Crawler with approximately 50 URLs, such as this one but did not specify a maxCrawlDepth. Additionally, the crawler automatically enqueued other links, including those with query parameters, which might have increased the scope of the crawl.

We understand this may have been an oversight. We've reimbursed $6 to your account as a compensation.

When you are starting the Actor with a list of URLs that you want to scrape, always set maxCrawlDepth=0

I'll go ahead and close this issue for now, but please feel free to ask another questions. Jiri

Add comment

Developer

Apify

Actor Metrics

5.5k monthly users
999 bookmarks
>99% runs succeeded
1.1 days response time
Created in Mar 2023
Modified 14 days ago

Categories

Fast Website Content Crawler

6sigmag/fast-website-content-crawler

A high-performance web scraper that rapidly extracts and analyzes content from multiple websites simultaneously. Perfect for competitive research, content aggregation, and website structure analysis.

David Deng

290

Deep Website Content Crawler

6sigmag/deep-website-content-crawler

Scrape Failed Killer! A high-performance web scraper that rapidly extracts and analyzes content from multiple websites simultaneously. Perfect for competitive research, content aggregation, and website structure analysis.

David Deng

164

AI Website Content Markdown Scraper

quaking_pail/ai-website-content-markdown-scraper

This Apify Actor, "Website Content Crawler with Markdown Extraction," is designed to perform a comprehensive crawl of specified websites, extract their text content, convert it into Markdown format, and store it in a structured dataset. The extracted content is suitable for feeding LLMs.

AI_Builder

332

Sing a page 🎶

josef.prochazka/sing-a-page

This Actor allows you to listen to a song of your favorite genre with lyrics generated from a page you provide.

Josef Procházka

Example Website Screenshot Crawler

dz_omar/example-website-screenshot-crawler

Automated website screenshot crawler using Pyppeteer and Apify. This open-source actor captures screenshots from specified URLs, uploads them to the Apify Key-Value Store, and provides easy access to the results, making it ideal for monitoring website changes and archiving web content.

Abdlhakim hefaia

Web Scraper

apify/web-scraper

Crawls arbitrary websites using the Chrome browser and extracts structured data from web pages using a provided JavaScript function. The Actor supports both recursive crawling and lists of URLs, and automatically manages concurrency for maximum performance.

Apify

76.1k

456

Video Link Crawler

infoweaver/video-link-crawler

Effortlessly discover and extract video links from any website with our powerful Video Link Crawler within few seconds. Starting from a specified URL, it navigates through web pages, identifies video content, and compiles structured datasets.! Try it Now!

InfoWeaver

News Website Crawler & Article Extractor

xtech/news-source-crawler

Scrape all articles from any news website. Extract full text, metadata, keywords, and summaries. Ideal for content analysis, research, and news aggregation.

Xtech

Web Crawler

rigelbytes/webcrawler

This web crawler is designed to provide users with complete flexibility by allowing them to use their **own proxies**. The scraper collects all pages from the website and returns extracts the **MetaData**, **Title**, and **Content** of the page in MarkDown.

Rigel Bytes

Google Maps Scraper

compass/crawler-google-places

Extract data from thousands of Google Maps locations and businesses. Get Google Maps data including reviews, reviewer details, images, contact info, opening hours, location, prices & more. Export scraped data, run the scraper via API, schedule and monitor runs, or integrate with other tools.