No credit card required

Website Content Crawler

apify/website-content-crawler

No credit card required

Crawl websites and extract text content to feed AI models, LLM applications, vector databases, or RAG pipelines. The Actor supports rich formatting using Markdown, cleans the HTML, downloads files, and integrates well with 🦜🔗 LangChain, LlamaIndex, and the wider LLM ecosystem.

Back to issues Create new issue

Hung crawl

Closed

mcantrell-owner opened this issue

The normal job should take around 45s but this one had to be aborted after 4.5m. I had to abort this and I'm not really sure what the cause was by looking at the logs.

If I hadn't caught this, it seems like it could have easily used our entire quota. As we're evaluating the framework, this is a bit of a red flag.

Jiří Spilka (jiri.spilka)

Hi,
Thank you for trying the Website Content Crawler.

I attempted to replicate the issue but did not encounter the problem. My run completed successfully in 34 seconds.

This should not happen. We'll investigate further to debug it.

Please let me know if there's anything else I can assist with or if anything is unclear. I'm happy to help.

Best regards,
Jiri

Jiří Spilka (jiri.spilka)

Hi, I ran this internally, but we were unable to reproduce the issue.

That said, each Actor has a timeout option, which hard-stops the Actor to prevent it from consuming more credits than intended.
I'm sorry I couldn't be of more help. I'll close this issue now, but feel free to reach out with any other questions. Jiri

Add comment

Developer

Apify

Actor Metrics

5.5k monthly users
999 bookmarks
>99% runs succeeded
1.1 days response time
Created in Mar 2023
Modified 14 days ago

Categories

Fast Website Content Crawler

6sigmag/fast-website-content-crawler

A high-performance web scraper that rapidly extracts and analyzes content from multiple websites simultaneously. Perfect for competitive research, content aggregation, and website structure analysis.

David Deng

290

Deep Website Content Crawler

6sigmag/deep-website-content-crawler

Scrape Failed Killer! A high-performance web scraper that rapidly extracts and analyzes content from multiple websites simultaneously. Perfect for competitive research, content aggregation, and website structure analysis.

David Deng

164

AI Website Content Markdown Scraper

quaking_pail/ai-website-content-markdown-scraper

This Apify Actor, "Website Content Crawler with Markdown Extraction," is designed to perform a comprehensive crawl of specified websites, extract their text content, convert it into Markdown format, and store it in a structured dataset. The extracted content is suitable for feeding LLMs.

AI_Builder

332

Sing a page 🎶

josef.prochazka/sing-a-page

This Actor allows you to listen to a song of your favorite genre with lyrics generated from a page you provide.

Josef Procházka

Example Website Screenshot Crawler

dz_omar/example-website-screenshot-crawler

Automated website screenshot crawler using Pyppeteer and Apify. This open-source actor captures screenshots from specified URLs, uploads them to the Apify Key-Value Store, and provides easy access to the results, making it ideal for monitoring website changes and archiving web content.

Abdlhakim hefaia

Web Scraper

apify/web-scraper

Crawls arbitrary websites using the Chrome browser and extracts structured data from web pages using a provided JavaScript function. The Actor supports both recursive crawling and lists of URLs, and automatically manages concurrency for maximum performance.

Apify

76.1k

456

Video Link Crawler

infoweaver/video-link-crawler

Effortlessly discover and extract video links from any website with our powerful Video Link Crawler within few seconds. Starting from a specified URL, it navigates through web pages, identifies video content, and compiles structured datasets.! Try it Now!

InfoWeaver

News Website Crawler & Article Extractor

xtech/news-source-crawler

Scrape all articles from any news website. Extract full text, metadata, keywords, and summaries. Ideal for content analysis, research, and news aggregation.

Xtech

Web Crawler

rigelbytes/webcrawler

This web crawler is designed to provide users with complete flexibility by allowing them to use their **own proxies**. The scraper collects all pages from the website and returns extracts the **MetaData**, **Title**, and **Content** of the page in MarkDown.

Rigel Bytes

Google Maps Scraper

compass/crawler-google-places

Extract data from thousands of Google Maps locations and businesses. Get Google Maps data including reviews, reviewer details, images, contact info, opening hours, location, prices & more. Export scraped data, run the scraper via API, schedule and monitor runs, or integrate with other tools.