No credit card required

Website Content Crawler

apify/website-content-crawler

No credit card required

Crawl websites and extract text content to feed AI models, LLM applications, vector databases, or RAG pipelines. The Actor supports rich formatting using Markdown, cleans the HTML, downloads files, and integrates well with 🦜🔗 LangChain, LlamaIndex, and the wider LLM ecosystem.

Back to issues Create new issue

Poor results

Closed

Digital_Mole opened this issue

I can't seem to get good results from simple documentation sites like OpenAI - most pages return partial info

Jiří Spilka (jiri.spilka)

Hi, I apologize for the delayed response.

Unfortunately, I've tried to resolve the issue, but was not very successful. I’ve reached out internally, and we’ll try to figure this out.

Thank you for your understanding. Jiri

Jiří Spilka (jiri.spilka)

Hi, thank you for your patience, and I apologize for the earlier delay in responding.

Credit to @jindrich.bar for this analysis

It appears the OpenAI documentation website behaves a bit unusually. However, he has identified some adjustments that can help improve the results:

Ensure content is fully loaded: Use "waitForSelector": ".anchor-heading" to make sure there’s content on the page before extraction begins.
Wait for dynamic content: Set "dynamicContentWaitSecs": 30" to allow sufficient time for the content to load.
Avoid dropping any content: Use "htmlTransformer": "none" to prevent content from being stripped during processing.

Here’s an example run for the OpenAI documentation where these settings were applied: example run.

Additionally, he noticed that many URLs in the OpenAI docs point to different sections of the same page (e.g., the API reference). As a result, the extracted results may include the entire API reference text multiple times. To optimize costs and performance, you may want to reduce the number of URLs being processed, as the current approach with a 30-second wait can become expensive.

Please let me know if you’d like further clarification. Jiri

Jiří Spilka (jiri.spilka)

I’ll go ahead and close this issue for now. However, feel free to ask additional questions or create a new issue if needed.

Add comment

Developer

Apify

Actor Metrics

5.5k monthly users
999 bookmarks
>99% runs succeeded
1.1 days response time
Created in Mar 2023
Modified 14 days ago

Categories

Fast Website Content Crawler

6sigmag/fast-website-content-crawler

A high-performance web scraper that rapidly extracts and analyzes content from multiple websites simultaneously. Perfect for competitive research, content aggregation, and website structure analysis.

David Deng

290

Deep Website Content Crawler

6sigmag/deep-website-content-crawler

Scrape Failed Killer! A high-performance web scraper that rapidly extracts and analyzes content from multiple websites simultaneously. Perfect for competitive research, content aggregation, and website structure analysis.

David Deng

164

AI Website Content Markdown Scraper

quaking_pail/ai-website-content-markdown-scraper

This Apify Actor, "Website Content Crawler with Markdown Extraction," is designed to perform a comprehensive crawl of specified websites, extract their text content, convert it into Markdown format, and store it in a structured dataset. The extracted content is suitable for feeding LLMs.

AI_Builder

332

Sing a page 🎶

josef.prochazka/sing-a-page

This Actor allows you to listen to a song of your favorite genre with lyrics generated from a page you provide.

Josef Procházka

Example Website Screenshot Crawler

dz_omar/example-website-screenshot-crawler

Automated website screenshot crawler using Pyppeteer and Apify. This open-source actor captures screenshots from specified URLs, uploads them to the Apify Key-Value Store, and provides easy access to the results, making it ideal for monitoring website changes and archiving web content.

Abdlhakim hefaia

Web Scraper

apify/web-scraper

Crawls arbitrary websites using the Chrome browser and extracts structured data from web pages using a provided JavaScript function. The Actor supports both recursive crawling and lists of URLs, and automatically manages concurrency for maximum performance.

Apify

76.1k

456

Video Link Crawler

infoweaver/video-link-crawler

Effortlessly discover and extract video links from any website with our powerful Video Link Crawler within few seconds. Starting from a specified URL, it navigates through web pages, identifies video content, and compiles structured datasets.! Try it Now!

InfoWeaver

News Website Crawler & Article Extractor

xtech/news-source-crawler

Scrape all articles from any news website. Extract full text, metadata, keywords, and summaries. Ideal for content analysis, research, and news aggregation.

Xtech

Web Crawler

rigelbytes/webcrawler

This web crawler is designed to provide users with complete flexibility by allowing them to use their **own proxies**. The scraper collects all pages from the website and returns extracts the **MetaData**, **Title**, and **Content** of the page in MarkDown.

Rigel Bytes

Google Maps Scraper

compass/crawler-google-places

Extract data from thousands of Google Maps locations and businesses. Get Google Maps data including reviews, reviewer details, images, contact info, opening hours, location, prices & more. Export scraped data, run the scraper via API, schedule and monitor runs, or integrate with other tools.