No credit card required

Website Content Crawler

apify/website-content-crawler

No credit card required

Crawl websites and extract text content to feed AI models, LLM applications, vector databases, or RAG pipelines. The Actor supports rich formatting using Markdown, cleans the HTML, downloads files, and integrates well with 🦜🔗 LangChain, LlamaIndex, and the wider LLM ecosystem.

Back to issues Create new issue

scraped data is redundant

Closed

visable opened this issue

can you please explain the scrapped data and how to scrape this website properly?

Jiří Spilka (jiri.spilka)

Hi (again),
That’s a very good question. For the targeted page, you need to select a country first; otherwise, the crawler won’t be able to retrieve the content.

You can achieve this by first using your browser to select the country, then copying the cookies. For example, you can use the Copy Cookies Google Chrome extension.

Once you have the cookies, paste them into the Website Content Crawler settings under the initialCookies field.
Here’s my example run, which I aborted after confirming that the crawling was working.

As with the previous issue, I noticed that the data is structured. You might get better results using a custom Web Scraper if you have some coding experience.
I hope this helps! Please let me know if it works for you. Jiri

Jiří Spilka (jiri.spilka)

I'll go ahead and close this issue now, but feel free to ask any questions or raise a new issue.

Add comment

Developer

Apify

Actor Metrics

5.5k monthly users
999 bookmarks
>99% runs succeeded
1.1 days response time
Created in Mar 2023
Modified 14 days ago

Categories

Fast Website Content Crawler

6sigmag/fast-website-content-crawler

A high-performance web scraper that rapidly extracts and analyzes content from multiple websites simultaneously. Perfect for competitive research, content aggregation, and website structure analysis.

David Deng

290

Deep Website Content Crawler

6sigmag/deep-website-content-crawler

Scrape Failed Killer! A high-performance web scraper that rapidly extracts and analyzes content from multiple websites simultaneously. Perfect for competitive research, content aggregation, and website structure analysis.

David Deng

164

AI Website Content Markdown Scraper

quaking_pail/ai-website-content-markdown-scraper

This Apify Actor, "Website Content Crawler with Markdown Extraction," is designed to perform a comprehensive crawl of specified websites, extract their text content, convert it into Markdown format, and store it in a structured dataset. The extracted content is suitable for feeding LLMs.

AI_Builder

332

Sing a page 🎶

josef.prochazka/sing-a-page

This Actor allows you to listen to a song of your favorite genre with lyrics generated from a page you provide.

Josef Procházka

Example Website Screenshot Crawler

dz_omar/example-website-screenshot-crawler

Automated website screenshot crawler using Pyppeteer and Apify. This open-source actor captures screenshots from specified URLs, uploads them to the Apify Key-Value Store, and provides easy access to the results, making it ideal for monitoring website changes and archiving web content.

Abdlhakim hefaia

Web Scraper

apify/web-scraper

Crawls arbitrary websites using the Chrome browser and extracts structured data from web pages using a provided JavaScript function. The Actor supports both recursive crawling and lists of URLs, and automatically manages concurrency for maximum performance.

Apify

76.1k

456

Video Link Crawler

infoweaver/video-link-crawler

Effortlessly discover and extract video links from any website with our powerful Video Link Crawler within few seconds. Starting from a specified URL, it navigates through web pages, identifies video content, and compiles structured datasets.! Try it Now!

InfoWeaver

News Website Crawler & Article Extractor

xtech/news-source-crawler

Scrape all articles from any news website. Extract full text, metadata, keywords, and summaries. Ideal for content analysis, research, and news aggregation.

Xtech

Web Crawler

rigelbytes/webcrawler

This web crawler is designed to provide users with complete flexibility by allowing them to use their **own proxies**. The scraper collects all pages from the website and returns extracts the **MetaData**, **Title**, and **Content** of the page in MarkDown.

Rigel Bytes

Google Maps Scraper

compass/crawler-google-places

Extract data from thousands of Google Maps locations and businesses. Get Google Maps data including reviews, reviewer details, images, contact info, opening hours, location, prices & more. Export scraped data, run the scraper via API, schedule and monitor runs, or integrate with other tools.