No credit card required

BeautifulSoup Scraper

apify/beautifulsoup-scraper

No credit card required

Crawls websites using raw HTTP requests. It parses the HTML with the BeautifulSoup library and extracts data from the pages using Python code. Supports both recursive crawling and lists of URLs. This Actor is a Python alternative to Cheerio Scraper.

Do you want to learn more about this Actor?

Get a demo

You can access the BeautifulSoup Scraper programmatically from your own applications by using the Apify API. You can choose the language preference from below. To use the Apify API, you’ll need an Apify account and your API token, found in Integrations settings in Apify Console.

1# Set API token
2API_TOKEN=<YOUR_API_TOKEN>
3
4# Prepare Actor input
5cat > input.json << 'EOF'
6{
7  "startUrls": [
8    {
9      "url": "https://crawlee.dev"
10    }
11  ],
12  "maxCrawlingDepth": 1,
13  "requestTimeout": 10,
14  "linkSelector": "a[href]",
15  "linkPatterns": [
16    ".*crawlee\\.dev.*"
17  ],
18  "pageFunction": "from typing import Any\n\n# See the context section in readme to find out what fields you can access \n# https://apify.com/vdusek/beautifulsoup-scraper#context    \ndef page_function(context: Context) -> Any:\n    url = context.request['url']\n    title = context.soup.title.string if context.soup.title else None\n    return {'url': url, 'title': title}\n",
19  "soupFeatures": "html.parser",
20  "proxyConfiguration": {
21    "useApifyProxy": true
22  }
23}
24EOF
25
26# Run the Actor using an HTTP API
27# See the full API reference at https://docs.apify.com/api/v2
28curl "https://api.apify.com/v2/acts/apify~beautifulsoup-scraper/runs?token=$API_TOKEN" \
29  -X POST \
30  -d @input.json \
31  -H 'Content-Type: application/json'

BeautifulSoup Scraper API

Below, you can find a list of relevant HTTP API endpoints for calling the BeautifulSoup Scraper Actor. For this, you’ll need an Apify account. Replace <YOUR_API_TOKEN> in the URLs with your Apify API token, which you can find under Integrations in Apify Console. For details, see the API reference .

Run Actor

POST

https://api.apify.com/v2/acts/apify~beautifulsoup-scraper/runs?token=<YOUR_API_TOKEN>

Note: By adding the method=POST query parameter, this API endpoint can be called using a GET request and thus used in third-party webhooks. Please refer to our Run Actor API documentation .

Run Actor synchronously and get dataset items

POST

https://api.apify.com/v2/acts/apify~beautifulsoup-scraper/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>

Note: This endpoint supports both POST and GET request methods. However, only the POST method allows you to pass input data. For more information, please refer to our Run Actor synchronously and get dataset items API documentation .

Get Actor

GET

https://api.apify.com/v2/acts/apify~beautifulsoup-scraper?token=<YOUR_API_TOKEN>

For more information, please refer to our Get Actor API documentation .

Actors can be used to scrape web pages, extract data, or automate browser tasks. Use the BeautifulSoup Scraper API programmatically via the Apify API.

You can choose from:

BeautifulSoup Scraper API in Python

BeautifulSoup Scraper API in JavaScript

BeautifulSoup Scraper API through CLI

You can start BeautifulSoup Scraper with the Apify API by sending an HTTP POST request to the Run Actor endpoint. An Actor’s input and its content type can be passed as a payload of the POST request, and additional options can be specified using URL query parameters. The BeautifulSoup Scraper is identified within the API by its ID, which is the creator’s username and the name of the Actor.

When the BeautifulSoup Scraper run finishes you can list the data from its default dataset (storage) via the API or you can preview the data directly on Apify Console .

Developer

Apify

Actor metrics

17 monthly users
4 stars
99.2% runs succeeded
Created in Jul 2023
Modified 4 months ago

Categories

Developer tools

For creators

Amazon search scraper

logical_scrapers/amazon-search-scraper

Fastest amazon search scraper that scrapes product listings from Amazon search results. This scraper extracts comprehensive product information including titles, prices (both sale and original), ratings, review counts and product URLs. Built with pagination support and robust error handling.

Goldmine

Lamudi.ph Real Estate Listings Scraper

pixelperfekt/lamudi-ph-real-estate-listings-scraper

This Python script uses Selenium and BeautifulSoup to scrape real estate listings from Lamudi.ph . It extracts details like property titles, prices, locations, bedrooms, and land size, storing the data in Apify's dataset for easy analysis.

PixelPerfekt

Website Content Crawler

apify/website-content-crawler

Crawl websites and extract text content to feed AI models, LLM applications, vector databases, or RAG pipelines. The Actor supports rich formatting using Markdown, cleans the HTML, downloads files, and integrates well with 🦜🔗 LangChain, LlamaIndex, and the wider LLM ecosystem.

Apify

24.6k

584

Web Scraper

apify/web-scraper

Crawls arbitrary websites using the Chrome browser and extracts data from pages using JavaScript code. The Actor supports both recursive crawling and lists of URLs and automatically manages concurrency for maximum performance. This is Apify's basic tool for web crawling and scraping.

Apify

70.4k

198

Cheerio Scraper

apify/cheerio-scraper

Crawls websites using raw HTTP requests, parses the HTML with the Cheerio library, and extracts data from the pages using a Node.js code. Supports both recursive crawling and lists of URLs. This actor is a high-performance alternative to apify/web-scraper for websites that do not require JavaScript.

Apify

5.4k

Puppeteer Scraper

apify/puppeteer-scraper

Crawls websites with the headless Chrome and Puppeteer library using a provided server-side Node.js code. This crawler is an alternative to apify/web-scraper that gives you finer control over the process. Supports both recursive crawling and list of URLs. Supports login to website.

Apify

4.4k

Legacy PhantomJS Crawler

apify/legacy-phantomjs-crawler

Replacement for the legacy Apify Crawler product with a backward-compatible interface. The actor uses PhantomJS headless browser to recursively crawl websites and extract data from them using a piece of front-end JavaScript code.

Apify

1.6k

Extended GPT Scraper

drobnikj/extended-gpt-scraper

Extract data from any website and feed it into GPT via the OpenAI API. Use ChatGPT to proofread content, analyze sentiment, summarize reviews, extract contact details, and much more.

Jakub Drobník

1.1k

Monitoring Reporter Dashboard

apify/monitoring-reporter-dashboard

The monitoring reporter dashboard is a part of the Apify Monitoring Suite (apify/monitoring). See its readme for more information and how to use this.

Apify

Playwright Scraper

apify/playwright-scraper

Crawls websites with the headless Chromium, Chrome, or Firefox browser and Playwright library using a provided server-side Node.js code. Supports both recursive crawling and a list of URLs. Supports login to a website.

Apify

808

Python web scraping tutorial (Step-by-step guide for 2024)

Web scraping with Python Requests

Web scraping with JavaScript vs. Python in 2024

Build new tools

Are you a developer? Build your own Actors and run them on Apify.

Learn more

Get a custom solution

Get a custom web scraping or RPA solution.

Book a demo

BeautifulSoup Scraper

BeautifulSoup Scraper

BeautifulSoup Scraper API

Run Actor

Run Actor synchronously and get dataset items

Get Actor

Amazon search scraper

Lamudi.ph Real Estate Listings Scraper

Website Content Crawler

Web Scraper

Cheerio Scraper

Puppeteer Scraper

Legacy PhantomJS Crawler

Extended GPT Scraper

Monitoring Reporter Dashboard

Playwright Scraper

Related articles

Where next?

Build new tools

Get a custom solution

BeautifulSoup Scraper

BeautifulSoup Scraper API

Run Actor

Run Actor synchronously and get dataset items

Get Actor

You might also like these Actors

Amazon search scraper

Lamudi.ph Real Estate Listings Scraper

Website Content Crawler

Web Scraper

Cheerio Scraper

Puppeteer Scraper

Legacy PhantomJS Crawler

Extended GPT Scraper

Monitoring Reporter Dashboard

Playwright Scraper

Related articles

Where next?

Build new tools

Get a custom solution