# Real Estate Actor Suite (`scrapeai/realestate-actor-suite`) Actor

Scrape public property, accommodation, and marketplace listing pages through isolated portal adapters with detailed structured output.

- **URL**: https://apify.com/scrapeai/realestate-actor-suite.md
- **Developed by:** [ScrapeAI](https://apify.com/scrapeai) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.49 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Real Estate Actor Suite

### What does Real Estate Actor Suite do?

This Actor runs isolated adapters for every portal in the supplied catalog, including Zillow, Auction.com, LoopNet, Otodom, Cian, Funda, SeLoger, Immobiliare, Immowelt, ImmoScout24, Australian portals, UAE portals, UK portals, Canadian portals, Idealista, Apartments.com, Bayut, Zoopla, Realtor.com, accommodation portals, and the public-page Skip Trace adapter. Each adapter follows the same Jobright-style Crawlee and Playwright structure while keeping its own portal configuration in `src/actor-configs.js`.

The suite extracts only publicly visible or publicly embedded page data. It does not invent values. Missing source values are represented by non-null empty strings, zero values, or empty arrays and are listed in `extractionWarnings`.

### Why use this Actor?

- Run one adapter, a selected group, or all configured adapters.
- Start with a configured public default URL for each portal, or override it with portal-specific URLs through `actorInputs`.
- Open detail pages for enrichment and retain structured JSON-LD, metadata, scoped card data, and selected JSON API responses.
- Deduplicate records per actor by canonical listing URL and identifier.
- Capture per-actor counters, block diagnostics, and failures in `RUN_SUMMARY`.
- Use public HTTP HTML for Centris, HotPads, Trulia, Zillow, Rightmove and Trip.com; use the browser for JavaScript-rendered portals such as Agoda. Embedded JSON and property-specific selectors supply listing details.
- Access the dataset through the Apify API, schedule runs, and monitor output from the Console.

### Standalone per-portal Actors

Each catalog entry is also available as an independent Jobright-style Actor under [`actors/`](C:/xampp/htdocs/scraper/realestate/actors/README.md). Every child folder contains its own `.actor` metadata and schemas, `src` runtime, `storage` results and diagnostics, `output` exports, `README.md`, Dockerfile, package files, and dataset validator. Run one child folder with `apify run`; regenerate the complete layout with `node scripts/scaffold-standalone-actors.mjs`.

### What data can it extract?

Each record includes the originating actor, portal, canonical URL, listing ID, title, description, transaction type, property type, price, address, bedrooms, bathrooms, area, year built, agent/company information, public contact fields, photos, amenities, raw structured data, quality flags, and extraction warnings.

### How to scrape property portals

1. Select the required actor names in the Input tab, or keep `all` to run every adapter.
2. The shared `startUrls` input already contains one public default URL per configured portal. Replace those URLs when you need a different location/search, or provide portal-specific URLs under `actorInputs`.
3. Set `maxItemsPerActor`, `maxPagesPerActor`, and `includeDetails`.
4. Run the Actor and inspect the `Detailed Listings from All Actors` dataset view.
5. Use the `actorName` and `sourcePortal` fields to split results by portal.

For local development, install dependencies with `npm ci`, validate the schema with `npm run validate`, and run with `apify run`. The `testMode` input runs a fast offline fixture smoke test for every selected adapter.

### Input

See the Input tab for all configuration options. For different URLs per actor, use this shape:

```json
{
  "actors": ["zillow-scrape-address-url-zpid", "realestate-com-au-scraper"],
  "actorInputs": {
    "zillow-scrape-address-url-zpid": {
      "startUrls": [{ "url": "https://www.zillow.com/" }]
    },
    "realestate-com-au-scraper": {
      "startUrls": [{ "url": "https://www.realestate.com.au/buy" }]
    }
  },
  "maxItemsPerActor": 10,
  "includeDetails": true,
  "includeRawData": true,
  "proxyEnabled": false,
  "maxRequestRetries": 3,
  "sameDomainDelaySecs": 3
}
```

`proxyEnabled` defaults to `false`, using the verified public HTTP route where configured. Set it to `true` to opt into the browser with an authorized Apify proxy or user-supplied `proxyUrls` inside `proxyConfiguration`. Proxy access is account- and plan-dependent and may incur costs. Unavailable access is reported; CAPTCHA and authentication requirements are not solved automatically. Fixture tests do not establish live portal coverage.

The default `startUrls` cover 31 portal adapters. The Skip Trace adapter has no fabricated default people-search URL: it requires authorized public `startUrls` supplied by the user.

For the recovered actors, use `scripts/live-http-input.json` and `scripts/live-recovery-input.json` with `apify run --input-file`. Keep separate relative `APIFY_LOCAL_STORAGE_DIR` and `CRAWLEE_STORAGE_DIR` values to preserve earlier runs. Raw platform datasets retain extraction warnings; the delivery exporter omits unknown values and never invents missing details or inserts error rows into data.

The Skip Trace adapter intentionally requires user-supplied public URLs. It does not execute a default person lookup.

### Output

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. Every emitted row has no JSON `null` values. When a portal does not publish a field, the row remains valid and records the field name in `extractionWarnings`.

```json
{
  "actorName": "realestate-com-au-scraper",
  "sourcePortal": "Realestate.com.au",
  "title": "Example property listing",
  "listingUrl": "https://www.realestate.com.au/property-example",
  "price": { "display": "$750,000", "value": 750000, "currency": "AUD", "period": "", "raw": "$750,000" },
  "address": { "full": "Example suburb", "street": "", "locality": "Example suburb", "region": "", "postalCode": "", "country": "AU", "latitude": 0, "longitude": 0 },
  "features": { "bedrooms": 3, "bathrooms": 2, "area": { "display": "", "value": 0, "unit": "" }, "yearBuilt": 0, "floor": "", "totalFloors": "", "parking": "", "furnished": "" },
  "extractionWarnings": ["area", "yearBuilt", "photos"]
}
```

### Cost, compliance, and support

Compute-unit and proxy costs depend on the number of portals, pages, detail requests, retries, and proxy type. Start with a small `maxItemsPerActor` value and increase it after checking coverage and blocking. Respect each portal's terms, robots directives, rate limits, and applicable privacy laws. The Skip Trace adapter can encounter personal data; use it only with a legitimate, documented purpose and appropriate authorization. Public data is not automatically unrestricted data.

For troubleshooting, inspect `RUN_SUMMARY`, per-actor `ACTOR_SUMMARY_<actor-name>` keys, and blocked-page diagnostics. Use the Issues tab for selector changes or portal-specific feedback.

# Actor input Schema

## `actors` (type: `array`):

Actor names to run. Use \["all"] to run every configured adapter.

## `startUrls` (type: `array`):

Default public listing/search URLs for the configured portals. URLs are routed to the matching portal adapter. Replace them with authorized public URLs for the locations and searches you need. Skip Trace intentionally requires user-supplied authorized public URLs.

## `query` (type: `string`):

Optional query metadata passed to adapters. Use explicit public startUrls when the portal requires a concrete route or search identifier.

## `location` (type: `string`):

Optional city, postcode, state, country, or neighbourhood.

## `listingType` (type: `string`):

Preferred transaction or accommodation mode where supported.

## `propertyType` (type: `string`):

Optional portal-specific property type filter.

## `maxItemsPerActor` (type: `integer`):

Maximum detailed records emitted by each selected adapter.

## `maxPagesPerActor` (type: `integer`):

Maximum search/list pages followed per actor.

## `includeDetails` (type: `boolean`):

Open each listing detail page to enrich the output.

## `includeRawData` (type: `boolean`):

Keep selected JSON-LD, metadata, card data, and JSON responses under rawData.

## `scrollCount` (type: `integer`):

Number of incremental scrolls used for lazy-loaded results.

## `requestTimeoutSecs` (type: `integer`):

Maximum seconds allowed for one page.

## `maxConcurrency` (type: `integer`):

Concurrent browser pages. Keep low for public portals.

## `maxRequestRetries` (type: `integer`):

Retries for transient HTTP blocks and navigation failures. Retries use session rotation and backoff.

## `sameDomainDelaySecs` (type: `integer`):

Minimum delay between requests to the same portal, in seconds.

## `proxyEnabled` (type: `boolean`):

Use an authorized Apify or user-supplied proxy route when available. This does not bypass access controls and may incur proxy costs.

## `proxyUrls` (type: `array`):

Optional custom proxy URLs (http://user:pass@host:port) to route requests through.

## `serpApiKey` (type: `string`):

Optional SerpApi API key used as fallback data provider when portals are blocked.

## `captchaApiKey` (type: `string`):

Optional CapSolver, 2Captcha, or Anti-Captcha API key to automatically solve challenge pages.

## `headless` (type: `boolean`):

Run Chromium without a visible window.

## `testMode` (type: `boolean`):

Run an offline fixture smoke test for every selected adapter instead of visiting websites.

## `actorInputs` (type: `object`):

Optional per-actor overrides. Explicit startUrls are filtered to that adapter's configured source hosts. Example: {"zillow-scrape-address-url-zpid": {"startUrls": \[{"url": "https://www.zillow.com/..."}], "maxItemsPerActor": 20}}.

## `proxyConfiguration` (type: `object`):

Optional Apify proxy configuration used when proxyEnabled is true.

## Actor input object example

```json
{
  "actors": [
    "all"
  ],
  "startUrls": [
    {
      "url": "https://www.zillow.com/san-francisco-ca/"
    },
    {
      "url": "https://www.auction.com/residential"
    },
    {
      "url": "https://www.loopnet.com/for-sale/"
    },
    {
      "url": "https://www.otodom.pl/pl/wyniki/sprzedaz/mieszkanie/cala-polska"
    },
    {
      "url": "https://www.cian.ru/cat.php?deal_type=sale&object_type%5B0%5D=1&offer_type=flat"
    },
    {
      "url": "https://www.funda.nl/zoeken/koop/"
    },
    {
      "url": "https://www.seloger.com/list.htm?projects=2&types=1"
    },
    {
      "url": "https://www.immobiliare.it/vendita-case/"
    },
    {
      "url": "https://www.immowelt.de/suche/kaufen/wohnung/berlin/ad04de11"
    },
    {
      "url": "https://www.immobilienscout24.de/Suche/de/wohnung-kaufen.html"
    },
    {
      "url": "https://flatmates.com.au/rooms"
    },
    {
      "url": "https://www.domain.com.au/sale/"
    },
    {
      "url": "https://hotpads.com/san-francisco-ca/apartments-for-rent"
    },
    {
      "url": "https://www.propertyfinder.ae/en/search?l=1&c=2&t=1"
    },
    {
      "url": "https://www.centris.ca/en/condos~for-sale~montreal"
    },
    {
      "url": "https://www.homes.com/"
    },
    {
      "url": "https://www.99.co/singapore/sale"
    },
    {
      "url": "https://www.trulia.com/CA/San_Francisco/"
    },
    {
      "url": "https://www.trip.com/hotels/list?city=359"
    },
    {
      "url": "https://www.agoda.com/city/bangkok-th.html"
    },
    {
      "url": "https://www.tripadvisor.com/Hotels"
    },
    {
      "url": "https://www.vrbo.com/en-ca/vacation-rentals/united-states/california/san-francisco-county/san-francisco"
    },
    {
      "url": "https://www.rightmove.co.uk/property-for-sale/London.html"
    },
    {
      "url": "https://www.realtor.ca/on/toronto/real-estate"
    },
    {
      "url": "https://dubai.dubizzle.com/en/property-for-sale/residential/apartment/"
    },
    {
      "url": "https://www.realestate.com.au/buy"
    },
    {
      "url": "https://www.idealista.com/en/"
    },
    {
      "url": "https://www.apartments.com/san-francisco-ca/"
    },
    {
      "url": "https://www.bayut.com/for-sale/apartments/dubai/"
    },
    {
      "url": "https://www.zoopla.co.uk/for-sale/property/london/"
    },
    {
      "url": "https://www.realtor.com/realestateandhomes-search/San-Francisco_CA"
    }
  ],
  "query": "",
  "location": "",
  "listingType": "sale",
  "propertyType": "",
  "maxItemsPerActor": 10,
  "maxPagesPerActor": 2,
  "includeDetails": true,
  "includeRawData": true,
  "scrollCount": 3,
  "requestTimeoutSecs": 45,
  "maxConcurrency": 2,
  "maxRequestRetries": 3,
  "sameDomainDelaySecs": 3,
  "proxyEnabled": false,
  "proxyUrls": [],
  "serpApiKey": "",
  "captchaApiKey": "",
  "headless": true,
  "testMode": false,
  "actorInputs": {},
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `dataset` (type: `string`):

Detailed, non-null records from every selected adapter.

## `files` (type: `string`):

Run summary, per-actor counters, and blocked-page diagnostics.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapeai/realestate-actor-suite").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("scrapeai/realestate-actor-suite").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call scrapeai/realestate-actor-suite --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapeai/realestate-actor-suite"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/sYxsQUC3rHGC0mSpB/builds/IrgrFvXUWlcGW3gOs/openapi.json
