# Facebook Page Scraper (Public, No Login) (`bgfc97/facebook-page-posts-scraper`) Actor

Scrape public Facebook Pages with no login: identity (name, category, followers, website, email, phone) plus recent public posts with text, permalink, reactions, comments and shares. Pure HTTP, no browser, fast and cheap. Never logs in.

- **URL**: https://apify.com/bgfc97/facebook-page-posts-scraper.md
- **Developed by:** [Bruno](https://apify.com/bgfc97) (community)
- **Categories:** Social media, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$4.00 / 1,000 item scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Facebook Public Page Scraper (no login)

Scrapes **public Facebook Pages** over plain HTTP — **no browser** (no Playwright/Chromium) and
**never logs in**. No account, no cookies, no email/password. It sends a realistic browser
TLS/HTTP2 fingerprint (`got-scraping`) so logged-out Facebook answers with the full page, then
parses the HTML with `cheerio`. Because it is HTTP-only, it runs fast and cheap on an Apify
**DATACENTER** proxy.

For each page it returns the full **public identity** plus the recent **public posts** Facebook
still serves to logged-out visitors.

### Input

```json
{
  "pageUrls": ["nasa", "https://www.facebook.com/CocaColaUS", "@natgeo"],
  "maxPosts": 20,
  "proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["BUYPROXIES94952"] },
  "timeoutSecs": 45
}
```

- **pageUrls** — usernames or full page URLs (with or without `@`/`https://`).
- **maxPosts** — cap on posts per page (see the note below on the logged-out limit).
- **proxyConfiguration** — DATACENTER (`BUYPROXIES94952`) is the cheap default; switch to
  `RESIDENTIAL` only if a page comes back blocked.
- **timeoutSecs** — per-request timeout (15–120s).

### Output (one dataset item per page)

```json
{
  "pageUrl": "https://www.facebook.com/nasa/",
  "loggedIn": false,
  "method": "http",
  "resolvedVia": "www.facebook.com",
  "name": "NASA - National Aeronautics and Space Administration",
  "category": "Government organization",
  "followers": "28,731,417",
  "likes": null,
  "intro": "NASA ... 28,731,417 followers · ...",
  "website": null,
  "email": "public-inquiries@hq.nasa.gov",
  "phone": null,
  "profilePic": "https://scontent.../....png",
  "isVerified": true,
  "postsFound": 1,
  "posts": [
    {
      "text": "Tune in now to watch ...",
      "timestamp_iso": "2026-09-26T16:59:18.000Z",
      "permalink": "https://www.facebook.com/NASA/posts/pfbid02...",
      "reactions": 308,
      "comments": 28,
      "shares": 14,
      "media": ["https://scontent.../image.jpg"]
    }
  ],
  "note": "OK: identity + 1 public post(s) ...",
  "attempts": [ ... ]
}
```

### How it works

- **www.facebook.com** is the primary host: the logged-out desktop HTML embeds the page identity
  in its `og:`/meta tags AND the recent public post(s) as JSON inside
  `<script type="application/json">` Relay caches (real message text, permalink, `creation_time`,
  reaction/comment/share counts). That JSON is parsed directly — no DOM rendering needed.
- **m.facebook.com** is a light identity-only fallback used if www yields nothing.
- **mbasic.facebook.com** is intentionally not used — it now hard-redirects logged-out visitors
  to a login page, so it yields no public data.

### Important limitation (be honest)

Logged out, Facebook only exposes a **small number of recent public posts** — in practice the
pinned/top post (typically 1–4 posts). Getting the full historical feed of a page requires a
logged-in session, which this actor **never uses**. When Facebook exposes no posts, the actor
returns identity only with `posts: []` — it never fabricates or pads posts.

# Actor input Schema

## `pageUrls` (type: `array`):

List of PUBLIC Facebook Page usernames or full URLs to scrape, e.g. "nasa", "@BillGates", or "https://www.facebook.com/BillGates". Each one is fetched over plain HTTP (no browser, no login) and the public identity plus the recent public posts Facebook still serves to logged-out visitors are parsed from the page. No account, cookies, email or password are ever used.

## `maxPosts` (type: `integer`):

Maximum number of recent public posts to extract per page. IMPORTANT: logged-out Facebook only exposes a small number of recent public posts (typically 1-4) — the full historical feed requires a logged-in session, which this actor never uses. Fewer posts (or zero) are reported honestly, never padded with fake data.

## `proxyConfiguration` (type: `object`):

This actor sends a realistic browser TLS/HTTP2 fingerprint over plain HTTP (no browser), so a cheap Apify DATACENTER proxy is enough for logged-out Facebook and is used by default. If a page ever comes back blocked, switch to the RESIDENTIAL group.

## `timeoutSecs` (type: `integer`):

Timeout for each HTTP request (www.facebook.com is tried first, with m.facebook.com as a light identity fallback). Between 15 and 120 seconds.

## Actor input object example

```json
{
  "pageUrls": [
    "nasa",
    "https://www.facebook.com/CocaColaUS"
  ],
  "maxPosts": 20,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "BUYPROXIES94952"
    ]
  },
  "timeoutSecs": 45
}
```

# Actor output Schema

## `dataset` (type: `string`):

One row per Facebook Page: public identity (name, category, followers, website, email, phone) plus recent public posts with text, permalink and engagement.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pageUrls": [
        "nasa",
        "https://www.facebook.com/CocaColaUS"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "BUYPROXIES94952"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("bgfc97/facebook-page-posts-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "pageUrls": [
        "nasa",
        "https://www.facebook.com/CocaColaUS",
    ],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["BUYPROXIES94952"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("bgfc97/facebook-page-posts-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pageUrls": [
    "nasa",
    "https://www.facebook.com/CocaColaUS"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "BUYPROXIES94952"
    ]
  }
}' |
apify call bgfc97/facebook-page-posts-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bgfc97/facebook-page-posts-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/KhNjkh7HMO63bwa5i/builds/AGtXlh5y3R9xuRhlG/openapi.json
