# Facebook Page Monitor - New Posts Scraper & Tracker (`eiv/facebook-page-monitor`) Actor

Watch Facebook Pages and get the posts they publish. Remembers what it already delivered, so a scheduled run returns only what is new - with post text, permalink, images, links and the time it was posted. No login and no API key.

- **URL**: https://apify.com/eiv/facebook-page-monitor.md
- **Developed by:** [Eimantas V](https://apify.com/eiv) (community)
- **Stats:** 1 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.00 / 1,000 post scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Facebook Page Monitor

Watch Facebook Pages and get the posts they publish. It remembers what it already
delivered, so a scheduled run returns **only what is new** - and only charges for that.

No login. No API key. No cookie of yours.

### What it is for

Put your competitors, your clients, or the accounts in your sector on a schedule and get
their posts as rows. The first run collects a Page's recent posts; every run after it
picks up where it left off, which for most Pages is one request and a handful of posts.

If you want a Page's whole back catalogue in one go, this is the wrong tool - see
**How deep it goes** below, and read it before you buy.

### What you get for every post

| Field | Notes |
|---|---|
| `text` | The post copy. Present on **99%** of posts measured |
| `postedAt` | ISO 8601, from Facebook's own timestamp |
| `url` | The post's permalink, so any row can be checked by hand |
| `authorName` | The Page as Facebook spells it |
| `attachmentType` | `photo`, `video`, `link`, or empty for a text-only post |
| `imageUrls` | The photo, the link's preview image, or a video's thumbnail. **97%** of posts carry at least one |
| `linkUrl` | Where a link post points, with Facebook's `l.php` redirector **unwrapped** so you get the real destination |
| `linkTitle` | The headline Facebook shows on the link card |
| `videoUrl`, `videoDurationMs` | A video's page on Facebook and its length |
| `altText` | Facebook's own description of the image, useful for classifying a photo without downloading it |

Each Page also gets a summary row: `postsFetched`, how many requests it took, the
watermark before and after, and `stoppedOn` - so a short delivery always says why.

### Two things it does not give you, and why

**No reaction, comment or share counts.** Facebook does not publish them to a logged-out
visitor. Measured across 72 posts on six Pages: not one carried a single engagement
number. Rather than ship a column of nulls - or worse, zeros, which read as "nobody
engaged" - there is no such field.

**No downloadable video file.** A video post gives you the video's *page* on Facebook and
its duration, not an `.mp4`. Twenty video posts were measured and none carried a playable
URL. `videoUrl` is named for what it is.

### How deep it goes

Facebook serves about **two posts per request** to a logged-out visitor, and asking
for more does not change that - it was measured at 50 and still returned three. So depth
costs requests, and this Actor is built as a monitor rather than an archiver:

- a **first** run on a Page takes up to your *Max posts per Page*, newest first;
- **every run after that** stops at the newest post it already delivered.

For a Page posting a few times a day, that is one or two requests per run. For a full
historical archive it would be hundreds, which is why the ceiling is set where it is.

If a run stops on a limit *before* reaching the post it last delivered, it says so in the
log and in `stoppedOn` - that means a gap, and the gap will not be picked up later.

### Use the residential proxy

**This Actor needs the RESIDENTIAL proxy group, and the input defaults to it.** Facebook
rate-limits Apify's shared datacenter pool: six different datacenter addresses were tried
and every one was refused on its first request. The same run on residential returned 60
posts. If you switch the proxy to datacenter you will get nothing, and the run will tell
you why rather than reporting an empty Page.

That is also why checking a Page is priced the way it is. Loading a Facebook Page in a
real browser costs about **4 MB** of residential traffic, and that happens whether or not
anything new has been posted - so watching ten Pages every hour is a real bill, while
watching ten Pages once a day is a small one. **Check no more often than you will act on
the answer.**

### Running more than one schedule

Each Page's watermark is stored under a name you control. Two schedules sharing that name
will hide each other's new posts - the first one to run consumes them. If you want two
independent watches over the same Pages, give each its own **State store name**.

### Example

```json
{
  "pages": ["bbcnews", "https://www.facebook.com/nasa", "natgeo"],
  "mode": "new",
  "maxPostsPerPage": 30,
  "maxAgeDays": 90,
  "proxyConfig": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}
```

### Honest limits

- **Public Pages only.** Private profiles, groups and anything behind a login are out of
  scope, and the Actor will tell you rather than return an empty dataset.
- **`robots.txt`.** Facebook's `robots.txt` ends with `User-agent: *` / `Disallow: /`;
  only a named allowlist of search-engine crawlers is granted access. This Actor reads
  publicly visible Pages without logging in, but you should decide whether that fits your
  own policy and your jurisdiction before you schedule it.
- **Zero is a normal answer.** A run that finds nothing new has worked, and reports
  success. `stoppedOn: "caught-up"` is the healthy state, not an error.

# Actor input Schema

## `pages` (type: `array`):

One Page per line. A username, a numeric id, or a full facebook.com URL - all three work, and a pasted address bar is fine.

## `mode` (type: `string`):

Only new posts is what a schedule wants: the Actor remembers the newest post it delivered for each Page and stops there next time, so you are not charged again for posts you already have. Latest posts ignores that and takes the newest every run.

## `maxPostsPerPage` (type: `integer`):

The ceiling for one Page. It bounds a Page's FIRST run, which has no watermark to stop at. Later runs normally stop far below it. Facebook serves about 2 posts per request, so a high number here costs requests and residential traffic.

## `maxTotalPosts` (type: `integer`):

A ceiling across every Page in the run.

## `maxAgeDays` (type: `integer`):

Stops a first run walking years back into a Page's history.

## `stateStoreName` (type: `string`):

The named key-value store holding each Page's watermark. Leave it alone unless you want two schedules watching the same Pages independently - give those a store name each, or they will hide each other's new posts.

## `requestDelayMs` (type: `integer`):

Raise this if runs start being refused.

## `proxyConfig` (type: `object`):

Use the RESIDENTIAL group. Facebook rate-limits Apify shared datacenter pool: six different datacenter addresses were tried and every one was refused on its first request. Residential works - the same run returned 60 posts. Leaving this on the default datacenter group will return nothing.

## Actor input object example

```json
{
  "pages": [
    "nasa"
  ],
  "mode": "new",
  "maxPostsPerPage": 30,
  "maxTotalPosts": 500,
  "maxAgeDays": 90,
  "stateStoreName": "facebook-page-monitor-state",
  "requestDelayMs": 800,
  "proxyConfig": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `posts` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pages": [
        "nasa"
    ],
    "proxyConfig": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("eiv/facebook-page-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "pages": ["nasa"],
    "proxyConfig": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("eiv/facebook-page-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pages": [
    "nasa"
  ],
  "proxyConfig": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call eiv/facebook-page-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,eiv/facebook-page-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/tblmNluLVGRbFAUK1/builds/WuOdNfVjQhhBfJUc8/openapi.json
