Facebook Page Monitor - New Posts Scraper & Tracker avatar

Facebook Page Monitor - New Posts Scraper & Tracker

Pricing

from $5.00 / 1,000 post scrapeds

Go to Apify Store
Facebook Page Monitor - New Posts Scraper & Tracker

Facebook Page Monitor - New Posts Scraper & Tracker

Watch Facebook Pages and get the posts they publish. Remembers what it already delivered, so a scheduled run returns only what is new - with post text, permalink, images, links and the time it was posted. No login and no API key.

Pricing

from $5.00 / 1,000 post scrapeds

Rating

0.0

(0)

Developer

Eimantas V

Eimantas V

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

6 days ago

Last modified

Categories

Share

Facebook Page Monitor

Watch Facebook Pages and get the posts they publish. It remembers what it already delivered, so a scheduled run returns only what is new - and only charges for that.

No login. No API key. No cookie of yours.

What it is for

Put your competitors, your clients, or the accounts in your sector on a schedule and get their posts as rows. The first run collects a Page's recent posts; every run after it picks up where it left off, which for most Pages is one request and a handful of posts.

If you want a Page's whole back catalogue in one go, this is the wrong tool - see How deep it goes below, and read it before you buy.

What you get for every post

FieldNotes
textThe post copy. Present on 99% of posts measured
postedAtISO 8601, from Facebook's own timestamp
urlThe post's permalink, so any row can be checked by hand
authorNameThe Page as Facebook spells it
attachmentTypephoto, video, link, or empty for a text-only post
imageUrlsThe photo, the link's preview image, or a video's thumbnail. 97% of posts carry at least one
linkUrlWhere a link post points, with Facebook's l.php redirector unwrapped so you get the real destination
linkTitleThe headline Facebook shows on the link card
videoUrl, videoDurationMsA video's page on Facebook and its length
altTextFacebook's own description of the image, useful for classifying a photo without downloading it

Each Page also gets a summary row: postsFetched, how many requests it took, the watermark before and after, and stoppedOn - so a short delivery always says why.

Two things it does not give you, and why

No reaction, comment or share counts. Facebook does not publish them to a logged-out visitor. Measured across 72 posts on six Pages: not one carried a single engagement number. Rather than ship a column of nulls - or worse, zeros, which read as "nobody engaged" - there is no such field.

No downloadable video file. A video post gives you the video's page on Facebook and its duration, not an .mp4. Twenty video posts were measured and none carried a playable URL. videoUrl is named for what it is.

How deep it goes

Facebook serves about two posts per request to a logged-out visitor, and asking for more does not change that - it was measured at 50 and still returned three. So depth costs requests, and this Actor is built as a monitor rather than an archiver:

  • a first run on a Page takes up to your Max posts per Page, newest first;
  • every run after that stops at the newest post it already delivered.

For a Page posting a few times a day, that is one or two requests per run. For a full historical archive it would be hundreds, which is why the ceiling is set where it is.

If a run stops on a limit before reaching the post it last delivered, it says so in the log and in stoppedOn - that means a gap, and the gap will not be picked up later.

Use the residential proxy

This Actor needs the RESIDENTIAL proxy group, and the input defaults to it. Facebook rate-limits Apify's shared datacenter pool: six different datacenter addresses were tried and every one was refused on its first request. The same run on residential returned 60 posts. If you switch the proxy to datacenter you will get nothing, and the run will tell you why rather than reporting an empty Page.

That is also why checking a Page is priced the way it is. Loading a Facebook Page in a real browser costs about 4 MB of residential traffic, and that happens whether or not anything new has been posted - so watching ten Pages every hour is a real bill, while watching ten Pages once a day is a small one. Check no more often than you will act on the answer.

Running more than one schedule

Each Page's watermark is stored under a name you control. Two schedules sharing that name will hide each other's new posts - the first one to run consumes them. If you want two independent watches over the same Pages, give each its own State store name.

Example

{
"pages": ["bbcnews", "https://www.facebook.com/nasa", "natgeo"],
"mode": "new",
"maxPostsPerPage": 30,
"maxAgeDays": 90,
"proxyConfig": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}

Honest limits

  • Public Pages only. Private profiles, groups and anything behind a login are out of scope, and the Actor will tell you rather than return an empty dataset.
  • robots.txt. Facebook's robots.txt ends with User-agent: * / Disallow: /; only a named allowlist of search-engine crawlers is granted access. This Actor reads publicly visible Pages without logging in, but you should decide whether that fits your own policy and your jurisdiction before you schedule it.
  • Zero is a normal answer. A run that finds nothing new has worked, and reports success. stoppedOn: "caught-up" is the healthy state, not an error.