Facebook Page Scraper (Public, No Login) avatar

Facebook Page Scraper (Public, No Login)

Pricing

$4.00 / 1,000 item scrapeds

Go to Apify Store
Facebook Page Scraper (Public, No Login)

Facebook Page Scraper (Public, No Login)

Scrape public Facebook Pages with no login: identity (name, category, followers, website, email, phone) plus recent public posts with text, permalink, reactions, comments and shares. Pure HTTP, no browser, fast and cheap. Never logs in.

Pricing

$4.00 / 1,000 item scrapeds

Rating

0.0

(0)

Developer

Bruno

Bruno

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Facebook Public Page Scraper (no login)

Scrapes public Facebook Pages over plain HTTP — no browser (no Playwright/Chromium) and never logs in. No account, no cookies, no email/password. It sends a realistic browser TLS/HTTP2 fingerprint (got-scraping) so logged-out Facebook answers with the full page, then parses the HTML with cheerio. Because it is HTTP-only, it runs fast and cheap on an Apify DATACENTER proxy.

For each page it returns the full public identity plus the recent public posts Facebook still serves to logged-out visitors.

Input

{
"pageUrls": ["nasa", "https://www.facebook.com/CocaColaUS", "@natgeo"],
"maxPosts": 20,
"proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["BUYPROXIES94952"] },
"timeoutSecs": 45
}
  • pageUrls — usernames or full page URLs (with or without @/https://).
  • maxPosts — cap on posts per page (see the note below on the logged-out limit).
  • proxyConfiguration — DATACENTER (BUYPROXIES94952) is the cheap default; switch to RESIDENTIAL only if a page comes back blocked.
  • timeoutSecs — per-request timeout (15–120s).

Output (one dataset item per page)

{
"pageUrl": "https://www.facebook.com/nasa/",
"loggedIn": false,
"method": "http",
"resolvedVia": "www.facebook.com",
"name": "NASA - National Aeronautics and Space Administration",
"category": "Government organization",
"followers": "28,731,417",
"likes": null,
"intro": "NASA ... 28,731,417 followers · ...",
"website": null,
"email": "public-inquiries@hq.nasa.gov",
"phone": null,
"profilePic": "https://scontent.../....png",
"isVerified": true,
"postsFound": 1,
"posts": [
{
"text": "Tune in now to watch ...",
"timestamp_iso": "2026-09-26T16:59:18.000Z",
"permalink": "https://www.facebook.com/NASA/posts/pfbid02...",
"reactions": 308,
"comments": 28,
"shares": 14,
"media": ["https://scontent.../image.jpg"]
}
],
"note": "OK: identity + 1 public post(s) ...",
"attempts": [ ... ]
}

How it works

  • www.facebook.com is the primary host: the logged-out desktop HTML embeds the page identity in its og:/meta tags AND the recent public post(s) as JSON inside <script type="application/json"> Relay caches (real message text, permalink, creation_time, reaction/comment/share counts). That JSON is parsed directly — no DOM rendering needed.
  • m.facebook.com is a light identity-only fallback used if www yields nothing.
  • mbasic.facebook.com is intentionally not used — it now hard-redirects logged-out visitors to a login page, so it yields no public data.

Important limitation (be honest)

Logged out, Facebook only exposes a small number of recent public posts — in practice the pinned/top post (typically 1–4 posts). Getting the full historical feed of a page requires a logged-in session, which this actor never uses. When Facebook exposes no posts, the actor returns identity only with posts: [] — it never fabricates or pads posts.