Facebook Pages Scraper avatar

Facebook Pages Scraper

Pricing

from $3.00 / 1,000 results

Go to Apify Store
Facebook Pages Scraper

Facebook Pages Scraper

Discover Facebook pages through search engines and extract structured page data including IDs, names, categories, addresses, follower counts, phone numbers, emails, websites, business hours, ratings, and media.

Pricing

from $3.00 / 1,000 results

Rating

0.0

(0)

Developer

Crawler Bros

Crawler Bros

Maintained by Community

Actor stats

0

Bookmarked

83

Total users

7

Monthly active users

13 hours

Issues response

11 days ago

Last modified

Share

Discover and extract structured data from Facebook Pages at scale. This actor discovers relevant Facebook Pages via Facebook's own search (most reliable, requires session cookies) plus DuckDuckGo and Google as supplementary sources, then extracts business details, engagement metrics, contact info, and more — all using residential IPs for reliability.

What You'll Get

Each discovered page is returned as a structured row with fields including:

  • Page identity: name, ID, canonical URL, categories
  • Engagement: follower count, likes, check-ins, talking-about
  • Contact: email, phone, address, website, Messenger link
  • Media: profile picture, cover photo
  • Business info: price range, business hours, rating, ad library reference

Features

  • Search-powered discovery: finds pages via Facebook's own search (when cookies are provided) plus DuckDuckGo and Google as supplementary sources
  • Scale to 1K+ results: combine multiple search terms × locations and increase maxItems (each combination adds up to ~40 pages to the pool)
  • Residential IPs always, with automatic session rotation on failure: all Facebook page requests use Apify residential proxy — no need to configure proxy settings. Any retry after a failed request automatically switches to a fresh residential IP session rather than hammering the same possibly-flagged one
  • Cookie authentication: when you provide session cookies the scraper bypasses most login-wall blocks
  • Smart fallback chain: HTTP → Playwright browser → Playwright with cookies (each step only activates when the previous one fails)
  • Dual extraction: regex over JSON-LD, Open Graph, inline scripts, and intro cards for maximum field coverage
  • Clean output: unavailable/deleted pages are detected and skipped — they never consume your maxItems budget or appear as junk rows in the output dataset

Input

FieldTypeDefaultDescription
startUrlsArray of {url}[]Direct list of Facebook page URLs to scrape — skips discovery for these. Either this or searchTerms is required
searchTermsArray of strings[]Plain search terms for page discovery (e.g. ["Coffee Shop", "Bakery"]). Either this or startUrls is required
locationsArray of strings[]Optional locations combined with each search term (e.g. ["London", "Manchester"])
maxItemsInteger10Maximum number of page rows to output
includeReviewsBooleantrueFetch and include real reviews per page (adds one extra request per page)
cookiesString""Facebook session cookies — strongly recommended. See Cookies below

Example Input — direct URLs

{
"startUrls": [
{ "url": "https://www.facebook.com/McDonalds" },
{ "url": "https://www.facebook.com/BBCNews" }
],
"cookies": "..."
}

Example Input — keyword + location discovery

{
"searchTerms": ["Coffee Shop", "Cafe", "Bakery"],
"locations": ["London", "Manchester", "Birmingham"],
"maxItems": 300,
"cookies": "..."
}

That single input produces 9 search combinations (3 terms × 3 cities), each yielding up to ~40 discovered URLs — giving a pool of up to ~360 candidates to fill 300 results from. startUrls and searchTerms can also be combined in the same run — results from both are merged into the same output dataset.

Cookies (Authentication)

Strongly recommended for any run.

Without session cookies, Facebook serves a compact login-wall stub for many discovered pages — you'll get some data but contact fields (phone, address, website) are gated. Providing cookies lets the scraper retry login-wall pages with an authenticated browser, dramatically increasing field coverage and yield.

How to export cookies from your browser:

  1. Log in to Facebook in Chrome or Firefox
  2. Open DevTools → Application → Cookies → facebook.com
  3. Export the cookies for .facebook.com using a browser extension like EditThisCookie (export as JSON) or Cookie Quick Manager on Firefox
  4. Paste the exported JSON array into the cookies field

Minimum cookies needed: c_user, xs, datr, fr, sb.

Accepted formats:

  1. JSON array (recommended — export directly from your browser extension):

    [
    { "name": "c_user", "value": "100074...", "domain": ".facebook.com", "path": "/", "expires": 1815212390, "httpOnly": false, "secure": true },
    { "name": "xs", "value": "8%3A2vaT...", "domain": ".facebook.com", "path": "/", "httpOnly": true, "secure": true },
    { "name": "datr", "value": "7JhPar...", "domain": ".facebook.com", "path": "/", "httpOnly": true, "secure": true }
    ]
  2. Cookie header string: "c_user=100074...; xs=8%3A...; datr=7JhP..."

Cookies expire (typically 90 days). If you see high skip rates, refresh your cookies.

Scaling to 1K+ Results

The actor is designed to scale — here is the fastest path to large datasets:

  1. Use multiple search terms and locations. Each (term, location) pair runs its own discovery pass. Example: 5 terms × 10 cities = 50 combinations × up to ~40 pages each = up to ~2,000 candidates in the pool.

  2. Set maxItems to your target. The actor over-discovers ~2× maxItems URLs to absorb login-walls and fetch failures without leaving you short.

  3. Provide session cookies. This is essential at scale — cookies enable Facebook's own search as the primary discovery source (far more reliable than external search engines) and let login-wall pages be retried with an authenticated browser.

  4. Vary your search terms. Generic terms (e.g. "Restaurant") produce many duplicate pages across locations. More specific terms (e.g. "Italian Restaurant", "Sushi Bar") yield more unique pages.

Example for 1K results:

{
"searchTerms": ["Restaurant", "Cafe", "Bar", "Pub", "Bistro"],
"locations": ["London", "Manchester", "Birmingham", "Glasgow", "Leeds", "Bristol", "Edinburgh", "Liverpool", "Sheffield", "Cardiff"],
"maxItems": 1000,
"cookies": "..."
}

50 search combinations, each yielding up to ~40 unique URLs — enough to fill 1K+ results (with cookies provided).

Output

Page Identity

FieldTypeDescription
facebookUrlStringFacebook page URL
pageUrlStringNormalized page URL
pageNameStringPage name or slug
pageIdString or nullFacebook page identifier
facebookIdString or nullFacebook ID (usually matches pageId)
titleString or nullFull page title with context

Classification

FieldTypeDescription
categoriesArray or nullPage categories
categoryString or nullPrimary category
infoArray or nullShort info lines from the page
introString or nullPage introduction or description

Contact & Business

FieldTypeDescription
phoneString or nullPublic phone number
emailString or nullPublic contact email
messengerString or nullMessenger link
addressString or nullAddress or coordinates
addressUrlString or nullGoogle Maps link for the address
websiteString or nullPrimary linked website
websitesArray or nullAll detected website links
priceRangeString or nullPrice range (e.g. "££")

Engagement

FieldTypeDescription
followersNumber or nullFollower count
likesNumber or nullLike count
followingsNumber or nullFollowing count
checkinsNumber or nullCheck-in count ("were here")
talkingAboutCountNumber or nullTalking-about count

Media

FieldTypeDescription
profilePictureUrlString or nullProfile picture URL
coverPhotoUrlString or nullCover photo URL

Business Details

FieldTypeDescription
business_hoursString or nullBusiness hours info
ratingOverallNumber or nullOverall rating score, synthesized to a 0-5 scale for compatibility (Facebook Pages actually use a recommend %, not 1-5 stars — see recommendPercentage)
ratingCountNumber or nullReview/rating count
ratingTextString or nullHuman-readable rating text
recommendPercentageNumber or nullRaw recommend percentage (0-100) as shown on the page — the authentic Facebook Pages rating metric
ownerOrganizationString or nullConfirmed owning organization/business ("X is responsible for this Page"), when Facebook discloses it
pageAdLibraryObject or nullAd library metadata or link
reviewsArray or nullReal reviews when available: [{reviewerName, recommends, reviewText, reviewDate}]. Controlled by the includeReviews input (default on)

Sample Output

{
"facebookUrl": "https://www.facebook.com/ozonecoffeeuk",
"pageUrl": "https://www.facebook.com/ozonecoffeeuk",
"pageName": "ozonecoffeeuk",
"pageId": "191234567",
"facebookId": "191234567",
"title": "Ozone Coffee Roasters | Facebook",
"categories": ["Coffee Shop"],
"category": "Coffee Shop",
"info": null,
"intro": "Specialty coffee roaster and cafe in London.",
"phone": null,
"email": null,
"messenger": "https://m.me/191234567",
"priceRange": null,
"address": null,
"addressUrl": null,
"website": null,
"websites": null,
"followers": 9100,
"followings": 12,
"checkins": null,
"talkingAboutCount": null,
"profilePictureUrl": "https://scontent.xx.fbcdn.net/v/...",
"coverPhotoUrl": "https://scontent.xx.fbcdn.net/v/...",
"ratingOverall": null,
"ratingCount": null,
"ratingText": null,
"recommendPercentage": null,
"ownerOrganization": null,
"business_hours": null,
"pageAdLibrary": "https://www.facebook.com/ads/library/?active_status=all&ad_type=all&country=ALL&view_all_page_id=191234567&search_type=page",
"reviews": [
{
"reviewerName": "Jane Smith",
"recommends": true,
"reviewText": "Great coffee and friendly staff!",
"reviewDate": "2025-11-02"
}
]
}

Unavailable Pages

When a discovered or provided URL turns out to be unavailable or deleted, it is detected and skipped entirely — it does not appear in the output dataset and does not consume your maxItems budget. This keeps the dataset free of junk rows; check the run log if you need to know which URLs were skipped and why.

Use Cases

  • Local business discovery: find businesses in specific cities or regions
  • Competitive intelligence: build datasets of competitor pages in your industry
  • Lead generation: collect pages with contact info for outreach
  • Market research: survey the Facebook page landscape for a category
  • Monitoring: run repeatedly to track new pages for a search term

FAQ

Do I need a Facebook account to use this?

No, but it's strongly recommended. Without cookies the scraper falls back to DuckDuckGo/Google discovery and extracts public HTML only. Providing session cookies from a logged-in Facebook account enables the most reliable discovery method (Facebook's own search), unlocks retry capability for login-wall pages, and gives access to contact fields that aren't in the public HTML.

Why are some fields null?

A null field means Facebook did not expose that information in the page's public HTML, the page owner did not fill it in, or the field layout for that page type doesn't include it. Fields like phone, address, and website are only populated when the page owner has listed them on their Facebook page. Adding cookies increases the chance of extracting these for login-gated pages.

Why did I get fewer results than maxItems?

The most common reasons:

  • Login walls: Facebook gates some pages behind a login session. Without cookies, these are skipped. Provide session cookies to retry login-wall pages with an authenticated browser.
  • No cookies provided: without session cookies, discovery relies solely on DuckDuckGo/Google, which are less reliable than Facebook's own search — add cookies for the best results.
  • Narrow search terms: each (term, location) pair yields a limited pool. Add more search terms or locations.

Can I get 1,000+ results?

Yes — see Scaling to 1K+ Results above. The short version: combine multiple search terms × multiple locations and set maxItems to your target. The actor handles all pagination and over-discovers URLs automatically.

How often do I need to refresh my cookies?

Facebook session cookies typically last 90 days. If you notice a sudden drop in results or many login-wall skips in the logs, export fresh cookies from your browser.

Does this work for Facebook Groups or personal profiles?

No. The scraper targets public Facebook Pages (business pages, community pages, public figures). Facebook Groups and personal profiles have different URL structures and content restrictions.

Is residential proxy required?

Residential proxy is used automatically for all Facebook page fetches and is already included — you do not need to configure or pay for it separately. The proxy cost is covered by your Apify usage credits.