# Changelog of Facebook Ads Library Scraper v2 (`prodiger/facebook-ads-library-scraper-v2`) Actor

- **URL**: https://apify.com/prodiger/facebook-ads-library-scraper-v2/changelog.md
- **Full Actor documentation**: https://apify.com/prodiger/facebook-ads-library-scraper-v2.md

## Changelog

All notable changes to this actor will be documented in this file.

### \[0.4.0] — 2026-09-29

#### Fixed — page URLs and `@handles` now return the page's own ads

- **Root cause:** a Facebook page URL (e.g. `https://www.facebook.com/weareallbirds`) or `@handle` was turned into a *keyword* search on the vanity name (`q=weareallbirds&search_type=keyword_unordered`). Few advertisers mention their own vanity handle, so Meta suppressed the search-results query, every bootstrap attempt ended in "no search-results connection captured (soft shadow-ban)", and the run returned 0 ads. When the keyword did match, results included unrelated advertisers.
- **Fix:** page URLs and `@handles` are now resolved to the numeric Ad Library page ID by fetching the public page over the (residential) proxy and reading `delegate_page.id` / `associated_page_id` (legacy pages: `pageID`, `fb://page/?id=`). The actor then scrapes `…/ads/library/?view_all_page_id=<id>&search_type=page`, so every row's `page_id` equals the page. The profile ID (`fb://profile/100…`) is deliberately ignored — the Ad Library does not key on it.
- Lookups retry up to 4 times on a fresh IP when Meta serves a login wall. If all fail, a numeric ID from the input (`profile.php?id=`, `/pages/Name/<id>`) is used, otherwise the actor falls back to the old keyword search with a warning.
- The resolved page ID and name are logged (`Resolved page "…" → page_id=…`).
- Unchanged: input schema, output fields, pricing. Plain keywords (including bare words without `@`) and full Ad Library URLs behave exactly as before. The lookup never writes to the dataset, so it adds no pay-per-event charges.
- New `src/page-resolver.ts` with 15 unit tests (85 tests total).

### \[0.3.0] — 2026-08-21

#### Fixed — the actor now actually returns ads by default

- **Default proxy is now RESIDENTIAL (was datacenter).** This was the cause of a ~95% failure rate (58 FAILED / 61 runs over the preceding 30 days). `useApifyProxy: true` with no explicit group resolves to Apify's shared **datacenter** pool, which Meta blocks on the Ad Library: every bootstrap attempt came back "Meta rate-limit", the run burned all 6 IP rotations and then died on `ERR_TUNNEL_CONNECTION_FAILED`.

  Verified head-to-head on 2026-08-21 with an identical query, changing only the proxy:

  | Proxy | Result |
  | --- | --- |
  | `RESIDENTIAL` (explicit) | **SUCCEEDED** — 5/5 ads in 104 s |
  | default (datacenter) | **FAILED** — 0 ads after 375 s |

  When the user picks no proxy group we now select `RESIDENTIAL` (and `US`, as before). Explicit `apifyProxyGroups` / `apifyProxyCountry` are always respected.
- **Graceful fallback when residential is unavailable.** Residential access is plan-gated; if we injected `RESIDENTIAL` ourselves and the account can't use it, the run falls back to the user's original proxy config with an explanatory warning instead of hard-failing at startup.
- **Input schema default and README updated** so the Console pre-selects residential US, and the no-proxy warning now states plainly that datacenter runs will likely return no ads.
- Proxy defaulting extracted into `src/proxy.ts` as a pure function with 7 unit tests (70 tests total).

### \[0.2.0] — 2026-05-19

#### Pagination

- **Replaced 30-cap DOM-scroll path with GraphQL cursor pagination.** The previous implementation called `page.mouse.wheel()` (which fired at default mouse position (0,0), not over Meta's virtualized scroll container) plus a 3-iteration stable-count early-exit. The result was a hard cap at ~30 ads per query regardless of `maxAdsPerQuery`. The new path opens one Playwright page per query, intercepts the first `/api/graphql/` request to harvest `lsd` / `fb_dtsg` / `doc_id` / cookies, closes the browser, and walks `AdLibrarySearchPaginationQuery` over HTTP using `got-scraping` until `has_next_page === false` or the user-supplied cap is reached. `maxAdsPerQuery: 500` now actually returns up to 500 ads for high-volume advertisers.
- **`doc_id` sniffed at bootstrap, not hardcoded.** Meta rotates query identifiers on every build; we read whatever is current from the live request rather than embedding a value that goes stale.
- **`forwardCursor` / `cursor` treated as opaque.** The captured `variables` object is reused verbatim with only the cursor key mutated per page.
- **Retry-with-exponential-backoff on 5xx / 429 / FB rate-limit error codes** (1675004 / 1357004). Three attempts per cursor, capped at 30s backoff, then a hard error so the run fails loudly instead of silently truncating.
- **800–2000ms jitter between pagination calls** to stay under Meta's anonymous-session throttle.
- **`nextPageCursor` surfaced on the final dataset row** when the run was stopped by `maxAdsPerQuery` (or `maxAds`) before the connection ended. Opaque resume token; absent when pagination ran to completion.

#### Removed

- `extractAdObjects()`, `findJsonObjectEnd()`, `scrapeRenderedAds()`, `MAX_SCROLLS_PER_SOURCE`, `SCROLL_WAIT_MILLIS` — DOM HTML-grep and scroll-loop helpers no longer needed.

#### Test infrastructure

- Added vitest with 41 tests covering URL parsing (`test/url.test.ts`), token extraction (`test/token-extraction.test.ts`), paginator cursor loop / retry / dedup / schema drift (`test/paginator.test.ts`), and runner orchestration (`test/runner.test.ts`). Live Meta calls remain a manual smoke test via `apify run`.

***

### 2026-04-29 — input simplification

- Added `query` as the headline input — accepts keywords (`crochet`), page handles (`@ZapierApp`), page URLs, or full Ad Library URLs.
- Renamed flat fields: `country`, `activeStatus`, `adType`, `mediaType`, `period`, `sortBy`, `maxAds`, `maxAdsPerQuery`. The old `scrapePageAds.*`, `urls`, `count`, and `limitPerSource` keys are still accepted for backwards compatibility.
- Added `mediaType` and `adType` filters that map to Ad Library URL parameters.
- `period` filter now applies to keyword searches too (not only page-ad scrapes).
- Each output row now includes a `query` field so multi-search runs can be split apart later.

### 2026-04-29 — initial

- Initial compatible actor for `curious_coder/facebook-ads-library-scraper`.
- Copied the public input schema, memory sizing, Store metadata, dataset shape, and pay-per-ad pricing event.
- Added a direct Playwright scrape path for rendered Meta Ad Library results.
