# Changelog of LinkedIn Posts Scraper (`renovative_basilisk/linkedin-posts-scraper`) Actor

- **URL**: https://apify.com/renovative\_basilisk/linkedin-posts-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/renovative\_basilisk/linkedin-posts-scraper.md

## Changelog

All notable changes to this Actor are documented here.
The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

### \[1.1.2] - 2026-09-05

#### Fixed

- **A search with a time window returned nothing for a niche query.** The
  engines were asked for their all-time ranking of the query and `postedLimit`
  was applied afterwards, from the activity id. For a topic with a long tail
  that ranking is the posts that have gathered the most links — the oldest —
  so a month-limited run for `hopital americain de paris` found ten links,
  dropped all ten as out of window, and reported that discovery had
  succeeded. The window is now passed to each engine's own date filter as
  well (DuckDuckGo `df`, Bing's epoch-day range, Mojeek `since`, Google
  `tbs`), so the result page holds posts the run can use. The engine
  restriction is never narrower than the window — presets are used only when
  they span it whole, exact ranges are padded by a day at each end — so the
  activity-id check remains the one that decides.
- **DuckDuckGo's bot challenge was read as an empty result page.** It is
  served with HTTP 202, which Crawlee does not treat as a block, so the page
  was parsed, yielded no links, and the run moved on — on an IP that a
  rotation would have fixed. The challenge is now recognised, by status and
  by markup, and sent down the same path a 403 takes: a fresh proxy session,
  at most ten of them, then a counted failure. Google's "unusual traffic"
  interstitial and Mojeek's "automated queries" refusal are handled the same
  way.
- **The first result page is requested the way a browser requests it.**
  DuckDuckGo and Mojeek were asked for page one with an explicit `s=0`, which
  no browser ever sends; on the platform DuckDuckGo's challenge met that page
  on every run while page two got through on most. The offset is now omitted
  on the first page, and search requests no longer overlay hand-written
  Chrome-style `Accept` headers on the HTTP client's Firefox impersonation —
  the mismatch is what a bot check keys on. With both, DuckDuckGo's first page
  came back with results on the platform.
- **The Output tab in Console failed with "Expected property `clean` to be of
  type `boolean`".** The output schema linked the dataset as
  `…/items?clean=true`, and Console hands a template's query string to its
  dataset viewer as strings, where `clean` must be a boolean. The link is now
  the documented bare `…/items`; the viewer applies its own options.
- **The Overview and Engagement views showed only the post and its URL.**
  Their column lists named the nested objects (`author`, `postedAt`,
  `engagement`) while also flattening them, and the platform filters columns
  after flattening, when those names no longer exist. The lists now name the
  flattened columns the views display.
- **The log names the newest post the window dropped.** "11 posts fall
  outside the window" could mean a window one day too narrow or an engine
  index years stale, and the two need opposite fixes. The message now gives
  the newest dropped post's date and says which it is.

#### Added

- **Yahoo as a search engine, on by default.** It serves Bing's index as
  plain HTML with `site:` honoured, a working past-day/week/month filter and
  no challenge page — everything Bing itself now withholds from a client that
  is not a full browser, which answers `site:linkedin.com/posts` with general
  web results. Seven results a page, so `scrapePages` buys less of it.

#### Changed

- **Mojeek is no longer a default engine.** It answers 403 "automated
  queries" to datacentre and residential proxies and to a home connection
  alike. Still selectable.
- **Google is no longer a default engine.** Since early 2025 it answers every
  client that does not execute JavaScript — from any IP, with any User-Agent —
  with an "enable JavaScript" or "update your browser" page, so each Google
  page cost a proxied fetch and returned nothing. The defaults are now
  DuckDuckGo, Bing and Mojeek. `google` remains selectable; a run that selects
  it is told once, in the log, when that wall is served.

### \[1.1.1] - 2026-09-05

#### Fixed

- **A query written without accents returned nothing.** The keyword filter is
  applied to every post *after* it has been fetched, and it compared raw
  characters: `hopital americain de paris` did not match a post body that says
  "l’Hôpital Américain de Paris". Discovery was finding those posts and the
  run was paying to fetch them, then discarding every one and reporting an
  empty result set — the worst shape a failure can take, since it costs full
  price and looks like "there are no such posts". Matching now folds accents,
  ligatures and typographic punctuation on both sides, so either spelling
  finds the other and a curly apostrophe or quote reads as its ASCII form.
  Whole-word anchoring and phrase order are unchanged: `hopital` still does
  not match `hospital`. This affected every non-English query — French,
  Spanish, German and Portuguese most of all.
- **The margin guard killed runs that were earning.** It counted work on the
  cost side that it refused to count on the revenue side, stopping a healthy
  fifty-page discovery at attempt 40 with nothing delivered. See the commit
  for the four instances.

### \[1.1.0] - 2026-09-04

Monetization readiness. Every change here is on the money path: what gets
charged, what does not, and what a run is allowed to spend.

#### Fixed

- **`maxComments: 0` returned one comment.** The per-post limit was applied
  after the comment was appended rather than before, so a run asking for no
  comments got one anyway — and because the charged event is selected by
  whether the record carries comments at all, that single comment billed
  **every post of the run at the with-comments rate**.
- **Blocked pages were charged as results.** LinkedIn answers a throttled or
  signed-out request with its sign-in shell under HTTP 200, so nothing retried
  it, and the activity ID was still recoverable from the request URL — leaving
  a record with no body and no author that passed the keyword filter (which
  deliberately accepts empty text) and was pushed and charged. Such records are
  now discarded before the filter, with the slot returned to the `maxPosts`
  budget and a warning naming the likely cause.
- **A run that could not produce anything still charged for starting.** Clearing
  `searchEngines` without naming authors or post URLs left nothing to crawl; the
  run charged `actor-start`, crawled nothing and exited SUCCEEDED. It is now
  rejected during input validation, before the start charge. `targetUrls` is
  resolved rather than counted, so a list holding only unusable entries no
  longer passes the check.
- **A comment container that rendered empty was billed as a comment.** A
  skeleton `.comment` node — a throttled response, or markup that never
  hydrated — produced an entry whose every field was `null`, which was enough to
  select the with-comments event. Contentless entries are now dropped before
  they are counted, so a real comment further down the list keeps its slot.
- **Author pages were not covered by any spend limit.** The discovery cap
  applied only to search; 5,000 author handles with `maxPosts: 10` still
  enqueued 5,000 proxied fetches. Both routes now share one budget, spent on
  author pages first because they do not depend on a search engine.
- **Any run needing a second extraction round crashed on the platform.**
  `crawler.run()` purges the request queue by default, and Apify's queue client
  cannot purge — it drops and recreates the queue, changing its id under the
  running crawler, so the next round died with `Request Queue was not found`.
  A second round happens whenever the keyword filter rejects enough posts for
  the budget to top itself up, which is the normal case. Found by a real
  platform run; it cannot reproduce locally, because the memory queue client
  purges cleanly.
- **A blocked post could be fetched ten times and charged nothing.** Crawlee
  classifies `401`, `403` and `429` as a session block, checks for that before
  it checks the status code, and exempts session errors from
  `max_request_retries` entirely — so a single candidate that LinkedIn was
  refusing burned up to ten proxied fetches, the default `max_session_rotations`,
  and then reached the failed-request handler, which released the slot and
  charged nothing. Observed in a real run: 19 rotations across 9 discovery
  requests, two of them failing after exactly nine. Rotations are now bounded,
  on the reasoning that a host which has refused several fresh sessions in a row
  will refuse the next one too.
- **A failed post fetch was free.** The failed-request handler returned the
  slot and charged nothing, so the most expensive outcome — every retry spent,
  nothing to show — was the one outcome that billed zero. A post page that was
  fetched and could not be delivered is now charged as `post-filtered`, once,
  through the same accounting as every other terminal outcome.
- **A charge cap too low to cover the run was accepted anyway.** The charge
  limit stop could only fire once a charge was *attempted*, which is after the
  whole discovery phase has been crawled and paid for out of `actor-start`; a
  caller capping `maxTotalChargeUsd` at the $0.01 platform minimum therefore got
  a full discovery crawl for $0.005, on a healthy run rather than a pathological
  one. Nothing read `maxTotalChargeUsd` at all. The cap is now priced against the
  discovery the run has actually planned, before the start charge, and a run that
  cannot be attempted for the money it is allowed to charge fails immediately
  with the figure it would need — about $0.0115 for a default single-query run,
  $0.0805 at the 200-page discovery ceiling. Refusing to start is cheaper for
  everyone than a scrape delivered at a loss.
- **A run could spend past what it could ever be charged for.** Counting charges
  cannot govern this Actor, because its first chargeable event comes after
  discovery is over and paid for: every one of nine measured runs was
  $0.004–$0.005 underwater by then. A margin guard now counts proxied *attempts*
  — which is what money is actually spent on, and what a request-level counter
  collapsing a blocked page into one entry cannot see — and stops the crawl when
  the fetches still allowed could not close the gap. It reaches that judgement
  from four things adversarial review showed it could not do without:
  - **Discovery had to become governable.** Projecting revenue from `maxPosts`
    alone credited a blocked, candidate-less discovery with a hundred fetches'
    worth of future earnings, so the guard never fired on the exact run it was
    written for. What discovery has actually found is now a ceiling on what
    extraction can be projected to earn.
  - **Discovery needed its own warm-up.** Gating solely on post attempts left the
    phase that makes no post attempts completely ungoverned: a run whose every
    search page was blocked placed 400 proxied attempts, found nothing, reached
    no post attempt, and was never looked at.
  - **The cost of the next fetch is measured, not assumed.** Crediting every
    remaining fetch with one attempt made the guard most optimistic exactly when
    it was being blocked — and because remaining fetches scale with `maxPosts`, a
    larger cap became a larger licence to keep losing. Measured: at
    `maxPosts: 10,000` the guard missed tripping by four cents on a run that
    ended $1.247 down. Attempts still in flight are subtracted first, so a
    healthy run at full concurrency is not mistaken for a blocked one.
  - **Revenue buffered by `sortBy: date` counts as earned.** A matched post is
    held until the round ends, so for the length of a round the guard saw
    attempts climbing against nothing — and killed runs that were $0.036 in the
    black at the moment it killed them.
- **The stop message asserted blocking whatever had actually happened.** The
  guard fires for several unrelated reasons and two of them are runs in which
  every fetch succeeded — a keyword that matches nothing, or a repost chain in
  which ten permalinks canonicalise onto one post. Telling that customer to
  switch on residential proxies sends them to fix a setting that was never the
  problem and hides the one that was. The terminal status message and the log
  now name what was observed: no candidates found, fetches never returning,
  fetches being retried *N* times each, pages fetched but rejected by the keyword
  filter, or discovery simply costing more than the posts it found repay — each
  with the input to change.
- **A run with pay-per-event pricing not in effect scraped for free.** The guard
  added for unpriced events lived inside the `if is_pay_per_event` branch, so
  the one case it could not catch was the platform reporting no pay-per-event
  pricing at all — the run then did the entire job and charged nothing, looking
  healthy throughout. That case now fails loudly, on the same reasoning as the
  unpriced-event check: failing the first run is strictly cheaper than
  delivering every run free.
- **`pytest` could not collect the suite.** Nothing put the repo root on
  `sys.path`, so the bare `pytest` that CI runs failed at import while
  `python -m pytest` succeeded — meaning the new CI would have been red on its
  first run and every gate protecting the money path skipped.

#### Added

- **Startup verification of pay-per-event pricing.** Charging an event the
  Actor has no configured price for is silently free: the SDK falls back to a
  price of zero, logs one line, and still reports the charge as successful — so
  a typo or an incomplete monetization tab would deliver every paid record for
  nothing while looking like a healthy run. All four event names are now
  checked against the platform's pricing before the first charge — a post event
  priced at zero counts as missing, since giving the product away cannot be
  deliberate — and a run fails loudly rather than scraping for free.

- **A ceiling on discovery spend.** Search discovery fans out as queries ×
  authors × engines × pages, and `maxPosts` capped only the results, never that
  product — 100 queries with 10 authors at 20 pages each meant 80,000 proxied
  search fetches for a run that could return at most 100 posts. The number of
  search-result pages is now capped in proportion to `maxPosts`, truncating
  round-robin so the budget is spread across engines and queries instead of
  being spent entirely on the first of each. Proportional turned out not to be
  the same as affordable, so an absolute ceiling of **200 discovery pages** now
  binds on top of it: discovery earns nothing at all, and tying the cap to
  `maxPosts` alone let a large `maxPosts` authorise a crawl no plausible
  extraction could repay — precisely the run the margin guard then has to kill
  halfway through. `searchQueries` is capped at 10 entries in the input schema,
  author lists at 50 and `targetUrls` at 200.

- **Input caps that bound a retry storm.** `maxConcurrency`'s schema maximum
  drops from 20 to 10: `max_tasks_per_minute` is `maxConcurrency * 20`, so the
  old ceiling let a caller quadruple the rate at which blocked requests were
  retried against a host that was already refusing them — more `429`s, not more
  throughput. `maxPosts: 0` still means "no limit" so existing callers keep
  working, but it no longer resolves to the 10,000-post maximum: it resolves to
  **1,000**. Nobody typing 0 has costed the run out, and the old mapping handed
  that one keystroke a 5,000-page discovery crawl — $2.50 of unrecoverable proxy
  spend against the $0.004 net that `actor-start` pays. A caller who genuinely
  wants ten thousand can still type ten thousand, and having typed it has
  accepted the bill.

- Continuous integration (`.github/workflows/ci.yml`): lint, formatting, the
  offline suite and a real image build on every push and pull request. Nothing
  pinned the resolved `apify`/`crawlee` version inside the allowed range before.

- A populated `fields` schema in `.actor/dataset_schema.json`, describing every
  column the dataset views reference.

- Search-engine result page fixtures for all four engines, covering each one's
  click-tracking wrapper end to end — including Bing's base64 payload and the
  visible-citation fallback used when that payload will not decode. Nothing
  exercised a whole result page before. They are hand-built to each engine's
  documented format, not captured live; a real capture would be worth more.

- **A fourth chargeable event, `post-filtered` ($0.001).** Charged for a post
  page that was fetched and examined but not delivered — rejected by the keyword
  filter, or returned by LinkedIn without a post in it. The proxied fetch is the
  expensive part of a run and costs the same whether the post turns out to
  match, and a rejected post frees its slot to pull another candidate in behind
  it, so a query matching nothing was previously the cheapest run to ask for and
  the most expensive to serve: unbounded fetches against a single $0.005 start
  charge. Measured, a hundred-post run matching nothing cost $0.92 and earned
  $0.004. It now earns $1.20 against the same cost.

- **One fetch is billed, and freed, at most once.** `ResultBudget.release` is
  now keyed by activity ID and idempotent. A handler that raises is retried by
  Crawlee and, on exhausting its retries, also reaches the failed-request
  handler — so a single transient error while charging could release one post's
  slot several times, and every freed-but-never-reserved slot let the run claim
  more work than `maxPosts` allowed. Reproduced before the fix: four charges and
  four proxied fetches for one post, and a run delivering four posts on a
  `maxPosts` of three.

- **A delivered post's slot can no longer be freed.** `release` consulted only
  the rejected set, so a post that was written and charged and whose request
  then failed — a handler that delivers and is cancelled or times out still
  reaches the failed-request handler — had its slot handed back. The freed slot
  pulled another candidate in, and the run delivered and billed past `maxPosts`.
  Reproduced against a real crawler: `maxPosts: 2` producing six rows and six
  `post-scraped` charges, while the summary read "Scraped 0 post(s)". The same
  bug counted one post in two mutually exclusive tallies.

- **A refused delivery returns its slot.** Two discovered permalinks can
  canonicalise to one post — a repost resolves to the original — so a duplicate
  is reachable without any retry. It previously consumed a slot and was never
  delivered, so the run stopped short of `maxPosts` with candidates queued. It
  is not charged: a duplicate is this Actor's own discovery counted twice.

- **A post is written to the dataset at most once.** The same retry hazard
  existed on the delivery side and was worse there: `push_data` writes the item
  and only then charges, so a failure at the charging step left the row in place,
  escaped the handler, and had Crawlee re-fetch, re-write and re-charge — at
  `post-scraped`, the dearest event. Reproduced against a real crawler: four
  fetches, four duplicate rows and four charges for one post. Delivery is now
  claimed before the write and refused on a repeat, and neither the handler path
  nor the date-sorted flush can raise out of a push.

- **A charging failure can no longer fail the fetch.** Charging goes over the
  network and can fail for reasons that have nothing to do with the post. Left
  to propagate it made Crawlee retry the whole request, and each retry re-ran
  the charging path — turning one transient error into several paid fetches.
  Losing a single charge is the cheaper failure, and it is logged.

- **A ceiling on post fetches, at five times `maxPosts`.** The counterpart to the
  event above: because rejected posts free their slots, extraction would
  otherwise drain the whole candidate pool, and a narrow keyword could turn a run
  sized at a hundred posts into a bill for a thousand fetches. Five leaves a
  query matching one post in five room to fill the run, and caps what a caller
  can be surprised by. The run logs why it stopped and how to search more
  narrowly.

#### Changed

- Pricing set from measured cost rather than a competitor's list price:
  `post-scraped` is **$0.00175** ($1.75 per 1,000 posts) and
  `post-scraped-with-comments` is **$0.003**, `post-filtered` **$0.001**.
  `actor-start` stays at $0.005 and is documented as covering the discovery
  phase, which runs before any post is fetched. The README's comparison figures
  were wrong and have been corrected.
- **Every customer-facing figure regenerated from the code.** The README and the
  input schema were written alongside the limits rather than from them, and by
  the time the limits settled they described a product that did not exist: 0 as
  the 10,000-post maximum, 50,000 charged fetches, a 5,000-page discovery crawl,
  and no mention that discovery is capped at 200 pages for *every* `maxPosts`,
  not just 0. Documentation that overstates a spending ceiling is worse than
  none, because it is the number a customer sizes a run against. Both files are
  now derived from the constants, the four worked pricing rows recomputed against
  the live event prices, and the two unfilled placeholders — the session-rotation
  cap and the minimum viable `maxTotalChargeUsd` — filled in. The margin guard is
  described as what it is: a stop, not a guarantee, with the run's normal
  successful terminal state on one side and an outright pre-flight failure on the
  other, which the previous copy ran together.
- Test suite grown from 159 to 318 tests, covering each of the above — including
  a harness that drives the real `main()` against a local server, which is what
  found the margin guard's projection and warm-up gaps.

### \[1.0.0] - 2026-08-06

First public release.

#### Added

- Search public LinkedIn posts by keyword, with LinkedIn's Boolean syntax
  (`AND`, `OR`, `NOT`, quoted phrases, parentheses) evaluated against every
  post's text — so a returned post really does mention what you asked for.
- Two discovery routes, merged and de-duplicated by activity ID:
  - **Author pages** — the public profile (`/in/…`) or company page
    (`/company/…`) of any author you name. LinkedIn-native and independent of
    any third party.
  - **Web search** — DuckDuckGo, Bing, Google and Mojeek, restricted to
    `linkedin.com/posts`. This is what makes an open keyword search possible;
    LinkedIn's own post search is behind the login wall.
- Direct scraping of individual post URLs passed in `targetUrls`.
- Full post data from LinkedIn's signed-out rendering: body text, author
  (name, public identifier, URL, avatar, follower count), exact publication
  timestamp, reaction and comment counts, reaction types, media, and any linked
  LinkedIn article.
- Optional comments (`scrapeComments`) with text, author identity, exact
  timestamp and per-comment like count, merged from the JSON-LD island and the
  rendered markup because neither carries all of it.
- Exact publication times recovered from the activity ID itself, so
  `postedLimit` and `postedLimitDate` are applied **before** a post is fetched —
  an out-of-window post costs nothing — and `sortBy: date` orders the run rather
  than just the output.
- `maxPosts` cap enforced through a reservation-based budget that releases a
  slot when a post fails the keyword filter, so filtered-out posts do not
  consume the cap. Extraction runs in rounds and tops up from the remaining
  candidates, so a run whose first `maxPosts` candidates all fail the filter
  still returns results instead of an empty dataset.
- Embed-fragment fallback for a post discovered as a bare activity ID: LinkedIn
  404s a permalink whose slug and hash were not the original ones, so one cannot
  be synthesised.
- Pay-per-event charging: `actor-start`, `post-scraped`,
  `post-scraped-with-comments`. The run stops cleanly at a charge limit.
- Apify Proxy support, defaulting to the `RESIDENTIAL` group.
- `respectRobotsTxt` input, off by default, documenting the conflict between
  LinkedIn's `robots.txt` and this Actor's function.
- Stable dataset columns: records always carry the full field set, with
  unavailable values as `null`, so CSV and Excel exports keep one shape.
- Dataset views for an overview table, post text, and engagement.
- Offline test suite (159 tests) running against captured fixtures, including
  unit coverage of the charge-accounting and budget decisions.

#### Notes

- Reposts are collapsed onto the post being shared. A reshare card carries the
  reshare's ID and the original's ID separately; the permalink points at the
  original, so using the outer ID would scrape the same post twice.
- Per-reaction identities (who reacted, and with what) are **not** available to
  a signed-out visitor. Aggregate counts and the reaction types displayed are.
- Repost counts are not rendered to signed-out visitors either; `shares` is
  always `null` rather than guessed.
- No login, cookies, session tokens or member contact data are involved at any
  point.
