# Changelog of LinkedIn Company Posts Scraper - Bulk Export \[NO LOGIN] ✅ (`unseenuser/company-posts`) Actor

- **URL**: https://apify.com/unseenuser/company-posts/changelog.md
- **Full Actor documentation**: https://apify.com/unseenuser/company-posts.md

## Changelog

### \[1.4.3] - 2026-09-24

#### Changed - streaming output

- **Posts now stream to the dataset as each HarvestAPI page arrives**,
  instead of buffering the whole per-URL pool until every page is
  fetched. Users see rows appear in the dataset within seconds of
  starting the run, and aborted runs preserve every row already
  pushed instead of losing everything.
- **Applies automatically** when the run does not need the full pool:
  `sort_by` is `newest` (default) or empty, and none of
  `include_post_extras`, `include_company_report`, or
  `include_strategy_summary` are on.
- **Batch mode preserved** when the pool is required: `sort_by=top`
  (ranks across the fetched pool), `sort_by=oldest` (reverses the
  pool), or any of the aggregating enrichments. Behavior in batch
  mode is unchanged.

#### Trade-off in streaming mode

- The `total_posts_for_company` column is `null` on every streamed
  row (the total is unknown until the last page arrives, and Apify
  datasets are append-only so pushed rows cannot be updated).
  Consumers who need the total can count rows per `company_url`
  after the run. Every other field is populated identically to
  batch mode.

#### Backwards compatibility

- Every field key, default, output shape (except the null
  `total_posts_for_company` note above), and pricing event is
  unchanged.
- Every existing saved Task, schedule, and API caller keeps
  running. Runs that hit batch mode (`sort_by=top` or any
  aggregating enrichment) behave exactly as before.
- All 7 pay-per-event slugs fire identically. `url-processed`
  charges once per URL, `post-scraped` charges per delivered post,
  enrichment events charge per unit as before.

### \[1.4.2] - 2026-09-24

#### Changed - free-tier gating (breaking for free-tier users only)

- **Per-run post cap on the free Apify plan reduced from 50 posts to
  5 posts.** Applies only to free-tier accounts
  (`APIFY_USER_IS_PAYING != "1"`). Paid tier is unaffected.
- **All opt-in enrichments are now disabled on the free Apify plan.**
  When a free-tier run turns on `include_author_profile`,
  `include_post_extras` (or the deprecated
  `include_post_analytics` / `include_media_details` aliases),
  `include_comments`, `include_company_report`, or
  `include_strategy_summary`, the toggles are silently overridden to
  `false` and a status line names which were skipped. Paid tier is
  unaffected.

#### Why

Apify's Actor.charge does not collect revenue from free-tier
accounts, but the Actor still pays HarvestAPI (posts, author
profiles, comments) and users still pay Anthropic (AI summary) on
every run. On the previous 50-post cap with enrichments allowed,
a single free-tier run with author profile enrichment could cost
the publisher ~$0.11 in HarvestAPI usage with $0.00 in revenue. The
new gate makes the free tier a strict preview: up to 5 posts, no
enrichments, ~$0.002 cost per URL.

#### Backwards compatibility

- Every field key, default, output column, and pricing event is
  unchanged.
- Every existing saved Task, schedule, and API caller keeps running.
  Paid-tier callers behave identically. Free-tier callers see fewer
  rows (5 instead of 50) and their enrichment toggles are ignored
  with a status message.
- No pricing events changed. Free-tier runs simply do not fire the
  enrichment events (they were $0 to the publisher anyway).

### \[1.4] - 2026-09-09

#### Added

- **Post analytics enrichment** (`include_post_analytics: true`) -
  adds `engagement_rate` (likes+comments+shares / author\_followers),
  `post_performance_percentile` (0-100 rank within the fetched pool),
  `posted_day_of_week`, and `posted_hour_utc` columns to every post
  row. Billed per post via `post-analytics-added`.
- **Media details enrichment** (`include_media_details: true`) -
  adds `media_details_json` column with a JSON array of per-image
  and per-video metadata (URL, dimensions, order, video duration,
  article thumbnail). Billed per post that has media via
  `media-details-added`.
- **Comment scraping** (`include_comments: true` +
  `max_comments_per_post`) - populates `top_comments_json` +
  `top_comments_count` columns using HarvestAPI's
  `/linkedin/post-comments` endpoint. Default cap 0 = fetch every
  available comment on each post (no hard ceiling). Billed per
  comment fetched via `comment-scraped`.
- **Company report** (`include_company_report: true`) - populates
  `company_report_json` column on the first row per company URL
  with posting cadence, best posting day and hour, top hashtags/
  mentions, and company metadata (headline, industry, location,
  followers). Billed per unique company URL via
  `company-report-added`.
- **AI strategy summary** (`include_strategy_summary: true` +
  `anthropic_api_key`) - populates `strategy_summary` column on the
  first row per company URL with a 3-paragraph AI-generated content
  strategy summary. Uses Claude Haiku 4.5 via the user's own
  Anthropic API key (~$0.015/company in tokens paid directly to
  Anthropic). Billed per company via `strategy-summary-added`.

#### Compatibility

- No field key renamed or removed. All new inputs default to
  `false` / `0` / empty; existing scheduled Tasks are unaffected.
- New output columns default to `null` on rows where the toggle is
  off, so existing consumers are unaffected.
- All new billing events are opt-in and only fire when the
  corresponding toggle is on. Runs with no toggles set incur only
  the existing `url-processed` + `post-scraped` charges.
- 5 new billing events: `post-analytics-added` ($0.001/post),
  `media-details-added` ($0.001/post), `comment-scraped`
  ($0.001/comment), `company-report-added` ($1.00/company),
  `strategy-summary-added` ($2.00/company). Configure on Apify's
  Pricing tab before use.

All notable changes to this Actor are documented here. Format follows
Keep-a-Changelog. Dates are UTC. No field key is ever renamed or removed
between versions - existing saved Tasks, API integrations, and schedules
keep working.

### \[1.3] - 2026-09-06

#### Added

- **`hashtags` output column.** Array of hashtags extracted from each
  post's content. Populated via HarvestAPI's `contentAttributes`
  structured entities (`type: "HASHTAG"`), with a regex fallback for
  any tags HarvestAPI did not tag structurally. Case-insensitive
  dedupe. Empty array on posts with no hashtags. No competitor in
  the category returns this as a first-class column.
- **`mentions` output column.** Array of company/person names
  @-mentioned in each post's content. Extracted from
  `contentAttributes` entries flagged `COMPANY_NAME` or
  `PERSON_NAME`. Case-insensitive dedupe. Empty array on posts with
  no mentions. Directly answers competitor complaint
  `more-data--mentions`.
- **`content_contains` input.** Optional case-insensitive keyword
  filter on post content. Enter one or more keywords; only posts
  whose text contains at least one are kept. Empty array (default)
  means no filter. Directly answers competitor complaint
  `company-posts-search`. Note: media-only posts with no text are
  dropped when any keyword is set.
- **`sort_by` input.** Optional select field with values `newest`
  (default, matches previous behavior), `oldest`, and `top`. When
  `top` is chosen the Actor fetches up to 3x the requested pool
  before ranking by likes + comments + shares descending, so the
  returned "top N" is meaningful rather than "top N of the newest
  N". Default empty string preserves the previous newest-first
  ordering exactly - no existing scheduled Task or API integration
  changes behavior.
- **"What's new" callout** above the fold in the README - top 4
  v1.3 features (author profile enrichment, bare-handle input,
  absolute date ranges, per-URL failure summary) surfaced right after
  the version pill.
- **Workflow recipes section** - 5 copy-paste JSON blocks for common
  jobs (weekly competitor monitor, 100-company bulk export, deep
  engagement analysis, range window, original-content-only
  benchmarking). Every recipe runs on the current build with no
  editing.
- **Integrations section** - one-paragraph setups for Google Sheets,
  Airtable, Zapier, Make, and n8n.
- **API and CLI usage examples** - working snippets for curl, Python
  (apify-client), and Node.js (apify-client). Cover start-and-wait
  runs and dataset iteration.
- **Troubleshooting table** - 10 common symptoms mapped to root cause
  and one-liner fix (empty dataset, mixed-case slugs, 429 rate
  limits, missing API key, unexpected billing, cap at 50 posts,
  dedupe collapse, bare-handle edge cases).
- **Full output column reference table** - all 32 output columns with
  type, when-null behavior, and a concrete example value per column.
  Turns the row-shape claim into a spec analysts can code against.
- **Maintenance signals** rolled into the README: version pill, "Last
  verified working" date, 24-hour business-day support SLA, `0` open
  issues, and a Changelog section linking to this file with the last
  two releases summarised inline.
- **[SUPPORT.md](./SUPPORT.md)** documenting where to reach me, the
  response SLA, what is in and out of scope, and the "no field key
  renamed or removed" compatibility guarantee.
- **[SECURITY.md](./SECURITY.md)** documenting the responsible
  disclosure policy, secret-handling behavior, and the
  no-cookies / no-login guarantee.
- **`.github/REPO_METADATA.md`** with the About-blurb, website URL,
  and topic tags to paste into the GitHub repo Settings for
  discoverability.
- **"Last verified" annotations** on the Technical Details, Output
  Reference, Output Schemas, and Maintenance sections so buyers can
  see when each core claim was last re-checked.
- **`author_public_id` output column** - LinkedIn public identifier
  extracted from the author URL. Cheap join key for pairing post rows
  with author profile data from other tools.
- **`include_author_profile` input toggle** - when on, resolves each
  unique author's HarvestAPI profile once (cached across the run), and
  adds `author_headline`, `author_industry`, `author_location`, and
  `author_followers` columns to every post row. Optional pay-per-event
  hook: `author-profile-fetched` (no-op until configured on Apify).
- **Per-URL failure summary in end-of-run status message** - up to 5
  failed URLs listed by URL + HTTP status so the run's status bar
  surfaces which URLs to inspect without opening the log.
- **Cost preview on run start** - logs an up-front estimate: URLs at
  $/URL, up to N posts at $/post, plus any extras enabled. Actual
  billing is per real event delivered.

#### Removed

- **`scrape_reactions`, `max_reactions`, `scrape_comments`,
  `max_comments`, `comments_posted_limit` input fields.** These were
  the "deep data (extra cost per post)" toggles. Removed because:
  - `scrape_reactions` was redundant with the always-populated
    `reactions` column, which already contains the same per-type
    breakdown as a compact string.
  - `scrape_comments` did not work on HarvestAPI's
    `/linkedin/company-posts` endpoint: the response's comment array
    came back empty regardless of the parameter, so the feature had
    no effect for users.
- **`reactions_breakdown` and `comments_details` output columns**
  removed for the same reason. The `reactions` column keeps the
  per-type counts as a string ("LIKE:142, PRAISE:38, ...").
- **`reactions-fetched` and `comments-fetched` billing hooks**
  removed. Only `url-processed`, `post-scraped`, and
  `author-profile-fetched` remain.

Existing input JSON that still carries any of the removed keys is
silently ignored - no error - so saved Tasks that pin those values
continue to run without a schema change.

#### Fixed

- **Author profile enrichment now works for company authors too.**
  The previous build only enriched individual (/in/) authors, so runs
  on company URLs (where the author is the company itself) returned
  all-null enrichment columns. `include_author_profile: true` now
  additionally calls HarvestAPI's /linkedin/company for each unique
  company/school/showcase author and populates author\_headline
  (company tagline), author\_industry, author\_location (headquarters),
  and author\_followers.
- **`is_repost` detection** now catches HarvestAPI reshares that have
  a populated `repostedBy` object even when `repostId` is absent
  (common on company-page reshares). Before, `include_reposts: false`
  silently kept these rows because the detection missed them; now it
  correctly filters them out.
- **Case-insensitive URL dedupe** now normalizes `www.` and protocol
  in addition to case and trailing slash. Previously
  `https://linkedin.com/company/Google` and
  `https://www.linkedin.com/company/google` were treated as different
  URLs, which resulted in duplicate rows in the output.
- **`reactions_breakdown` column** now populates from the reactions
  array HarvestAPI returns at `engagement.reactions` when
  `scrape_reactions=true`. Previously the extraction looked at
  top-level `raw.reactionsDetail`/`raw.reactions` which HarvestAPI
  does not populate, so the column stayed null even with the toggle
  on.
- **`comments_details` column** now checks additional HarvestAPI
  response field candidates (`commentsList`, `postComments`, plus
  variants inside `engagement`). Same root cause as reactions.

#### Changed

- **Store title** now `LinkedIn Company Posts Scraper - Bulk Export [NO LOGIN] ✅` (57 chars). Swapped in `- Bulk Export` in place of
  the `[NO COOKIES &` prefix to capture the high-intent "bulk
  export" search term no competitor in the category claims. Kept the
  `[NO LOGIN]` differentiator (higher search value than `[NO
  COOKIES]`) and the ✅ marker. `no cookies` still appears in the
  store description and the README H1 so buyers searching for that
  term still hit the actor.
- **Store description** rewritten around the new title: names the
  three v1.3 columns (hashtags, mentions, author profile
  enrichment) that no competitor in the subcategory returns,
  advertises the >99% success rate up-front, and keeps `no login,
  no cookies` explicit.
- **Store description rewritten** for SEO and to reflect current
  features: names bulk scraping, engagement metrics, author profile
  enrichment, bare-handle input, and the no-login / no-cookies
  differentiators inside a single sentence.
- **Keyword tags expanded** from 6 to 14 to cover primary keyword
  variations (`linkedin company posts scraper`,
  `linkedin posts scraper`, `linkedin scraper`), differentiators
  (`linkedin no login`, `linkedin no cookies`), and intent categories
  (`competitor monitoring`, `content benchmarking`, `abm`, `b2b data`).
- **Retry policy hardcoded at 3 attempts** with 1s/2s/4s exponential
  backoff (worst-case 7s wait per URL). This is a deliberate choice to
  keep the user-facing surface simple and to cap HarvestAPI billing
  exposure - a per-run tunable would let mis-configured runs burn
  many extra calls with no upside.
- **Case-insensitive URL dedupe** in `normalizeUrls`. Users pasting the
  same LinkedIn URL in different casings (or with a trailing slash) no
  longer end up with duplicate rows.
- Retry list now covers `500` in addition to `429/502/503/504`. Retry
  delays follow exponential backoff based on `max_retries`.
- Retry log lines include per-URL context ("attempt 1/3 \[3/12]") so
  bulk runs show which URL each retry is for.

#### Compatibility

- No field key renamed or removed.
- All existing saved Tasks continue to run identically. `max_retries`
  and `include_author_profile` are optional; existing input JSON that
  omits them behaves as before (3 retries, profile enrichment off).
- New output columns (`author_public_id`, `author_headline`,
  `author_industry`, `author_location`, `author_followers`) default to
  `null` for existing consumers - no downstream schema needs to change.

### \[1.2] - 2026-09-06

#### Added

- 7 new input fields, all optional (no existing Task input needs to change):
  - `posted_after_date` (datepicker) - absolute lower cutoff
  - `posted_before_date` (datepicker) - absolute upper cutoff
  - `include_quote_posts` (boolean, default `true`)
  - `include_reposts` (boolean, default `true`)
  - `scrape_reactions` (boolean, default `false`)
  - `max_reactions` (integer, default `0` = all)
  - `scrape_comments` (boolean, default `false`)
  - `max_comments` (integer, default `0` = all)
  - `comments_posted_limit` (enum) - comment age window
- Output row gained 3 columns: `is_quote_post` (boolean),
  `reactions_breakdown` (JSON string, populated when
  `scrape_reactions=true`), `comments_details` (JSON string, populated
  when `scrape_comments=true`). Empty rows keep the same shape.
- `date_filter` gained `Last 2 years` and `Last 3 years` window options.
- `companyUrls` now accepts bare LinkedIn handles (e.g. `google`,
  `openai`) and partial URLs (e.g. `linkedin.com/company/google`) in
  addition to full URLs. Bare values auto-expand to
  `https://www.linkedin.com/company/<handle>`.
- Actor title now carries `[NO COOKIES] [NO LOGIN]` explicitly to
  match store-listing convention.
- Optional pay-per-event hooks: `reactions-fetched` and
  `comments-fetched`. No-op until configured on the Apify pricing tab.

#### Changed

- Numeric limit fields now follow the `0 = fetch everything` convention:
  - `max_posts` default `0` (unchanged), schema `maximum` removed
  - `max_reactions` default `100` -> `0`, schema `maximum` removed
  - `max_comments` default `50` -> `0`, schema `maximum` removed
    No hard ceiling in code either. Whatever positive number the user
    enters is what runs.
- Absolute date range wins over `date_filter`: if the user sets
  `posted_after_date` or `posted_before_date`, `date_filter` is ignored
  for that run and a warning is logged.
- Input schema fields grouped by `sectionCaption`:
  - top: `companyUrls`, `max_posts`, `date_filter`
  - `Date range (absolute cutoffs)`: `posted_after_date`,
    `posted_before_date`
  - `Include and exclude`: `include_quote_posts`, `include_reposts`,
    `skip_empty_rows`
  - `Deep data (extra cost per post)`: `scrape_reactions`,
    `max_reactions`, `scrape_comments`, `max_comments`,
    `comments_posted_limit`
- Titles and descriptions rewritten in plain English with a concrete
  example in each field.

#### Compatibility

- No input field key was renamed or removed.
- Every new input is optional; input JSON that only carries
  `companyUrls` runs identically to earlier versions.
- Existing saved Tasks that pin `max_reactions` or `max_comments` to
  a specific number continue to work; that number is still respected.

### \[1.1] - 2026-09-05

#### Added

- Per-URL pay-per-event billing (`url-processed`) alongside per-post
  (`post-scraped`), so URLs that return zero posts are not billed.
- `skip_empty_rows` input.

#### Changed

- Free-plan cap is now counted in posts (not dataset rows), still 50
  posts per run.
- API-key scrubbing centralised in `src/secret.ts` so keys never leak
  to logs or dataset rows.

### \[1.0] - 2026-09-02

#### Initial release

- LinkedIn Company Posts Scraper via HarvestAPI
  (`/linkedin/company-posts`).
- Inputs: `companyUrls`, `max_posts`, `date_filter`.
- One flat row per post, 26 top-level columns.
- Pay-per-event billing scaffold.
- Structured error rows (`type` / `message` / `fix`) instead of thrown
  errors, so a broken run still delivers a readable dataset.
