# Changelog of LinkedIn Scraper: Profiles, Companies, Posts & Jobs (`data_forge_org/linkedin-scraper`) Actor

- **URL**: https://apify.com/data\_forge\_org/linkedin-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/data\_forge\_org/linkedin-scraper.md

## Changelog

All notable changes to this Actor are documented here. Versions follow the Actor's `MAJOR.MINOR` build versions.

### 0.3 (2026-09-28)

LinkedIn jobs. The title is now "LinkedIn Scraper: Profiles, Companies, Posts & Jobs".

- **Job URLs** (`jobUrls`): `/jobs/view/<id>` (with or without the title), `/comm/jobs/view/…`, any jobs URL with `currentJobId=`, or a bare job id → one `job` item with its details, read from LinkedIn's 23–80 KB guest job page: description, employment type, industry, seniority, job function, applicants, pay range (parsed into `salary_min`, `salary_max`, `salary_currency`, `salary_period`), Easy Apply, `closed` for jobs that no longer accept applications, numeric company id, and the job poster when shown. An unknown id gives `NOT_FOUND`.
- **Jobs search URLs** (`jobSearchUrls`): LinkedIn jobs search URLs with any filters, job listing pages (`/jobs/<slug>`) and a company's *See jobs* link → one row per job (title, company, location, `posted_at`, logo), 10 jobs per request, up to `maxJobs` (1–1,000, default 25). A `/jobs/search` URL without a location covers the whole world; a job listing page is searched in the location LinkedIn picks for it (for example the United States), so add `location=` to it or use a `/jobs/search` URL for another location. Only the search parameters of a URL are used (keywords, location, `f_*` filters, distance, sort order): tracking and job-alert sign-in tokens never reach LinkedIn or `source_url`, and two alerts of the same saved search are one list. `/jobs/collections/…` gives `INVALID_INPUT`.
- **Company open jobs** (`includeCompanyJobs`): the jobs behind the company page's *See jobs* button (operation `companies.jobs`). A company without open jobs has no button and costs no extra request.
- **Job details for lists** (`includeJobDetails`): one more small request per job, and the row's `posted_at` is kept. A detail that fails every retry (or answers 404) still saves its row, with the list page as its `source_url`, plus the error item.
- **Paging**: up to 3 list pages fetched ahead per search; no page past `maxJobs`, the exact total or LinkedIn's 1,000 results. A list page blocked on every retry gives a `BLOCKED` item (a server or network error gives `INTERNAL_ERROR`) and the list goes on with the pages it would have fetched next (1, or up to 3 after a job listing page), so one failure doesn't drop the pages after it; it stops when the search's first list page fails (a job listing page's continuation goes on), when that next page fails too, and at once after an error LinkedIn would repeat (HTTP 4xx), so persistent failures cost a few requests, not the whole list.
- **Deduplication**: a job found by several searches or companies of a run is output and charged once (kept across migrations: a job is recorded only once its item is saved, and the record is saved again after the pages in flight finish), and a job listed in `jobUrls` is never also output as a row.
- **Pay-per-event**: new events `job` (a job with details, $1.00 per 1,000 down to $0.70 on Gold) and `job_card` (a list row without details, $0.30 down to $0.21). Each job is charged one event, never both.
- **Standby**: `GET /v1/jobs/enrich?li_job_url=` and `GET /v1/jobs/search?li_job_search_url=&page=&page_size=&include_job_details=` (a JSON array; request pages until `[]`), also accepting `linkedin_job_url` and `linkedin_job_search_url` as parameter names; `max_results=n` returns the first n jobs (1–100) in one call. `/v1/companies/enrich` takes `include_jobs`, `max_jobs` (1–100) and `include_job_details` and returns `jobs`. Job lists fetch up to 5 pages at a time and stop at the end of the list once a page shows it. A partial result, on any endpoint, answers 200 with what it has and the headers `X-Results-Incomplete` and `X-Results-Incomplete-Reason` (`blocked`, `error`, `budget`, `timeout`). The run's maximum cost is checked before anything is downloaded: a call it can't pay for answers 402 without any request to LinkedIn, and a job list stops at the jobs that fit. `INVALID_INPUT` messages name the API's endpoints and parameters.
- **Schemas**: dataset schema with the `job` entity, the `jobs.enrich`, `jobs.search` and `companies.jobs` operations, the job fields, a *Jobs* view and a *Job details* view with the description, salary and poster (`title` and `linkedin_job_url` added to *Overview*); input schema *Jobs* section; output schema `jobs` and `job_details` links; OpenAPI 0.3.0 with the job routes, `Job` schema and incomplete-result headers.
- **README**: LinkedIn jobs scraper sections (job data, how to scrape jobs, company jobs, details and salaries, daily monitoring), job output examples, Standby job endpoints, prices and FAQs.

### 0.2 (2026-09-27)

Fixes from the pre-release review.

- **Name**: the Actor is now `linkedin-scraper` (title "LinkedIn Scraper: Profiles, Companies & Posts"). API paths and the Standby URL use the new name: `<username>~linkedin-scraper`, `https://<username>--linkedin-scraper.apify.actor`.

- **README**: rewritten for search: keyword-first intro, question-style headings, Python and JavaScript sections, a comparison table, prices and a longer FAQ.

- **Standby**: each call now uses only its own query parameters. Before, the URLs saved in the run's input (for example a Task's prefilled examples) could be scraped and charged instead of the requested one.

- **Standby**: charged only after the answer is sent; a disconnect or timeout stops the call at once. A 402 now carries `error_label: LIMIT_REACHED`.

- **Max cost per run**: all events share one budget, reserved before each save, so pages finishing in parallel or mixing events (a company with its posts) can no longer save or charge beyond the limit. The final status message says when the run stopped on the limit.

- **Input**: a stray `%` in a handle no longer fails the whole run (it becomes an `INVALID_INPUT` item). Handles and URL prefixes are case-insensitive (`JBalada` = `jbalada`, fetched once). An encoded slash (`/in/x%2Fy`) no longer resolves to a different profile. `INVALID_INPUT` items carry the input value as `source_url` instead of `""`.

- **Posts**: `urn:li:share:` and `urn:li:ugcPost:` URLs are read from the ~24 KB embed page when comments are not requested. Plain reposts keep the counts shown on their card; a comment with only LinkedIn's hidden like counter gets `comment_like_count: 0`.

- **Memory**: default 512 MB, minimum 256 MB (measured peak for profile batches at 10 parallel pages was above 256 MB).

- **Docs and schemas**: company fields `featured_employees`, `affiliates`, `li_employees_url` and `li_job_search_url` added to the dataset and OpenAPI schemas; fields never produced (`repost_count`, `comment_reply_count`) removed; the OpenAPI spec now matches the server's responses (including 402, 405, 504).

### 0.1 (2026-09-27)

Initial release.

- **Scrapes the LinkedIn URLs you provide**, logged out: profiles (`/in/…`), company pages (`/company/…`) and posts (`/posts/…`, `/feed/update/…`, `/embed/…`, `/pulse/…` or a bare activity id). No account, no cookies, no search-engine discovery. School, directory and activity-feed pages are out of scope.
- **Fast and cheap**: plain HTTP with a browser TLS fingerprint (no headless browser), each page parsed once, 256 MB default memory, one request per URL. Posts without comments are read from the ~24 KB embed page.
- **From the same page, no extra request**: recent posts of a profile or company (`includeProfilePosts`, `includeCompanyPosts`), the first public comments of a post (`includeComments`), featured employees and similar pages of a company.
- **Output**: flat snake\_case field names (`li_*` for people and companies, `linkedin_*` for posts and comments) plus provenance fields (`entity_type`, `operation`, `input`, `source_url`, `scraped_at`). Unknown fields are omitted, never `null`. Failures produce `error` items (`INVALID_INPUT`, `NOT_FOUND`, `BLOCKED`, `INTERNAL_ERROR`).
- **Blocks handled**: HTTP 999, sign-in walls and 429s are retried on a fresh residential IP.
- **Standby API**: `GET /v1/people/enrich`, `/v1/companies/enrich`, `/v1/posts/enrich` with a JSON error body (`error_label`, `message`, `status_code`); charged only when the answer is delivered.
- **Pay-per-event** events: `person`, `company`, `post`, `comment`. Errors are never charged.
