# LinkedIn Post Detail Scraper (`xtracto/linkedin-post-detail-scraper`) Actor

Extract public LinkedIn posts without cookies or an account: full text, author, publish date, reaction and comment counts, the comments shown to anonymous visitors, and attached media. Accepts activity IDs or post URLs. HTTP-only, no login, no browser.

- **URL**: https://apify.com/xtracto/linkedin-post-detail-scraper.md
- **Developed by:** [Farhan Febrian Nauval](https://apify.com/xtracto) (community)
- **Categories:** Lead generation
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.70 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## LinkedIn Post Detail Scraper

Extract public LinkedIn posts — **no cookies, no account, no browser**. Give it activity IDs
or post URLs and it returns the full text, author, publish date, reaction and comment counts,
the comments LinkedIn shows anonymous visitors, and any attached media.

Pairs with the **LinkedIn Profile Scraper**: feed its `activityIds` output straight into this
actor's `posts` input.

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `posts` | array | *required* | Activity IDs (`7501466755261820928`), `/feed/update/urn:li:activity:<id>` URLs, or `/posts/…-activity-<id>-<hash>` URLs |
| `includeComments` | boolean | `true` | Include the comments rendered for anonymous visitors |
| `maxAttempts` | integer | `3` | Retries on throttling, each with a fresh exit IP |
| `maxConcurrency` | integer | `4` | Posts in parallel. Keep it low |
| `proxyConfiguration` | object | RESIDENTIAL | **Residential strongly recommended** |

All three input forms resolve to the same post — verified: a bare ID and the matching
`/posts/…` URL returned identical reaction and comment counts.

### Output

One row per input, always — including failures, so downstream joins never silently lose a key.

```jsonc
{
  "_input": "7501466755261820928",
  "_source": "S1-jsonld",
  "_scrapedAt": "2026-09-06T16:39:10Z",
  "_host": "uk.linkedin.com",
  "_attempts": 1,
  "activityId": "7501466755261820928",
  "postUrl": "https://www.linkedin.com/posts/williamhgates_fda-approves-…",

  // --- raw schema.org node, upstream field names kept verbatim ---
  "@type": "SocialMediaPosting",
  "headline": "This is great news in the fight against Alzheimer's.",
  "articleBody": "…full post text…",
  "datePublished": "2026-09-04T02:30:44.966Z",
  "author": { "name": "Bill Gates", "url": "https://www.linkedin.com/in/williamhgates" },
  "comment": [ /* schema.org Comment nodes: text, datePublished, creator, likes */ ],
  "image": { "url": "https://media.licdn.com/…" },

  // --- stable names across post types (see below) ---
  "postType": "SocialMediaPosting",
  "authorName": "Bill Gates",
  "authorUrl": "https://www.linkedin.com/in/williamhgates",
  "likeCount": 2908,
  "commentCount": 322,
  "commentsReturned": 10
}
```

#### Two post types, two field names

LinkedIn renders a post under one of two schema.org types depending on its media, and the two
disagree on field names:

| | text / image post | video post |
|---|---|---|
| `@type` | `SocialMediaPosting` | `VideoObject` |
| author field | `author` | `creator` |
| body field | `articleBody` | `description` |
| extra fields | `image`, `hasPart` | `contentUrl`, `duration`, `thumbnailUrl`, `transcript` |

The raw node is passed through unchanged, so nothing upstream is lost. On top of it the actor
adds `postType`, `authorName`, `authorUrl`, `likeCount` and `commentCount`, which mean the same
thing for both — use those and you never have to branch on post type.

### How it works

The `<script type="application/ld+json">` block on the guest post page is the whole source —
it survives redesigns that break CSS selectors.

The post node is picked by matching the **activity ID in its `@id`**, not by taking the first
node of a matching type: a post page also embeds sidebar nodes of the same type for "related
posts", so picking by type alone would return a neighbour's post.

#### Block detection

Two things are deliberately **not** used:

- **Status code alone.** LinkedIn serves its guest sign-in wall as a `200` with a full-size
  body, and answers a throttled request with `999` — its own status code, not an HTTP one.
- **Body substrings.** Every LinkedIn page carries the LiX flag
  `data-recaptcha-v3-integration-lix-value`, so a substring test for `"captcha"` marks good
  renders as blocked.

What is used: the **final URL** (a redirect to `/authwall`, `/checkpoint/`, `/uas/login`) and a
**positive data marker** — the presence of a post node carrying the right activity ID.

#### Deleted posts vs. blocked requests

LinkedIn does not `404` a post that is gone; it serves the same sign-in wall it serves for a
private one. The two are indistinguishable from the outside, so the actor reports a persistent
wall as `not_found_or_private` rather than `blocked` — telling you the proxy failed when the
post was simply deleted sends you debugging the wrong thing. A genuine block appears as `999`
and keeps the `blocked` code.

### Known limits

| Limit | Detail |
|---|---|
| Comments are capped | LinkedIn renders roughly the **top 10** comments to anonymous visitors. `commentCount` is the true total, `commentsReturned` is what you actually got. There is no anonymous pagination past that |
| No reactor identities | Reaction *counts* are returned; *who* reacted is not exposed to anonymous visitors |
| Rate limiting is aggressive | A single IP is throttled to `999` within a few dozen requests. **Use a residential proxy** |
| Deleted vs private | Not distinguishable — both report `not_found_or_private` |

### Related

- **LinkedIn Profile Scraper** — profiles, and the `activityIds` that feed this actor
- **LinkedIn Company Profile Scraper** — company overviews
- **LinkedIn Jobs Scraper** — public job postings

# Actor input Schema

## `posts` (type: `array`):

Activity IDs (7501466755261820928), /feed/update/urn:li:activity:<id> URLs, or /posts/...-activity-<id>-<hash> URLs. Feed the activityIds output of the LinkedIn Profile Scraper straight in. One row is emitted per entry, including for failures.

## `includeComments` (type: `boolean`):

Include the comments LinkedIn renders for anonymous visitors (the top 10, not the full thread - see the README). Costs no extra request.

## `maxAttempts` (type: `integer`):

Retries on throttling, each with a fresh exit IP and a different country subdomain.

## `maxConcurrency` (type: `integer`):

Posts fetched in parallel. Keep this low - LinkedIn throttles aggressively and returns HTTP 999 to a hot IP.

## `proxyConfiguration` (type: `object`):

Residential proxy is strongly recommended. LinkedIn rate-limits a single IP to HTTP 999 within a few dozen requests.

## Actor input object example

```json
{
  "posts": [
    "7501466755261820928"
  ],
  "includeComments": true,
  "maxAttempts": 3,
  "maxConcurrency": 4,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `posts` (type: `string`):

Dataset items shown in the 'Posts' view.

## `items` (type: `string`):

Every record this run produced, with all fields, as JSON.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "posts": [
        "https://www.linkedin.com/feed/update/urn:li:activity:7501466755261820928"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("xtracto/linkedin-post-detail-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "posts": ["https://www.linkedin.com/feed/update/urn:li:activity:7501466755261820928"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("xtracto/linkedin-post-detail-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "posts": [
    "https://www.linkedin.com/feed/update/urn:li:activity:7501466755261820928"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call xtracto/linkedin-post-detail-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,xtracto/linkedin-post-detail-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9oBSakmgUC7BwEYk9/builds/uvSgWmkX1qOb4404T/openapi.json
