LinkedIn Post Detail Scraper avatar

LinkedIn Post Detail Scraper

Pricing

from $1.70 / 1,000 results

Go to Apify Store
LinkedIn Post Detail Scraper

LinkedIn Post Detail Scraper

Extract public LinkedIn posts without cookies or an account: full text, author, publish date, reaction and comment counts, the comments shown to anonymous visitors, and attached media. Accepts activity IDs or post URLs. HTTP-only, no login, no browser.

Pricing

from $1.70 / 1,000 results

Rating

0.0

(0)

Developer

Farhan Febrian Nauval

Farhan Febrian Nauval

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

1

Monthly active users

6 days ago

Last modified

Categories

Share

Extract public LinkedIn posts — no cookies, no account, no browser. Give it activity IDs or post URLs and it returns the full text, author, publish date, reaction and comment counts, the comments LinkedIn shows anonymous visitors, and any attached media.

Pairs with the LinkedIn Profile Scraper: feed its activityIds output straight into this actor's posts input.

Input

FieldTypeDefaultDescription
postsarrayrequiredActivity IDs (7501466755261820928), /feed/update/urn:li:activity:<id> URLs, or /posts/…-activity-<id>-<hash> URLs
includeCommentsbooleantrueInclude the comments rendered for anonymous visitors
maxAttemptsinteger3Retries on throttling, each with a fresh exit IP
maxConcurrencyinteger4Posts in parallel. Keep it low
proxyConfigurationobjectRESIDENTIALResidential strongly recommended

All three input forms resolve to the same post — verified: a bare ID and the matching /posts/… URL returned identical reaction and comment counts.

Output

One row per input, always — including failures, so downstream joins never silently lose a key.

{
"_input": "7501466755261820928",
"_source": "S1-jsonld",
"_scrapedAt": "2026-09-06T16:39:10Z",
"_host": "uk.linkedin.com",
"_attempts": 1,
"activityId": "7501466755261820928",
"postUrl": "https://www.linkedin.com/posts/williamhgates_fda-approves-…",
// --- raw schema.org node, upstream field names kept verbatim ---
"@type": "SocialMediaPosting",
"headline": "This is great news in the fight against Alzheimer's.",
"articleBody": "…full post text…",
"datePublished": "2026-09-04T02:30:44.966Z",
"author": { "name": "Bill Gates", "url": "https://www.linkedin.com/in/williamhgates" },
"comment": [ /* schema.org Comment nodes: text, datePublished, creator, likes */ ],
"image": { "url": "https://media.licdn.com/…" },
// --- stable names across post types (see below) ---
"postType": "SocialMediaPosting",
"authorName": "Bill Gates",
"authorUrl": "https://www.linkedin.com/in/williamhgates",
"likeCount": 2908,
"commentCount": 322,
"commentsReturned": 10
}

Two post types, two field names

LinkedIn renders a post under one of two schema.org types depending on its media, and the two disagree on field names:

text / image postvideo post
@typeSocialMediaPostingVideoObject
author fieldauthorcreator
body fieldarticleBodydescription
extra fieldsimage, hasPartcontentUrl, duration, thumbnailUrl, transcript

The raw node is passed through unchanged, so nothing upstream is lost. On top of it the actor adds postType, authorName, authorUrl, likeCount and commentCount, which mean the same thing for both — use those and you never have to branch on post type.

How it works

The <script type="application/ld+json"> block on the guest post page is the whole source — it survives redesigns that break CSS selectors.

The post node is picked by matching the activity ID in its @id, not by taking the first node of a matching type: a post page also embeds sidebar nodes of the same type for "related posts", so picking by type alone would return a neighbour's post.

Block detection

Two things are deliberately not used:

  • Status code alone. LinkedIn serves its guest sign-in wall as a 200 with a full-size body, and answers a throttled request with 999 — its own status code, not an HTTP one.
  • Body substrings. Every LinkedIn page carries the LiX flag data-recaptcha-v3-integration-lix-value, so a substring test for "captcha" marks good renders as blocked.

What is used: the final URL (a redirect to /authwall, /checkpoint/, /uas/login) and a positive data marker — the presence of a post node carrying the right activity ID.

Deleted posts vs. blocked requests

LinkedIn does not 404 a post that is gone; it serves the same sign-in wall it serves for a private one. The two are indistinguishable from the outside, so the actor reports a persistent wall as not_found_or_private rather than blocked — telling you the proxy failed when the post was simply deleted sends you debugging the wrong thing. A genuine block appears as 999 and keeps the blocked code.

Known limits

LimitDetail
Comments are cappedLinkedIn renders roughly the top 10 comments to anonymous visitors. commentCount is the true total, commentsReturned is what you actually got. There is no anonymous pagination past that
No reactor identitiesReaction counts are returned; who reacted is not exposed to anonymous visitors
Rate limiting is aggressiveA single IP is throttled to 999 within a few dozen requests. Use a residential proxy
Deleted vs privateNot distinguishable — both report not_found_or_private
  • LinkedIn Profile Scraper — profiles, and the activityIds that feed this actor
  • LinkedIn Company Profile Scraper — company overviews
  • LinkedIn Jobs Scraper — public job postings