LinkedIn Post Detail Scraper
Pricing
from $1.70 / 1,000 results
LinkedIn Post Detail Scraper
Extract public LinkedIn posts without cookies or an account: full text, author, publish date, reaction and comment counts, the comments shown to anonymous visitors, and attached media. Accepts activity IDs or post URLs. HTTP-only, no login, no browser.
Pricing
from $1.70 / 1,000 results
Rating
0.0
(0)
Developer
Farhan Febrian Nauval
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
6 days ago
Last modified
Categories
Share
Extract public LinkedIn posts — no cookies, no account, no browser. Give it activity IDs or post URLs and it returns the full text, author, publish date, reaction and comment counts, the comments LinkedIn shows anonymous visitors, and any attached media.
Pairs with the LinkedIn Profile Scraper: feed its activityIds output straight into this
actor's posts input.
Input
| Field | Type | Default | Description |
|---|---|---|---|
posts | array | required | Activity IDs (7501466755261820928), /feed/update/urn:li:activity:<id> URLs, or /posts/…-activity-<id>-<hash> URLs |
includeComments | boolean | true | Include the comments rendered for anonymous visitors |
maxAttempts | integer | 3 | Retries on throttling, each with a fresh exit IP |
maxConcurrency | integer | 4 | Posts in parallel. Keep it low |
proxyConfiguration | object | RESIDENTIAL | Residential strongly recommended |
All three input forms resolve to the same post — verified: a bare ID and the matching
/posts/… URL returned identical reaction and comment counts.
Output
One row per input, always — including failures, so downstream joins never silently lose a key.
{"_input": "7501466755261820928","_source": "S1-jsonld","_scrapedAt": "2026-09-06T16:39:10Z","_host": "uk.linkedin.com","_attempts": 1,"activityId": "7501466755261820928","postUrl": "https://www.linkedin.com/posts/williamhgates_fda-approves-…",// --- raw schema.org node, upstream field names kept verbatim ---"@type": "SocialMediaPosting","headline": "This is great news in the fight against Alzheimer's.","articleBody": "…full post text…","datePublished": "2026-09-04T02:30:44.966Z","author": { "name": "Bill Gates", "url": "https://www.linkedin.com/in/williamhgates" },"comment": [ /* schema.org Comment nodes: text, datePublished, creator, likes */ ],"image": { "url": "https://media.licdn.com/…" },// --- stable names across post types (see below) ---"postType": "SocialMediaPosting","authorName": "Bill Gates","authorUrl": "https://www.linkedin.com/in/williamhgates","likeCount": 2908,"commentCount": 322,"commentsReturned": 10}
Two post types, two field names
LinkedIn renders a post under one of two schema.org types depending on its media, and the two disagree on field names:
| text / image post | video post | |
|---|---|---|
@type | SocialMediaPosting | VideoObject |
| author field | author | creator |
| body field | articleBody | description |
| extra fields | image, hasPart | contentUrl, duration, thumbnailUrl, transcript |
The raw node is passed through unchanged, so nothing upstream is lost. On top of it the actor
adds postType, authorName, authorUrl, likeCount and commentCount, which mean the same
thing for both — use those and you never have to branch on post type.
How it works
The <script type="application/ld+json"> block on the guest post page is the whole source —
it survives redesigns that break CSS selectors.
The post node is picked by matching the activity ID in its @id, not by taking the first
node of a matching type: a post page also embeds sidebar nodes of the same type for "related
posts", so picking by type alone would return a neighbour's post.
Block detection
Two things are deliberately not used:
- Status code alone. LinkedIn serves its guest sign-in wall as a
200with a full-size body, and answers a throttled request with999— its own status code, not an HTTP one. - Body substrings. Every LinkedIn page carries the LiX flag
data-recaptcha-v3-integration-lix-value, so a substring test for"captcha"marks good renders as blocked.
What is used: the final URL (a redirect to /authwall, /checkpoint/, /uas/login) and a
positive data marker — the presence of a post node carrying the right activity ID.
Deleted posts vs. blocked requests
LinkedIn does not 404 a post that is gone; it serves the same sign-in wall it serves for a
private one. The two are indistinguishable from the outside, so the actor reports a persistent
wall as not_found_or_private rather than blocked — telling you the proxy failed when the
post was simply deleted sends you debugging the wrong thing. A genuine block appears as 999
and keeps the blocked code.
Known limits
| Limit | Detail |
|---|---|
| Comments are capped | LinkedIn renders roughly the top 10 comments to anonymous visitors. commentCount is the true total, commentsReturned is what you actually got. There is no anonymous pagination past that |
| No reactor identities | Reaction counts are returned; who reacted is not exposed to anonymous visitors |
| Rate limiting is aggressive | A single IP is throttled to 999 within a few dozen requests. Use a residential proxy |
| Deleted vs private | Not distinguishable — both report not_found_or_private |
Related
- LinkedIn Profile Scraper — profiles, and the
activityIdsthat feed this actor - LinkedIn Company Profile Scraper — company overviews
- LinkedIn Jobs Scraper — public job postings