Instagram Post Scraper avatar

Instagram Post Scraper

Pricing

from $1.49 / 1,000 results

Go to Apify Store
Instagram Post Scraper

Instagram Post Scraper

Pricing

from $1.49 / 1,000 results

Rating

0.0

(0)

Developer

Fertech

Fertech

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

1

Monthly active users

9 hours ago

Last modified

Categories

Share

Instagram Post & Reel Scraper

Give it Instagram post, reel or IGTV URLs. Get back structured data with exact like and comment counts — not the rounded "102K" the page shows.

No login, no cookies, no session tokens.

One flat rate, whatever Apify plan you are on. No plan ladder — see the pricing section of this listing for the current rate.

Why this one

Exact numbers, not rounded ones. Instagram's page says "24.8K". The payload behind it says 24,815 — and that is what this Actor reads.

No guessing. Where Instagram publishes no number, the field is null: never a substituted 0, never a different metric wearing the wrong name. videoPlayCount stays empty rather than being filled with the view count, because plays and views are not the same thing.

What you get

Feed it a list of Instagram post, reel or IGTV URLs, get one structured record per post.

  • Engagement: exact likes, exact comments, video views
  • Post: caption, hashtags, mentions, type, product type
  • Media: display dimensions, cover image, duration
  • Owner: username, id
  • Music: track metadata when Instagram publishes it

Input

{
"directUrls": [
"https://www.instagram.com/reel/CxAbc123DeF/",
"https://www.instagram.com/p/CyDef456GhI/"
],
"maxAttemptsPerUrl": 3
}

Share links with ?igsh=, ?igsi= or ?stkn= tracking parameters work as given — the parameters are stripped, so the same post submitted in two forms is fetched, delivered and charged once.

URLs prefixed with the owner's username — the form Instagram serves when you open a post from a profile, e.g. https://www.instagram.com/<username>/p/<code>/ or https://www.instagram.com/<username>/reel/<code>/ — are also accepted; the username is dropped and the post is normalised to its bare /p/ or /reel/ form.

Output

One record per URL, with every field below always present. Field names, types and shapes are exactly what the Actor emits; the values are illustrative:

{
"inputUrl": "https://www.instagram.com/reel/CxAbc123DeF",
"id": "3123456789012345678",
"type": "Video",
"shortCode": "CxAbc123DeF",
"caption": "sunset over the harbour 🌅 #travel #photography #goldenhour",
"hashtags": ["travel", "photography", "goldenhour"],
"mentions": [],
"url": "https://www.instagram.com/p/CxAbc123DeF/",
"commentsCount": 312,
"firstComment": "",
"latestComments": [],
"dimensionsWidth": 640,
"dimensionsHeight": 1136,
"originalWidth": null,
"originalHeight": null,
"displayUrl": "https://scontent.cdninstagram.com/v/t51.71878-15/000000000_0000000000000000_0000000000000000000_n.jpg?_nc_cat=…&oe=…",
"images": [],
"videoUrl": "",
"alt": "",
"likesCount": 24815,
"videoViewCount": 247282,
"timestamp": null,
"childPosts": [],
"ownerFullName": "",
"ownerUsername": "example_user",
"ownerId": "1234567890",
"productType": "clips",
"videoDuration": 18.7,
"musicInfo": { "uses_original_audio": true },
"isCommentsDisabled": false,
"videoPlayCount": null,
"paidPartnership": false
}

displayUrl is shortened above for readability. Real output carries the full signed CDN URL — several hundred characters, and valid for a limited time, so download the image rather than storing the link.

Nine fields are always empty. They are listed in the table below and explained under What this Actor does not return.

What the values mean

FieldNotes
likesCount, commentsCountExact numbers from the page payload — not the rounded "102K" the page displays.
dimensionsWidth / HeightThe display-size cover image (640×1136 above).
originalWidth / HeightAlways null — the source resolution is not published on the surface this Actor reads. dimensions* above is the display size.
videoDurationSeconds, fractional. Present for videos and reels only.
typeImage, Video, Sidecar (carousel), or Unknown.
captionRaw caption text. hashtags and mentions are parsed out of it.
altAlways "" — the accessibility description is not published.
musicInfoTrack metadata (artist, song, audio id) when Instagram publishes it. When it publishes none, { "uses_original_audio": null } — the Actor does not guess.
firstComment, latestCommentsAlways "" and [] — comment bodies are not published. commentsCount is still the true total.
timestampAlways null — the post's publication date is not published.
videoUrlAlways "" — no playable media URL is published. displayUrl (the cover image) is.
ownerFullNameAlways "" — only ownerUsername and ownerId are published.
isCommentsDisabledAlways false. Not published, so this is a default rather than a reading.
images, childPostsAlways [] — carousel children are not expanded.
videoViewCountA video's view count. null for images and carousels — Instagram publishes no count for those.
videoPlayCountAlways null — see below.
paidPartnershipAlways false — not present in the payload.

Views are not plays

videoViewCount is Instagram's view count. It is not the same as a play count, which counts replays — for one reel Instagram reported 429,413 plays against 247,282 views. Instagram publishes no play count on any public surface, so videoPlayCount is always null and this Actor will not substitute the view count for it. A number in the wrong field is worse than an empty one.

A number that is missing is null, never 0. A zero here always means a real zero. If Instagram did not publish a figure, this Actor says so rather than reporting a plausible-looking number you cannot distinguish from data.

Errors

A URL that cannot be scraped still produces a record, carrying error and errorDescription:

errorMeaningRetried
invalid-urlNot a post, reel or IGTV URLno
not-foundDeleted, private, or never existedno
unsupported-pageThe response was not an Instagram post page at allno — re-fetching returns the same page
networkConnection, TLS, timeout or proxy failureyes, up to maxAttemptsPerUrl
blockedInstagram served a page but withheld the post data from every IP tried, or returned HTTP 401/403/429yes, up to maxAttemptsPerUrl; the session rotates once it has failed enough times

Instagram intermittently serves a real page with the post data stripped out, depending on which IP asks. It is transient: the Actor detects it, switches to a different IP and retries, so it rarely reaches your dataset. If you do see blocked records, raising maxAttemptsPerUrl gives each URL more IPs to try.

Pricing: every submitted URL is charged once

See the pricing section of this listing for the current rate.

One URL in, one charge out.

Posts get deleted, accounts go private, and Instagram sometimes withholds a post's data. When a URL can't be scraped you get an error record instead of a post record, so you always know what happened to it — and it costs the same as a delivered one.

Retries and duplicates are not charged on top. You pay once per URL you submit, however many attempts it takes.

Requirements

Residential proxies are recommended. Instagram rate-limits repeated requests from one IP, and sometimes serves a real page with the post data stripped out depending on which IP asks.

The Actor uses Apify residential proxies by default, so there is nothing to configure — no proxy picker in the input form, though an API caller can still override it.

Residential proxies are not included in the Apify Free plan. On the Free plan this Actor stops with a clear message rather than spending your credit on requests that mostly cannot succeed.

Limitations

  • Single post URLs only. Profile, hashtag and location feeds are not supported.
  • Carousel child media is not expanded.

What this Actor does not return

Nine fields are always empty, because the surface this Actor reads does not publish them: timestamp, videoUrl, originalWidth, originalHeight, firstComment, latestComments, ownerFullName, alt and isCommentsDisabled.

They are present in every record, empty, rather than dropped — so existing integrations keep parsing. If you need the post date or a playable media URL, this Actor cannot give them to you.

Reading the run summary

Each run writes a summary to the log and to the key-value store, under RUN_SUMMARY:

requested 200 | delivered 187 (93.5%) | unpaid 0 | notfound 8 | failed 5 | blocked-responses 42 | callbackErrors 0
FieldWhat it counts
requestedURLs you submitted
deliveredURLs that produced a post record
notfoundDeleted, private or non-existent posts
failedURLs that ended in any other error record
blocked-responsesResponses where Instagram withheld the data — not URLs
unpaidURLs processed after the charge budget ran out
callbackErrorsRows that could not be written. Should always be 0

These do not sum to requested, and that is deliberate. Every URL ends as exactly one of delivered, notfound or failed. blocked-responses sits alongside them: one post withheld twice and then fetched adds 1 to delivered and 2 to blocked-responses.

Using the API

from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("YOUR_USERNAME/instagram-post-scraper").call(run_input={
"directUrls": ["https://www.instagram.com/reel/CxAbc123DeF/"],
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["likesCount"], item["ownerUsername"])

This Actor collects only publicly available data — information anyone can see without logging in. It does not access private accounts or login-walled content. You remain responsible for how you use the data, particularly regarding personal information and the GDPR. If you plan to process personal data, seek your own legal advice first.

Issues and requests

Found a bug, or want profiles, hashtags or full comment threads supported? Open an issue on the Actor's Issues tab — feature demand genuinely drives what gets built next.