Instagram Scraper — Profiles, Posts & Comments | from $2/1K avatar

Instagram Scraper — Profiles, Posts & Comments | from $2/1K

Pricing

from $1.99 / 1,000 profile items

Go to Apify Store
Instagram Scraper — Profiles, Posts & Comments | from $2/1K

Instagram Scraper — Profiles, Posts & Comments | from $2/1K

**Scrape public Instagram profiles, posts, and comments** without login. Returns followers, captions, like/comment counts, media URLs, timestamps, and `parse_confidence` drift signal in every record. Uses Instagram's stable mobile API — no doc_id rotation, no silent failures.

Pricing

from $1.99 / 1,000 profile items

Rating

0.0

(0)

Developer

Vitalii Bondarev

Vitalii Bondarev

Maintained by Community

Actor stats

0

Bookmarked

6

Total users

3

Monthly active users

10 hours ago

Last modified

Share

Instagram Scraper — Profiles, Posts & Comments

Scrape public Instagram profiles, posts, and comments using Instagram's stable mobile API — no doc_id rotation, no maintenance windows, no login required for profiles and posts. Returns 15+ fields per record including followers, captions, like counts, media URLs, timestamps, and parse_confidence in every record.

Most Instagram scrapers use GraphQL endpoints with doc_id parameters that rotate every 2–4 weeks, causing silent failures. This actor calls Instagram's mobile API endpoints (/api/v1/users/web_profile_info/, /api/v1/feed/user/{id}/) that have been stable for years — the same endpoints used by the official iOS app.


What you get

Profile mode (mode: "profile")

FieldDescription
usernameInstagram handle
user_idNumeric user ID
full_nameDisplay name
followersFollower count
followingFollowing count
posts_countTotal post count
is_verifiedBlue checkmark
is_privatePrivate account flag
bioProfile biography text
external_urlLink in bio
profile_pic_urlProfile picture URL
parse_confidenceData quality score (0.0–1.0)
warningsMachine-readable quality codes
scraped_atISO-8601 timestamp

Posts mode (mode: "posts" / "post_details")

FieldDescription
shortcodePost code (e.g. DYhkH24lf3j)
media_pkNumeric media ID
usernameAuthor handle
user_idAuthor numeric ID
captionPost caption text
media_typephoto / video / carousel
media_type_codeRaw code: 1=photo, 2=video, 8=carousel
like_countLike count
comment_countComment count
play_countVideo view count (None for photos)
taken_atPost timestamp (ISO-8601 UTC)
display_urlBest image URL
video_urlVideo file URL (videos only)
location_nameLocation tag name
location_cityLocation city
parse_confidenceData quality score (0.0–1.0)
warningsMachine-readable quality codes
scraped_atISO-8601 timestamp

Comments mode (mode: "comments")

Requires sessionId — Instagram blocks anonymous comment requests. Provide your own IG sessionid cookie. See input options below.

FieldDescription
comment_pkNumeric comment ID
media_pkParent post ID
textComment text
authorCommenter handle
author_idCommenter numeric ID
like_countComment like count
created_atComment timestamp (ISO-8601 UTC)
parse_confidenceData quality score (0.0–1.0)
warningsMachine-readable quality codes
scraped_atISO-8601 timestamp

Input

ParameterTypeDefaultDescription
modeprofile/posts/post_details/commentspostsWhat to scrape
usernamesstring[]—Handles to scrape (profile/posts modes)
postUrlsstring[]—Post URLs for comments mode
shortcodesstring[]—Post shortcodes for comments mode
maxItemsinteger50Max records per username/post
sessionIdstring—Your IG sessionid cookie (comments only)
proxyConfigurationobjectRESIDENTIALApify proxy config

Example input — profiles

{
"mode": "profile",
"usernames": ["natgeo", "nasa", "instagram"],
"maxItems": 50
}

Example input — posts

{
"mode": "posts",
"usernames": ["natgeo"],
"maxItems": 50
}

Example input — comments

{
"mode": "comments",
"postUrls": ["https://www.instagram.com/p/DYhkH24lf3j/"],
"maxItems": 100,
"sessionId": "YOUR_SESSIONID_COOKIE"
}

  1. Log in to Instagram in your browser (Chrome/Firefox)
  2. Open DevTools → Application → Cookies → https://www.instagram.com
  3. Find the cookie named sessionid — copy its Value
  4. Paste it into the sessionId input field

Important: Use your own account. Sessions last ~90 days before expiring. This actor never stores or shares your cookie — it is only used for the current run.


Why RESIDENTIAL proxy is required

Instagram blocks datacenter IP ranges immediately. When running on Apify cloud (datacenter IPs), RESIDENTIAL proxy is mandatory. The default proxy config (useApifyProxy: true, apifyProxyGroups: ["RESIDENTIAL"]) handles this automatically — the proxy cost is billed to your Apify account at standard rates.


Why this scraper beats the alternatives

This scraper@data-slayer/instagram-post-details ($2,648/mo)@krazee_kaushik (profile+comments)
Price~$1.50–2.00/1khigherhigher
No doc_id rotation maintenance✓✗✗
parse_confidence in every record✓✗✗
Graceful degradation without sessionId✓✗partial
Both feed + profile node shapes parsed✓partialpartial
play_count (video views)✓partial✗

Key advantage: Most competitors use GraphQL endpoints with doc_id parameters that rotate every 2–4 weeks, causing silent failures until manually patched. This actor uses Instagram's stable mobile API endpoints — no doc_id rotation, no maintenance window.


Technical notes

  • Uses lightweight HTTP requests with proper browser headers (requests/httpx are blocked by Instagram)
  • Profile + posts: anonymous access via /api/v1/users/web_profile_info/ and /api/v1/feed/user/{id}/
  • Comments: requires sessionid cookie via /api/v1/media/{pk}/comments/
  • Proxy: Apify RESIDENTIAL (required for cloud runs; buyer-paid)
  • parse_confidence (0.0–1.0) in every record — schema drift is visible in the dataset
  • Rate-limit aware: exponential backoff on 429, retries on transient errors
  • Shortcode ↔ media_pk conversion via base64url (A-Za-z0-9 + -_)

Pricing

~$1.50–2.00 per 1,000 records (PPE — pay per result, no per-run fees).


Use with AI agents (MCP)

This scraper is callable as a tool by AI agents (Claude, Cursor, n8n, CrewAI) via Apify's MCP server. Minimal agent call:

{
"mode": "posts",
"usernames": ["natgeo"],
"maxItems": 10
}

parse_confidence (0.0–1.0) in every record lets agents filter low-quality rows without manual inspection.


Not affiliated with Instagram or Meta. For personal use and research.

Integrations

Built for social-listening and influencer-marketing teams pulling profile stats, post engagement, and comment data at scale — the JSON/dataset output drops into the tools you already run, no glue code:

  • n8n / Make / Zapier — trigger a run or pipe every new dataset item into 500+ apps (Google Sheets, Airtable, Slack, HubSpot, your database) with no code: n8n, Make, Zapier.
  • Webhooks — fire your own endpoint the moment a run finishes, to push results straight into your pipeline (docs).
  • MCP server — expose this actor as a tool to Claude, Cursor, or any MCP client so an AI agent can pull this data mid-conversation (guide).
  • API & SDKs — fetch the dataset as JSON, CSV, or Excel through the Apify REST API or the Python / JS SDKs.

See all Apify integrations.

Usage statistics

This Actor creates a small, content-free summary at the end of each run. It is used only to monitor reliability and improve this Actor. A copy is saved as USAGE_STATS in your own Apify key-value store, so you can see the exact record created for your run.

Set disableUsageStats to true in the input to opt out. Nothing is sent then; your USAGE_STATS record only says that statistics were disabled.

Only these fields are recorded:

  • schema version, Actor name and build number;
  • UTC start and finish hour (not a precise timestamp);
  • run duration, number of results and time to the first result, each as a coarse range;
  • whether the result was empty, the end status, and an error type from a fixed list;
  • memory setting and counts of charged events;
  • names of the input fields you set, never their values;
  • the selected option for input fields that offer a fixed list of choices (for example a sort order).

We do not collect input text, search terms, URLs, domains, usernames, email addresses, names, proxy credentials, tokens, scraped records, output items, raw error messages, stack traces, or hashes of any of those values. Records are kept for no longer than 13 months, used only as aggregated operational statistics, and never sold or shared.

Additional fields (Phase 2)

This Actor also records your Apify user ID, whether Apify marks the account as paying, the size range of list inputs, the selected country when the input offers a fixed list of countries, and one category from a fixed Actor taxonomy. We use these fields only for aggregate reliability, repeat-use and cross-Actor analysis; reports suppress any cell with fewer than five distinct users.

The same disableUsageStats: true input flag turns these fields off too. The user ID is removed after 13 months; we do not export, sell, share, or attempt to re-identify this data.

Run-outcome signals (v2)

To learn whether a run did what it was asked to do, the record also holds a few more coarse ranges and yes/no flags. None of them contains content:

  • the result limit you asked for (a range, when the input has one) and what share of it was delivered;
  • results delivered per input item you listed (a range);
  • output quality as ranges: how fully the result fields were filled, the share of rows that look like errors, the share of duplicate rows, and how many different fields appeared. These are counted in memory while results are saved; no result content is kept;
  • how the run was started (console, API, schedule, webhook, another Actor);
  • how it ended: stopped by you, timed out, reached the requested limit, stopped by the charge limit, and how many times the platform moved the run;
  • if this Actor reports it: how many items to process worked or failed (ranges) and one failure reason from a fixed list;
  • a short code made from the names of the input fields you set, never their values.

Repeat-run fingerprint (v2)

When your Apify user ID is recorded (see above), the record also holds an 8-character one-way code made from your input (proxy settings left out) and this Actor's name. It only lets us see that the same account ran the same input again soon after an unsatisfying run; we never see the input itself. It is stored only in the database, never published, and reports use it in aggregate with the same five-user minimum. It is the one exception to the statement above that no hashes are collected, and disableUsageStats: true turns it off.