Following Scraper: Audience Overlap & Segment Breakdown
Pricing
$19.99/month + usage
Following Scraper: Audience Overlap & Segment Breakdown
π± Instagram Following Scraper captures the full list of accounts a public profile follows β usernames, handles, profile URLs, bios, verification & follower counts. β‘ Ideal for influencer discovery, competitor analysis, audience research & lead gen. π Fast, accurate, export-ready.
Pricing
$19.99/month + usage
Rating
0.0
(0)
Developer
Scraper Engine
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
2 days ago
Last modified
Categories
Share
Instagram Audience Overlap Scraper β Overlap Tiers and Segments
Following Scraper: Audience Overlap & Segment Breakdown pools who each of your seed Instagram profiles follows, keeps only accounts shared by at least N seeds, and returns every one as structured JSON with seedOverlapCount, segmentOverlapTier, segmentPrivacy, and segmentReelActivity already computed. No manual cross-referencing or spreadsheet formulas β each run also writes a per-segment roll-up summary. Add two or more seed profiles and a session cookie below to start mapping your shared audience.
What is Following Scraper: Audience Overlap & Segment Breakdown?
Following Scraper: Audience Overlap & Segment Breakdown is an Apify Actor that pools the Following lists of several Instagram seed profiles, computes how many seeds follow each account, and segments the result by verification status, privacy, cross-seed overlap tier, and reel activity. It returns typed JSON rows plus a run-level roll-up summary β no spreadsheet math required afterward. Instagram's Following list is login-gated, so the Actor needs a valid sessionId cookie from a logged-in Instagram session; without one, Instagram returns a login wall and no accounts are collected. It's built for growth marketers, competitive researchers, and data/AI engineers who need to map the audience two or more Instagram accounts share.
What Instagram following data is publicly available to scrape?
Logged-out visitors to a public Instagram profile see the username, name, bio, and follower/following counts β never the actual list of accounts someone follows, which sits behind Instagram's login wall.
| Data category | Publicly available | Restricted |
|---|---|---|
| Username, full name, verified badge | Yes | β |
| Numeric follower/following counts | Yes | β |
| Enumerated Following list | No | sessionId cookie (login-gated API) |
| Private accounts' Following lists | No | Login + approved follow request |
| Recent reel/story activity signal | No | Same login-gated API call |
| Cross-seed audience-overlap analysis | Not an Instagram feature | Computed by this Actor, not Instagram |
This Actor reproduces only what your logged-in session can already see β it doesn't bypass authentication or reach private accounts you're not approved to follow.
What data can I extract with Following Scraper: Audience Overlap & Segment Breakdown?
Every row combines Instagram's raw following-payload fields with overlap and segmentation metrics the Actor computes locally.
| Field | Description |
|---|---|
username | Instagram handle of the followed account |
full_name | Display name |
pk / pk_id / id | Instagram's internal numeric user ID (three aliases the API returns interchangeably) |
profile_pic_id / profile_pic_url | Profile photo identifiers |
has_anonymous_profile_picture | Whether the account still uses Instagram's default avatar |
account_badges | Instagram badge list (empty array if none) |
profileUrl | Profile link the Actor constructs from username |
scrapedAt | ISO timestamp the row was recorded |
seedOverlapCount | How many of your seed profiles follow this account |
followedBySeeds | List of seed usernames that follow it |
sourceProfile | Comma-joined seed usernames β clarified provenance |
segment | Nested object: { verified, privacy, overlapTier, reelActivity } |
segmentOverlapTier / segmentPrivacy / segmentReelActivity | Flat mirrors of segment, added so the dataset table view renders them without dot-notation |
is_private | Whether the account is private |
is_verified | Whether the account has Instagram's verified badge |
is_favorite | Whether your own session account has favorited this account (session-relative, not a public signal) |
latest_reel_media | Raw timestamp Instagram returns if the account has recent reel/story activity |
followed_by | Instagram's own value for this field (typically null; not seed provenance) |
fbid_v2 | Meta cross-app account ID |
third_party_downloads_enabled | Instagram media-download permission flag |
strong_id__ | Instagram's internal strong identifier |
Identity and profile fields
username, full_name, pk/pk_id/id, profile_pic_id, profile_pic_url, has_anonymous_profile_picture, account_badges, profileUrl, scrapedAt.
Overlap and segmentation fields
seedOverlapCount, followedBySeeds, sourceProfile, segment, segmentOverlapTier, segmentPrivacy, segmentReelActivity.
Raw account metadata
is_private, is_verified, is_favorite, latest_reel_media, followed_by, fbid_v2, third_party_downloads_enabled, strong_id__.
π€ Add-on: Need additional Instagram data?
Pair this Actor with Instagram Profile Scraper for full bio and profile detail on any account this Actor surfaces, or Instagram Followers Scraper With Bio Contact Enrichment if you need contact details enriched onto the follower side instead of the following side. To track how a shared audience changes over time, Instagram Followers β New Followers & Unfollows Checker monitors follower churn on a single account.
How does Following Scraper: Audience Overlap & Segment Breakdown differ from the official Instagram API?
Meta's Instagram Graph API has no endpoint that lists the accounts a user follows β only a numeric following count for the API-owning Business/Creator account β and no feature for computing audience overlap across accounts.
| Feature | Instagram Graph API | Following Scraper: Audience Overlap & Segment Breakdown |
|---|---|---|
| Following-list access | No endpoint returns followed accounts, only a count | Returns the full list for any public profile as a seed |
| Cross-account audience overlap | Not available in any form | Computes seedOverlapCount, followedBySeeds, overlap tiers |
| Account scope | Business/Creator accounts you own, linked to a Page | Works on any Instagram profile as a seed |
| Approval process | Meta App Review and use-case justification | None β supply a sessionId cookie and run |
| Segmentation/analytics | None built in | Verified, privacy, overlap-tier, reel-activity buckets with roll-ups |
| Setup | OAuth app registration, Page linking, permission review | Paste a session cookie into the input form |
Use the Graph API for aggregate stats on accounts you own inside Meta's approved app ecosystem. Use this Actor for the actual following list and cross-account overlap analysis on profiles you don't own.
How to use Following Scraper: Audience Overlap & Segment Breakdown
This Actor runs on Apify's platform β no separate Instagram developer account or app approval is needed, only an Apify account and an Instagram session cookie.
- Open Following Scraper: Audience Overlap & Segment Breakdown on the Apify Store and start a new run.
- Add two or more seed profiles under Audience seed profiles (
audienceSeeds) β URLs,@handles, or plain usernames. The Actor requires at least one seed to run at all, and needs 2+ to make overlap meaningful. - Paste your Instagram
sessionIdcookie value β required for the Actor to return any accounts. - Optionally set
sampleSizePerSeedandminSeedOverlapto control sampling depth and the overlap threshold. - Start the run, then download or stream the dataset as JSON or CSV from the Apify Console or API.
How to scale to bulk Instagram audience overlap extraction
audienceSeeds already accepts any number of seeds as a single array, and overlap is computed jointly across every seed in that one run β that is the Actor's built-in batching mechanism, not a separate bulk mode. For separate campaigns that each need their own overlap threshold or seed set, run the Actor again with a different audienceSeeds list, or loop multiple runs through the Apify API β one call per campaign or client.
What can you do with Instagram audience overlap data?
- A brand partnerships manager comparing two complementary accounts uses
segmentOverlapTier == "core"andseedOverlapCountto find the audience followed by every seed β strong co-marketing partner candidates. - A competitive growth marketer runs two competitors as seeds with
minSeedOverlap = 2and readssourceProfileto see exactly which seeds share each account, informing positioning against a rival's audience. - A talent scout filters
segmentPrivacy == "public",is_verified == true, andsegmentReelActivity == "active"to shortlist reachable, active, verified accounts worth an influencer outreach message. - A community manager reads the run's roll-up summary (
byOverlapTier,byPrivacy,byReelActivity) to judge how concentrated versus diffuse a niche audience is across several seed accounts. - An AI engineer feeds
username,sourceProfile, andsegmentinto a RAG pipeline or agent tool so an LLM can auto-tag prospects by overlap tier and reel activity before generating outreach copy.
How does Following Scraper: Audience Overlap & Segment Breakdown handle rate limits and blocking?
The Actor escalates through three connection stages when Instagram pushes back: it starts direct, falls back to an Apify datacenter proxy on the first block, and finally to a sticky residential proxy β retrying up to 3 times with exponential backoff once on residential. A block is detected from HTTP 401/403/429 responses or from response text containing markers like "checkpoint", "challenge", "login_required", or "please wait". Once escalated to residential, it rotates to a fresh residential IP before each new seed, since a single sticky IP that IG blocks after one seed's pagination would otherwise fail every seed that follows. Requests to Instagram are also spaced with randomized 1β2 second delays between pages and between seeds. If a seed's following list still can't be collected after this escalation, the Actor logs it as a failed seed, counts it in the run's seedsFailed total, and continues on to the remaining seeds rather than stopping the whole run. It does not solve CAPTCHAs.
β¬οΈ Input
| Parameter | Required | Type | Description | Example value |
|---|---|---|---|---|
audienceSeeds | No | array | Instagram profile links, @handle, or plain usernames β one per line. Add 2+ seeds to make the overlap filter meaningful. | ["https://www.instagram.com/nike/", "https://www.instagram.com/adidas/"] |
sampleSizePerSeed | No | integer | How many followed accounts to read from EACH seed. 0 = the full list (slower, higher block risk). Start with 50β200. | 50 |
minSeedOverlap | No | integer | Keep an account only if it is followed by at least this many seeds. 1 = keep everything (union). 2 = only accounts shared by β₯2 seeds. Default 1. | 2 |
sessionId | No | string | Copy the sessionid cookie value from a browser logged into Instagram and paste it here. Without it, Instagram returns a login wall and no accounts are collected. | 58023926432%3AeXaMpLeCookieValue%3A9 |
proxyConfiguration | No | object | Default runs direct and only escalates to Apify Proxy on block. Toggle Apify Proxy here if your network needs it. | { "useApifyProxy": false } |
No parameter is marked required in the input schema, but audienceSeeds needs at least one entry for the run to start, and sessionId is functionally required β without it Instagram's login wall returns zero accounts.
Example input
{"audienceSeeds": ["https://www.instagram.com/nike/","https://www.instagram.com/adidas/"],"sampleSizePerSeed": 50,"minSeedOverlap": 2,"sessionId": "58023926432%3AeXaMpLeCookieValue%3A9","proxyConfiguration": { "useApifyProxy": false }}
β¬οΈ Output
Each dataset row is one unique account, deduplicated across every seed, carrying Instagram's raw following-payload fields alongside the overlap and segmentation fields the Actor computes. Results export directly from the Apify Console or API as JSON, CSV, or Excel. Every row pushed to the dataset is billed under the row_result charged event; the Actor never pushes a separate uncharged accounting row. The run also writes a roll-up summary object (totalAccounts, byOverlapTier, byPrivacy, byReelActivity, verifiedCount, privateCount, bySeedOverlapCount) to the run's key-value store under OUTPUT.
Example output
{"pk": "1234567890","pk_id": "1234567890","id": "1234567890","full_name": "Sarah Kim","is_private": false,"fbid_v2": "17841400000000000","third_party_downloads_enabled": 0,"strong_id__": "1234567890","profile_pic_id": "3234567890123456789_1234567890","profile_pic_url": "https://scontent.cdninstagram.com/v/t51.2885-19/example_150x150.jpg","is_verified": false,"username": "sarahkim.fit","has_anonymous_profile_picture": false,"account_badges": [],"latest_reel_media": 1737936000,"is_favorite": false,"followed_by": null,"followedBySeeds": ["nike", "adidas"],"seedOverlapCount": 2,"sourceProfile": "nike, adidas","profileUrl": "https://www.instagram.com/sarahkim.fit/","scrapedAt": "2026-07-25T14:32:10.000Z","segment": {"verified": false,"privacy": "public","overlapTier": "core","reelActivity": "active"},"segmentOverlapTier": "core","segmentPrivacy": "public","segmentReelActivity": "active"}
How does it work?
The Actor first loads each seed profile's public HTML page to extract Instagram's internal APP_ID and numeric user ID, then calls Instagram's own friendships/{user_id}/following/ endpoint page by page using your session cookie. When Instagram blocks a request, it automatically escalates from a direct connection to a datacenter proxy and finally to a sticky residential proxy, rotating a fresh residential IP for every new seed so one blocked IP doesn't cascade into failures for later seeds. Once every seed is collected, overlap counting, threshold filtering, and segmentation all run locally against the already-collected data β no extra requests to Instagram. Only data your session can already see is returned, and the output field names stay the same run to run regardless of Instagram's UI changes, since the Actor talks to Instagram's underlying API responses rather than parsing rendered pages.
Integrations
Following Scraper: Audience Overlap & Segment Breakdown runs on the Apify platform, so it works with anything that can call the Apify API, plus Apify's own MCP server for AI agents.
Calling Following Scraper: Audience Overlap & Segment Breakdown programmatically
from apify_client import ApifyClientclient = ApifyClient("<APIFY_API_TOKEN>")run = client.actor("Scraper-Engine/instagram-following-scraper-audience-overlap-segment-breakdown").call(run_input={"audienceSeeds": ["https://www.instagram.com/nike/", "https://www.instagram.com/adidas/"],"minSeedOverlap": 2,"sessionId": "<YOUR_SESSIONID_COOKIE>",})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
Works in Go, Ruby, Node.js, cURL β any language that can make an HTTP request.
MCP integration for AI agents
This Actor is reachable through Apify's official @apify/actors-mcp-server. Register it with an MCP-compatible client (Claude Desktop, Claude Code, Cursor) using:
APIFY_TOKEN=<your_token> npx -y @apify/actors-mcp-server --actors Scraper-Engine/instagram-following-scraper-audience-overlap-segment-breakdown
No-code tools (n8n, Make, LangChain)
In n8n, use the HTTP Request node (or the community Apify node) pointed at this Actor's run endpoint with your API token. In Make, the Apify app's "Run an Actor" module triggers a run and passes results into your scenario. In LangChain, load results with an Apify dataset loader after triggering the run via the apify-client SDK.
Is it legal to scrape Instagram following/audience data?
Scraping publicly accessible Instagram data is generally lawful; this Actor returns only what your own authenticated Instagram session can already see, and does not defeat CAPTCHAs, bypass paywalls, or access private accounts you haven't been approved to follow. Because usernames, names, and verification status identify real individuals, the accounts and overlap data this Actor collects are personal data under GDPR and CCPA β you need a lawful basis to store and further process them, and should be able to honor access or deletion requests. Storing this data at scale, particularly for profiling or targeting, increases your compliance obligations under both Instagram's Terms of Use and applicable data-protection law. Consult legal counsel if your use case involves bulk storage of personal data.
Frequently asked questions
What Instagram following fields does Following Scraper: Audience Overlap & Segment Breakdown return?
It returns username, full_name, seedOverlapCount, segmentOverlapTier, and segmentPrivacy on every row, plus every raw following-payload field Instagram provides. See the data fields table above for the full list.
Does Following Scraper: Audience Overlap & Segment Breakdown require an Instagram account or login?
Yes. The Following list is login-gated, so you must supply a valid sessionId cookie value from a browser logged into Instagram. No field is marked "required" in the input schema, but without a session cookie Instagram returns a login wall and zero accounts are collected.
How many accounts can I extract in one run?
There's no fixed cap in the input schema. sampleSizePerSeed controls how many followed accounts are read per seed (0 reads the full list), and every seed's pool is merged before the minSeedOverlap filter is applied β so total output size depends on your seed count, sample size, and threshold.
What happens if a seed profile is private, deleted, or the sessionId is invalid?
The Actor logs that seed as failed, adds it to the run's seedsFailed count, and continues collecting the remaining seeds rather than stopping the whole run. If no seeds return any accounts and no session cookie was supplied, the run raises an error telling you to set a valid sessionId.
Can I scrape multiple Instagram following lists at once?
Yes. audienceSeeds accepts any number of profile URLs or usernames in a single run, and overlap plus segmentation are computed jointly across all of them together β that's the actual point of the Actor.
How is the audience overlap and segment breakdown actually computed?
Entirely locally, after all seeds are collected: the Actor deduplicates accounts across seeds by their Instagram user ID, counts how many seeds follow each one into seedOverlapCount/followedBySeeds, applies your minSeedOverlap threshold, and then buckets each surviving account into verified, privacy, overlapTier, and reelActivity segments β plus a per-segment roll-up count. No extra network calls to Instagram are made for this step; it's pure computation over the following-payload fields already collected.
Can I use Following Scraper: Audience Overlap & Segment Breakdown without managing proxies or browser infrastructure?
Yes. The Actor automatically escalates from a direct connection to Apify datacenter and residential proxies when Instagram blocks a request, and rotates residential IPs between seeds β you don't configure any of this yourself unless you want to.
Does Following Scraper: Audience Overlap & Segment Breakdown work with Claude, ChatGPT, and other AI agent tools?
Yes. It's callable as an HTTP endpoint by any agent framework through the Apify API, and reachable through Apify's official MCP server for MCP-compatible clients such as Claude Desktop and Claude Code.
How does Following Scraper: Audience Overlap & Segment Breakdown compare to other Instagram following scrapers?
Checked on the Apify Store on 2026-07-25, a comparable listing (Louis Deconinck's Instagram Following Scraper) returns unfiltered per-profile following data with no login required and no cross-account overlap or segmentation β sorting and bucketing are left to the user. This Actor requires a session cookie precisely because it needs authenticated access to the Following list, but in exchange it pools multiple seeds in one run and ships seedOverlapCount, overlap tiers, and segment buckets already computed.
What happens when Instagram changes its layout or anti-bot system?
This Actor is maintained, and its output field names stay stable across updates since it reads Instagram's underlying API responses rather than parsing rendered HTML. No specific update turnaround time is published.
Which Instagram following fields work best for AI training data and RAG indexing?
For RAG, index username, full_name, and sourceProfile as the descriptive text fields. For structured training data, seedOverlapCount, segmentOverlapTier, and segmentReelActivity are the most consistently structured fields across every record β all return as typed primitives (strings, integers, booleans).
Related scrapers
| Scraper name | What it extracts |
|---|---|
| Instagram Profile Scraper | Core profile fields β bio, counts, verification |
| Instagram Followers Scraper With Bio Contact Enrichment | Follower lists enriched with bio and contact details |
| Instagram Followers β New Followers & Unfollows Checker | Follower gains and unfollows on one account over time |
| Instagram Followers Count Scraper With Engagement Quality Score | Follower counts with an engagement quality score |
| Instagram Related Person Scraper | Instagram's own "related accounts" suggestions |
Your feedback
Found a bug or think a field is missing? Let us know so we can fix it. Reach the Scraper-Engine team through the Issues tab on this Actor's Apify Store page β we actively maintain this Actor as Instagram's API responses and anti-bot behavior change.