Following Scraper: Audience Overlap & Segment Breakdown avatar

Following Scraper: Audience Overlap & Segment Breakdown

Pricing

$19.99/month + usage

Go to Apify Store
Following Scraper: Audience Overlap & Segment Breakdown

Following Scraper: Audience Overlap & Segment Breakdown

πŸ“± Instagram Following Scraper captures the full list of accounts a public profile follows β€” usernames, handles, profile URLs, bios, verification & follower counts. ⚑ Ideal for influencer discovery, competitor analysis, audience research & lead gen. πŸ”Ž Fast, accurate, export-ready.

Pricing

$19.99/month + usage

Rating

0.0

(0)

Developer

Scraper Engine

Scraper Engine

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

2 days ago

Last modified

Share

Instagram Audience Overlap Scraper β€” Overlap Tiers and Segments

Following Scraper: Audience Overlap & Segment Breakdown pools who each of your seed Instagram profiles follows, keeps only accounts shared by at least N seeds, and returns every one as structured JSON with seedOverlapCount, segmentOverlapTier, segmentPrivacy, and segmentReelActivity already computed. No manual cross-referencing or spreadsheet formulas β€” each run also writes a per-segment roll-up summary. Add two or more seed profiles and a session cookie below to start mapping your shared audience.

What is Following Scraper: Audience Overlap & Segment Breakdown?

Following Scraper: Audience Overlap & Segment Breakdown is an Apify Actor that pools the Following lists of several Instagram seed profiles, computes how many seeds follow each account, and segments the result by verification status, privacy, cross-seed overlap tier, and reel activity. It returns typed JSON rows plus a run-level roll-up summary β€” no spreadsheet math required afterward. Instagram's Following list is login-gated, so the Actor needs a valid sessionId cookie from a logged-in Instagram session; without one, Instagram returns a login wall and no accounts are collected. It's built for growth marketers, competitive researchers, and data/AI engineers who need to map the audience two or more Instagram accounts share.

What Instagram following data is publicly available to scrape?

Logged-out visitors to a public Instagram profile see the username, name, bio, and follower/following counts β€” never the actual list of accounts someone follows, which sits behind Instagram's login wall.

Data categoryPublicly availableRestricted
Username, full name, verified badgeYesβ€”
Numeric follower/following countsYesβ€”
Enumerated Following listNosessionId cookie (login-gated API)
Private accounts' Following listsNoLogin + approved follow request
Recent reel/story activity signalNoSame login-gated API call
Cross-seed audience-overlap analysisNot an Instagram featureComputed by this Actor, not Instagram

This Actor reproduces only what your logged-in session can already see β€” it doesn't bypass authentication or reach private accounts you're not approved to follow.

What data can I extract with Following Scraper: Audience Overlap & Segment Breakdown?

Every row combines Instagram's raw following-payload fields with overlap and segmentation metrics the Actor computes locally.

FieldDescription
usernameInstagram handle of the followed account
full_nameDisplay name
pk / pk_id / idInstagram's internal numeric user ID (three aliases the API returns interchangeably)
profile_pic_id / profile_pic_urlProfile photo identifiers
has_anonymous_profile_pictureWhether the account still uses Instagram's default avatar
account_badgesInstagram badge list (empty array if none)
profileUrlProfile link the Actor constructs from username
scrapedAtISO timestamp the row was recorded
seedOverlapCountHow many of your seed profiles follow this account
followedBySeedsList of seed usernames that follow it
sourceProfileComma-joined seed usernames β€” clarified provenance
segmentNested object: { verified, privacy, overlapTier, reelActivity }
segmentOverlapTier / segmentPrivacy / segmentReelActivityFlat mirrors of segment, added so the dataset table view renders them without dot-notation
is_privateWhether the account is private
is_verifiedWhether the account has Instagram's verified badge
is_favoriteWhether your own session account has favorited this account (session-relative, not a public signal)
latest_reel_mediaRaw timestamp Instagram returns if the account has recent reel/story activity
followed_byInstagram's own value for this field (typically null; not seed provenance)
fbid_v2Meta cross-app account ID
third_party_downloads_enabledInstagram media-download permission flag
strong_id__Instagram's internal strong identifier

Identity and profile fields

username, full_name, pk/pk_id/id, profile_pic_id, profile_pic_url, has_anonymous_profile_picture, account_badges, profileUrl, scrapedAt.

Overlap and segmentation fields

seedOverlapCount, followedBySeeds, sourceProfile, segment, segmentOverlapTier, segmentPrivacy, segmentReelActivity.

Raw account metadata

is_private, is_verified, is_favorite, latest_reel_media, followed_by, fbid_v2, third_party_downloads_enabled, strong_id__.

πŸ€– Add-on: Need additional Instagram data?

Pair this Actor with Instagram Profile Scraper for full bio and profile detail on any account this Actor surfaces, or Instagram Followers Scraper With Bio Contact Enrichment if you need contact details enriched onto the follower side instead of the following side. To track how a shared audience changes over time, Instagram Followers β€” New Followers & Unfollows Checker monitors follower churn on a single account.

How does Following Scraper: Audience Overlap & Segment Breakdown differ from the official Instagram API?

Meta's Instagram Graph API has no endpoint that lists the accounts a user follows β€” only a numeric following count for the API-owning Business/Creator account β€” and no feature for computing audience overlap across accounts.

FeatureInstagram Graph APIFollowing Scraper: Audience Overlap & Segment Breakdown
Following-list accessNo endpoint returns followed accounts, only a countReturns the full list for any public profile as a seed
Cross-account audience overlapNot available in any formComputes seedOverlapCount, followedBySeeds, overlap tiers
Account scopeBusiness/Creator accounts you own, linked to a PageWorks on any Instagram profile as a seed
Approval processMeta App Review and use-case justificationNone β€” supply a sessionId cookie and run
Segmentation/analyticsNone built inVerified, privacy, overlap-tier, reel-activity buckets with roll-ups
SetupOAuth app registration, Page linking, permission reviewPaste a session cookie into the input form

Use the Graph API for aggregate stats on accounts you own inside Meta's approved app ecosystem. Use this Actor for the actual following list and cross-account overlap analysis on profiles you don't own.

How to use Following Scraper: Audience Overlap & Segment Breakdown

This Actor runs on Apify's platform β€” no separate Instagram developer account or app approval is needed, only an Apify account and an Instagram session cookie.

  1. Open Following Scraper: Audience Overlap & Segment Breakdown on the Apify Store and start a new run.
  2. Add two or more seed profiles under Audience seed profiles (audienceSeeds) β€” URLs, @handles, or plain usernames. The Actor requires at least one seed to run at all, and needs 2+ to make overlap meaningful.
  3. Paste your Instagram sessionId cookie value β€” required for the Actor to return any accounts.
  4. Optionally set sampleSizePerSeed and minSeedOverlap to control sampling depth and the overlap threshold.
  5. Start the run, then download or stream the dataset as JSON or CSV from the Apify Console or API.

How to scale to bulk Instagram audience overlap extraction

audienceSeeds already accepts any number of seeds as a single array, and overlap is computed jointly across every seed in that one run β€” that is the Actor's built-in batching mechanism, not a separate bulk mode. For separate campaigns that each need their own overlap threshold or seed set, run the Actor again with a different audienceSeeds list, or loop multiple runs through the Apify API β€” one call per campaign or client.

What can you do with Instagram audience overlap data?

  • A brand partnerships manager comparing two complementary accounts uses segmentOverlapTier == "core" and seedOverlapCount to find the audience followed by every seed β€” strong co-marketing partner candidates.
  • A competitive growth marketer runs two competitors as seeds with minSeedOverlap = 2 and reads sourceProfile to see exactly which seeds share each account, informing positioning against a rival's audience.
  • A talent scout filters segmentPrivacy == "public", is_verified == true, and segmentReelActivity == "active" to shortlist reachable, active, verified accounts worth an influencer outreach message.
  • A community manager reads the run's roll-up summary (byOverlapTier, byPrivacy, byReelActivity) to judge how concentrated versus diffuse a niche audience is across several seed accounts.
  • An AI engineer feeds username, sourceProfile, and segment into a RAG pipeline or agent tool so an LLM can auto-tag prospects by overlap tier and reel activity before generating outreach copy.

How does Following Scraper: Audience Overlap & Segment Breakdown handle rate limits and blocking?

The Actor escalates through three connection stages when Instagram pushes back: it starts direct, falls back to an Apify datacenter proxy on the first block, and finally to a sticky residential proxy β€” retrying up to 3 times with exponential backoff once on residential. A block is detected from HTTP 401/403/429 responses or from response text containing markers like "checkpoint", "challenge", "login_required", or "please wait". Once escalated to residential, it rotates to a fresh residential IP before each new seed, since a single sticky IP that IG blocks after one seed's pagination would otherwise fail every seed that follows. Requests to Instagram are also spaced with randomized 1–2 second delays between pages and between seeds. If a seed's following list still can't be collected after this escalation, the Actor logs it as a failed seed, counts it in the run's seedsFailed total, and continues on to the remaining seeds rather than stopping the whole run. It does not solve CAPTCHAs.

⬇️ Input

ParameterRequiredTypeDescriptionExample value
audienceSeedsNoarrayInstagram profile links, @handle, or plain usernames β€” one per line. Add 2+ seeds to make the overlap filter meaningful.["https://www.instagram.com/nike/", "https://www.instagram.com/adidas/"]
sampleSizePerSeedNointegerHow many followed accounts to read from EACH seed. 0 = the full list (slower, higher block risk). Start with 50–200.50
minSeedOverlapNointegerKeep an account only if it is followed by at least this many seeds. 1 = keep everything (union). 2 = only accounts shared by β‰₯2 seeds. Default 1.2
sessionIdNostringCopy the sessionid cookie value from a browser logged into Instagram and paste it here. Without it, Instagram returns a login wall and no accounts are collected.58023926432%3AeXaMpLeCookieValue%3A9
proxyConfigurationNoobjectDefault runs direct and only escalates to Apify Proxy on block. Toggle Apify Proxy here if your network needs it.{ "useApifyProxy": false }

No parameter is marked required in the input schema, but audienceSeeds needs at least one entry for the run to start, and sessionId is functionally required β€” without it Instagram's login wall returns zero accounts.

Example input

{
"audienceSeeds": [
"https://www.instagram.com/nike/",
"https://www.instagram.com/adidas/"
],
"sampleSizePerSeed": 50,
"minSeedOverlap": 2,
"sessionId": "58023926432%3AeXaMpLeCookieValue%3A9",
"proxyConfiguration": { "useApifyProxy": false }
}

⬆️ Output

Each dataset row is one unique account, deduplicated across every seed, carrying Instagram's raw following-payload fields alongside the overlap and segmentation fields the Actor computes. Results export directly from the Apify Console or API as JSON, CSV, or Excel. Every row pushed to the dataset is billed under the row_result charged event; the Actor never pushes a separate uncharged accounting row. The run also writes a roll-up summary object (totalAccounts, byOverlapTier, byPrivacy, byReelActivity, verifiedCount, privateCount, bySeedOverlapCount) to the run's key-value store under OUTPUT.

Example output

{
"pk": "1234567890",
"pk_id": "1234567890",
"id": "1234567890",
"full_name": "Sarah Kim",
"is_private": false,
"fbid_v2": "17841400000000000",
"third_party_downloads_enabled": 0,
"strong_id__": "1234567890",
"profile_pic_id": "3234567890123456789_1234567890",
"profile_pic_url": "https://scontent.cdninstagram.com/v/t51.2885-19/example_150x150.jpg",
"is_verified": false,
"username": "sarahkim.fit",
"has_anonymous_profile_picture": false,
"account_badges": [],
"latest_reel_media": 1737936000,
"is_favorite": false,
"followed_by": null,
"followedBySeeds": ["nike", "adidas"],
"seedOverlapCount": 2,
"sourceProfile": "nike, adidas",
"profileUrl": "https://www.instagram.com/sarahkim.fit/",
"scrapedAt": "2026-07-25T14:32:10.000Z",
"segment": {
"verified": false,
"privacy": "public",
"overlapTier": "core",
"reelActivity": "active"
},
"segmentOverlapTier": "core",
"segmentPrivacy": "public",
"segmentReelActivity": "active"
}

How does it work?

The Actor first loads each seed profile's public HTML page to extract Instagram's internal APP_ID and numeric user ID, then calls Instagram's own friendships/{user_id}/following/ endpoint page by page using your session cookie. When Instagram blocks a request, it automatically escalates from a direct connection to a datacenter proxy and finally to a sticky residential proxy, rotating a fresh residential IP for every new seed so one blocked IP doesn't cascade into failures for later seeds. Once every seed is collected, overlap counting, threshold filtering, and segmentation all run locally against the already-collected data β€” no extra requests to Instagram. Only data your session can already see is returned, and the output field names stay the same run to run regardless of Instagram's UI changes, since the Actor talks to Instagram's underlying API responses rather than parsing rendered pages.

Integrations

Following Scraper: Audience Overlap & Segment Breakdown runs on the Apify platform, so it works with anything that can call the Apify API, plus Apify's own MCP server for AI agents.

Calling Following Scraper: Audience Overlap & Segment Breakdown programmatically

from apify_client import ApifyClient
client = ApifyClient("<APIFY_API_TOKEN>")
run = client.actor("Scraper-Engine/instagram-following-scraper-audience-overlap-segment-breakdown").call(
run_input={
"audienceSeeds": ["https://www.instagram.com/nike/", "https://www.instagram.com/adidas/"],
"minSeedOverlap": 2,
"sessionId": "<YOUR_SESSIONID_COOKIE>",
}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)

Works in Go, Ruby, Node.js, cURL β€” any language that can make an HTTP request.

MCP integration for AI agents

This Actor is reachable through Apify's official @apify/actors-mcp-server. Register it with an MCP-compatible client (Claude Desktop, Claude Code, Cursor) using:

APIFY_TOKEN=<your_token> npx -y @apify/actors-mcp-server --actors Scraper-Engine/instagram-following-scraper-audience-overlap-segment-breakdown

No-code tools (n8n, Make, LangChain)

In n8n, use the HTTP Request node (or the community Apify node) pointed at this Actor's run endpoint with your API token. In Make, the Apify app's "Run an Actor" module triggers a run and passes results into your scenario. In LangChain, load results with an Apify dataset loader after triggering the run via the apify-client SDK.

Scraping publicly accessible Instagram data is generally lawful; this Actor returns only what your own authenticated Instagram session can already see, and does not defeat CAPTCHAs, bypass paywalls, or access private accounts you haven't been approved to follow. Because usernames, names, and verification status identify real individuals, the accounts and overlap data this Actor collects are personal data under GDPR and CCPA β€” you need a lawful basis to store and further process them, and should be able to honor access or deletion requests. Storing this data at scale, particularly for profiling or targeting, increases your compliance obligations under both Instagram's Terms of Use and applicable data-protection law. Consult legal counsel if your use case involves bulk storage of personal data.

Frequently asked questions

What Instagram following fields does Following Scraper: Audience Overlap & Segment Breakdown return?

It returns username, full_name, seedOverlapCount, segmentOverlapTier, and segmentPrivacy on every row, plus every raw following-payload field Instagram provides. See the data fields table above for the full list.

Does Following Scraper: Audience Overlap & Segment Breakdown require an Instagram account or login?

Yes. The Following list is login-gated, so you must supply a valid sessionId cookie value from a browser logged into Instagram. No field is marked "required" in the input schema, but without a session cookie Instagram returns a login wall and zero accounts are collected.

How many accounts can I extract in one run?

There's no fixed cap in the input schema. sampleSizePerSeed controls how many followed accounts are read per seed (0 reads the full list), and every seed's pool is merged before the minSeedOverlap filter is applied β€” so total output size depends on your seed count, sample size, and threshold.

What happens if a seed profile is private, deleted, or the sessionId is invalid?

The Actor logs that seed as failed, adds it to the run's seedsFailed count, and continues collecting the remaining seeds rather than stopping the whole run. If no seeds return any accounts and no session cookie was supplied, the run raises an error telling you to set a valid sessionId.

Can I scrape multiple Instagram following lists at once?

Yes. audienceSeeds accepts any number of profile URLs or usernames in a single run, and overlap plus segmentation are computed jointly across all of them together β€” that's the actual point of the Actor.

How is the audience overlap and segment breakdown actually computed?

Entirely locally, after all seeds are collected: the Actor deduplicates accounts across seeds by their Instagram user ID, counts how many seeds follow each one into seedOverlapCount/followedBySeeds, applies your minSeedOverlap threshold, and then buckets each surviving account into verified, privacy, overlapTier, and reelActivity segments β€” plus a per-segment roll-up count. No extra network calls to Instagram are made for this step; it's pure computation over the following-payload fields already collected.

Can I use Following Scraper: Audience Overlap & Segment Breakdown without managing proxies or browser infrastructure?

Yes. The Actor automatically escalates from a direct connection to Apify datacenter and residential proxies when Instagram blocks a request, and rotates residential IPs between seeds β€” you don't configure any of this yourself unless you want to.

Does Following Scraper: Audience Overlap & Segment Breakdown work with Claude, ChatGPT, and other AI agent tools?

Yes. It's callable as an HTTP endpoint by any agent framework through the Apify API, and reachable through Apify's official MCP server for MCP-compatible clients such as Claude Desktop and Claude Code.

How does Following Scraper: Audience Overlap & Segment Breakdown compare to other Instagram following scrapers?

Checked on the Apify Store on 2026-07-25, a comparable listing (Louis Deconinck's Instagram Following Scraper) returns unfiltered per-profile following data with no login required and no cross-account overlap or segmentation β€” sorting and bucketing are left to the user. This Actor requires a session cookie precisely because it needs authenticated access to the Following list, but in exchange it pools multiple seeds in one run and ships seedOverlapCount, overlap tiers, and segment buckets already computed.

What happens when Instagram changes its layout or anti-bot system?

This Actor is maintained, and its output field names stay stable across updates since it reads Instagram's underlying API responses rather than parsing rendered HTML. No specific update turnaround time is published.

Which Instagram following fields work best for AI training data and RAG indexing?

For RAG, index username, full_name, and sourceProfile as the descriptive text fields. For structured training data, seedOverlapCount, segmentOverlapTier, and segmentReelActivity are the most consistently structured fields across every record β€” all return as typed primitives (strings, integers, booleans).

Scraper nameWhat it extracts
Instagram Profile ScraperCore profile fields β€” bio, counts, verification
Instagram Followers Scraper With Bio Contact EnrichmentFollower lists enriched with bio and contact details
Instagram Followers β€” New Followers & Unfollows CheckerFollower gains and unfollows on one account over time
Instagram Followers Count Scraper With Engagement Quality ScoreFollower counts with an engagement quality score
Instagram Related Person ScraperInstagram's own "related accounts" suggestions

Your feedback

Found a bug or think a field is missing? Let us know so we can fix it. Reach the Scraper-Engine team through the Issues tab on this Actor's Apify Store page β€” we actively maintain this Actor as Instagram's API responses and anti-bot behavior change.