Instagram Profile Scraper With Bio Link & Email Extraction avatar

Instagram Profile Scraper With Bio Link & Email Extraction

Pricing

$14.99/month + usage

Go to Apify Store
Instagram Profile Scraper With Bio Link & Email Extraction

Instagram Profile Scraper With Bio Link & Email Extraction

Scrapes full data from any public Instagram profile, capturing bio, username, follower and following counts, posts, Reels, highlights, profile image, engagement stats, and URLs. Ideal for influencer research, competitor analysis, brand insights, and large-scale Instagram profiling

Pricing

$14.99/month + usage

Rating

0.0

(0)

Developer

Scrapio

Scrapio

Maintained by Community

Actor stats

0

Bookmarked

46

Total users

5

Monthly active users

3 days ago

Last modified

Share

Instagram Profile Scraper With Bio Link & Email Extraction pulls structured profile data, expanded bio links, and contact details — emails, phone numbers, and social handles — from any public Instagram profile, returned as typed JSON. Unlike scraping frameworks that return raw HTML, it hands back ready-to-use fields for your CRM, database, or LLM pipeline without any parsing step. Paste a username or profile URL and it fetches follower counts, bio text, and business metadata, expands Linktree- and Beacons-style aggregator links, and mines the bio for real, non-fabricated contact data. This guide covers every input option, output field, and deployment pattern for running it in production.

Instagram Profile Scraper With Bio Link & Email Extraction is a Playwright-driven Actor that loads an Instagram profile page, intercepts the same web_profile_info response Instagram's own web client uses, and turns it into one structured JSON row per profile. No Instagram account is required for public profiles, though Instagram has increasingly login-walled logged-out profile views since mid-2026 — an optional sessionId cookie is supported for when that happens. On top of the base profile fields, this variant adds two features: contact extraction from the bio and external link, and destination-link expansion for link-aggregator pages.

  • Full profile metadata: username, full name, biography, follower/following counts, posts count, verified and business-account flags, business category
  • Contact mining: emails[], phoneNumbers[], and socialHandles{} (WhatsApp, Telegram, Linktree handle) extracted from the bio and external link
  • Link-aggregator expansion: destination URLs behind Linktree, Beacons, linkin.bio, campsite.bio, msha.ke, Taplink, AllMyLinks, and other aggregator hosts
  • Optional About data: join date and country when includeAboutSection is enabled
  • Related profiles, latest IGTV videos, and latest posts (captions, hashtags, mentions, media) returned alongside the profile
  • Honest nulls: emails, phones, and expanded links are only ever emitted when actually found — never fabricated placeholders

⚡ Features & Capabilities

The Actor's capabilities split into the base profile fetch, the contact-mining layer built on top of it, and how it sits alongside the rest of the Scrapio Instagram lineup.

Core features

  • Interception-based fetch — reads Instagram's own web_profile_info JSON response via Playwright instead of parsing rendered HTML, so field names map directly to Instagram's API shape (full_name, edge_followed_by.count, is_business_account, etc.)
  • Regex-based contact extraction (contacts.py) — pulls emails[] via an RFC-shaped email pattern, phoneNumbers[] via an 8–15 digit E.164-plausible pattern, and socialHandles{} for WhatsApp (wa.me/api.whatsapp.com links or plain-text "WhatsApp: +...") , Telegram (t.me/telegram.me), and Linktree usernames
  • Link-aggregator detection and expansion — recognizes 25 aggregator hosts (linktr.ee, beacons.ai, linkin.bio, campsite.bio, msha.ke, taplink.cc, allmylinks.com, solo.to, bio.link, carrd.co, and more) and fetches the aggregator page itself, parsing embedded JSON and anchor tags to return real destination links in expandedLinks
  • Optional sessionId authentication — a Playwright cookie injection (sessionid + derived ds_user_id) to fetch profile data as a logged-in session when Instagram's logged-out login-wall blocks the request
  • Residential proxy by defaultsetup_proxy() defaults to Apify's RESIDENTIAL proxy group unless the caller supplies their own proxyConfiguration, with 3 retry attempts (RETRY_ATTEMPTS) on timeout or error per profile
  • Per-profile fault isolation — one profile failing (login-wall, timeout, invalid username) is logged and skipped; it does not abort the rest of the run

This Actor covers profile data, bio-link expansion, and bio-derived contacts for a known list of usernames or URLs. If you need to discover leads rather than enrich a known list, use Instagram Phone Number Scraper & Email Lead Finder or Instagram B2B Email Scraper: Business Type Leads, which search Google or Instagram's hashtag/mentions surfaces for candidate accounts first. For follower lists and shared-audience overlap, use Instagram Followers Scraper: Multi-Profile Analysis. For a scored, outreach-ready lead list instead of raw contact fields, use Instagram Profile Phone Number Scraper: Outreach Lead Scorer.

Why do developers and data teams scrape Instagram?

🏢 Lead generation and sales outreach teams

Sales and growth teams already have a list of relevant Instagram usernames — competitor followers, hashtag search results, a conference attendee list — and need to know which ones are reachable. Feed profileTargets with that list and read back hasContact, emails, phoneNumbers, and socialHandles per profile: hasContact == true rows go straight into an outreach queue, while businessCategoryName and isBusinessAccount help route business accounts differently from personal ones. Because expandLinks resolves Linktree/Beacons pages first, contacts hidden behind a link-in-bio page are captured too, not just what's visible in the raw biography text.

📊 AI training data and RAG indexing

The biography field and the caption, hashtags, and mentions sub-fields on each latestPosts entry are the highest-information text fields for language tasks — bios carry a subject's self-description in a few dense sentences, and captions carry contextual, informal language tied to real media. For RAG enrichment, indexing biography plus businessCategoryName gives a retrieval system enough context to answer "who is this account and what do they do" without a separate summarization pass. For training data, emails, phoneNumbers, and socialHandles are useful as a labeled contact-detection dataset — every field arrives as a typed primitive (string, array, or boolean), so no HTML cleanup or regex re-extraction is needed before use.

📱 Competitive and market intelligence

Track a competitor's Instagram presence over time by re-running the Actor against the same profileTargets on a schedule and diffing followersCount, followsCount, postsCount, and externalUrl between runs. A changed externalUrl often signals a new campaign or product launch before it's announced elsewhere, and a jump in followersCount combined with new latestPosts entries flags an active push worth investigating.

🔬 Research and academic use

Public bio text, business-category distribution, and contact-availability rates (hasContact) across a sampled set of profiles support studies on social-commerce adoption, influencer marketing structure, or platform self-presentation norms. This Actor only returns data from profiles that are public at request time — private accounts return no usable data, keeping the dataset within publicly accessible scope.

🎥 Product and SaaS development

Contact-enrichment products, influencer directories, and outreach-scoring tools can build directly on top of emails, phoneNumbers, socialHandles, and expandedLinks without writing their own regex or link-aggregator parser — both are already implemented and reused as-is from src/contacts.py.

🍚 Input Parameters

No input parameter is required — the Actor falls back gracefully but produces no rows if profileTargets is empty.

ParameterRequiredTypeDescriptionExample Value
profileTargetsNoarrayInstagram profiles to mine for contact data. Accepts a full URL (https://www.instagram.com/username/), a bare username, @username, or instagram.com/username. The legacy startUrls array ([{"url": "..."}]) is also accepted as a fallback.["gymshark", "https://www.instagram.com/natgeo/"]
expandLinksNoboolean (default true)When on, if a profile's externalUrl is a known link-aggregator page, the Actor fetches it and returns its destination links in expandedLinks.true
includeAboutSectionNoboolean (default false)When on, adds an about object with joinedDate and country, sourced from Instagram's own response.false
sessionIdNostring (secret)Optional Instagram sessionid cookie value (format <numeric_user_id>:xxxxx…), used to fetch profile data as a logged-in session when logged-out access is blocked. Leave empty to try logged-out first.(left empty)
proxyConfigurationNoobjectApify proxy configuration. Defaults to residential proxy with automatic retries; supply your own Apify proxy groups or custom proxy URLs to override.{"useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"]}
{
"profileTargets": ["gymshark", "https://www.instagram.com/natgeo/", "@nike"],
"expandLinks": true,
"includeAboutSection": false,
"sessionId": "",
"proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}

Supported URL types and input formats

  • Full profile URLhttps://www.instagram.com/gymshark/ — parsed via urlparse, taking the last non-empty path segment as the username.
  • Bare usernamenatgeo — used directly once @ and slashes are stripped.
  • @username@nike — the leading @ is stripped before the profile URL is built.
  • Legacy object form{"url": "https://www.instagram.com/gymshark/"} — still read from a startUrls array if you supply that key instead of profileTargets.

One honest limitation: pasting a bare domain path with no scheme (e.g. instagram.com/username, without https://) is not parsed the same way as a full URL in the current version — the username extractor treats it as a literal string rather than a URL, which can produce an incorrect username. Use a full https:// URL, @username, or a bare username instead.

📦 Output Format

Every successfully scraped profile is pushed as one dataset row combining the base profile fields with the contact-extraction fields added by this variant. Output is a standard Apify dataset — typed JSON per row, with no schema drift between runs beyond Instagram adding or removing a field on its side.

Output for Profiles

{
"inputUrl": "https://www.instagram.com/gymshark",
"url": "https://www.instagram.com/gymshark",
"id": "174424367",
"username": "gymshark",
"fullName": "Gymshark",
"biography": "The Original. 💪 Est. 2012 in a garage.\nShop the drop below 👇",
"externalUrls": [
{ "url": "https://linktr.ee/gymshark", "link_type": "external", "title": "" }
],
"externalUrl": "https://linktr.ee/gymshark",
"externalUrlShimmed": "https://l.instagram.com/?u=https%3A%2F%2Flinktr.ee%2Fgymshark",
"followersCount": 6500000,
"followsCount": 210,
"hasChannel": false,
"highlightReelCount": 24,
"isBusinessAccount": true,
"joinedRecently": false,
"businessCategoryName": "Sporting Goods",
"private": false,
"verified": true,
"profilePicUrl": "https://scontent.cdninstagram.com/v/t51.2885-19/gymshark_profile.jpg",
"profilePicUrlHD": "https://scontent.cdninstagram.com/v/t51.2885-19/gymshark_profile_hd.jpg",
"igtvVideoCount": 0,
"postsCount": 6100,
"about": { "joinedDate": null, "country": null },
"relatedProfiles": [
{ "id": "1798450931", "full_name": "Gymshark Women", "is_verified": false, "profile_pic_url": "https://scontent.cdninstagram.com/gymshark_women.jpg", "username": "gymsharkwomen", "is_private": false }
],
"latestIgtvVideos": [],
"latestPosts": [
{
"id": "3401234567890123456",
"type": "Sidecar",
"shortCode": "C9xY2zAOabc",
"caption": "New drop is live. Link in bio.",
"hashtags": ["gymshark", "trainlikeagymshark"],
"mentions": [],
"url": "https://www.instagram.com/p/C9xY2zAOabc/",
"commentsCount": 412,
"dimensionsHeight": 1350,
"dimensionsWidth": 1080,
"displayUrl": "https://scontent.cdninstagram.com/gymshark_post.jpg",
"alt": "Photo by Gymshark",
"likesCount": 28450,
"timestamp": "2026-07-20T09:15:00.000Z",
"ownerUsername": "gymshark",
"ownerId": "174424367",
"images": ["https://scontent.cdninstagram.com/gymshark_post.jpg"],
"childPosts": [],
"taggedUsers": []
}
],
"emails": ["support@gymshark.com"],
"phoneNumbers": [],
"socialHandles": { "linktree": "gymshark" },
"expandedLinks": [
"https://gymshark.com/collections/new",
"https://gymshark.com/pages/gymshark-training-app"
],
"hasContact": true
}

These fields are embedded in the same row above but are worth calling out separately since they are this variant's core differentiator over a plain profile scrape:

{
"externalUrl": "https://linktr.ee/gymshark",
"expandedLinks": [
"https://gymshark.com/collections/new",
"https://gymshark.com/pages/gymshark-training-app"
],
"emails": ["support@gymshark.com"],
"phoneNumbers": [],
"socialHandles": { "linktree": "gymshark" },
"hasContact": true
}

expandedLinks is null (not []) when expandLinks is off, the external link isn't a recognized aggregator, or the aggregator fetch failed — [] means the aggregator page was fetched but no destination links were found. socialHandles is null when none of WhatsApp, Telegram, or a Linktree handle were found in the text.

Schema stability and export options

Field names stay stable across runs because they come from this Actor's own USER_FIELDS mapping over Instagram's web_profile_info API response, not from parsing rendered page markup — a redesign of Instagram's front end doesn't break the mapping unless Instagram also changes the underlying API's field names. Results are stored in a standard Apify dataset and can be exported as JSON, CSV, Excel, XML, RSS, or an HTML table from the Apify Console or API, or pulled programmatically with the Apify API/SDK. Only rows that are successfully built and pushed are billed under the row_result charged event — profiles that fail to load or return no user data are logged and skipped without being pushed, so they are never charged.

🎯 Strategy 1: Real-time enrichment pipeline

Trigger a run whenever your team adds a new Instagram username to a CRM or spreadsheet: call the Actor with profileTargets set to just that username. Read back hasContact, emails, phoneNumbers, socialHandles, and expandedLinks, and write them onto the existing lead record keyed by username. Because hasContact is a single boolean, it's a fast filter for a downstream "needs outreach" queue without inspecting every array field individually.

🎯 Strategy 2: Scheduled monitoring and alerting

For a watchlist of accounts (competitors, partners, or influencers), set up an Apify Console Schedule to re-run the Actor daily or weekly against the same profileTargets. Store each run's output keyed by username and diff followersCount, externalUrl, and biography against the previous run. Alert on a changed externalUrl (often a new campaign link) or a followersCount jump, rather than alerting on every run regardless of whether anything changed.

🎯 Strategy 3: Bulk dataset build

For a research dataset or a large enrichment backlog, split your full username list into batches and launch one Actor run per batch via the Apify API rather than passing the entire list to a single run — the Actor processes profiles sequentially within a run (one Playwright browser session per profile, not in parallel), so several smaller concurrent runs finish a large list faster than one large run. Aggregate each run's dataset into your warehouse or a CSV export once all runs report a terminal status.

Strategy comparison at a glance

StrategyBest forRun patternOutput format
Real-time enrichmentSingle new leads as they appearOn-demand, single profileTargets entryJSON row appended to CRM/record
Scheduled monitoringWatchlists tracked over timeApify Console Schedule, recurringJSON diffed against prior run
Bulk dataset buildLarge one-off listsMultiple parallel runs, batched inputAggregated JSON/CSV export
Scraper NameWhat it extracts
Instagram Phone Number Scraper & Email Lead FinderEmails and phone numbers discovered via Google-indexed Instagram pages, no login needed
Instagram B2B Email Scraper: Business Type LeadsBusiness/personal emails via hashtag search, mentions mining, or keyword SERP discovery
Instagram Profile Phone Number Scraper: Outreach Lead ScorerSame SERP-based contact mining, plus a 0–100 outreach lead score and grade
Instagram Related Person By Niche Content Creator SearchRelated/suggested profiles around a seed account, filtered and ranked as niche influencers
Instagram Followers Scraper: Multi-Profile AnalysisFollower lists per profile plus shared-follower overlap across multiple profiles
Instagram Highlights Scraper: Content ExtractionHighlight metadata and, with a session cookie, full story-item content per highlight
LinkedIn B2B Emails Scraper By Phone & Email FinderThe same contact-mining pattern applied to public LinkedIn profiles instead of Instagram
Extract Emails Contacts Socials: Verified Phone & Email ListMulti-platform (including Instagram) contact cards — emails, phones, and social profile links
Instagram Hashtag & Engagement ScraperPosts and reels by hashtag, ranked by derived engagement metrics
Instagram Comments Scraper With Engagement AnalyticsComments and replies on a post/reel with per-comment engagement scoring

Instagram Profile Scraper With Bio Link & Email Extraction works with any language or tool that can call the Apify API — there is no bespoke API surface beyond standard Actor run and dataset endpoints.

Python

import csv
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
profiles = ["gymshark", "natgeo", "nike"]
run_input = {
"profileTargets": profiles,
"expandLinks": True,
"includeAboutSection": False,
}
run = client.actor(
"Scrapio/instagram-profile-scraper-with-bio-link-email-extraction"
).call(run_input=run_input)
rows = list(client.dataset(run["defaultDatasetId"]).iterate_items())
with open("instagram_contacts.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
writer.writerow(["username", "hasContact", "emails", "phoneNumbers", "expandedLinks"])
for row in rows:
writer.writerow([
row.get("username"),
row.get("hasContact"),
";".join(row.get("emails") or []),
";".join(row.get("phoneNumbers") or []),
";".join(row.get("expandedLinks") or []) if row.get("expandedLinks") else "",
])
print(f"Wrote {len(rows)} profile rows to instagram_contacts.csv")

Node.js

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor(
'Scrapio/instagram-profile-scraper-with-bio-link-email-extraction'
).call({
profileTargets: ['gymshark', 'natgeo'],
expandLinks: true,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
const withContact = items.filter((row) => row.hasContact);
console.log(`${withContact.length} of ${items.length} profiles returned a contact`);

Async and scheduled pipelines

There is no webhook-based delivery built into this Actor's own code, so the supported pattern is polling: start a run via the API, poll client.run(runId).get() (or its Node equivalent) until the status is terminal, then read the dataset. For recurring monitoring, use an Apify Console Schedule to fire the same profileTargets on a cron cadence instead of looping a script.

🏢 Sales and outreach teams

A sales team building a target list of boutique fitness brands feeds their Instagram usernames into profileTargets and reads back emails, phoneNumbers, and businessCategoryName to prioritize which accounts have a reachable contact before a single message is sent.

📊 AI and RAG engineering teams

Teams building a retrieval index over creator or brand accounts pull biography, businessCategoryName, and latestPosts[].caption as high-density text fields, and use emails/phoneNumbers/socialHandles as a structured, pre-labeled contact-detection example set — no additional cleaning required.

📱 Competitive intelligence analysts

An analyst tracking three competitor brand accounts re-runs the Actor weekly against the same profileTargets and watches followersCount, externalUrl, and postsCount for the week-over-week deltas that signal a new campaign or product push.

🔬 Researchers

Academic researchers studying influencer marketing or business self-presentation on Instagram sample public profiles for biography text, businessCategoryName distribution, and contact-availability rate (hasContact) — scope stays limited to what is public at request time.

🎥 Product and SaaS developers

Developers building a contact-enrichment or influencer-directory product call this Actor as their Instagram enrichment layer, using emails, phoneNumbers, socialHandles, and expandedLinks directly rather than building and maintaining their own bio-parsing and link-aggregator logic.

Scraping publicly accessible Instagram data is generally lawful in the United States: in hiQ Labs, Inc. v. LinkedIn Corp. (9th Cir. 2019, cert. denied 2022), the court held that scraping data a website makes publicly available does not violate the Computer Fraud and Abuse Act. That precedent addresses access, not Instagram's own Terms of Service — scraping in violation of Instagram's ToS can still expose you to civil claims (breach of contract, account suspension) even where no criminal statute is implicated. Separately, because this Actor extracts profile bios, external links, and derived emails/phone numbers, it returns personal data for individuals and business contacts alike, which brings GDPR (EU/UK) and CCPA/CPRA (California) obligations into play for anyone storing or processing it — lawful basis, retention limits, and deletion requests are the operator's responsibility, not something the Actor manages. Instagram Profile Scraper With Bio Link & Email Extraction returns only publicly accessible data. What you do with that data is your responsibility — consult legal counsel for commercial applications involving personal data.

❓ Frequently asked questions

Yes, for most public profiles — the Actor first attempts a logged-out fetch of Instagram's web_profile_info endpoint. Instagram has increasingly login-walled this endpoint for logged-out traffic since mid-2026, so if a profile returns no data, supply an optional sessionId cookie from a logged-in browser session to fetch it as an authenticated request.

How does it handle Instagram's anti-scraping measures?

It uses a real Chromium browser via Playwright (not a raw HTTP client), routes requests through Apify's residential proxy pool by default, and retries up to 3 times on a timeout or transient error before giving up on a given profile. An optional session cookie handles Instagram's login-wall specifically.

Can I run it at scale without getting blocked?

Profiles within a single run are processed one at a time, each with its own browser session, rather than in parallel — there is no documented concurrency or rate limit built into the Actor's code beyond the 3-attempt retry per profile. For large lists, splitting input across multiple parallel runs is the documented approach in the Strategy Guide above, rather than assuming unlimited throughput from one run.

How fresh is the data it returns?

Fully live per run — every profile is fetched fresh from Instagram's web_profile_info response at the moment the Actor runs. Nothing is cached or reused between runs; re-running against the same username always issues a new request.

Yes, in the sense that hiQ Labs, Inc. v. LinkedIn Corp. (9th Cir. 2019) establishes that accessing publicly available web data is not a CFAA violation — but that doesn't override Instagram's Terms of Service, and it doesn't remove your obligations under data protection law once you're handling personal contact data. See "Is it legal to scrape Instagram?" above for the full breakdown.

Which fields work best for AI training and RAG indexing?

For RAG: biography, businessCategoryName, and latestPosts[].caption carry the most descriptive natural-language content per profile. For training data: emails, phoneNumbers, and socialHandles are the most structurally consistent fields across records — each returns as a typed array, object, or null, with no HTML entities or inconsistent formatting to normalize before use.

Does this Actor return personal data, and who is responsible for handling it lawfully?

Yes — biographies, external links, and any extracted emails, phone numbers, or social handles are personal data under GDPR and CCPA/CPRA when they identify an individual. The Actor returns only publicly available data at the time of the run; establishing a lawful basis for storing, processing, or contacting people from that data is the responsibility of whoever runs the Actor, not the Actor itself.

Does it work with Claude, ChatGPT, and other AI agent tools?

There is no MCP server for this Actor in its current configuration, but it is callable as a standard HTTP endpoint through the Apify API by any agent framework that can make a REST call — every response is typed JSON, so no HTML parsing is needed before passing results into an LLM context window.

expandedLinks is returned as null, not an empty array or a guess — the Actor only expands a fixed, known list of aggregator hosts (Linktree, Beacons, linkin.bio, and others). A personal website or a direct e-commerce link is left as-is in externalUrl and is not expanded.

What happens if a profile is private or doesn't exist?

The Actor logs the failure (invalid username, login-wall, or Instagram returning no user object) and skips that profile — it does not stop the rest of the run, and no row is pushed or charged for a profile that fails to resolve.

ℹ️ Disclaimer

Instagram Profile Scraper With Bio Link & Email Extraction extracts only publicly available data from Instagram. This tool is intended for lawful use cases only. Users are responsible for complying with Instagram's terms of service and applicable data protection laws in their jurisdiction.