Influencer Profile Scraper
Pricing
from $6.00 / 1,000 profile reads
Influencer Profile Scraper
Reads public creator profiles from a list of handles or URLs, one or a batch per run. One flat row per profile with followers, post count, bio, bio link, public business email, engagement rate, verified flag, source URL, and read date. A private or missing account returns a labeled error row.
Pricing
from $6.00 / 1,000 profile reads
Rating
0.0
(0)
Developer
Mamba Labs
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 hours ago
Last modified
Categories
Share
👤 What can Influencer Profile Scraper do?
Handles or profile URLs in, profile detail out. One flat row per profile with followers, following, post count, engagement rate and the method that computed it, verified flag, bio, bio link, location and country guess, and a public business email or manager email when one is written in the bio. Every row carries its source URL and read time. A private or missing account returns a labeled error row, not a blank.
| 📦 What you get | ⚙️ Features and integrations |
|---|---|
| 👥 Followers, following, post count, exact where the platform shows it 📊 Engagement rate with the method that computed it ✅ Verified flag, bio, bio link, location, country guess 📧 Business or manager email when it is written in the bio 🧾 57 flat fields, snake_case, the suite's shared row | 🌐 Seven platforms, TikTok, Instagram, YouTube, Pinterest, Twitch, Threads, podcasts 🔁 Datacenter first, one residential retry on a block 🧭 Labeled error rows, a private or missing account is never a blank 🧪 Clay ready, one handle per row lands on the same path ⬇️ Export to JSON, CSV, Excel, HTML or XML |
Bought by teams that sell to creators, newsletter and creator tool companies, talent and brand partnership teams, and anyone turning a handle list into a contactable, sized creator table.
🚫 This is not a post or comment scraper, and it does not log in. It reads the public profile page once and returns the profile row. Feeds, comments, and anything behind a login stay unread; YouTube's business email button is one of those, see Known limits.
💡 Why use Influencer Profile Scraper?
| If you sell | Read these fields |
|---|---|
| Newsletter or email tools to creators | bio_link, business_email, followers, then the Link in Bio Scraper and Newsletter Detector |
| Brand partnerships or sponsorships | followers, engagement_rate, engagement_method, verified |
| Creator tools, courses, or memberships | post_count, bio, country_guess |
| Talent representation or management | business_email, manager_email, bio |
| Anything, as a disqualifier | row_status, error_reason, country_guess |
🧭 One route per platform, measured before the price was set
Five public profiles per platform, read at 1, 2, 3, and 5 in flight. USD per profile is transfer at the account's datacenter rate ($0.20 per GB) or residential rate ($8.00 per GB) plus an estimated compute share (1 GB, wall seconds over 3,600, at $0.30 per compute unit). The default concurrency sits below the point where rows started dropping, or two below the largest wave tested when nothing dropped.
| Platform | Route | Fields returned | Per profile | Ok of 5 at 1, 2, 3, 5 | Default concurrency |
|---|---|---|---|---|---|
| TikTok | Profile page rehydration JSON over datacenter; one residential retry on a block | name, followers (exact from statsV2), following, video count, engagement (lifetime likes per video over followers), verified, bio, bio link, business email from the bio | 372 KB, 1.5 to 3.0 s, $0.00007 (a residential recovery costs $0.0030) | 5, 5, 5, 5 minus one block per wave at 5 | 3 |
Profile embed widget over datacenter (counts, name, verified, 6 recent posts); bio, bio link, and following from web_profile_info on about 1 in 4 datacenter attempts, otherwise from the profile page over residential when escalate_on_block is on (event instagram-bio-fetch, $0.010, only when that page came back readable) | name, followers (exact), post count, verified, private flag, engagement (median likes plus comments of recent posts over followers), bio, bio link, business email from the bio | embed 303 KB and $0.00006; residential bio page 720 to 830 KB, up to $0.0066 by decompressed bytes and $0.0012 by the platform's metered transfer (run o6u6d7x7zfElI9WQm, 2026-09-22); the residential page ran on 7 of 10 reads on 2026-09-22 and read on 6 | 5, 5, 5, 5 | 4 | |
| YouTube | /@handle/about page ytInitialData over datacenter; innertube fallback; residential only after every datacenter route blocks | name, subscribers (rounded by YouTube), video count, description, channel links, country (platform region), verified, engagement (average recent views over subscribers) | 2.2 MB, 2.0 s, $0.0005 | 5, 5, 5, 5 | 3 |
| Threads | Public profile page Relay payload over datacenter; residential only on a block signature | name, followers (exact), bio, bio link, verified, business email from the bio | 0.93 MB, 5.0 s, $0.0003 | 5, 5, 5, 5 | 3 |
Profile page __PWS_INITIAL_PROPS__ over datacenter; residential retry on a real block | name, followers, following, pin count, verified merchant, about, website, location, business email, other social profiles | 1.39 MB, 1.8 to 3.0 s, $0.0003 | 5, 5, 5, 5 | 3 | |
| Twitch | Public GQL endpoint with the site's own web client id over datacenter; channel page JSON-LD fallback; your own Helix client id and app token when supplied | display name, followers, video count, partner flag, description, social links from the channel panels, business email in the description | 1.2 KB, 1.2 to 1.6 s, under $0.00001 | 5, 5, 5, 5 | 3 |
| Podcast | iTunes lookup plus the RSS feed, direct (no proxy bytes); Spotify show page over datacenter | show name, episode count, description, site link, owner email from the feed (business_email_source feed), author, genres, country | feed 0.5 to 5.4 MB direct, compute only, 0.8 to 2.9 s | 5, 5, 5, 5 | 3 |
Discovery per keyword, measured the same day: TikTok 7 handles over three Google pages (Google drops the site: operator on some exit countries, so a page without a handle is fetched once more); Instagram 1 to 8; YouTube 20 from the channel filtered results page with no search engine; Threads 11 to 15 over two Google pages; Twitch 11 from its own search; Pinterest 50 from its user search resource; podcasts 50 from the iTunes Search API. A Google page costs $0.0025 on the runner's account.
What each platform does not expose without login, and what the row says instead: YouTube's business email sits behind "View email address" and a verification step, so it is not read and business_email is filled only from an address written in the public description; Instagram has no location field, so country_guess comes from the bio; Threads and podcasts have no per post engagement; podcasts have no follower count; TikTok's anonymous render carries no region, so the country comes from the bio.
📋 What data can Influencer Profile Scraper extract?
57 fields per row, the same 57 in the same order on every actor in this suite. The ones this actor fills:
| Field | What it holds |
|---|---|
creator_id, creator_id_method, creator_id_confidence | One ID per creator across platforms and runs; see The shared row below |
platform, handle, profile_url, display_name | The profile read, with the canonical public URL |
followers, following, post_count | Counts as the platform shows them; TikTok rounds above about ten thousand, YouTube to three significant figures |
engagement_rate, engagement_method | Percentage with two decimals and the formula that produced it; null where the platform gives no per post figures without login |
verified | The platform's mark; false means read and not verified, null means not readable |
bio, bio_link | The bio text and the outbound link in it |
location_text, country_guess, country_guess_method | Location as written, an ISO country guess, and what it came from |
business_email, business_email_source | A public email in the bio or the feed, and where it was found |
other_social_profiles_json | Other social profile URLs the profile links to |
source_url, read_at, row_status, error_reason, billable_events_json | Where the row came from, when, whether it is usable, and what it cost |
⚠️
falseandnullare not the same thing, and an empty cell is not a failed read. A column this actor does not own isnull, never missing. On a boolean,falsemeans the page was read and the signal was not there;nullmeans nothing read it. On an Influencer Profile Scraper row,verified: falseis a profile without the mark andverified: nullis a profile that could not be read. Readrow_statusbefore any other column; the paragraph below says how.
Every actor in this suite writes the same 57 flat columns, in the same order, and fills the ones it owns. A column an actor does not own is null, never missing. Nested data lives only in columns ending _json, as a JSON string, so a Clay column reads one cell.
| Group | Columns | Filled by |
|---|---|---|
| Identity | creator_id, creator_id_method, creator_id_confidence, platform, handle, profile_url, display_name | Influencer Finder, Influencer Profile Scraper |
| Discovery | search_keyword, niche, similar_creators, similar_creators_method | Influencer Finder |
| Profile | followers, following, post_count, engagement_rate, engagement_method, verified, bio, bio_link, location_text, country_guess, country_guess_method | Influencer Profile Scraper |
| Contact | business_email, business_email_source, website_email, website_email_source_url, manager_name, manager_email, manager_source_url, agency_name, agency_domain, agency_match_method | Influencer Profile Scraper, Link in Bio Scraper and Newsletter Detector, Influencer Talent Agency Lookup |
| Links | link_in_bio_platform, outbound_links_json, other_social_profiles_json, own_website | Link in Bio Scraper and Newsletter Detector; other_social_profiles_json also by the Influencer Profile Scraper |
| Newsletter | newsletter_status, newsletter_platform, newsletter_url, newsletter_evidence_url, newsletter_check_method | Link in Bio Scraper and Newsletter Detector |
| Sells | sells_course, sells_coaching, sells_digital_product, sells_merch, sells_membership, brand_deals_visible, discount_codes_visible | Link in Bio Scraper and Newsletter Detector |
| Change | change_type, change_from, change_to, previous_run_at | Influencer Change Monitor |
| Run | source_url, read_at, row_status, error_reason, billable_events_json | All six |
The Influencer Lead List Builder runs the Influencer Finder, the Influencer Profile Scraper, and the Link in Bio Scraper and Newsletter Detector in one run and fills what they fill. The Influencer Change Monitor and the Influencer Talent Agency Lookup read the profile (and the Influencer Change Monitor reads the link page too when check_links is on), so their rows carry the identity, profile, contact, and link columns as well.
Read row_status first. ok is a normal row. error carries the reason in error_reason (not_found: user banned, private, blocked: bot detection page, outside launch country scope) and charges nothing. partial means the page was read but a step after it failed, and the reason says which.
The creator ID. creator_id is cr_ plus 16 hex characters of the SHA-1 of platform:handle of the creator's primary profile, so the same creator gets the same ID in every actor and every run. When one run sees the same creator on two platforms, the rows share the ID and each row says how it was linked: bio_link_match (0.9, one profile links to the other), website_match (0.8, same own website), handle_match (0.5, same handle on two platforms, which collides on common words). A lone row is seed_handle at 1.
Country scope. Launch scope is US creators. country_guess comes from the platform's region field, the location text, the bio, or a flag emoji, and country_guess_method says which. With us_only on (the default) a row whose guess is a known non US country is returned as an error row saying so. A row with no signal is kept, because most Instagram and TikTok profiles carry none.
🛠️ How to read an influencer profile from a handle
- Open the Input tab and put a profile URL,
platform:@handle, or a bare@handleinhandles, one per line. - If any entry is a bare
@handle, pick theplatformsit should be looked up on. - Leave
escalate_on_blockon if you want Instagram bio fields filled on the rows where the datacenter API does not answer; turn it off to never payinstagram-bio-fetch. - Click Start. One row per profile lands in the dataset as each wave finishes.
- Read
row_statusfirst, thenfollowers,engagement_rate,bio_link, andbusiness_email. - Export from the Output tab, or pull the rows through the API.
🧪 Using it in Clay
Add an Enrichment > Apify column, pick this actor, and map your handle or profile URL column to handles. Every input is accepted as a string, which is what Clay sends. The 57 fields land as one flat row with no reshaping. Gate the run on a follower band or a niche column so you only spend the event on creators you would actually work.
📚 Batch or single
One handle is one run; a list is a batch. Both shapes reach the same code path, and so does the shape the platform produces when Clay sends a top level array against an object schema. Input is deduplicated before any fetch. Rows are pushed as each wave finishes, so a run that hits its timeout keeps every row it already resolved. A row that throws becomes an error row with the reason and the run continues.
Concurrency (batch_size) defaults to the measured per platform value in the table above. Above the measured point results start dropping; that is why the default is not higher.
💵 How much does it cost to read a creator profile?
Pay per event. You are charged for output, never for input.
| Event | Fires when | Price |
|---|---|---|
actor-start | Once per run, on start. | $0.001 |
profile-read | Once per profile row where the public profile page was read and at least the follower count or the bio came back. A private, missing, or blocked profile returns a labeled error row and does not charge. | $0.006 |
instagram-bio-fetch | Once per Instagram profile row when the bio, bio link, and following were not on the embed widget or the datacenter API and the profile page was read over the residential proxy and came back readable. Only when escalate_on_block is on; uncheck it to avoid the charge and accept empty bio fields on about half of Instagram rows. Never on the embed or datacenter reads, never on another platform, never on a blocked page, never on an error row. | $0.010 |
Free Apify plan users get 180 results per calendar month, reset monthly. Upgrade to any paid Apify plan for unlimited use: https://apify.com/pricing?fpr=mamba
💳 What you are billed for.
profile-readfires once per profile row where the public page was read and at least the follower count or the bio came back. A private, missing, or blocked profile returns a labeled error row and charges nothing.instagram-bio-fetchfires only withescalate_on_blockon, only on Instagram, and only when the residential page came back readable; uncheck the option and it can never fire.
Apify bills its own platform usage on top of the event prices.
⌨️ Input
Everything is on the Input tab. The options worth explaining:
| Field | What it does |
|---|---|
handles | One per line: a profile URL on any supported platform, platform:@handle, or a bare @handle (which needs platforms and is looked up on each). Duplicates are removed before any fetch. |
platforms | Where a bare handle is looked up. |
us_only | Launch scope. A row whose country_guess is a known non US country comes back as an error row saying so; a row with no signal is kept. |
escalate_on_block | A profile fetch that returns a bot detection page is retried once over the residential proxy. On Instagram this also fetches the bio page over residential when the datacenter API does not answer; a readable page charges instagram-bio-fetch ($0.010) on top of the profile read. Off: no instagram-bio-fetch is ever charged, every read stays on datacenter, and the Instagram bio, bio link, following, and therefore business email stay empty on about half of the rows. |
batch_size | Rows in flight per platform. Leave empty for the measured default in the table above. |
twitch_client_id, twitch_app_token | Optional. Your own Helix credentials; with both set, Twitch reads use the official API. |
📤 Output
Exports to JSON, CSV, Excel, HTML or XML. One flat, snake_case row per creator per platform, 57 columns, with null rather than a missing key. No nested objects, so it drops straight into Clay, a spreadsheet or a warehouse table without a flattening step. Nested data lives only in columns ending _json, as a JSON string, so a Clay column reads one cell.
Influencer Profile Scraper fills the identity, profile, contact, and run columns; the link, newsletter, sells, and change columns are null until the Link in Bio Scraper and Newsletter Detector or the Influencer Change Monitor fills them. The row below is trimmed to the columns this actor fills.
{"creator_id": "cr_1aef6e69476cfbad","creator_id_method": "seed_handle","creator_id_confidence": 1,"platform": "youtube","handle": "mkbhd","profile_url": "https://www.youtube.com/@mkbhd","display_name": "Marques Brownlee","followers": 21300000,"following": null,"post_count": 1851,"engagement_rate": 38.26,"engagement_method": "avg_recent_views_over_subscribers","verified": true,"bio": "MKBHD: Quality Tech Videos | YouTuber | Geek | Consumer Electronics | Tech Head | Internet Personality!\n\nbusiness@MKBHD.com\n\nNYC","bio_link": "http://twitter.com/MKBHD","location_text": "United States","country_guess": "US","country_guess_method": "platform_region","business_email": "business@mkbhd.com","business_email_source": "bio","other_social_profiles_json": "[\"http://instagram.com/MKBHD\",\"http://youtube.com/c/TheStudio\",\"http://reddit.com/r/MKBHD\",\"http://discord.gg/MKBHD\"]","source_url": "https://www.youtube.com/@mkbhd/about","read_at": "2026-09-22T08:19:07.956Z","row_status": "ok","error_reason": null,"billable_events_json": "[{\"event\":\"profile-read\",\"count\":1}]"}
💡 Tips
- Pass profile URLs rather than bare handles where you have them. A URL names the platform, so nothing is looked up twice.
- Read
engagement_methodnext toengagement_rate. A TikTok lifetime figure and an Instagram recent post median are not the same number. - Leave
batch_sizeempty. The default is the measured per platform concurrency in the table above; above it, rows start dropping. - Feed
bio_linkto the Link in Bio Scraper and Newsletter Detector. That read is included in itslinks-checkedevent, so nothing is paid twice.
⚠️ Known limits
Business email coverage. Business email coverage is partial. The actor fills business_email only when a creator publishes one: in a bio, a channel description, or a podcast feed; the link page is read by the Link in Bio Scraper and Newsletter Detector, not here. YouTube keeps its business email button behind a sign-in and a CAPTCHA, which the actor does not bypass; in testing, 2 of 5 channels published an address in their description instead. Empty business_email means no public address was found, not that none exists.
Instagram bio fields need the residential page on about half of the rows. The embed widget carries counts, name, and verified. Bio, bio link, and following come from the datacenter API on about 1 in 4 attempts, otherwise from the residential profile page when escalate_on_block is on, which is the instagram-bio-fetch event. With it off, those fields stay empty on about half of Instagram rows.
Some fields are not exposed without login. What each platform does not expose without login, and what the row says instead: YouTube's business email sits behind "View email address" and a verification step, so it is not read and business_email is filled only from an address written in the public description; Instagram has no location field, so country_guess comes from the bio; Threads and podcasts have no per post engagement; podcasts have no follower count; TikTok's anonymous render carries no region, so the country comes from the bio.
Counts are what the platform shows. TikTok rounds above about ten thousand and YouTube rounds to three significant figures; the row carries the shown figure.
Country scope is US at launch. A row whose country_guess is a known non US country comes back as an error row when us_only is on; a row with no signal is kept, because most Instagram and TikTok profiles carry none.
Not X, not Facebook pages, not LinkedIn. Seven platforms are read, listed above.
What is never done. No login. No CAPTCHA solving. No LinkedIn fetch. No message to anyone. No key of ours is used on your run; this actor has no AI option, and the Twitch credentials you can supply are never logged or written to a row.
❓ FAQ
Why is business_email empty on a YouTube channel that has one?
YouTube keeps its business email button behind a sign-in and a CAPTCHA, which the actor does not bypass. The field is filled only from an address written in the public description.
Why is bio empty on an Instagram row?
The datacenter API did not answer and escalate_on_block was off, or the residential page came back blocked. Turn escalate_on_block on to fetch the bio page over residential, at $0.010 per readable page.
What does a private account return?
A labeled error row (row_status error, error_reason private) that charges nothing.
Why is engagement_rate null on Threads and podcasts?
Neither exposes per post figures without login, so no rate is computed and engagement_method is null.
Can I bring my own Twitch credentials?
Yes. With twitch_client_id and twitch_app_token set, Twitch reads use the official Helix API.
🧩 Want other GTM data?
Mamba Labs builds a fleet of GTM enrichment actors that share one flat,
Clay-ready output convention, so their rows join on company_domain or
creator_id with no cleaning step:
Every actor in the suite takes a domain, a company, or a creator and returns one flat row, so they stack in the same Clay table without reshaping anything.
🛠️ Need something custom built for you or your team? Tell us what you are trying to find and we will build it. Talk to Mamba Labs.
🆘 Support
Issues, field requests and bug reports: open an issue on the actor's Issues tab. Mamba Labs reads every one.
ℹ️ Sourcing and legal. Every field comes from pages the platforms and the creators publish to anyone without a login, read directly or through a proxy, with no login, no CAPTCHA solving, and no LinkedIn fetch. The row records what was public at
read_at. Nothing is assessed or scored; a class, a flag, or a match method says what was read and where. You are responsible for how you use the output, including under applicable data protection and platform terms.
Built by Mamba Labs.