Link in Bio Scraper and Newsletter Detector avatar

Link in Bio Scraper and Newsletter Detector

Pricing

from $8.00 / 1,000 links checkeds

Go to Apify Store
Link in Bio Scraper and Newsletter Detector

Link in Bio Scraper and Newsletter Detector

Follows a creator's link-in-bio page and website. One flat row with every outbound link, other social profiles, newsletter status and platform, what they sell, and manager contact. Each classification carries its evidence URL. Contributes to a shared creator and agency pool, on by default.

Pricing

from $8.00 / 1,000 links checkeds

Rating

0.0

(0)

Developer

Mamba Labs

Mamba Labs

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

21 minutes ago

Last modified

Share

Follows a creator's link-in-bio page and website. One flat row with every outbound link, other social profiles, the link-in-bio service, the creator's own website, newsletter status and platform, what they sell, brand deal and discount code signals, and manager contact. Each classification carries its evidence URL, so a weak call is visible rather than silent.

📦 What you get⚙️ Features and integrations
📰 Newsletter status, four classes, each with its evidence URL
🛍️ What they sell, course, coaching, digital product, merch, membership
🔗 Every outbound link, other social profiles, own website
📧 Manager and business email with the page it came from
🧾 57 flat fields, snake_case, the suite's shared row
🌐 Linktree, Stan Store, linkin.bio, and the creator's own site
🖥️ Headless render for pages that return an empty shell, opt in
🤖 AI check on your own key, opt in, never ours
🏢 Agency match against 143 agencies, opt in
⬇️ Export to JSON, CSV, Excel, HTML or XML

Bought by newsletter and email platforms selling to creators, course and membership tools, brand partnership teams checking what a creator already sells, and talent teams looking for the manager email.

🚫 This is not a link-in-bio page builder or a mass page crawler. It reads the pages a creator owns, starting from the bio link, and returns one classified row. Pages that are not the creator's (an affiliate storefront, a brand's site) are listed in outbound_links_json and never used as evidence.

If you sellRead these fields
Newsletter or email platformsnewsletter_status, newsletter_platform, newsletter_evidence_url
Course, coaching, or membership toolssells_course, sells_coaching, sells_membership
Commerce, merch, or digital product toolssells_digital_product, sells_merch, discount_codes_visible
Brand partnershipsbrand_deals_visible, outbound_links_json, own_website
Talent representationmanager_email, manager_source_url, agency_name
Anything, as a disqualifierrow_status, newsletter_check_method, link_in_bio_platform

🧭 Four classes, and an evidence URL on every one

newsletter_statusMeans
has_newsletterA Substack, beehiiv, Ghost, or Buttondown publication of the creator's; or an email platform link or an owned page that names a newsletter, a weekly email, an email list, or a Subscribe navigation item.
email_capture_onlyAn email platform form or an email input with list wording, but no newsletter named.
noneNothing found on readable owned pages.
unknown_fetch_failedNo bio link page rendered readable text. Turn on render_unreadable_pages for these.

A signal counts only from a page judged to belong to the creator: the link-in-bio page, a domain sharing a name token with the handle or display name, a newsletter platform page whose identity shares a token, or a single non affiliate bio link. newsletter_evidence_url is the page or link that decided it.

Sells flags come from link text and owned page text (course, coaching, digital product, merch, membership) plus affiliate and ambassador wording (brand_deals_visible) and discount code wording (discount_codes_visible). Manager contact is an email labeled management, mgmt, managed by, manager, or booking in the bio, on a link page, or on the website, with manager_source_url.

📏 Accuracy

The pre-research rule classifier measured 5 errors in 30 (17%) on a hand spot check. The same 30 rows on this build measured 0 errors in 30 on 2026-09-21, and two runs over them on 2026-09-22 returned the same status, platform, and evidence URL on every row. Three of the 30 drove rule 7 (cadence wording, footer only headings, email platform paths named /newsletter), so they are in sample; before that rule the build measured 3 errors in 30 (10%). Residual errors run in both directions, which is why every row carries its evidence URL.

57 fields per row, the same 57 in the same order on every actor in this suite. The ones this actor fills:

FieldWhat it holds
link_in_bio_platformlinktr.ee, stan.store, linkin.bio, and so on, when the bio link is a link-in-bio service
outbound_links_json, other_social_profiles_json, own_websiteEvery outbound link on the pages read (capped at 100), the other social profiles found, and the creator's own site
newsletter_status, newsletter_platform, newsletter_urlThe class, the platform (substack, beehiiv, kit, and so on), and the newsletter's own URL when exposed
newsletter_evidence_url, newsletter_check_methodThe page or link that decided the status, and whether rules, rules_plus_ai, or browser decided it
sells_course, sells_coaching, sells_digital_product, sells_merch, sells_membershipWhat is sold on an owned page
brand_deals_visible, discount_codes_visibleAffiliate or ambassador wording, and a promo code on an owned page
business_email, business_email_source, manager_name, manager_email, manager_source_urlPublic contact, and where it was found
website_email, website_email_source_urlFilled by the website scan add-on
agency_name, agency_domain, agency_match_methodFilled by the agency match add-on
source_url, read_at, row_status, error_reason, billable_events_jsonWhere the row came from, when, whether it is usable, and what it cost

⚠️ false and null are not the same thing, and an empty cell is not a failed read. A column this actor does not own is null, never missing. On a boolean, false means the page was read and the signal was not there; null means nothing read it. On an Influencer Profile Scraper row, verified: false is a profile without the mark and verified: null is a profile that could not be read. Read row_status before any other column; the paragraph below says how.

Every actor in this suite writes the same 57 flat columns, in the same order, and fills the ones it owns. A column an actor does not own is null, never missing. Nested data lives only in columns ending _json, as a JSON string, so a Clay column reads one cell.

GroupColumnsFilled by
Identitycreator_id, creator_id_method, creator_id_confidence, platform, handle, profile_url, display_nameInfluencer Finder, Influencer Profile Scraper
Discoverysearch_keyword, niche, similar_creators, similar_creators_methodInfluencer Finder
Profilefollowers, following, post_count, engagement_rate, engagement_method, verified, bio, bio_link, location_text, country_guess, country_guess_methodInfluencer Profile Scraper
Contactbusiness_email, business_email_source, website_email, website_email_source_url, manager_name, manager_email, manager_source_url, agency_name, agency_domain, agency_match_methodInfluencer Profile Scraper, Link in Bio Scraper and Newsletter Detector, Influencer Talent Agency Lookup
Linkslink_in_bio_platform, outbound_links_json, other_social_profiles_json, own_websiteLink in Bio Scraper and Newsletter Detector; other_social_profiles_json also by the Influencer Profile Scraper
Newsletternewsletter_status, newsletter_platform, newsletter_url, newsletter_evidence_url, newsletter_check_methodLink in Bio Scraper and Newsletter Detector
Sellssells_course, sells_coaching, sells_digital_product, sells_merch, sells_membership, brand_deals_visible, discount_codes_visibleLink in Bio Scraper and Newsletter Detector
Changechange_type, change_from, change_to, previous_run_atInfluencer Change Monitor
Runsource_url, read_at, row_status, error_reason, billable_events_jsonAll six

The Influencer Lead List Builder runs the Influencer Finder, the Influencer Profile Scraper, and the Link in Bio Scraper and Newsletter Detector in one run and fills what they fill. The Influencer Change Monitor and the Influencer Talent Agency Lookup read the profile (and the Influencer Change Monitor reads the link page too when check_links is on), so their rows carry the identity, profile, contact, and link columns as well.

Read row_status first. ok is a normal row. error carries the reason in error_reason (not_found: user banned, private, blocked: bot detection page) and charges nothing. partial means the page was read but a step after it failed, and the reason says which.

The creator ID. creator_id is cr_ plus 16 hex characters of the SHA-1 of platform:handle of the creator's primary profile, so the same creator gets the same ID in every actor and every run. When one run sees the same creator on two platforms, the rows share the ID and each row says how it was linked: bio_link_match (0.9, one profile links to the other), website_match (0.8, same own website), handle_match (0.5, same handle on two platforms, which collides on common words). A lone row is seed_handle at 1.

Country scope. This actor applies no country filter. A row started from a handle carries country_guess and country_guess_method from the profile read (the platform's region field, the location text, the bio, or a flag emoji); a row started from a bare bio_links entry has no profile read, so both are null.

  1. Open the Input tab and put creator handles or profile URLs in handles, or link-in-bio and website URLs in bio_links, one per line.
  2. Turn on render_unreadable_pages if your list carries Stan Store, linkin.bio, or Typeform pages; a plain fetch returns an empty shell on those.
  3. Turn on scan_website_for_email and match_agencies if you want the add-ons.
  4. Click Start. One row per creator lands in the dataset as each wave finishes.
  5. Read row_status first, then newsletter_status with newsletter_evidence_url next to it.
  6. Export from the Output tab, or pull the rows through the API.

🧪 Using it in Clay

Add an Enrichment > Apify column, pick this actor, and map your handle column to handles or your bio link column to bio_links. Every input is accepted as a string, which is what Clay sends. Filter on newsletter_status before filtering on anything else; an unknown_fetch_failed row carries no information about the creator and must not be scored as no newsletter.

🎯 Best input

Pass creator handles where you have them. A handle lets the actor match the creator's own site by name. A bare website link works, but a site named differently from the creator can read one class lower, for example email_capture_only instead of has_newsletter.

📚 Batch or single

One handle or one bio link is one run; a list is a batch. Both shapes reach the same code path, and so does the shape the platform produces when Clay sends a top level array against an object schema. Input is deduplicated before any fetch. Rows are pushed as each wave finishes, so a run that hits its timeout keeps every row it already resolved. A row that throws becomes an error row with the reason and the run continues.

Concurrency (batch_size) is creators in flight, default 3, capped at 10. Renders run one at a time whatever the setting.

Pay per event. You are charged for output, never for input.

EventFires whenPrice
actor-startOnce per run, on start.$0.002
links-checkedOnce per creator whose bio link or link-in-bio page was fetched and classified. A creator with no bio link returns newsletter_status none and does not charge this event.$0.008
browser-renderOnce per page rendered in the headless browser because the plain fetch returned no readable content, and the rendered page came back readable: 120 characters of visible text or 3 links off the page's host, and no block page. A render that returns the same empty shell is not charged, and a charged render always reaches the classifier (a link grid with little text, such as a Linkin.bio page, counts as readable by its links). Only when render_unreadable_pages is on.$0.004
website-scanOnce per creator whose own website was scanned for an email. Only when scan_website_for_email is on.$0.005
agency-matchOnce per row where a manager or business email domain matched the bundled agency list and an agency name came back.$0.003
instagram-bio-fetchOnce per Instagram profile row when the bio, bio link, and following were not on the embed widget or the datacenter API and the profile page was read over the residential proxy and came back readable. Only when escalate_on_block is on; uncheck it to avoid the charge and accept empty bio fields on about half of Instagram rows. Never on the embed or datacenter reads, never on another platform, never on a blocked page, never on an error row.$0.010

Free Apify plan users get 140 results per calendar month, reset monthly. Upgrade to any paid Apify plan for unlimited use: https://apify.com/pricing?fpr=mamba

💳 What you are billed for. links-checked fires once per creator whose bio link page was fetched and classified; a creator with no bio link returns none and does not charge it. browser-render, website-scan, and agency-match fire only when their option is on and only when they produced something. The profile read that finds the bio link is not charged here (links-checked covers it), but an Instagram handle whose bio needs the residential page is charged instagram-bio-fetch. Start from bio_links to skip the profile read altogether.

Apify bills its own platform usage on top of the event prices.

How a render runs: the page is polled once a second until it is readable, returned 1.5 seconds after that, and given up at 25 seconds; a session whose document never arrives or whose proxy tunnel drops a script gets one retry on a fresh session inside a 40 second budget. Renders run one at a time. Stan Store pages carry their content in the HTML and read at domcontentloaded (5 to 18 seconds over the proxy, measured 2026-09-22); Linkin.bio fills its link grid 1 to 2 seconds after that. At 1 GB the worst case (both attempts spent) costs about $0.0033 of compute; a typical render costs $0.0009 to $0.0015, so the $0.004 price covers it.

⌨️ Input

Everything is on the Input tab. The options worth explaining:

FieldWhat it does
handlesHandles or profile URLs. The profile is read for the bio link first; that read is part of the links-checked event.
bio_linksLink-in-bio or website URLs to start from directly, skipping the profile read. When one input item carries both a handle and bio links, the links are read as that creator's links and the actor returns one row for that creator, not one row per link; to get one row per link, pass the links in bio_links alone.
render_unreadable_pagesOff by default. Stan Store, linkin.bio, and Typeform shells return an empty page to a plain fetch and are classified unknown_fetch_failed (157 of 1,034 creators in the pre-research). On, those pages are rendered in a headless browser and charged per render.
ai_check, ai_provider, ai_api_keyOff by default. With your own Anthropic or OpenAI key, a model reads the rule classifier's evidence and page text and rules on each row; newsletter_check_method becomes rules_plus_ai. The key is never stored, logged, or written to a row, and no Mamba Labs key is ever used.
scan_website_for_emailOff by default. For creators with their own website (not a link-in-bio page), reads the home, contact, and about pages and the footer for an email and records where it was found in website_email_source_url. Charged per creator scanned.
match_agenciesMatches the domain of a manager or business email against the bundled talent agency list (143 agencies from 13 public directories, 93 with a public contact). Charged per matched row.
escalate_on_blockOn by default. A profile fetch that returns a bot detection page is retried once over the residential proxy, and on Instagram the bio page is fetched over residential when the datacenter API does not answer; a readable page charges instagram-bio-fetch ($0.010). Off: no instagram-bio-fetch is ever charged, and the Instagram bio link stays empty on about half of the rows, which leaves those rows with no page to follow. Ignored for bio_links.

📤 Output

Exports to JSON, CSV, Excel, HTML or XML. One flat, snake_case row per creator per platform, 57 columns, with null rather than a missing key. No nested objects, so it drops straight into Clay, a spreadsheet or a warehouse table without a flattening step. Nested data lives only in columns ending _json, as a JSON string, so a Clay column reads one cell.

This row started from a bare bio_links entry, so platform is null and creator_id_method is bio_link_only; a row started from a handle carries the platform and profile columns too. The row below is trimmed to the columns this actor fills.

{
"creator_id": "cr_1f22152aa4fd5c19",
"creator_id_method": "bio_link_only",
"creator_id_confidence": 0.3,
"platform": null,
"handle": "lifewithnancy1",
"bio_link": "https://stan.store/lifewithnancy1",
"business_email": "coachnancybookings@gmail.com",
"business_email_source": "link_page",
"manager_email": null,
"link_in_bio_platform": "stan.store",
"outbound_links_json": "[{\"url\":\"https://www.instagram.com/lifewithnancy1\",\"from\":\"https://stan.store/lifewithnancy1\"},{\"url\":\"https://www.tiktok.com/@lifewithnancy1\",\"from\":\"https://stan.store/lifewithnancy1\"},{\"url\":\"https://www.youtube.com/@LifewithNancy1-cn9te\",\"from\":\"https://stan.store/lifewithnancy1\"}]",
"other_social_profiles_json": "[\"https://www.instagram.com/lifewithnancy1\",\"https://www.tiktok.com/@lifewithnancy1\",\"https://www.youtube.com/@LifewithNancy1-cn9te\"]",
"own_website": null,
"newsletter_status": "none",
"newsletter_platform": null,
"newsletter_url": null,
"newsletter_evidence_url": "https://stan.store/lifewithnancy1",
"newsletter_check_method": "browser",
"sells_course": false,
"sells_coaching": false,
"sells_digital_product": false,
"sells_merch": false,
"sells_membership": false,
"brand_deals_visible": false,
"discount_codes_visible": false,
"source_url": "https://stan.store/lifewithnancy1",
"read_at": "2026-09-22T08:22:28.675Z",
"row_status": "ok",
"error_reason": null,
"billable_events_json": "[{\"event\":\"browser-render\",\"count\":1},{\"event\":\"links-checked\",\"count\":1}]"
}

💡 Tips

  • Pass handles where you have them, not bare links. See Best input above.
  • Turn render_unreadable_pages on for a list with Stan Store or linkin.bio pages, and read newsletter_check_method to see which rows the browser decided.
  • Read newsletter_evidence_url on every has_newsletter row you act on. A weak call is visible there, not silent.
  • Bring your own Anthropic or OpenAI key with ai_check when a rule verdict alone is not enough for your list; newsletter_check_method then reads rules_plus_ai.

⚠️ Known limits

Business email coverage. Business email coverage is partial. The actor fills business_email only when a creator publishes one: in a bio, a link-in-bio page, or a YouTube channel description. YouTube keeps its business email button behind a sign-in and a CAPTCHA, which the actor does not bypass; in testing, 2 of 5 channels published an address in their description instead. Empty business_email means no public address was found, not that none exists.

An unreadable page is unknown_fetch_failed, not none. Stan Store, linkin.bio, and Typeform shells return an empty page to a plain fetch (157 of 1,034 creators in the pre-research). With render_unreadable_pages off, those rows say unknown_fetch_failed and carry no information about the newsletter.

A signal counts only from a page judged to belong to the creator. A site named differently from the creator can read one class lower, for example email_capture_only instead of has_newsletter; pass the handle so the site can be matched by name.

The rule classifier is not perfect. The pre-research version measured 5 errors in 30; this build measured 0 in the same 30, three of which drove rule 7 and are in sample. Residual errors run in both directions, which is why every row carries its evidence URL.

The agency list is partial by construction. 143 agencies from 13 public directories; a creator represented by an agency outside the list returns no match, never a guess.

Not X, not Facebook pages, not LinkedIn. A bio link to one of those is listed in outbound_links_json and not read.

What is never done. No login. No CAPTCHA solving. No LinkedIn fetch. No message to anyone. No key of ours is used on your run; the AI check runs only on the key you supply, and the key is never logged or written to a row.

What this actor shares

This run contributes the records it finds to a shared creator and agency pool that all users of this actor read from. What one run finds, the next run can read.

The toggle is contribute_to_shared_pool. It is a boolean input and it is on by default. Turn it off and the run still reads the pool and writes nothing to it.

What this actor contributes. The manager contacts it finds, the agency it matched them to, and any manager email domain the pool has not seen before, as a candidate agency for a later run to confirm.

Only public data that is already in your own output. Every field written to the pool is a field this run returned to you, read from a page the platform or the creator publishes to anyone without a login. Nothing from your Apify account, your input list, your API keys, or your own notes is sent. A contribution never deletes anything from the pool.

What a contribution is labeled with. The actor ID, the run ID, the pool key issued to the actor build, and a hash of the calling IP address, used for the rate limit and nothing else. Your Apify account and your user ID are not recorded.

Contributing is free. No event is charged for a write to the pool. If the pool is unreachable the run finishes as normal, the rows are dropped, and the run log says so.

❓ FAQ

I passed one handle and three links and got one row. Where are the other two? When one input item carries both a handle and bio links, the links are read as that creator's links and the actor returns one row for that creator. To get one row per link, pass the links in bio_links alone.

What is the difference between none and unknown_fetch_failed? none means readable owned pages were checked and nothing was found. unknown_fetch_failed means no bio link page rendered readable text, so nothing is known. Only none is about the creator.

Why was a render not charged? A render that returns the same empty shell as the plain fetch is not charged. Only a render that came back readable (120 characters of visible text or 3 links off the page's host, and no block page) fires browser-render, and a charged render always reaches the classifier.

Does the AI check use a Mamba Labs key? No. It runs only on the key you supply, and the key is never stored, logged, or written to a row.

Is the profile read charged here? No. The profile read that finds the bio link is not charged here (links-checked covers it), but an Instagram handle whose bio needs the residential page is charged instagram-bio-fetch.

🧩 Want other GTM data?

Mamba Labs builds a fleet of GTM enrichment actors that share one flat, Clay-ready output convention, so their rows join on company_domain or creator_id with no cleaning step:

🕵️ Agent Accessibility Auditor🤖 AI Tooling Detector
📡 B2B Buying Signals Aggregator🚀 Prospect Engine
📝 Publishing Frequency Tracker🦋 Bluesky Brand Presence Mapper
Sequencer Lead Push🔄 Company Change-Event Feed
📇 Company Contact Details Extractor🧭 Company Discovery List Builder
🏢 Company Firmographic Enricher🪪 Company Identity Resolver
🌐 Company Social Presence Mapper🏷️ Contact Classifier
📈 Influencer Change Monitor🔎 Influencer Finder
🧾 Influencer Lead List Builder👤 Influencer Profile Scraper
📬 Domain Deliverability Checker🔗 Domain to LinkedIn URL Resolver
🛒 Ecommerce Platform Profiler✉️ Work Email Waterfall Finder
🎪 Event Presence Index💵 Funding Record and Filings
💰 Funding and Press Signal Scanner🐙 GitHub Organization Signal Scanner
🧑‍💼 GTM Hiring Signal Scraper📋 Job Posting Monitor
🧱 Tech Stack Detector🎯 ICP Fit Scorer
🔑 Job Board Keyword Scanner⚖️ Legal Entity Resolver
💼 LinkedIn Company Page Mapper💬 LinkedIn Post Tracker and Comment Capture
📢 Meta Ad Library Monitor📸 Instagram and Facebook Brand Mapper
📮 Outbound Stack Detector📄 Page Finder and Extractor
👤 People Finder and Email Verifier📌 Pinterest Brand Presence Mapper
🏛️ Government Contract Award Monitor📅 Public Company Reporting Window Finder
👽 Reddit Brand and Mention MonitorTrustpilot Reputation Enricher
👥 Team Page People Extractor🎵 TikTok Brand Presence Mapper
📈 Website Traffic Rank Estimator🏅 Workplace Program Detector
✖️ X Twitter Brand Presence Mapper▶️ YouTube Channel Stats Extractor

Every actor in the suite takes a domain, a company, or a creator and returns one flat row, so they stack in the same Clay table without reshaping anything.

🛠️ Need something custom built for you or your team? Tell us what you are trying to find and we will build it. Talk to Mamba Labs.

🆘 Support

Issues, field requests and bug reports: open an issue on the actor's Issues tab. Mamba Labs reads every one.

ℹ️ Sourcing and legal. Every field comes from pages the platforms and the creators publish to anyone without a login, read directly or through a proxy, with no login, no CAPTCHA solving, and no LinkedIn fetch. The row records what was public at read_at. Nothing is assessed or scored; a class, a flag, or a match method says what was read and where. You are responsible for how you use the output, including under applicable data protection and platform terms.

Built by Mamba Labs.