Link in Bio Scraper and Newsletter Detector
Pricing
from $8.00 / 1,000 links checkeds
Link in Bio Scraper and Newsletter Detector
Follows a creator's link-in-bio page and website. One flat row with every outbound link, other social profiles, newsletter status and platform, what they sell, brand deal signals, and manager contact. Each classification carries its evidence URL, so a weak call is visible rather than silent.
Pricing
from $8.00 / 1,000 links checkeds
Rating
0.0
(0)
Developer
Mamba Labs
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
7 hours ago
Last modified
Categories
Share
🔗 What can Link in Bio Scraper and Newsletter Detector do?
Follows a creator's link-in-bio page and website. One flat row with every outbound link, other social profiles, the link-in-bio service, the creator's own website, newsletter status and platform, what they sell, brand deal and discount code signals, and manager contact. Each classification carries its evidence URL, so a weak call is visible rather than silent.
| 📦 What you get | ⚙️ Features and integrations |
|---|---|
| 📰 Newsletter status, four classes, each with its evidence URL 🛍️ What they sell, course, coaching, digital product, merch, membership 🔗 Every outbound link, other social profiles, own website 📧 Manager and business email with the page it came from 🧾 57 flat fields, snake_case, the suite's shared row | 🌐 Linktree, Stan Store, linkin.bio, and the creator's own site 🖥️ Headless render for pages that return an empty shell, opt in 🤖 AI check on your own key, opt in, never ours 🏢 Agency match against 103 agencies, opt in ⬇️ Export to JSON, CSV, Excel, HTML or XML |
Bought by newsletter and email platforms selling to creators, course and membership tools, brand partnership teams checking what a creator already sells, and talent teams looking for the manager email.
🚫 This is not a link-in-bio page builder or a mass page crawler. It reads the pages a creator owns, starting from the bio link, and returns one classified row. Pages that are not the creator's (an affiliate storefront, a brand's site) are listed in
outbound_links_jsonand never used as evidence.
💡 Why use Link in Bio Scraper and Newsletter Detector?
| If you sell | Read these fields |
|---|---|
| Newsletter or email platforms | newsletter_status, newsletter_platform, newsletter_evidence_url |
| Course, coaching, or membership tools | sells_course, sells_coaching, sells_membership |
| Commerce, merch, or digital product tools | sells_digital_product, sells_merch, discount_codes_visible |
| Brand partnerships | brand_deals_visible, outbound_links_json, own_website |
| Talent representation | manager_email, manager_source_url, agency_name |
| Anything, as a disqualifier | row_status, newsletter_check_method, link_in_bio_platform |
🧭 Four classes, and an evidence URL on every one
newsletter_status | Means |
|---|---|
has_newsletter | A Substack, beehiiv, Ghost, or Buttondown publication of the creator's; or an email platform link or an owned page that names a newsletter, a weekly email, an email list, or a Subscribe navigation item. |
email_capture_only | An email platform form or an email input with list wording, but no newsletter named. |
none | Nothing found on readable owned pages. |
unknown_fetch_failed | No bio link page rendered readable text. Turn on render_unreadable_pages for these. |
A signal counts only from a page judged to belong to the creator: the link-in-bio page, a domain sharing a name token with the handle or display name, a newsletter platform page whose identity shares a token, or a single non affiliate bio link. newsletter_evidence_url is the page or link that decided it.
Sells flags come from link text and owned page text (course, coaching, digital product, merch, membership) plus affiliate and ambassador wording (brand_deals_visible) and discount code wording (discount_codes_visible). Manager contact is an email labeled management, mgmt, managed by, manager, or booking in the bio, on a link page, or on the website, with manager_source_url.
📏 Accuracy
The pre-research rule classifier measured 5 errors in 30 (17%) on a hand spot check. The same 30 rows on this build measured 0 errors in 30 on 2026-09-21, and two runs over them on 2026-09-22 returned the same status, platform, and evidence URL on every row. Three of the 30 drove rule 7 (cadence wording, footer only headings, email platform paths named /newsletter), so they are in sample; before that rule the build measured 3 errors in 30 (10%). Residual errors run in both directions, which is why every row carries its evidence URL.
📋 What data can Link in Bio Scraper and Newsletter Detector extract?
57 fields per row, the same 57 in the same order on every actor in this suite. The ones this actor fills:
| Field | What it holds |
|---|---|
link_in_bio_platform | linktr.ee, stan.store, linkin.bio, and so on, when the bio link is a link-in-bio service |
outbound_links_json, other_social_profiles_json, own_website | Every outbound link on the pages read (capped at 100), the other social profiles found, and the creator's own site |
newsletter_status, newsletter_platform, newsletter_url | The class, the platform (substack, beehiiv, kit, and so on), and the newsletter's own URL when exposed |
newsletter_evidence_url, newsletter_check_method | The page or link that decided the status, and whether rules, rules_plus_ai, or browser decided it |
sells_course, sells_coaching, sells_digital_product, sells_merch, sells_membership | What is sold on an owned page |
brand_deals_visible, discount_codes_visible | Affiliate or ambassador wording, and a promo code on an owned page |
business_email, business_email_source, manager_name, manager_email, manager_source_url | Public contact, and where it was found |
website_email, website_email_source_url | Filled by the website scan add-on |
agency_name, agency_domain, agency_match_method | Filled by the agency match add-on |
source_url, read_at, row_status, error_reason, billable_events_json | Where the row came from, when, whether it is usable, and what it cost |
⚠️
falseandnullare not the same thing, and an empty cell is not a failed read. A column this actor does not own isnull, never missing. On a boolean,falsemeans the page was read and the signal was not there;nullmeans nothing read it. On an Influencer Profile Scraper row,verified: falseis a profile without the mark andverified: nullis a profile that could not be read. Readrow_statusbefore any other column; the paragraph below says how.
Every actor in this suite writes the same 57 flat columns, in the same order, and fills the ones it owns. A column an actor does not own is null, never missing. Nested data lives only in columns ending _json, as a JSON string, so a Clay column reads one cell.
| Group | Columns | Filled by |
|---|---|---|
| Identity | creator_id, creator_id_method, creator_id_confidence, platform, handle, profile_url, display_name | Influencer Finder, Influencer Profile Scraper |
| Discovery | search_keyword, niche, similar_creators, similar_creators_method | Influencer Finder |
| Profile | followers, following, post_count, engagement_rate, engagement_method, verified, bio, bio_link, location_text, country_guess, country_guess_method | Influencer Profile Scraper |
| Contact | business_email, business_email_source, website_email, website_email_source_url, manager_name, manager_email, manager_source_url, agency_name, agency_domain, agency_match_method | Influencer Profile Scraper, Link in Bio Scraper and Newsletter Detector, Influencer Talent Agency Lookup |
| Links | link_in_bio_platform, outbound_links_json, other_social_profiles_json, own_website | Link in Bio Scraper and Newsletter Detector; other_social_profiles_json also by the Influencer Profile Scraper |
| Newsletter | newsletter_status, newsletter_platform, newsletter_url, newsletter_evidence_url, newsletter_check_method | Link in Bio Scraper and Newsletter Detector |
| Sells | sells_course, sells_coaching, sells_digital_product, sells_merch, sells_membership, brand_deals_visible, discount_codes_visible | Link in Bio Scraper and Newsletter Detector |
| Change | change_type, change_from, change_to, previous_run_at | Influencer Change Monitor |
| Run | source_url, read_at, row_status, error_reason, billable_events_json | All six |
The Influencer Lead List Builder runs the Influencer Finder, the Influencer Profile Scraper, and the Link in Bio Scraper and Newsletter Detector in one run and fills what they fill. The Influencer Change Monitor and the Influencer Talent Agency Lookup read the profile (and the Influencer Change Monitor reads the link page too when check_links is on), so their rows carry the identity, profile, contact, and link columns as well.
Read row_status first. ok is a normal row. error carries the reason in error_reason (not_found: user banned, private, blocked: bot detection page) and charges nothing. partial means the page was read but a step after it failed, and the reason says which.
The creator ID. creator_id is cr_ plus 16 hex characters of the SHA-1 of platform:handle of the creator's primary profile, so the same creator gets the same ID in every actor and every run. When one run sees the same creator on two platforms, the rows share the ID and each row says how it was linked: bio_link_match (0.9, one profile links to the other), website_match (0.8, same own website), handle_match (0.5, same handle on two platforms, which collides on common words). A lone row is seed_handle at 1.
Country scope. This actor applies no country filter. A row started from a handle carries country_guess and country_guess_method from the profile read (the platform's region field, the location text, the bio, or a flag emoji); a row started from a bare bio_links entry has no profile read, so both are null.
🛠️ How to check a creator's link in bio page for a newsletter
- Open the Input tab and put creator handles or profile URLs in
handles, or link-in-bio and website URLs inbio_links, one per line. - Turn on
render_unreadable_pagesif your list carries Stan Store, linkin.bio, or Typeform pages; a plain fetch returns an empty shell on those. - Turn on
scan_website_for_emailandmatch_agenciesif you want the add-ons. - Click Start. One row per creator lands in the dataset as each wave finishes.
- Read
row_statusfirst, thennewsletter_statuswithnewsletter_evidence_urlnext to it. - Export from the Output tab, or pull the rows through the API.
🧪 Using it in Clay
Add an Enrichment > Apify column, pick this actor, and map your handle column to handles or your bio link column to bio_links. Every input is accepted as a string, which is what Clay sends. Filter on newsletter_status before filtering on anything else; an unknown_fetch_failed row carries no information about the creator and must not be scored as no newsletter.
🎯 Best input
Pass creator handles where you have them. A handle lets the actor match the creator's own site by name. A bare website link works, but a site named differently from the creator can read one class lower, for example email_capture_only instead of has_newsletter.
📚 Batch or single
One handle or one bio link is one run; a list is a batch. Both shapes reach the same code path, and so does the shape the platform produces when Clay sends a top level array against an object schema. Input is deduplicated before any fetch. Rows are pushed as each wave finishes, so a run that hits its timeout keeps every row it already resolved. A row that throws becomes an error row with the reason and the run continues.
Concurrency (batch_size) is creators in flight, default 3, capped at 10. Renders run one at a time whatever the setting.
💵 How much does it cost to check a link in bio page?
Pay per event. You are charged for output, never for input.
| Event | Fires when | Price |
|---|---|---|
actor-start | Once per run, on start. | $0.002 |
links-checked | Once per creator whose bio link or link-in-bio page was fetched and classified. A creator with no bio link returns newsletter_status none and does not charge this event. | $0.008 |
browser-render | Once per page rendered in the headless browser because the plain fetch returned no readable content, and the rendered page came back readable: 120 characters of visible text or 3 links off the page's host, and no block page. A render that returns the same empty shell is not charged, and a charged render always reaches the classifier (a link grid with little text, such as a Linkin.bio page, counts as readable by its links). Only when render_unreadable_pages is on. | $0.004 |
website-scan | Once per creator whose own website was scanned for an email. Only when scan_website_for_email is on. | $0.005 |
agency-match | Once per row where a manager or business email domain matched the bundled agency list and an agency name came back. | $0.003 |
instagram-bio-fetch | Once per Instagram profile row when the bio, bio link, and following were not on the embed widget or the datacenter API and the profile page was read over the residential proxy and came back readable. Only when escalate_on_block is on; uncheck it to avoid the charge and accept empty bio fields on about half of Instagram rows. Never on the embed or datacenter reads, never on another platform, never on a blocked page, never on an error row. | $0.010 |
Free Apify plan users get 140 results per calendar month, reset monthly. Upgrade to any paid Apify plan for unlimited use: https://apify.com/pricing?fpr=mamba
💳 What you are billed for.
links-checkedfires once per creator whose bio link page was fetched and classified; a creator with no bio link returnsnoneand does not charge it.browser-render,website-scan, andagency-matchfire only when their option is on and only when they produced something. The profile read that finds the bio link is not charged here (links-checkedcovers it), but an Instagram handle whose bio needs the residential page is chargedinstagram-bio-fetch. Start frombio_linksto skip the profile read altogether.
Apify bills its own platform usage on top of the event prices.
How a render runs: the page is polled once a second until it is readable, returned 1.5 seconds after that, and given up at 25 seconds; a session whose document never arrives or whose proxy tunnel drops a script gets one retry on a fresh session inside a 40 second budget. Renders run one at a time. Stan Store pages carry their content in the HTML and read at domcontentloaded (5 to 18 seconds over the proxy, measured 2026-09-22); Linkin.bio fills its link grid 1 to 2 seconds after that. At 1 GB the worst case (both attempts spent) costs about $0.0033 of compute; a typical render costs $0.0009 to $0.0015, so the $0.004 price covers it.
⌨️ Input
Everything is on the Input tab. The options worth explaining:
| Field | What it does |
|---|---|
handles | Handles or profile URLs. The profile is read for the bio link first; that read is part of the links-checked event. |
bio_links | Link-in-bio or website URLs to start from directly, skipping the profile read. When one input item carries both a handle and bio links, the links are read as that creator's links and the actor returns one row for that creator, not one row per link; to get one row per link, pass the links in bio_links alone. |
render_unreadable_pages | Off by default. Stan Store, linkin.bio, and Typeform shells return an empty page to a plain fetch and are classified unknown_fetch_failed (157 of 1,034 creators in the pre-research). On, those pages are rendered in a headless browser and charged per render. |
ai_check, ai_provider, ai_api_key | Off by default. With your own Anthropic or OpenAI key, a model reads the rule classifier's evidence and page text and rules on each row; newsletter_check_method becomes rules_plus_ai. The key is never stored, logged, or written to a row, and no Mamba Labs key is ever used. |
scan_website_for_email | Off by default. For creators with their own website (not a link-in-bio page), reads the home, contact, and about pages and the footer for an email and records where it was found in website_email_source_url. Charged per creator scanned. |
match_agencies | Matches the domain of a manager or business email against the bundled talent agency list (103 agencies from 9 public directories, 70 with a public contact). Charged per matched row. |
escalate_on_block | On by default. A profile fetch that returns a bot detection page is retried once over the residential proxy, and on Instagram the bio page is fetched over residential when the datacenter API does not answer; a readable page charges instagram-bio-fetch ($0.010). Off: no instagram-bio-fetch is ever charged, and the Instagram bio link stays empty on about half of the rows, which leaves those rows with no page to follow. Ignored for bio_links. |
📤 Output
Exports to JSON, CSV, Excel, HTML or XML. One flat, snake_case row per creator per platform, 57 columns, with null rather than a missing key. No nested objects, so it drops straight into Clay, a spreadsheet or a warehouse table without a flattening step. Nested data lives only in columns ending _json, as a JSON string, so a Clay column reads one cell.
This row started from a bare bio_links entry, so platform is null and creator_id_method is bio_link_only; a row started from a handle carries the platform and profile columns too. The row below is trimmed to the columns this actor fills.
{"creator_id": "cr_1f22152aa4fd5c19","creator_id_method": "bio_link_only","creator_id_confidence": 0.3,"platform": null,"handle": "lifewithnancy1","bio_link": "https://stan.store/lifewithnancy1","business_email": "coachnancybookings@gmail.com","business_email_source": "link_page","manager_email": null,"link_in_bio_platform": "stan.store","outbound_links_json": "[{\"url\":\"https://www.instagram.com/lifewithnancy1\",\"from\":\"https://stan.store/lifewithnancy1\"},{\"url\":\"https://www.tiktok.com/@lifewithnancy1\",\"from\":\"https://stan.store/lifewithnancy1\"},{\"url\":\"https://www.youtube.com/@LifewithNancy1-cn9te\",\"from\":\"https://stan.store/lifewithnancy1\"}]","other_social_profiles_json": "[\"https://www.instagram.com/lifewithnancy1\",\"https://www.tiktok.com/@lifewithnancy1\",\"https://www.youtube.com/@LifewithNancy1-cn9te\"]","own_website": null,"newsletter_status": "none","newsletter_platform": null,"newsletter_url": null,"newsletter_evidence_url": "https://stan.store/lifewithnancy1","newsletter_check_method": "browser","sells_course": false,"sells_coaching": false,"sells_digital_product": false,"sells_merch": false,"sells_membership": false,"brand_deals_visible": false,"discount_codes_visible": false,"source_url": "https://stan.store/lifewithnancy1","read_at": "2026-09-22T08:22:28.675Z","row_status": "ok","error_reason": null,"billable_events_json": "[{\"event\":\"browser-render\",\"count\":1},{\"event\":\"links-checked\",\"count\":1}]"}
💡 Tips
- Pass handles where you have them, not bare links. See Best input above.
- Turn
render_unreadable_pageson for a list with Stan Store or linkin.bio pages, and readnewsletter_check_methodto see which rows the browser decided. - Read
newsletter_evidence_urlon everyhas_newsletterrow you act on. A weak call is visible there, not silent. - Bring your own Anthropic or OpenAI key with
ai_checkwhen a rule verdict alone is not enough for your list;newsletter_check_methodthen readsrules_plus_ai.
⚠️ Known limits
Business email coverage. Business email coverage is partial. The actor fills business_email only when a creator publishes one: in a bio, a link-in-bio page, or a YouTube channel description. YouTube keeps its business email button behind a sign-in and a CAPTCHA, which the actor does not bypass; in testing, 2 of 5 channels published an address in their description instead. Empty business_email means no public address was found, not that none exists.
An unreadable page is unknown_fetch_failed, not none. Stan Store, linkin.bio, and Typeform shells return an empty page to a plain fetch (157 of 1,034 creators in the pre-research). With render_unreadable_pages off, those rows say unknown_fetch_failed and carry no information about the newsletter.
A signal counts only from a page judged to belong to the creator. A site named differently from the creator can read one class lower, for example email_capture_only instead of has_newsletter; pass the handle so the site can be matched by name.
The rule classifier is not perfect. The pre-research version measured 5 errors in 30; this build measured 0 in the same 30, three of which drove rule 7 and are in sample. Residual errors run in both directions, which is why every row carries its evidence URL.
The agency list is partial by construction. 103 agencies from 9 public directories; a creator represented by an agency outside the list returns no match, never a guess.
Not X, not Facebook pages, not LinkedIn. A bio link to one of those is listed in outbound_links_json and not read.
What is never done. No login. No CAPTCHA solving. No LinkedIn fetch. No message to anyone. No key of ours is used on your run; the AI check runs only on the key you supply, and the key is never logged or written to a row.
❓ FAQ
I passed one handle and three links and got one row. Where are the other two?
When one input item carries both a handle and bio links, the links are read as that creator's links and the actor returns one row for that creator. To get one row per link, pass the links in bio_links alone.
What is the difference between none and unknown_fetch_failed?
none means readable owned pages were checked and nothing was found. unknown_fetch_failed means no bio link page rendered readable text, so nothing is known. Only none is about the creator.
Why was a render not charged?
A render that returns the same empty shell as the plain fetch is not charged. Only a render that came back readable (120 characters of visible text or 3 links off the page's host, and no block page) fires browser-render, and a charged render always reaches the classifier.
Does the AI check use a Mamba Labs key? No. It runs only on the key you supply, and the key is never stored, logged, or written to a row.
Is the profile read charged here?
No. The profile read that finds the bio link is not charged here (links-checked covers it), but an Instagram handle whose bio needs the residential page is charged instagram-bio-fetch.
🧩 Want other GTM data?
Mamba Labs builds a fleet of GTM enrichment actors that share one flat,
Clay-ready output convention, so their rows join on company_domain or
creator_id with no cleaning step:
Every actor in the suite takes a domain, a company, or a creator and returns one flat row, so they stack in the same Clay table without reshaping anything.
🛠️ Need something custom built for you or your team? Tell us what you are trying to find and we will build it. Talk to Mamba Labs.
🆘 Support
Issues, field requests and bug reports: open an issue on the actor's Issues tab. Mamba Labs reads every one.
ℹ️ Sourcing and legal. Every field comes from pages the platforms and the creators publish to anyone without a login, read directly or through a proxy, with no login, no CAPTCHA solving, and no LinkedIn fetch. The row records what was public at
read_at. Nothing is assessed or scored; a class, a flag, or a match method says what was read and where. You are responsible for how you use the output, including under applicable data protection and platform terms.
Built by Mamba Labs.