Page Finder and Extractor Company Website Pages for Clay avatar

Page Finder and Extractor Company Website Pages for Clay

Pricing

from $3.40 / 1,000 page type locateds

Go to Apify Store
Page Finder and Extractor Company Website Pages for Clay

Page Finder and Extractor Company Website Pages for Clay

Give it a company domain and name the page you want. Finds pricing, careers, investor relations, trust centers, terms, contact and 40 more page types on that company's own site, in 11 languages, and returns the URL, how it was found and a confidence for that method. Structured fields on request.

Pricing

from $3.40 / 1,000 page type locateds

Rating

0.0

(0)

Developer

Mamba Labs

Mamba Labs

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Share

🧭 What can Page Finder and Extractor do?

Give it a company domain or a company name, and name the page you want. It finds that page on the company's own website, tells you how it found it and how confident it is in that method, and when you ask, reads the page and returns structured fields. One flat Clay ready row per input, every time.

📦 What you get⚙️ Features and integrations
📍 The page URL, for 46 named page types
🧪 The method that found it, plus a confidence for that method
🗒️ The evidence phrase, the anchor text or heading that identified it
🧾 Structured fields off the page, when you ask for them
🌍 Discovery vocabulary in 11 languages, not English with translations bolted on
🔗 Reads the link graph, not a list of guessed paths
🪪 Company name input, with identity resolution first
🧊 14 day cache, and export to JSON, CSV, Excel, HTML or XML

Built for GTM teams who hold a domain and need the answer that lives on one specific page of that company's site: the pricing page, the trust center, the careers page, the investor relations section, the terms page.

🚫 This is not a web crawler and it is not a scraper of everything. It does not dump a site. You name a page type, it finds that page type, and it tells you honestly when it could not.

🎯 Why use Page Finder and Extractor?

You want toRead this field
The page URL itself{type}_url, for example pricing_url
To know whether an empty answer means anything{type}_found, coverage, fetch_status
To trust only strong matches{type}_confidence, {type}_method
To prove the match before you use it{type}_evidence_phrase
The structured data on the pageThe findings dataset, one record per field
To know what got filtered out and whyfindings_rejected_count, rejected_by_filter

Finding a company's pricing page by trying /pricing works on American software companies and falls apart everywhere else. We measured it: across 2,207 European company websites, a nine path guess method failed to reach the investor section on 516 companies, and found the section but not the page below it on another 365. That is 881 companies, 39.9 percent, from link discovery alone. Every other failure put together, blocked plus JavaScript only plus dead domain plus robots refusals plus every odd status code, came to 312.

So this actor reads the homepage and footer link graph, follows sitemap.xml and its shards, matches anchor and heading vocabulary in eleven languages, follows one hop into a located section to find the page below it, and recognizes known third party hosts. Path guessing is the last thing it tries, and it is scored as the weak method it is.

📏 How often does it actually find the page

60.93 percent, for investor_relations, across 2,207 European listed company domains. That is the only page type measured at that scale, and the number is per page type, not a claim about the product as a whole.

CorpusEvery company on the fourteen main European venues that carries a domain, 2,186 distinct after deduplication
Page typeinvestor_relations only
Located1,332, 60.93%
Read and graded a real false456, 20.9%
Undetermined, returned null398, 18.2%

The 18.2 percent that came back null is the ceiling on this population, and it is not a bug. By status: 175 companies refused us outright, 98 were unreachable, 34 served nothing readable even to a browser, 14 disallowed us in robots.txt and 14 had dead domains. The remaining 63 are rows where the site read fine but the look did not complete: on 31 the request budget ran out, and on 32 the page the evidence pointed at refused us, which is a null and not a false however much of the site we read. A row that could not be read returns null, never false.

Other page types are measured on smaller corpora and the figures are in the limits section below. A rate for pricing on European industrials would be meaningless, because most of them do not have one.

🈳 Eleven languages, by default, not as an option

Of the rows the prior crawl found at all, English carried 76.7 percent. The other 23.3 percent was concentrated in Milan, Paris and Xetra, which are three of the largest venues in the population. Without German, French and Italian vocabulary most of that yield simply disappears.

Vocabulary ships in English, German, French, Italian, Spanish, Dutch, Swedish, Norwegian, Danish, Finnish and Portuguese, for all 46 page types. languageHints reorders that list. It never shortens it.

📊 What data can Page Finder and Extractor extract?

46 page types, in nine groups:

GroupPage types
Money and buyingpricing, demo_request, free_trial, procurement_vendor
Company factsabout, leadership_team, locations, investor_relations, annual_report, governance
Trust and risksecurity_trust_center, compliance_certifications, privacy_policy, terms_of_service, dpa_subprocessors, accessibility_statement, status_page
Hiringcareers, job_board, benefits, culture
Product and technicaldocumentation, api_reference, integrations, changelog, roadmap, developer_portal
Marketing and contentblog, press_newsroom, case_studies, customers_logos, resources_library, events_webinars, podcast, media_kit
Ecosystempartners, reseller_channel, affiliate_program, marketplace_listing, community
Reachcontact, support_help_center, login_app
ESGsustainability_esg, diversity_programs, giving_volunteering

In locate_and_extract mode you also get two groups of fields.

Available on any page type: copyright line and year, legal entity name, registration number, VAT number, emails, phone numbers, postal addresses, social links, meta title and description, page language, last modified date, canonical URL, hreflang set, schema.org types, forms and their fields, primary calls to action, and tech markers in the page source (analytics, chat widget, marketing automation, cookie vendor).

Bound to a specific page type, for seventeen of them:

Page typeWhat you get
pricingPlan names, prices, currency, billing period, seat versus usage, free tier, trial length, annual discount, enterprise contact, feature gates
investor_relationsFiling rows with report type, period label and publication date. Then the derived fields: fiscal year end, reporting cadence, publication lag. Plus ticker, ISIN, auditor, investor contact
security_trust_centerCertifications claimed, subprocessor list, data residency, pen test cadence, bug bounty
careersATS host and slug, open role count, titles, departments, locations, remote policy, salary bands where posted
aboutFounding year, stated headcount, named leadership, mission statement
press_newsroomLatest release date, publishing cadence, media contact, funding mentions
partnersNamed partners, tier structure, program present, integration count
customers_logosNamed logos, case study count, industries served
integrationsNamed integrations, categories, count
terms_of_serviceGoverning law, jurisdiction, arbitration clause, contracting entity
privacy_policyDPO contact, lawful basis stated, retention, residency
status_pageUptime figure, incident count, hosting provider
documentation, api_referenceAPI present, auth method, SDK languages, rate limits, versioning
changelogRelease cadence, last release date
sustainability_esgReports published, targets stated, frameworks named
contactContact block fields, support channels, stated response times

Page types without a bound field map still locate, and still return the whole page agnostic group.

🧮 The fiscal year end, and why it is hard

A stated fiscal year end appears on an investor relations landing page for 45 companies out of 2,207, which is 2.0 percent. It is almost never written down. So the actor derives it, by four routes, and tells you which one it used in year_end_source:

  1. explicit_ending_phrase, confidence 0.95. The page says financial year ended 31 December, esercizio chiuso al, exercice clos le. Legal phrasings, which is why Italy and France supply most of the ones that exist.
  2. stated_period_end, 0.88 for an annual report and 0.82 for an interim. The page says Interim Report First Half (as at 30 June 2026). The report type fixes the period length, the date fixes its end, and the fiscal year start follows. For a quarterly report the quarter number has to be readable, because Q2 as at 30 June is not the first quarter; where it is not readable this route declines rather than guessing.
  3. implicit_period_label, confidence up to 0.95. A label reading Interim Report, 1 January to 30 September 2026 fixes the year start directly. Several labels agreeing raise the confidence. This is the cleanest route and the rarest: it fires on 4 companies in 2,186.
  4. derived_from_report_type_and_publication, capped at 0.72. A calendar that states none of the above, only Nine-month results, 29 October 2026. The report type gives the period length and the publication date bounds when it ended, so every row on the calendar votes. This is the one to filter out.

Filter on year_end_source if you only want what a company wrote down. Routes 1 and 2 are what the company wrote down; 3 is arithmetic on a label; 4 is an inference from the shape of a calendar.

🎯 How accurate is it, and what does that number depend on

The actor filters nothing. It returns every result with a confidence, and you choose the threshold. That is a deliberate choice: the right threshold depends on what you are doing with the answer, and hiding rows to make a headline look better would be lying by omission.

So here is the accuracy at each threshold, measured against the SEC, which publishes the true fiscal year end for every US filer. 51 US companies where this actor derived one and EDGAR carries ground truth:

What you keepnAgreement with the SEC
confidence >= 0.8, the threshold this README recommends4197.6% within 7 days, 90.2% exact
Everything, every confidence5190.2% within 7 days, 84.3% exact
Only confidence < 0.81060.0%

Read all three rows. The 97.6 percent is conditioned on the filter, and the filter throws away a fifth of the results. If you keep everything you get more coverage at 90.2 percent. Neither number is the "real" one; they are the same data at two settings.

Accuracy differs sharply by which route derived the answer, which is exactly what year_end_source is for:

Routeyear_end_sourceConfidenceAgreementn
The company stated its year endexplicit_ending_phrase0.95100%36
The company stated a period endstated_period_end0.82 to 0.8880%5
Inferred from the shape of the calendarderived_from_report_type_and_publicationup to 0.7250%10

The inferred route sits below 0.8 on purpose, so the recommended threshold excludes it. It is weaker on US companies than European ones, because a US investor page tends to list earnings call dates rather than period-labelled filings. Treat it as a hint, not a fact.

The "within 7 days" column exists because a 52 or 53 week fiscal year genuinely ends on a fixed weekday near a month end: Papa John's declares 31 December to the SEC and closes on 27 December. Both numbers are published rather than the flattering one.

🔍 How to find a page on a company's website

Three ways in, and they cost different amounts.

You haveWhat you setWhat happensWhat it costs
A bare domaindomain or domainsDiscovery runs on that site. This is the normal path.A locate event per page type.
A company namecompanies, as {"name": "..."}Identity resolution runs first to turn the name into a domain, then discovery runs.A identity-resolved event at $0.007 on top, once per company, plus the locate events.
A name and a URL you already holdcompanies with url, or knownUrls keyed by page typeThe URL is taken as given. Discovery is skipped for that page type, so it is faster and exact.No identity event, and the located page bills at the same locate rate with method: known_url and confidence 1.0.

Supplying a domain, or a name alongside a URL, never triggers identity resolution and is never charged for it.

Two modes. locate returns the URL, the method, the confidence and the evidence phrase, and nothing else. locate_and_extract does all of that and then reads each located page, returning the structured fields into the findings dataset. Extraction adds a page-extracted event per page read.

  1. Open the Input tab and put a domain in domain, for example stripe.com.
  2. Pick your page types in pageTypes. Each one you add costs a locate event and takes more requests.
  3. Leave mode on locate for URLs only, or set locate_and_extract to read the pages.
  4. Run it. A cold locate run on one page type takes about 10 to 20 seconds per company.
  5. Read pricing_url for the answer and pricing_confidence for how much to trust it.

🧬 Using it in Clay

Add it as an enrichment on a company table and map domain to your domain column. Every input returns exactly one row, including the ones where nothing is found, so your table never loses accounts. Map {type}_url and {type}_confidence into columns and filter on {type}_confidence >= 0.8.

If you hold the URL already, pass it in knownUrls keyed by page type. Discovery is skipped for that type, which is faster and exact.

🎚️ Filtering for outreach quality

{type}_method tells you what actually found the page, and precision differs sharply between methods. known_url and known_host are near certain. homepage_anchor and footer_anchor mean the company's own navigation said what the page is. sitemap means a URL slug matched and nothing on the page confirmed it. path_guess means we guessed and the page happened to confirm it. Threshold at 0.8 and above for anything a customer will see.

💵 How much does it cost?

Pay per event, so you pay for what you asked for.

EventWhen it is charged
page-type-locatedPer page type looked for, per company, when the look actually happened
page-extractedPer page read, in locate_and_extract mode only
identity-resolvedOnly on the company name path, never when you supply a domain
apify-actor-startOnce per run, per gigabyte of memory

Every event at every Apify tier, in dollars:

EventFREEBRONZESILVERGOLDPLATINUMDIAMOND
page-type-located0.0040.00380.00360.00340.00340.0034
page-extracted0.0030.002850.00270.002550.002550.00255
identity-resolved0.0070.006650.00630.005950.005950.00595
apify-actor-start0.000050.000050.000050.000050.000050.00005

The three work events get cheaper as your Apify plan gets larger, down 15 percent at GOLD and flat from there. The actor start fee is set by Apify, not by us, and does not tier.

🧠 This actor runs a browser, so it is billed at 4 GB and the actor start fee is four times the fleet's. Apify charges the start event once per gigabyte of memory, and this actor's default is 4096 MB where most Mamba Labs actors sit at 256 or 512. So a run costs 4 x $0.00005 = $0.0002 to start, before any page is looked at.

What that means in practice:

How you call itStart feeWorkStart as a share
One company per run, the Clay shape$0.0002$0.0044.8%
100 companies in one run$0.0002$0.4000.05%
1,000 companies in one run$0.0002$4.0000.005%

If you are calling it one row at a time from Clay, batch where you can. The memory is not padding: rendering a JavaScript page needs a real browser, and dropping to 512 MB to save $0.00015 per run would trade out of memory crashes for a rounding error.

💳 A look that happens is billed, including the ones that come back empty. The actor reads robots.txt, the homepage, the link graph and the sitemap whether or not the page is there, and a documented false is a real answer that cost real work. A look that does not happen is not billed. If the domain is dead, refuses us, or is disallowed by robots.txt, nothing was looked at, the row comes back with found: null, and no locate event is charged for it.

Why that line and not "hits only". A readable site that returns nothing is the MOST expensive outcome this actor produces, not the cheapest. Measured across 2,207 European domains: a no_paths_found row takes a median of 64 seconds against 30 for a hit, because it is the only outcome that exhausts every method, the homepage link graph, the footer, the sitemap and its shards, a hop into any located section, and finally every path guess. A hit stops as soon as it finds something. Billing hits only would give away the dearest work and keep the cheapest, which is backwards. The rows that genuinely cost nothing, a dead domain at 1.4 seconds or a refusal at 8, are the ones that are free.

🪪 You are never charged for identity resolution unless you use the company name path. Supplying a domain, or supplying a URL alongside a name, skips it entirely.

⌨️ Input

Set everything on the Input tab. Nothing is required, and a run with no usable input still returns a row saying so.

FieldDefaultWhat it does
domainnoneOne company domain. Protocol and path are stripped.
domainsnoneA list of company domains.
companiesnoneCompanies by name, as objects with name and optionally country, isin, ticker, url. Identity resolution runs first.
pageTypes["pricing"]Which of the 46 page types to find.
modelocatelocate_and_extract also reads the pages.
extractionFieldsall 13The page agnostic menu.
extractPageTypeFieldstrueAlso run the field map bound to each page type.
knownUrlsnoneURLs you already hold, keyed by page type. Skips discovery for those.
maxPagesPerType4Candidate pages to open per page type. Lowering it is faster and finds less.
maxRequestsPerInput60Hard ceiling on requests to one company's site. Hitting it returns coverage: partial, never a false negative.
concurrency10How many companies to work on at once. Per company the actor is still strictly one request at a time.
allowRendertrueOpen a browser for pages that serve no readable HTML.
languageHintsnoneReorders the vocabulary. Never shortens it.
skipCachefalsetrue forces a fresh crawl instead of the 14 day cache.

📤 Output

Two datasets. The default dataset holds one flat row per input. The findings dataset holds one record per extracted field, keyed back by input_key, because eleven filing rows do not fit in one cell.

The per page type columns exist only for the page types you asked for. Ask for pricing and you get six pricing columns, not 270 empty ones.

This is a real row, copied unedited from a run of the shipped build on stripe.com asking for two page types in locate mode:

{
"input_key": "stripe.com",
"input_path": "domain",
"domain": "stripe.com",
"company_name": null,
"mode": "locate",
"page_types_requested": "pricing,security_trust_center",
"page_types_requested_count": 2,
"page_types_found": 2,
"page_types_not_found": 0,
"page_types_undetermined": 0,
"best_confidence": 0.93,
"identity_resolved": null,
"identity_method": null,
"identity_confidence": null,
"findings_count": 0,
"findings_agnostic_count": 0,
"findings_bound_count": 0,
"findings_rejected_count": 0,
"rejected_by_filter": null,
"coverage": "partial",
"fetch_status": "ok",
"render_mode": "http",
"pages_reached": 3,
"requests_made": 6,
"renders_made": 0,
"destination_mismatches": 0,
"render_failures": 0,
"render_broken": false,
"robots_crawl_delay_ms": 1000,
"robots_disallowed_paths": 0,
"blocked_pages": 0,
"javascript_only_pages": 0,
"homepage_links_seen": 92,
"sitemap_urls_seen": 0,
"fetch_error": null,
"cache_hit": false,
"run_date": "2026-08-16T08:03:49.406Z",
"pricing_url": "https://stripe.com/pricing",
"pricing_found": true,
"pricing_method": "homepage_anchor",
"pricing_confidence": 0.93,
"pricing_evidence_phrase": "pricing",
"pricing_pages_reached": 1,
"security_trust_center_url": "https://docs.stripe.com/security",
"security_trust_center_found": true,
"security_trust_center_method": "path_guess",
"security_trust_center_confidence": 0.6,
"security_trust_center_evidence_phrase": "Security at Stripe",
"security_trust_center_pages_reached": 1
}

Two things in that row are worth reading, because they are the row telling on itself. coverage is partial, not good, because sitemap_urls_seen is 0: the sitemap was not readable on this run, so one discovery method never ran. And the trust center came back at 0.6 on path_guess, meaning we guessed the path and the page confirmed it, against 0.93 for the pricing page which Stripe's own navigation named. Both are found, and they are not equally trustworthy. That is the point of shipping the method and the confidence next to the URL.

The findings dataset. In locate_and_extract mode the same domain produced 22 findings records off the pricing page. One record per field, so eleven filing rows or six pricing plans do not have to be crushed into one cell:

{
"input_key": "stripe.com",
"domain": "stripe.com",
"page_type": "pricing",
"source_url": "https://stripe.com/pricing",
"group": "bound",
"field": "plan_name",
"value": "Payments",
"value_type": "string",
"method": "plan_card",
"confidence": 0.85,
"detail_json": "{\"plan_index\":1}",
"is_personal_data": false,
"lawful_basis": null,
"page_status": "ok",
"run_date": "2026-08-16T08:04:19.535Z"
}

group is agnostic for the menu that runs on any page and bound for the fields tied to that page type. Join back to the flat row on input_key.

⚠️ false and null mean different things and are never collapsed. pricing_found: false means the homepage was read, its link graph was searched, the sitemap was searched, and the page is not reachable. pricing_found: null means not enough was readable to say. A blocked or JavaScript only site returns null, never false.

Three fields to read before you trust a false:

FieldWhat a good answer looks like
coveragegood means the homepage, its link graph and the sitemap were all read. partial means one was missing or the request budget ran out. poor means the homepage itself was not read, so the actual method never ran. none means nothing was readable and every found on the row is null.
pages_reachedHow many pages returned real readable content. A false on a row with pages_reached: 0 is not a negative, it is a row where nothing was seen.
fetch_statusok, partial and no_paths_found are real answers. Everything else is a failure to look, and none of the failures is billed. Full list below.

A false with coverage: good and pages_reached above zero is a negative you can act on. Anything else is an admission, not a finding.

Every fetch_status value, and what each one means for the row:

fetch_statusWhat happenedfound valuesBilled?
okThe site was read and at least one requested page type was located.Real true and false
partialThe site was read but the per company request ceiling ran out before every method finished. Treat a false here as unproven.true, or false you should not trust
no_paths_foundEverything was readable, every method ran to the end, and the page is not there. This is a real answer, not a failure.false
unreachableDNS resolved but the site did not answer: connection refused, TLS failure, or timeout.All null
blockedThe site refused us with 401, 403, 407, 429 or 451, or served a CAPTCHA interstitial. We do not retry and we do not work around it.All null
javascript_onlyThe page still had no readable text after a real browser rendered it. Usually a client side app behind a shell.All null
robots_disallowedThe company's own robots.txt disallows our user agent on the paths we would need. We stop.All null
domain_deadThe domain does not resolve at all.All null

The homepage decides the row. If the homepage comes back blocked, domain_dead, robots_disallowed, javascript_only or unreachable, that status is the row's status, because every discovery method depends on reading the homepage first.

💡 Tips

  • Ask for the page types you will actually use. Each one costs an event and adds requests.
  • Pass knownUrls for anything you already hold. It is faster, exact, and skips discovery for that type.
  • Filter on {type}_confidence >= 0.8 before putting a URL in front of a customer.
  • coverage: good with found: false is a real, usable negative. Anything else means the look was incomplete.
  • Re-run with skipCache: "true" after a site redesign.
  • For company names rather than domains, put the name in companies. Stockholm and Helsinki are domain starved, and the name path is the only way to reach most of those companies.

⚠️ Known limits

  • Some sites refuse us, and we do not argue. 401, 403, 429 and CAPTCHA interstitials come back as fetch_status: "blocked", unretried, with every field null. In the 2,186 rows of the shipped build's run that was 175 companies, the single largest reason for a null. Re-running does not help, because we are not trying to get around it.

  • robots.txt is honored, including Crawl-delay. Fourteen companies in that population disallowed the paths we needed and came back robots_disallowed. That is the correct outcome, not a bug.

  • A browser renders, it does not break in. Rendering is used for pages a company serves publicly that need JavaScript to read. It is never used against a block, a CAPTCHA, a login, or robots.txt. A page that is still unreadable after rendering reports javascript_only, which now means "a browser looked and there was nothing there".

  • security_trust_center is restricted on purpose, and it returns fewer pages than it could. For this page type only, an anchor that says "Security" is not enough: the URL has to agree. The last part of the path must itself be a trust token and the page must not sit under a product or content section, so acme.com/trust is accepted and acme.com/industries/security is not. A dedicated trust. subdomain or a known trust vendor host skips the check, because the host is the evidence.

    Why: at a security vendor the word Security in the navigation names a product. Before the restriction this page type measured 55.0 percent correct across two hand read samples, n=60. After it, 83.3 percent, n=30. The cost is coverage: it now finds a trust center on 25.0 percent of a US software corpus where it previously claimed 39.0 percent, so about a third of the old results are gone and most of them were wrong.

    Precision by method on the restricted build, n=30:

    security_trust_center_methodCorrectn
    footer_anchor10 of 1010
    known_subdomain, a trust. or security. host6 of 66
    path_guess3 of 44
    sitemap2 of 33
    homepage_anchor4 of 77

    homepage_anchor is still the weak one. If you want near certainty, keep known_subdomain and footer_anchor and drop the rest.

  • Precision by page type, hand read, every figure with its sample size: investor_relations 96.7% (n=30), careers 90.0% (n=30), contact 90.0% (n=30), about 83.3% (n=30), security_trust_center 83.3% (n=30, restricted build), pricing 80.0% (n=30). The other 40 page types have not been measured at this scale and no figure is published for them.

  • A page type with no bound field map returns only the page agnostic fields. That is deliberate. Seventeen page types have bound extractors; the other 29 locate and return the general group.

  • Landing pages are indexes. For investor_relations the actor walks one hop into the section to reach the calendar or reports page, because that is where filing tables live. It reads up to three such pages.

  • Derived fields are marked as derived. year_end_source tells you whether a fiscal year end was stated by the company or inferred from the shape of its calendar. The weakest inferred route, reading a fiscal year end off the shape of a filing calendar, is capped at 0.72 confidence and is easy to filter out. It measured 25 correct of 34 hand read, 73.5 percent, which is why the cap sits below the 0.8 threshold this README tells you to use.

  • One request at a time per domain, with a delay. This is polite by design and it sets the floor on how fast a large batch can run.

  • No proxy. The actor runs proxyless, so rate limiting surfaces in fetch_status rather than being hidden.

❓ FAQ

How is this different from just guessing /pricing? Guessing is one of seven methods here and it is the last one tried. On the 2,207 company European population, guessing paths accounted for 881 failures on its own. The link graph, the sitemap, the multilingual vocabulary and the one hop walk are what find the other pages.

Why did my company come back empty? Read coverage and fetch_status first. coverage: good with fetch_status: no_paths_found means a complete look found nothing, which is a real answer. blocked, javascript_only, robots_disallowed or domain_dead mean the look did not complete, and those rows are not billed a locate event.

Can I trust a path_guess result? Only when the page itself confirmed it, which is the only case where one is returned at all, and it still scores 0.6. Threshold at 0.8 to exclude them.

Does it work outside Europe? Yes. The vocabulary is European because that is where the measurement was done, and English covers the US and UK. Nothing about the method is region specific.

Why two datasets? Because a pricing page has four plans and an investor relations page has eleven filing rows, and neither fits in a flat cell. The flat row answers "where is it and how sure are you". The findings dataset answers "what was on it".

Can I get more page types? Adding a page type is vocabulary in eleven languages plus an optional field map. Open an issue on the Issues tab with the page type you want and what you would read off it.

🧩 Want other GTM data?

Mamba Labs builds custom actors for B2B go-to-market teams. The public versions of that work live here on the Store, so our users get the same tooling we build under contract.

🧑‍💼 GTM Hiring Signal Scraper🧱 Tech Stack Detector
📡 B2B Buying Signals Aggregator🔑 Job Board Keyword Scanner
🔗 Domain to LinkedIn URL Resolver🎯 ICP Fit Scorer
📋 Job Posting Monitor📬 Domain Deliverability Checker
🏢 Company Firmographic Enricher🌐 Company Social Presence Mapper
🪪 Company Identity Resolver💰 Funding and Press Signal Scanner
🔄 Company Change-Event Feed👤 People Finder and Email Verifier
🚀 Prospect Engine🤖 AI Tooling Detector
📮 Outbound Stack Detector📝 Publishing Frequency Tracker
✉️ Work Email Waterfall FinderSequencer Lead Push
🏅 Workplace Program Detector👥 Team Page People Extractor
🧭 Company Discovery List Builder🎪 Event Presence Index
⚖️ Legal Entity Resolver🏷️ Contact Classifier
💬 LinkedIn Post Tracker and Comment Capture🏛️ Government Contract Award Monitor
🕵️ Agent Accessibility Auditor

Every actor in the suite takes a domain or a company and returns one flat row, so they stack in the same Clay table without reshaping anything.

🛠️ Need something custom built for you or your team? Tell us what you are trying to find and we will build it. Talk to Mamba Labs.

🛟 Support

Issues, feature requests and data questions go through the Issues tab on this actor. Include the domain, the page type and the run ID and it can be traced.

👤 Personal data, stated plainly

This actor can return a named person, so read this section rather than assuming.

Two page types can return one: contact returns the contact block a company publishes on its own contact page, and about returns leadership names published on the company's own about page. That is the whole of it. Every other page type returns company information only.

  • It only ever returns what the company published on its own website. It does not search for people, does not look anyone up elsewhere, does not enrich, does not guess an email pattern, and does not join a name to any other source.
  • It infers no personal attribute. No seniority scoring, no gender, no location inference, no profile building. A name, a job title and a business contact method, exactly as the company printed them.
  • Every such field is flagged. Records carrying a person have is_personal_data: true and a lawful_basis in the findings dataset, so you can filter the entire class out with one predicate, or exclude it up front by leaving contact and about out of pageTypes.
  • Lawful basis. Legitimate interest in publicly published business contact information, under GDPR Article 6(1)(f). These are role holders acting in a business capacity, published by their employer for the purpose of being contacted.
  • Deletion. Open an issue on the Issues tab naming the page and the person, and the record is removed from our cache and added to a suppression list. We cannot remove it from the company's own website, which is where it is published.

If you want a roster of people at a company, this is the wrong actor. Use Team Page People Extractor to pull a team page, and Contact Classifier to classify contacts you already hold. This actor returns a contact block that happened to sit on a page it was asked to find, and nothing more.

Sourcing and legal. This actor reads publicly available pages on the company's own website. It reads and honors robots.txt per host, including Crawl-delay, makes one request at a time per domain with a delay, and sends a descriptive user agent that names it and links here. It does not impersonate a browser, does not sign headers, and does not retry to get around a block. Where a page is publicly served but needs JavaScript to read, a browser renders it; a browser is never used against an access control of any kind. Refusals are reported in the output rather than hidden.

Built by Mamba Labs.