# Page Finder and Extractor Company Website Pages for Clay (`mambalabs/page-finder-extractor`) Actor

Give it a company domain and name the page you want. Finds pricing, careers, investor relations, trust centers, terms, contact and 40 more page types on that company's own site, in 11 languages, and returns the URL, how it was found and a confidence for that method. Structured fields on request.

- **URL**: https://apify.com/mambalabs/page-finder-extractor.md
- **Developed by:** [Mamba Labs](https://apify.com/mambalabs) (community)
- **Categories:** Lead generation, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.40 / 1,000 page type locateds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

### 🧭 What can Page Finder and Extractor do?

Give it a **company domain** or a **company name**, and name the page you want. It finds that page on the company's own website, tells you **how** it found it and **how confident** it is in that method, and when you ask, reads the page and returns structured fields. One flat Clay ready row per input, every time.

| 📦 What you get | ⚙️ Features and integrations |
|---|---|
| 📍 **The page URL**, for 46 named page types<br>🧪 **The method that found it**, plus a confidence for that method<br>🗒️ **The evidence phrase**, the anchor text or heading that identified it<br>🧾 **Structured fields off the page**, when you ask for them | 🌍 **Discovery vocabulary in 11 languages**, not English with translations bolted on<br>🔗 **Reads the link graph**, not a list of guessed paths<br>🪪 **Company name input**, with identity resolution first<br>🧊 **14 day cache**, and export to JSON, CSV, Excel, HTML or XML |

Built for GTM teams who hold a domain and need the answer that lives on one specific page of that company's site: the pricing page, the trust center, the careers page, the investor relations section, the terms page.

> 🚫 **This is not a web crawler and it is not a scraper of everything.** It does not dump a site. You name a page type, it finds that page type, and it tells you honestly when it could not.

### 🎯 Why use Page Finder and Extractor?

| You want to | Read this field |
|---|---|
| The page URL itself | `{type}_url`, for example `pricing_url` |
| To know whether an empty answer means anything | `{type}_found`, `coverage`, `fetch_status` |
| To trust only strong matches | `{type}_confidence`, `{type}_method` |
| To prove the match before you use it | `{type}_evidence_phrase` |
| The structured data on the page | The findings dataset, one record per field |
| To know what got filtered out and why | `findings_rejected_count`, `rejected_by_filter` |

#### 🔗 It reads the link graph. That is the whole product.

Finding a company's pricing page by trying `/pricing` works on American software companies and falls apart everywhere else. We measured it: across **2,207 European company websites**, a nine path guess method failed to reach the investor section on **516 companies**, and found the section but not the page below it on another **365**. That is **881 companies, 39.9 percent**, from link discovery alone. Every other failure put together, blocked plus JavaScript only plus dead domain plus robots refusals plus every odd status code, came to **312**.

So this actor reads the homepage and footer link graph, follows `sitemap.xml` and its shards, matches anchor and heading vocabulary in eleven languages, follows one hop into a located section to find the page below it, and recognizes known third party hosts. Path guessing is the **last** thing it tries, and it is scored as the weak method it is.

#### 📏 How often does it actually find the page

**60.93 percent, for `investor_relations`, across 2,207 European listed company domains.** That is the only page type measured at that scale, and the number is per page type, not a claim about the product as a whole.

| | |
|---|---|
| Corpus | Every company on the fourteen main European venues that carries a domain, 2,186 distinct after deduplication |
| Page type | `investor_relations` only |
| Located | 1,332, **60.93%** |
| Read and graded a real `false` | 456, 20.9% |
| Undetermined, returned `null` | 398, 18.2% |

The 18.2 percent that came back `null` is the ceiling on this population, and it is not a bug. By status: **175 companies refused us outright**, 98 were unreachable, 34 served nothing readable even to a browser, 14 disallowed us in `robots.txt` and 14 had dead domains. The remaining 63 are rows where the site read fine but the look did not complete: on 31 the request budget ran out, and on 32 the page the evidence pointed at refused us, which is a `null` and not a `false` however much of the site we read. A row that could not be read returns `null`, never `false`.

**Other page types are measured on smaller corpora and the figures are in the limits section below.** A rate for `pricing` on European industrials would be meaningless, because most of them do not have one.

#### 🈳 Eleven languages, by default, not as an option

Of the rows the prior crawl found at all, English carried 76.7 percent. The other 23.3 percent was concentrated in Milan, Paris and Xetra, which are three of the largest venues in the population. Without German, French and Italian vocabulary most of that yield simply disappears.

Vocabulary ships in **English, German, French, Italian, Spanish, Dutch, Swedish, Norwegian, Danish, Finnish and Portuguese**, for all 46 page types. `languageHints` reorders that list. It never shortens it.

### 📊 What data can Page Finder and Extractor extract?

**46 page types**, in nine groups:

| Group | Page types |
|---|---|
| Money and buying | `pricing`, `demo_request`, `free_trial`, `procurement_vendor` |
| Company facts | `about`, `leadership_team`, `locations`, `investor_relations`, `annual_report`, `governance` |
| Trust and risk | `security_trust_center`, `compliance_certifications`, `privacy_policy`, `terms_of_service`, `dpa_subprocessors`, `accessibility_statement`, `status_page` |
| Hiring | `careers`, `job_board`, `benefits`, `culture` |
| Product and technical | `documentation`, `api_reference`, `integrations`, `changelog`, `roadmap`, `developer_portal` |
| Marketing and content | `blog`, `press_newsroom`, `case_studies`, `customers_logos`, `resources_library`, `events_webinars`, `podcast`, `media_kit` |
| Ecosystem | `partners`, `reseller_channel`, `affiliate_program`, `marketplace_listing`, `community` |
| Reach | `contact`, `support_help_center`, `login_app` |
| ESG | `sustainability_esg`, `diversity_programs`, `giving_volunteering` |

In `locate_and_extract` mode you also get two groups of fields.

**Available on any page type:** copyright line and year, legal entity name, registration number, VAT number, emails, phone numbers, postal addresses, social links, meta title and description, page language, last modified date, canonical URL, hreflang set, schema.org types, forms and their fields, primary calls to action, and tech markers in the page source (analytics, chat widget, marketing automation, cookie vendor).

**Bound to a specific page type**, for seventeen of them:

| Page type | What you get |
|---|---|
| `pricing` | Plan names, prices, currency, billing period, seat versus usage, free tier, trial length, annual discount, enterprise contact, feature gates |
| `investor_relations` | Filing rows with report type, period label and publication date. Then the derived fields: **fiscal year end**, reporting cadence, publication lag. Plus ticker, ISIN, auditor, investor contact |
| `security_trust_center` | Certifications claimed, subprocessor list, data residency, pen test cadence, bug bounty |
| `careers` | ATS host and slug, open role count, titles, departments, locations, remote policy, salary bands where posted |
| `about` | Founding year, stated headcount, named leadership, mission statement |
| `press_newsroom` | Latest release date, publishing cadence, media contact, funding mentions |
| `partners` | Named partners, tier structure, program present, integration count |
| `customers_logos` | Named logos, case study count, industries served |
| `integrations` | Named integrations, categories, count |
| `terms_of_service` | Governing law, jurisdiction, arbitration clause, contracting entity |
| `privacy_policy` | DPO contact, lawful basis stated, retention, residency |
| `status_page` | Uptime figure, incident count, hosting provider |
| `documentation`, `api_reference` | API present, auth method, SDK languages, rate limits, versioning |
| `changelog` | Release cadence, last release date |
| `sustainability_esg` | Reports published, targets stated, frameworks named |
| `contact` | Contact block fields, support channels, stated response times |

Page types without a bound field map still locate, and still return the whole page agnostic group.

#### 🧮 The fiscal year end, and why it is hard

A stated fiscal year end appears on an investor relations landing page for **45 companies out of 2,207**, which is 2.0 percent. It is almost never written down. So the actor derives it, by four routes, and tells you which one it used in `year_end_source`:

1. **`explicit_ending_phrase`**, confidence 0.95. The page says `financial year ended 31 December`, `esercizio chiuso al`, `exercice clos le`. Legal phrasings, which is why Italy and France supply most of the ones that exist.
2. **`stated_period_end`**, 0.88 for an annual report and 0.82 for an interim. The page says `Interim Report First Half (as at 30 June 2026)`. The report type fixes the period length, the date fixes its end, and the fiscal year start follows. For a quarterly report the quarter number has to be readable, because `Q2 as at 30 June` is not the first quarter; where it is not readable this route declines rather than guessing.
3. **`implicit_period_label`**, confidence up to 0.95. A label reading `Interim Report, 1 January to 30 September 2026` fixes the year start directly. Several labels agreeing raise the confidence. This is the cleanest route and the rarest: it fires on 4 companies in 2,186.
4. **`derived_from_report_type_and_publication`**, capped at 0.72. A calendar that states none of the above, only `Nine-month results, 29 October 2026`. The report type gives the period length and the publication date bounds when it ended, so every row on the calendar votes. This is the one to filter out.

Filter on `year_end_source` if you only want what a company wrote down. Routes 1 and 2 are what the company wrote down; 3 is arithmetic on a label; 4 is an inference from the shape of a calendar.

#### 🎯 How accurate is it, and what does that number depend on

**The actor filters nothing. It returns every result with a confidence, and you choose the threshold.** That is a deliberate choice: the right threshold depends on what you are doing with the answer, and hiding rows to make a headline look better would be lying by omission.

So here is the accuracy at each threshold, measured against the SEC, which publishes the true fiscal year end for every US filer. 51 US companies where this actor derived one and EDGAR carries ground truth:

| What you keep | n | Agreement with the SEC |
|---|---:|---:|
| **`confidence >= 0.8`**, the threshold this README recommends | 41 | **97.6%** within 7 days, 90.2% exact |
| Everything, every confidence | 51 | 90.2% within 7 days, 84.3% exact |
| Only `confidence < 0.8` | 10 | 60.0% |

**Read all three rows.** The 97.6 percent is conditioned on the filter, and the filter throws away a fifth of the results. If you keep everything you get more coverage at 90.2 percent. Neither number is the "real" one; they are the same data at two settings.

Accuracy differs sharply by which route derived the answer, which is exactly what `year_end_source` is for:

| Route | `year_end_source` | Confidence | Agreement | n |
|---|---|---:|---:|---:|
| The company stated its year end | `explicit_ending_phrase` | 0.95 | **100%** | 36 |
| The company stated a period end | `stated_period_end` | 0.82 to 0.88 | 80% | 5 |
| Inferred from the shape of the calendar | `derived_from_report_type_and_publication` | up to 0.72 | 50% | 10 |

The inferred route sits below 0.8 on purpose, so the recommended threshold excludes it. It is weaker on US companies than European ones, because a US investor page tends to list earnings call dates rather than period-labelled filings. **Treat it as a hint, not a fact.**

The "within 7 days" column exists because a 52 or 53 week fiscal year genuinely ends on a fixed weekday near a month end: Papa John's declares 31 December to the SEC and closes on 27 December. Both numbers are published rather than the flattering one.

### 🔍 How to find a page on a company's website

**Three ways in, and they cost different amounts.**

| You have | What you set | What happens | What it costs |
|---|---|---|---|
| A bare domain | `domain` or `domains` | Discovery runs on that site. This is the normal path. | A locate event per page type. |
| A company name | `companies`, as `{"name": "..."}` | **Identity resolution runs first** to turn the name into a domain, then discovery runs. | **A `identity-resolved` event at $0.007 on top**, once per company, plus the locate events. |
| A name and a URL you already hold | `companies` with `url`, or `knownUrls` keyed by page type | The URL is taken as given. **Discovery is skipped** for that page type, so it is faster and exact. | No identity event, and the located page bills at the same locate rate with `method: known_url` and confidence 1.0. |

Supplying a domain, or a name alongside a URL, never triggers identity resolution and is never charged for it.

**Two modes.** `locate` returns the URL, the method, the confidence and the evidence phrase, and nothing else. `locate_and_extract` does all of that and then reads each located page, returning the structured fields into the findings dataset. Extraction adds a `page-extracted` event per page read.

1. Open the **Input** tab and put a domain in `domain`, for example `stripe.com`.
2. Pick your page types in `pageTypes`. Each one you add costs a locate event and takes more requests.
3. Leave `mode` on `locate` for URLs only, or set `locate_and_extract` to read the pages.
4. Run it. A cold `locate` run on one page type takes about 10 to 20 seconds per company.
5. Read `pricing_url` for the answer and `pricing_confidence` for how much to trust it.

#### 🧬 Using it in Clay

Add it as an enrichment on a company table and map `domain` to your domain column. Every input returns exactly one row, including the ones where nothing is found, so your table never loses accounts. Map `{type}_url` and `{type}_confidence` into columns and filter on `{type}_confidence >= 0.8`.

If you hold the URL already, pass it in `knownUrls` keyed by page type. Discovery is skipped for that type, which is faster and exact.

#### 🎚️ Filtering for outreach quality

`{type}_method` tells you what actually found the page, and precision differs sharply between methods. `known_url` and `known_host` are near certain. `homepage_anchor` and `footer_anchor` mean the company's own navigation said what the page is. `sitemap` means a URL slug matched and nothing on the page confirmed it. `path_guess` means we guessed and the page happened to confirm it. Threshold at **0.8 and above** for anything a customer will see.

### 💵 How much does it cost?

Pay per event, so you pay for what you asked for.

| Event | When it is charged |
|---|---|
| `page-type-located` | Per page type looked for, per company, when the look actually happened |
| `page-extracted` | Per page read, in `locate_and_extract` mode only |
| `identity-resolved` | Only on the company name path, never when you supply a domain |
| `apify-actor-start` | Once per run, per gigabyte of memory |

**Every event at every Apify tier, in dollars:**

| Event | FREE | BRONZE | SILVER | GOLD | PLATINUM | DIAMOND |
|---|---:|---:|---:|---:|---:|---:|
| `page-type-located` | 0.004 | 0.0038 | 0.0036 | 0.0034 | 0.0034 | 0.0034 |
| `page-extracted` | 0.003 | 0.00285 | 0.0027 | 0.00255 | 0.00255 | 0.00255 |
| `identity-resolved` | 0.007 | 0.00665 | 0.0063 | 0.00595 | 0.00595 | 0.00595 |
| `apify-actor-start` | 0.00005 | 0.00005 | 0.00005 | 0.00005 | 0.00005 | 0.00005 |

The three work events get cheaper as your Apify plan gets larger, down 15 percent at GOLD and flat from there. The actor start fee is set by Apify, not by us, and does not tier.

> 🧠 **This actor runs a browser, so it is billed at 4 GB and the actor start fee is four times the fleet's.** Apify charges the start event once per gigabyte of memory, and this actor's default is 4096 MB where most Mamba Labs actors sit at 256 or 512. So a run costs **4 x $0.00005 = $0.0002 to start**, before any page is looked at.
>
> What that means in practice:
>
> | How you call it | Start fee | Work | Start as a share |
> |---|---:|---:|---:|
> | One company per run, the Clay shape | $0.0002 | $0.004 | **4.8%** |
> | 100 companies in one run | $0.0002 | $0.400 | 0.05% |
> | 1,000 companies in one run | $0.0002 | $4.000 | 0.005% |
>
> **If you are calling it one row at a time from Clay, batch where you can.** The memory is not padding: rendering a JavaScript page needs a real browser, and dropping to 512 MB to save $0.00015 per run would trade out of memory crashes for a rounding error.

> 💳 **A look that happens is billed, including the ones that come back empty.** The actor reads robots.txt, the homepage, the link graph and the sitemap whether or not the page is there, and a documented `false` is a real answer that cost real work. **A look that does not happen is not billed.** If the domain is dead, refuses us, or is disallowed by robots.txt, nothing was looked at, the row comes back with `found: null`, and no locate event is charged for it.
>
> **Why that line and not "hits only".** A readable site that returns nothing is the MOST expensive outcome this actor produces, not the cheapest. Measured across 2,207 European domains: a `no_paths_found` row takes a median of 64 seconds against 30 for a hit, because it is the only outcome that exhausts every method, the homepage link graph, the footer, the sitemap and its shards, a hop into any located section, and finally every path guess. A hit stops as soon as it finds something. Billing hits only would give away the dearest work and keep the cheapest, which is backwards. The rows that genuinely cost nothing, a dead domain at 1.4 seconds or a refusal at 8, are the ones that are free.

> 🪪 **You are never charged for identity resolution unless you use the company name path.** Supplying a domain, or supplying a URL alongside a name, skips it entirely.

### ⌨️ Input

Set everything on the **Input** tab. Nothing is required, and a run with no usable input still returns a row saying so.

| Field | Default | What it does |
|---|---|---|
| `domain` | none | One company domain. Protocol and path are stripped. |
| `domains` | none | A list of company domains. |
| `companies` | none | Companies by name, as objects with `name` and optionally `country`, `isin`, `ticker`, `url`. Identity resolution runs first. |
| `pageTypes` | `["pricing"]` | Which of the 46 page types to find. |
| `mode` | `locate` | `locate_and_extract` also reads the pages. |
| `extractionFields` | all 13 | The page agnostic menu. |
| `extractPageTypeFields` | `true` | Also run the field map bound to each page type. |
| `knownUrls` | none | URLs you already hold, keyed by page type. Skips discovery for those. |
| `maxPagesPerType` | `4` | Candidate pages to open per page type. Lowering it is faster and finds less. |
| `maxRequestsPerInput` | `60` | Hard ceiling on requests to one company's site. Hitting it returns `coverage: partial`, never a false negative. |
| `concurrency` | `10` | How many companies to work on at once. Per company the actor is still strictly one request at a time. |
| `allowRender` | `true` | Open a browser for pages that serve no readable HTML. |
| `languageHints` | none | Reorders the vocabulary. Never shortens it. |
| `skipCache` | `false` | `true` forces a fresh crawl instead of the 14 day cache. |

### 📤 Output

**Two datasets.** The default dataset holds one flat row per input. The `findings` dataset holds one record per extracted field, keyed back by `input_key`, because eleven filing rows do not fit in one cell.

The per page type columns exist only for the page types you asked for. Ask for `pricing` and you get six pricing columns, not 270 empty ones.

This is a real row, copied unedited from a run of the shipped build on `stripe.com` asking for two page types in `locate` mode:

```json
{
  "input_key": "stripe.com",
  "input_path": "domain",
  "domain": "stripe.com",
  "company_name": null,
  "mode": "locate",
  "page_types_requested": "pricing,security_trust_center",
  "page_types_requested_count": 2,
  "page_types_found": 2,
  "page_types_not_found": 0,
  "page_types_undetermined": 0,
  "best_confidence": 0.93,
  "identity_resolved": null,
  "identity_method": null,
  "identity_confidence": null,
  "findings_count": 0,
  "findings_agnostic_count": 0,
  "findings_bound_count": 0,
  "findings_rejected_count": 0,
  "rejected_by_filter": null,
  "coverage": "partial",
  "fetch_status": "ok",
  "render_mode": "http",
  "pages_reached": 3,
  "requests_made": 6,
  "renders_made": 0,
  "destination_mismatches": 0,
  "render_failures": 0,
  "render_broken": false,
  "robots_crawl_delay_ms": 1000,
  "robots_disallowed_paths": 0,
  "blocked_pages": 0,
  "javascript_only_pages": 0,
  "homepage_links_seen": 92,
  "sitemap_urls_seen": 0,
  "fetch_error": null,
  "cache_hit": false,
  "run_date": "2026-08-16T08:03:49.406Z",
  "pricing_url": "https://stripe.com/pricing",
  "pricing_found": true,
  "pricing_method": "homepage_anchor",
  "pricing_confidence": 0.93,
  "pricing_evidence_phrase": "pricing",
  "pricing_pages_reached": 1,
  "security_trust_center_url": "https://docs.stripe.com/security",
  "security_trust_center_found": true,
  "security_trust_center_method": "path_guess",
  "security_trust_center_confidence": 0.6,
  "security_trust_center_evidence_phrase": "Security at Stripe",
  "security_trust_center_pages_reached": 1
}
```

Two things in that row are worth reading, because they are the row telling on itself. `coverage` is `partial`, not `good`, because `sitemap_urls_seen` is 0: the sitemap was not readable on this run, so one discovery method never ran. And the trust center came back at **0.6 on `path_guess`**, meaning we guessed the path and the page confirmed it, against 0.93 for the pricing page which Stripe's own navigation named. Both are found, and they are not equally trustworthy. That is the point of shipping the method and the confidence next to the URL.

**The findings dataset.** In `locate_and_extract` mode the same domain produced 22 findings records off the pricing page. One record per field, so eleven filing rows or six pricing plans do not have to be crushed into one cell:

```json
{
  "input_key": "stripe.com",
  "domain": "stripe.com",
  "page_type": "pricing",
  "source_url": "https://stripe.com/pricing",
  "group": "bound",
  "field": "plan_name",
  "value": "Payments",
  "value_type": "string",
  "method": "plan_card",
  "confidence": 0.85,
  "detail_json": "{\"plan_index\":1}",
  "is_personal_data": false,
  "lawful_basis": null,
  "page_status": "ok",
  "run_date": "2026-08-16T08:04:19.535Z"
}
```

`group` is `agnostic` for the menu that runs on any page and `bound` for the fields tied to that page type. Join back to the flat row on `input_key`.

> ⚠️ **`false` and `null` mean different things and are never collapsed.** `pricing_found: false` means the homepage was read, its link graph was searched, the sitemap was searched, and the page is not reachable. `pricing_found: null` means not enough was readable to say. A blocked or JavaScript only site returns `null`, never `false`.

**Three fields to read before you trust a `false`:**

| Field | What a good answer looks like |
|---|---|
| `coverage` | `good` means the homepage, its link graph and the sitemap were all read. `partial` means one was missing or the request budget ran out. `poor` means the homepage itself was not read, so the actual method never ran. `none` means nothing was readable and every `found` on the row is `null`. |
| `pages_reached` | How many pages returned real readable content. A `false` on a row with `pages_reached: 0` is not a negative, it is a row where nothing was seen. |
| `fetch_status` | `ok`, `partial` and `no_paths_found` are real answers. Everything else is a failure to look, and none of the failures is billed. Full list below. |

A `false` with `coverage: good` and `pages_reached` above zero is a negative you can act on. Anything else is an admission, not a finding.

**Every `fetch_status` value, and what each one means for the row:**

| `fetch_status` | What happened | `found` values | Billed? |
|---|---|---|:--:|
| `ok` | The site was read and at least one requested page type was located. | Real `true` and `false` | ✅ |
| `partial` | The site was read but the per company request ceiling ran out before every method finished. Treat a `false` here as unproven. | `true`, or `false` you should not trust | ✅ |
| `no_paths_found` | Everything was readable, every method ran to the end, and the page is not there. **This is a real answer, not a failure.** | `false` | ✅ |
| `unreachable` | DNS resolved but the site did not answer: connection refused, TLS failure, or timeout. | All `null` | ❌ |
| `blocked` | The site refused us with 401, 403, 407, 429 or 451, or served a CAPTCHA interstitial. **We do not retry and we do not work around it.** | All `null` | ❌ |
| `javascript_only` | The page still had no readable text after a real browser rendered it. Usually a client side app behind a shell. | All `null` | ❌ |
| `robots_disallowed` | The company's own `robots.txt` disallows our user agent on the paths we would need. We stop. | All `null` | ❌ |
| `domain_dead` | The domain does not resolve at all. | All `null` | ❌ |

The homepage decides the row. If the homepage comes back `blocked`, `domain_dead`, `robots_disallowed`, `javascript_only` or `unreachable`, that status is the row's status, because every discovery method depends on reading the homepage first.

### 💡 Tips

- Ask for the page types you will actually use. Each one costs an event and adds requests.
- Pass `knownUrls` for anything you already hold. It is faster, exact, and skips discovery for that type.
- Filter on `{type}_confidence >= 0.8` before putting a URL in front of a customer.
- `coverage: good` with `found: false` is a real, usable negative. Anything else means the look was incomplete.
- Re-run with `skipCache: "true"` after a site redesign.
- For company names rather than domains, put the name in `companies`. Stockholm and Helsinki are domain starved, and the name path is the only way to reach most of those companies.

### ⚠️ Known limits

- **Some sites refuse us, and we do not argue.** 401, 403, 429 and CAPTCHA interstitials come back as `fetch_status: "blocked"`, unretried, with every field null. In the 2,186 rows of the shipped build's run that was 175 companies, the single largest reason for a `null`. Re-running does not help, because we are not trying to get around it.
- **robots.txt is honored, including `Crawl-delay`.** Fourteen companies in that population disallowed the paths we needed and came back `robots_disallowed`. That is the correct outcome, not a bug.
- **A browser renders, it does not break in.** Rendering is used for pages a company serves publicly that need JavaScript to read. It is never used against a block, a CAPTCHA, a login, or robots.txt. A page that is still unreadable after rendering reports `javascript_only`, which now means "a browser looked and there was nothing there".
- **`security_trust_center` is restricted on purpose, and it returns fewer pages than it could.** For this page type only, an anchor that says "Security" is not enough: the URL has to agree. The last part of the path must itself be a trust token and the page must not sit under a product or content section, so `acme.com/trust` is accepted and `acme.com/industries/security` is not. A dedicated `trust.` subdomain or a known trust vendor host skips the check, because the host is the evidence.

  Why: at a security vendor the word Security in the navigation names a product. Before the restriction this page type measured **55.0 percent correct across two hand read samples, n=60**. After it, **83.3 percent, n=30**. The cost is coverage: it now finds a trust center on 25.0 percent of a US software corpus where it previously claimed 39.0 percent, so **about a third of the old results are gone and most of them were wrong**.

  Precision by method on the restricted build, n=30:

  | `security_trust_center_method` | Correct | n |
  |---|---:|---:|
  | `footer_anchor` | 10 of 10 | 10 |
  | `known_subdomain`, a `trust.` or `security.` host | 6 of 6 | 6 |
  | `path_guess` | 3 of 4 | 4 |
  | `sitemap` | 2 of 3 | 3 |
  | `homepage_anchor` | 4 of 7 | 7 |

  `homepage_anchor` is still the weak one. If you want near certainty, keep `known_subdomain` and `footer_anchor` and drop the rest.
- **Precision by page type, hand read, every figure with its sample size:** `investor_relations` 96.7% (n=30), `careers` 90.0% (n=30), `contact` 90.0% (n=30), `about` 83.3% (n=30), `security_trust_center` 83.3% (n=30, restricted build), `pricing` 80.0% (n=30). The other 40 page types have not been measured at this scale and no figure is published for them.
- **A page type with no bound field map returns only the page agnostic fields.** That is deliberate. Seventeen page types have bound extractors; the other 29 locate and return the general group.
- **Landing pages are indexes.** For `investor_relations` the actor walks one hop into the section to reach the calendar or reports page, because that is where filing tables live. It reads up to three such pages.
- **Derived fields are marked as derived.** `year_end_source` tells you whether a fiscal year end was stated by the company or inferred from the shape of its calendar. The weakest inferred route, reading a fiscal year end off the shape of a filing calendar, is capped at 0.72 confidence and is easy to filter out. It measured 25 correct of 34 hand read, 73.5 percent, which is why the cap sits below the 0.8 threshold this README tells you to use.
- **One request at a time per domain, with a delay.** This is polite by design and it sets the floor on how fast a large batch can run.
- **No proxy.** The actor runs proxyless, so rate limiting surfaces in `fetch_status` rather than being hidden.

### ❓ FAQ

**How is this different from just guessing `/pricing`?**
Guessing is one of seven methods here and it is the last one tried. On the 2,207 company European population, guessing paths accounted for 881 failures on its own. The link graph, the sitemap, the multilingual vocabulary and the one hop walk are what find the other pages.

**Why did my company come back empty?**
Read `coverage` and `fetch_status` first. `coverage: good` with `fetch_status: no_paths_found` means a complete look found nothing, which is a real answer. `blocked`, `javascript_only`, `robots_disallowed` or `domain_dead` mean the look did not complete, and those rows are not billed a locate event.

**Can I trust a `path_guess` result?**
Only when the page itself confirmed it, which is the only case where one is returned at all, and it still scores 0.6. Threshold at 0.8 to exclude them.

**Does it work outside Europe?**
Yes. The vocabulary is European because that is where the measurement was done, and English covers the US and UK. Nothing about the method is region specific.

**Why two datasets?**
Because a pricing page has four plans and an investor relations page has eleven filing rows, and neither fits in a flat cell. The flat row answers "where is it and how sure are you". The findings dataset answers "what was on it".

**Can I get more page types?**
Adding a page type is vocabulary in eleven languages plus an optional field map. Open an issue on the **Issues** tab with the page type you want and what you would read off it.

### 🧩 Want other GTM data?

Mamba Labs builds custom actors for B2B go-to-market teams. The public versions
of that work live here on the Store, so our users get the same tooling we build
under contract.

| | |
|---|---|
| 🧑‍💼 [GTM Hiring Signal Scraper](https://apify.com/mambalabs/gtm-hiring-signal-scraper) | 🧱 [Tech Stack Detector](https://apify.com/mambalabs/gtm-tech-stack-signal-scraper) |
| 📡 [B2B Buying Signals Aggregator](https://apify.com/mambalabs/b2b-buying-signals-hiring-tech-stack-intent-for-clay) | 🔑 [Job Board Keyword Scanner](https://apify.com/mambalabs/job-board-keyword-signal-scanner) |
| 🔗 [Domain to LinkedIn URL Resolver](https://apify.com/mambalabs/domain-to-linkedin-url-resolver) | 🎯 [ICP Fit Scorer](https://apify.com/mambalabs/icp-account-lead-scoring-fit-scorer-0-100-for-clay) |
| 📋 [Job Posting Monitor](https://apify.com/mambalabs/gtm-job-discovery) | 📬 [Domain Deliverability Checker](https://apify.com/mambalabs/domain-deliverability-checker) |
| 🏢 [Company Firmographic Enricher](https://apify.com/mambalabs/company-firmographic-enricher) | 🌐 [Company Social Presence Mapper](https://apify.com/mambalabs/company-social-presence-mapper) |
| 🪪 [Company Identity Resolver](https://apify.com/mambalabs/company-identity-resolver) | 💰 [Funding and Press Signal Scanner](https://apify.com/mambalabs/funding-press-signal-scanner) |
| 🔄 [Company Change-Event Feed](https://apify.com/mambalabs/company-change-event-feed) | 👤 [People Finder and Email Verifier](https://apify.com/mambalabs/people-finder) |
| 🚀 [Prospect Engine](https://apify.com/mambalabs/b2b-prospect-engine) | 🤖 [AI Tooling Detector](https://apify.com/mambalabs/ai-tooling-detector) |
| 📮 [Outbound Stack Detector](https://apify.com/mambalabs/outbound-infrastructure-fingerprint) | 📝 [Publishing Frequency Tracker](https://apify.com/mambalabs/blog-publishing-frequency) |
| ✉️ [Work Email Waterfall Finder](https://apify.com/mambalabs/email-waterfall-orchestrator) | ⏩ [Sequencer Lead Push](https://apify.com/mambalabs/clay-to-instantly-smartlead-push) |
| 🏅 [Workplace Program Detector](https://apify.com/mambalabs/workplace-program-detector) | 👥 [Team Page People Extractor](https://apify.com/mambalabs/team-page-people-extractor) |
| 🧭 [Company Discovery List Builder](https://apify.com/mambalabs/company-discovery-list-builder) | 🎪 [Event Presence Index](https://apify.com/mambalabs/event-presence-index) |
| ⚖️ [Legal Entity Resolver](https://apify.com/mambalabs/legal-entity-resolver) | 🏷️ [Contact Classifier](https://apify.com/mambalabs/contact-classifier) |
| 💬 [LinkedIn Post Tracker and Comment Capture](https://apify.com/mambalabs/linkedin-post-engager-capture) | 🏛️ [Government Contract Award Monitor](https://apify.com/mambalabs/public-award-monitor) |
| 🕵️ [Agent Accessibility Auditor](https://apify.com/mambalabs/agent-accessibility-auditor) | |

> Every actor in the suite takes a domain or a company and returns one flat row,
> so they stack in the same Clay table without reshaping anything.

> 🛠️ **Need something custom built for you or your team?** Tell us what you are
> trying to find and we will build it. [Talk to Mamba Labs](https://mambabuilt.com/contact).

### 🛟 Support

Issues, feature requests and data questions go through the **Issues** tab on this actor. Include the domain, the page type and the run ID and it can be traced.

#### 👤 Personal data, stated plainly

**This actor can return a named person, so read this section rather than assuming.**

Two page types can return one: `contact` returns the contact block a company publishes on its own contact page, and `about` returns leadership names published on the company's own about page. That is the whole of it. Every other page type returns company information only.

- **It only ever returns what the company published on its own website.** It does not search for people, does not look anyone up elsewhere, does not enrich, does not guess an email pattern, and does not join a name to any other source.
- **It infers no personal attribute.** No seniority scoring, no gender, no location inference, no profile building. A name, a job title and a business contact method, exactly as the company printed them.
- **Every such field is flagged.** Records carrying a person have `is_personal_data: true` and a `lawful_basis` in the findings dataset, so you can filter the entire class out with one predicate, or exclude it up front by leaving `contact` and `about` out of `pageTypes`.
- **Lawful basis.** Legitimate interest in publicly published business contact information, under GDPR Article 6(1)(f). These are role holders acting in a business capacity, published by their employer for the purpose of being contacted.
- **Deletion.** Open an issue on the **Issues** tab naming the page and the person, and the record is removed from our cache and added to a suppression list. We cannot remove it from the company's own website, which is where it is published.

**If you want a roster of people at a company, this is the wrong actor.** Use [Team Page People Extractor](https://apify.com/mambalabs/team-page-people-extractor) to pull a team page, and [Contact Classifier](https://apify.com/mambalabs/contact-classifier) to classify contacts you already hold. This actor returns a contact block that happened to sit on a page it was asked to find, and nothing more.

> **Sourcing and legal.** This actor reads publicly available pages on the company's own website. It reads and honors `robots.txt` per host, including `Crawl-delay`, makes one request at a time per domain with a delay, and sends a descriptive user agent that names it and links here. It does not impersonate a browser, does not sign headers, and does not retry to get around a block. Where a page is publicly served but needs JavaScript to read, a browser renders it; a browser is never used against an access control of any kind. Refusals are reported in the output rather than hidden.

Built by [Mamba Labs](https://apify.com/mambalabs).

# Actor input Schema

## `domain` (type: `string`):

A single company domain, for example stripe.com. Protocol and path are stripped. Use this for one company; use Company domains for a list.

## `domains` (type: `array`):

A list of company domains. Every domain returns a row, including the ones where nothing is found.

## `companies` (type: `array`):

Companies you have a name for but not a domain. Each entry is an object with name, and optionally country, isin, ticker and url. Identity resolution runs first on this path and is charged as its own event. Send as JSON when calling from Clay.

## `pageTypes` (type: `array`):

Which page types to locate. 46 are available. Each one you add costs a locate event and takes more requests, so ask for what you will use.

## `mode` (type: `string`):

locate returns the page URL, the method that found it and a confidence. locate\_and\_extract also reads the page and returns structured fields. Locating alone is cheaper and is a complete product on its own.

## `extractionFields` (type: `array`):

The page agnostic menu. These are available on every page type. Only used in locate\_and\_extract mode. Leave empty for all of them.

## `extractPageTypeFields` (type: `string`):

true also runs the field map bound to each page type: pricing plans, filing rows and a derived fiscal year end, certifications, ATS host, governing law, and the rest. Only used in locate\_and\_extract mode.

## `knownUrls` (type: `object`):

URLs you already hold, keyed by page type, for example {"pricing": "https://stripe.com/pricing"}. Discovery is skipped for that page type, which is faster and more accurate. Send as JSON when calling from Clay.

## `maxPagesPerType` (type: `string`):

Between 1 and 12. How many candidate pages to open per page type before giving up. Lowering it is faster and finds less: at 1 the actor opens only its single best candidate, so a company whose real pricing page was the second best guess comes back not found. 4 is the measured default. Sent as a string so it works from Clay.

## `maxRequestsPerInput` (type: `string`):

Between 5 and 200. The hard ceiling on requests to one company's site, robots.txt and sitemap included. Hitting it returns coverage partial rather than a false negative. Sent as a string so it works from Clay.

## `concurrency` (type: `string`):

Between 1 and 30. How many companies to work on at once. Politeness is enforced per domain, not globally, so raising this never sends two requests to the same site at once. 10 is the default and it is what keeps the per company compute cost low enough to price. Sent as a string so it works from Clay.

## `allowRender` (type: `string`):

true opens a browser for pages that serve no readable HTML, which is most Nordic investor calendars. It is never used to get past a block, a CAPTCHA, a login or robots.txt: those come back refused, unretried. false skips rendering entirely and those pages report javascript\_only.

## `languageHints` (type: `array`):

Optional. Discovery vocabulary is multilingual by default in all 11 languages; this only reorders it so the languages you name are tried first. It never removes a language.

## `skipCache` (type: `string`):

false uses the 14 day cache. true forces a fresh crawl.

## `source_tag` (type: `string`):

Internal attribution tag set by Mamba Labs on published task examples. Not required, and nothing depends on it. Leave it empty.

## Actor input object example

```json
{
  "domain": "stripe.com",
  "domains": [
    "stripe.com",
    "vercel.com"
  ],
  "companies": [
    {
      "name": "Wartsila",
      "country": "FI"
    }
  ],
  "pageTypes": [
    "pricing",
    "security_trust_center"
  ],
  "mode": "locate",
  "extractionFields": [
    "copyright",
    "legal_entity",
    "emails",
    "phones",
    "addresses",
    "social_links",
    "meta",
    "last_modified",
    "canonical",
    "schema_org",
    "forms",
    "ctas",
    "tech_markers"
  ],
  "extractPageTypeFields": "true",
  "knownUrls": {
    "pricing": "https://stripe.com/pricing"
  },
  "maxPagesPerType": "4",
  "maxRequestsPerInput": "60",
  "concurrency": "10",
  "allowRender": "true",
  "languageHints": [
    "de",
    "fr"
  ],
  "skipCache": "false"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `findings` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domain": "stripe.com",
    "domains": [
        "stripe.com"
    ],
    "pageTypes": [
        "pricing"
    ],
    "mode": "locate",
    "extractionFields": [
        "copyright",
        "legal_entity",
        "emails",
        "phones",
        "addresses",
        "social_links",
        "meta",
        "last_modified",
        "canonical",
        "schema_org",
        "forms",
        "ctas",
        "tech_markers"
    ],
    "extractPageTypeFields": "true",
    "maxPagesPerType": "4",
    "maxRequestsPerInput": "60",
    "concurrency": "10",
    "allowRender": "true",
    "skipCache": "false"
};

// Run the Actor and wait for it to finish
const run = await client.actor("mambalabs/page-finder-extractor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domain": "stripe.com",
    "domains": ["stripe.com"],
    "pageTypes": ["pricing"],
    "mode": "locate",
    "extractionFields": [
        "copyright",
        "legal_entity",
        "emails",
        "phones",
        "addresses",
        "social_links",
        "meta",
        "last_modified",
        "canonical",
        "schema_org",
        "forms",
        "ctas",
        "tech_markers",
    ],
    "extractPageTypeFields": "true",
    "maxPagesPerType": "4",
    "maxRequestsPerInput": "60",
    "concurrency": "10",
    "allowRender": "true",
    "skipCache": "false",
}

# Run the Actor and wait for it to finish
run = client.actor("mambalabs/page-finder-extractor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domain": "stripe.com",
  "domains": [
    "stripe.com"
  ],
  "pageTypes": [
    "pricing"
  ],
  "mode": "locate",
  "extractionFields": [
    "copyright",
    "legal_entity",
    "emails",
    "phones",
    "addresses",
    "social_links",
    "meta",
    "last_modified",
    "canonical",
    "schema_org",
    "forms",
    "ctas",
    "tech_markers"
  ],
  "extractPageTypeFields": "true",
  "maxPagesPerType": "4",
  "maxRequestsPerInput": "60",
  "concurrency": "10",
  "allowRender": "true",
  "skipCache": "false"
}' |
apify call mambalabs/page-finder-extractor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mambalabs/page-finder-extractor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/TpurgcOZbVnknlaiC/builds/Blw3MToaobA1Vqh2P/openapi.json
