# Wellfound Scraper — Startup & Job Data (`b2b_leads/wellfound-real-time-data-scraper`) Actor

Turn Wellfound (AngelList) into clean, structured startup data. Get company profiles with hiring signals and open jobs with salaries, plus websites, emails, and socials when available. Rows stream to your dataset in real time — no setup, pay per result.

- **URL**: https://apify.com/b2b\_leads/wellfound-real-time-data-scraper.md
- **Developed by:** [Emmanuel](https://apify.com/b2b_leads) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Wellfound Real-Time Data Scraper — Startup Companies, Jobs & People at Scale

**Turn Wellfound (AngelList) into clean, structured startup data.** One run collects the live startup feed — company profiles with descriptions, markets, size, and locations, plus their open job listings — and public people profiles, with optional **lead details** (website, email, phone, social profiles) added to each company when available. Rows stream into your Apify dataset **as they are finalized**, so even long runs stay light and nothing is lost mid-run.

No API keys. No spreadsheets. No cleanup. Just press **Start**.

> ⚡ **Real-time streaming output** • 🏢 **Companies + jobs in one run** • 👤 **People profiles** • 🎯 **Lead details without filtering** • 💸 **Pay only for what you export**

***

### ✅ Pricing & free-plan note (read first)

This Actor uses Apify's **pay-per-result** model — you're charged per exported row (see [Pricing](#pricing)).

- **Apify free plan:** runs are limited to a **2-result sample** per run, and the log tells you to upgrade. This is an Apify platform plan restriction, applied transparently — it is not an error.
- **Any paid Apify plan (Bronze and above):** full, uncapped output. You get exactly what your input asks for.
- Every run also honors your **max run charge** (spending limit): the Actor finishes gracefully the moment your configured limit is reached — no more charges than you approved.

Details: [Free plan limitations](#free-plan-limitations) · [Pricing](#pricing)

***

### Why this Actor

| | Wellfound Real-Time Data Scraper | Typical scraper |
|---|---|---|
| **Output** | Flat, stable JSON field names | Messy, inconsistent |
| **Streaming** | Rows saved live during the run | Results only at the end |
| **Memory** | **512 MB** default | 2–4 GB+ |
| **Lead details** | Added when available — never filters | Filters and drops rows |
| **Setup** | Organized input UI — run immediately | Needs tuning |
| **Freshness** | Collected at run time ("real-time") | Cached or stale |

***

### What you get

Every row includes `featureType`, `recordType`, and `scrapedAt` so you can filter, join, and pipe into any workflow.

#### Company rows (`recordType: "company"`)

| Group | Fields |
|-------|--------|
| **Identity** | `companyId`, `name`, `profileUrl`, `logoUrl` |
| **About** | `description` (tagline, plus website description when found), `location`, `markets[]`, `companySize` |
| **Signals** | `badges[]` (e.g. actively hiring), `openJobsCount` (open roles when shown) |
| **Lead details** | `website`, `email`, `emails[]`, `phone`, `phones[]`, `linkedinUrl`, `twitterUrl`, `facebookUrl`, `instagramUrl`, `address` |
| **Structured view** | `leadDetails` (object with all lead fields), `searchQuery`, `source`, `scrapedAt` |

#### Job rows (`recordType: "job"`)

| Group | Fields |
|-------|--------|
| **Identity** | `jobId`, `title`, `jobUrl`, `companyName`, `companyProfileUrl` |
| **Compensation & location** | `salary` (e.g. `$120k – $160k • 0.1–0.3% equity`), `location`, `workplaceType` (ONSITE / REMOTE / HYBRID), `remote` |
| **Timing** | `postedAt`, `searchQuery`, `source`, `scrapedAt` |

#### People rows (`recordType: "person"`)

| Group | Fields |
|-------|--------|
| **Identity** | `name`, `title` (headline), `profileUrl`, `avatarUrl` |
| **Public links** | `linkedinUrl`, `twitterUrl`, `githubUrl`, `email` (when displayed) |
| **Structured view** | `leadDetails`, `source`, `scrapedAt` |

Missing values are `null` — field names are stable across runs, so your transforms won't break.

***

### Features

#### 🏢 Companies (on by default)

Collect startup company profiles from the paginated startup directory: name, description, markets, size, location, badges, and the company's open-jobs count. Directory pages rotate, so a run is not limited to the live feed's pass size — collect 10 or 10,000 companies. Cap the number with **Max companies this run** (`0` = keep going until the directory is exhausted).

#### 💼 Jobs (off by default)

Collect open job listings joined to their company: title, salary, workplace type (onsite / remote), location, and posting date. Listings are collected across role families and locations in rotating pages, so job volume scales with your cap instead of stopping at a fixed pass size — the live feed pass is only used as a top-up. Cap the number with **Max jobs this run**, or keep only remote roles with **Remote jobs only**. Job keywords that clearly name a job family (engineering, sales, design, data, support…) steer which role listings are collected first.

#### 🏷️ Keywords as context

Each section carries its own keywords, recorded on every exported row so you can tag and filter your own exports (e.g. `AI` for companies, `engineering` for jobs). For jobs, a keyword that names a job family also influences which listings are collected first; otherwise keywords label the data — they never filter it.

#### 🎯 Enable lead details (on by default) — enrichment, never a filter

Adds lead contact details to each company row — **website, email, phone, social profiles** — when they are available. The email finder covers every way companies publish addresses: contact and about pages, structured contact data, protected/encoded addresses, and human-written forms like `hello [at] company [dot] com`.

**Every company is still exported even when no lead details are found.** Nothing is ever dropped for lacking contact info, so the time and cost of a run stay predictable (this is deliberate: you always know what a run of N companies costs). Adds a little extra time per company.

#### 👤 People profiles

Paste public member profile URLs (`wellfound.com/u/...`) and collect people records as their own rows — name, headline, avatar, and the public social links the member displays. Ideal for founder and recruiting outreach lists.

#### ⚙️ Output controls

Cap companies and jobs independently, keep only remote jobs, cap total items, and deliver rows to a webhook in real time (see [Webhooks](#webhooks--real-time-delivery-to-your-stack)).

***

### Who it's for

- **Recruiters & talent teams** — companies that are hiring right now, with roles, salaries, and direct contact channels.
- **Sales & BD teams selling to startups** — filter by market, size, and location; reach founders and hiring managers with verified contact details.
- **Lead-generation agencies** — a steady stream of fresh startup leads with website, email, and phone when available, streamed into your CRM.
- **Growth & partnerships managers** — track fast-growing companies by badge-worthy signals (actively hiring, growing fast) in your space.
- **Investors & scouting analysts** — market maps of active startups with markets, size, location, and descriptions for deal screening.
- **Job seekers & career coaches** — structured lists of open startup roles with salary ranges and workplace types.
- **Market researchers & analysts** — joinable JSON across runs for startup-ecosystem studies.
- **AI & automation builders** — clean rows for LLM prompts, scoring agents, and enrichment pipelines via the Apify API and MCP.
- **CRM & data teams** — stable flat JSON that drops straight into Postgres, BigQuery, Airtable, or Sheets.

### Use cases

- **Hiring-company prospecting** — every run gives you companies with live job openings; combine with lead details for direct outreach.
- **Recruiting-fee hunting** — export roles with salary bands, find the hiring company's website and contact channels, pitch your services.
- **SaaS lead gen by market** — tag rows with your keywords (e.g. `AI`, `fintech`, `healthtech`), route to your CRM by webhook, and let sales filter by market and size.
- **Foundering-target monitoring** — schedule daily runs on your keywords; new companies appearing in the feed are fresh leads.
- **Remote-job tracking** — enable *Remote jobs only* and stream remote roles to Slack for a candidate community.
- **Investor deal flow** — collect startups in your thesis markets weekly; `markets[]`, `companySize`, and `description` feed your screener.
- **Sales-trigger monitoring** — companies that just posted roles are companies that are growing: feed them to your outbound sequences first.
- **Talent-market mapping** — compare salary ranges and workplace types across locations and markets over time.
- **Enrichment pipeline input** — take company rows into your own enrichment stack; website/socials are already filled when available.
- **AI company research** — feed descriptions, markets, and job data to an LLM to summarize and score startups automatically.

### Quick start

1. Open the Actor in Apify Console.
2. Leave the defaults — the startup feed with lead details is on.
3. Optionally add **keywords** to tag rows (e.g. `AI`, `fintech`).
4. Set **Max companies this run** (Companies section) and **Max jobs this run** (Jobs section) — defaults: 10 each. Turn either section off if you only need the other.
5. Press **Start** and watch rows stream into the dataset.
6. Export as **JSON / CSV / Excel**, pull via the **Apify API**, wire up a **webhook**, or use **MCP**.

#### Example — startup feed with lead details (default)

```json
{
  "enableCompanies": true,
  "companyKeywords": ["AI"],
  "feedMaxCompanies": 10,
  "enableJobs": true,
  "jobKeywords": ["engineering"],
  "feedMaxJobs": 10,
  "enableLeadDetails": true
}
```

#### Example — remote jobs only, larger run

```json
{
  "enableCompanies": false,
  "enableJobs": true,
  "feedRemoteOnly": true,
  "feedMaxJobs": 100
}
```

#### Example — companies with lead details only

```json
{
  "enableCompanies": true,
  "enableJobs": false,
  "feedMaxCompanies": 50,
  "enableLeadDetails": true
}
```

#### Example — people profiles from URLs

```json
{
  "enableCompanies": false,
  "enableJobs": false,
  "enablePeopleProfiles": true,
  "peopleUrls": [
    "https://wellfound.com/u/example-founder",
    "https://wellfound.com/u/example-cto"
  ]
}
```

#### Example — results to a webhook (Slack-friendly)

```json
{
  "enableCompanies": true,
  "feedMaxCompanies": 25,
  "webhookUrl": "https://hooks.slack.com/services/T000/B000/XXXX",
  "webhookFormat": "slack"
}
```

### Input reference

| Field | Type | Default | Description |
|---|---|---|---|
| `enableCompanies` | boolean | `true` | Companies section: collect startup company profiles from the paginated startup directory. |
| `companyKeywords` | string\[] | `[]` | Keywords recorded on every company row as context tags (e.g. `AI`, `fintech`). |
| `feedMaxCompanies` | integer | `10` | Cap company rows this run — the main cost lever for companies. `0` = no cap (keep paginating until the directory is exhausted). |
| `enableJobs` | boolean | `false` | Jobs section: collect open job listings joined to their companies. |
| `jobKeywords` | string\[] | `[]` | Keywords recorded on every job row as context tags (e.g. `engineering`, `sales`); job-family keywords also steer which listings are collected first. |
| `feedMaxJobs` | integer | `10` | Cap job rows this run — the main cost lever for jobs. `0` = no cap (keep collecting until listings are exhausted). |
| `feedRemoteOnly` | boolean | `false` | Keep only job rows marked remote (Jobs section). |
| `enablePeopleProfiles` | boolean | `false` | Collect public member profiles from `peopleUrls`. |
| `peopleUrls` | string\[] | `[]` | Public member profile URLs (`wellfound.com/u/...`). |
| `enableLeadDetails` | boolean | `true` | Add lead contact details to company rows when available. **Never filters** — every company is exported either way. Adds a little extra time per company. Free plan: capped at 2 results per run — see [Free plan limitations](#free-plan-limitations). |
| `webhookUrl` | string | `""` | Optional. Each row is saved to the dataset **and** delivered to this URL in real time. |
| `webhookFormat` | select | `json` | `json` (full record) or `slack` (message payload). |
| `proxyConfiguration` | proxy | US residential | Apify proxy settings (US residential on by default). |

***

### Output reference

Every exported record is a **flat JSON object** in the run's dataset, tagged by `recordType`:

| `recordType` | Row shape |
|---|---|
| `company` | Startup company from the feed (with lead details when available). |
| `job` | Job listing from the feed, joined to its company. |
| `person` | Public member profile from a profile URL. |

#### Company row — full field list

| Field | Type | Description |
|---|---|---|
| `featureType` | string | Always `company_profiles`. |
| `recordType` | string | Always `company`. |
| `companyId` | string | null | Platform company ID. |
| `name` | string | null | Company name. |
| `profileUrl` | string | null | Company profile URL. |
| `description` | string | null | Tagline / short description (longer website description when found). |
| `location` | string | null | Primary location(s), e.g. `New York City`. |
| `markets` | string\[] | Market / industry tags, e.g. `["Artificial Intelligence", "SaaS"]`. |
| `companySize` | string | null | Size band, e.g. `11-50`. |
| `badges` | string\[] | Public company signals, e.g. `ACTIVELY_HIRING`, `YC`, `GROWING_FAST`. |
| `openJobsCount` | number | null | Open roles the company lists on the platform, when shown. |
| `logoUrl` | string | null | Logo image URL. |
| `website` | string | null | Company website (lead details). |
| `email` | string | null | Public contact email (lead details). |
| `emails` | string\[] | Up to 3 public emails (lead details). |
| `phone` | string | null | Public phone in `(555) 555-5555` format (lead details). |
| `phones` | string\[] | Up to 5 public phones (lead details). |
| `linkedinUrl` / `twitterUrl` / `facebookUrl` / `instagramUrl` | string | null | Social profiles (lead details). |
| `address` | string | null | Public postal address (lead details). |
| `leadDetails` | object | null | All lead fields grouped in one object (same values as the flat fields). |
| `searchQuery` | string | null | Keywords this run tagged the row with. |
| `source` | string | Which stage produced the row. |
| `scrapedAt` | string | ISO 8601 UTC timestamp of collection. |

#### Job row — full field list

| Field | Type | Description |
|---|---|---|
| `featureType` / `recordType` | string | `jobs_feed` / `job` (`featureType` is a stable data tag; the input sections are Companies and Jobs). |
| `jobId` | string | null | Platform job ID. |
| `title` | string | null | Job title. |
| `jobUrl` | string | null | Job listing URL. |
| `companyName` | string | null | Hiring company name. |
| `companyProfileUrl` | string | null | Company profile URL. |
| `salary` | string | null | Compensation string, e.g. `$120k – $160k • 0.1–0.3% equity`. |
| `location` | string | null | Job location(s), e.g. `San Francisco, New York City`. |
| `workplaceType` | string | null | `ONSITE`, `REMOTE`, or `HYBRID`. |
| `remote` | boolean | null | True when the role is remote. |
| `postedAt` | string | null | ISO 8601 UTC posting date. |
| `searchQuery` / `source` / `scrapedAt` | — | Same as company rows. |

#### People row — full field list

| Field | Type | Description |
|---|---|---|
| `featureType` / `recordType` | string | `people_profiles` / `person`. |
| `name` | string | null | Member name. |
| `title` | string | null | Headline. |
| `profileUrl` | string | null | Profile URL you provided. |
| `avatarUrl` | string | null | Avatar image URL. |
| `linkedinUrl` / `twitterUrl` / `githubUrl` | string | null | Public social links displayed on the profile. |
| `email` | string | null | Public email displayed on the profile. |
| `leadDetails` / `source` / `scrapedAt` | — | Same as other rows. |

#### Run summary (`OUTPUT` key-value store)

Each run also writes a summary you can read via the Apify API:

```json
{
  "totalPushed": 120,
  "enabledFeatures": ["companies", "jobs", "lead_details"],
  "errors": [],
  "spendingLimitReached": false,
  "paywall": {
    "detected": true,
    "isPaying": true,
    "pricingTier": "SILVER",
    "blocked": false,
    "limited": false,
    "mode": "limit",
    "freeTierMaxItems": null
  },
  "finishedAt": "2026-09-25T12:00:00.000Z"
}
```

| Summary field | Meaning |
|---|---|
| `totalPushed` | Rows exported this run. |
| `enabledFeatures` | Features that ran (e.g. `companies`, `jobs`, `people_profiles`, `lead_details`). |
| `errors` | Non-fatal failures, expressed as fixed user-facing sentences. |
| `spendingLimitReached` | `true` when your max run charge was reached — the run then finishes gracefully. |
| `paywall` | Free-plan gate transparency: `detected`, `isPaying`, `pricingTier`, `blocked`, `limited`, `mode`, `freeTierMaxItems`. |
| `finishedAt` | When the run wrapped up. |

#### Dataset views

The dataset ships with ready-made views: **Overview**, **Companies**, **Jobs**, and **People** — switch between them in the Console's Dataset tab.

***

### Webhooks — real-time delivery to your stack

Every row is **always** saved to the run's dataset. Optionally, set **Webhook URL** and each row is *also* delivered to your URL the moment it is collected — perfect for CRMs, Slack, Zapier, Make, n8n, or Google Sheets.

| Setting | Values |
|---|---|
| `webhookUrl` | Your receiving URL (any service that accepts a POST). |
| `webhookFormat` | `json` = the full record; `slack` = Slack message payload. |

**Payload example (`json` format — company row):**

```json
{
  "featureType": "company_profiles",
  "recordType": "company",
  "name": "Example Startup",
  "description": "AI for supply chains",
  "location": "San Francisco",
  "markets": ["Artificial Intelligence"],
  "companySize": "11-50",
  "website": "https://examplestartup.com/",
  "email": "hello@examplestartup.com",
  "phone": "(555) 555-5555",
  "linkedinUrl": "https://www.linkedin.com/company/examplestartup",
  "profileUrl": "https://wellfound.com/company/example-startup",
  "scrapedAt": "2026-09-25T12:00:00.000Z"
}
```

**Payload example (`json` format — job row):**

```json
{
  "featureType": "jobs_feed",
  "recordType": "job",
  "title": "Senior Software Engineer",
  "companyName": "Example Startup",
  "salary": "$150k – $200k • 0.1–0.3% equity",
  "location": "New York City",
  "workplaceType": "REMOTE",
  "remote": true,
  "jobUrl": "https://wellfound.com/jobs/1234567",
  "scrapedAt": "2026-09-25T12:00:00.000Z"
}
```

**Payload example (`slack` format):**

```json
{
  "text": ":office: *Example Startup*\nAI for supply chains\n*Location:* San Francisco\n*Size:* 11-50\n*Contact:* hello@examplestartup.com • (555) 555-5555\n*Website:* https://examplestartup.com/"
}
```

Failed deliveries are logged as a single warning (`Webhook delivery failed`) and **never** interrupt the run or dataset writes.

**Apify's own webhooks** (run succeeded / failed / aborted) are configured separately under your run's *Storage* → *Webhooks* in the Console and work with any Actor.

***

### Integrations — API, MCP, and automations

#### Apify API

Pull results programmatically as JSON, CSV, or Excel:

- `GET https://api.apify.com/v2/datasets/{DATASET_ID}/items` — your rows.
- Full reference: [Apify API v2](https://docs.apify.com/api/v2)

Start and monitor runs with `POST /v2/acts/{actorId}/runs` — schedule them with [Schedules](https://docs.apify.com/platform/schedules) for daily lead refreshes.

#### MCP usage (AI assistants)

Use the [Apify MCP server](https://docs.apify.com/platform/integrations/mcp) so AI assistants (Claude, ChatGPT, Cursor, and others) can **run this Actor and read its results directly in chat**:

1. Connect the Apify MCP server to your assistant ([setup guide](https://docs.apify.com/platform/integrations/mcp)).
2. Ask in natural language — the assistant calls this Actor with the right input.
3. Results come back as dataset items the assistant can summarize, tabulate, or chain onward.

Example prompts:

```
"Collect 20 AI startups that are hiring and list each with its website and LinkedIn."
"Run the Wellfound actor for fintech companies and turn the top 10 into a CSV with name, size, and location."
"Get the latest remote startup jobs with salaries and summarize the market in 5 bullets."
```

Typical MCP flow:

```
User: "Find 10 healthtech startups hiring engineers and draft outreach to each"
→ MCP runs Actor with companyKeywords=["healthtech"], feedMaxCompanies=10
→ MCP reads dataset items (name, website, email, linkedinUrl, roles)
→ Assistant drafts the outreach
```

#### LLM & RAG pipelines

Output is stable, flat JSON — ideal for ChatGPT, Claude, Gemini, LangChain, and LlamaIndex:

```json
{
  "recordType": "company",
  "name": "Example Startup",
  "markets": ["Artificial Intelligence"],
  "companySize": "11-50",
  "website": "https://examplestartup.com/",
  "email": "hello@examplestartup.com",
  "scrapedAt": "2026-09-25T12:00:00.000Z"
}
```

Workflow: run the Actor → fetch dataset items via the API → pass records to your LLM or vector store.

#### Native integrations

Push rows straight from the dataset to **Google Sheets, Airtable, Dropbox, Google Drive, Zapier, Make, and more** via [Apify Integrations](https://docs.apify.com/platform/integrations) — no code required.

***

### Pricing

- **Pay per exported result** (pay-per-event). You are charged for rows actually delivered to your dataset — not for time, and not for work you don't receive.
- **Spending limit respected:** if you set a max charge for a run in Apify, the Actor stops gracefully the moment that limit is reached, with a clear status message — never an overrun. See [Apify's pay-per-event docs](https://docs.apify.com/actors/publishing/monetize/pay-per-event#respect-user-spending-limits).
- **Cost levers:** cap `feedMaxCompanies` / `feedMaxJobs`, shorten the `peopleUrls` list, and turn off **Enable lead details** for lighter runs.

#### Free plan limitations

Apify free-plan accounts are limited to a **2-result sample per run**. The run log shows an explicit notice to **upgrade to a paid Apify plan for full, unlimited data**, and the run's `paywall` summary object records the decision (`detected`, `isPaying`, `pricingTier`, `blocked`, `limited`). This is a transparent platform plan restriction — runs finish cleanly; nothing errors. **Paid plans (Bronze and above) always get the full, normal output with no cap.**

***

### Performance & scale

- **Streaming:** rows are written to the dataset as they are finalized — memory stays light and a long run never loses completed work.
- **512 MB** default memory (configurable up to 2 GB).
- **10,000-second** run timeout — supports very long scheduled runs.
- **Up to 50,000 rows** per run (bounded by the per-section caps).
- **US residential connectivity** enabled by default for consistent results.

***

### FAQ

**Do I need a Wellfound account or API key?**
No. Just press Start.

**What does "real-time" mean here?**
Data is collected when your run executes — not served from a stale cache. Schedule runs to keep it fresh.

**I'm on the Apify free plan — why only 2 results?**
That's the platform free-plan restriction for this Actor: 2 results per run plus an upgrade notice in the log. Upgrade to any paid Apify plan for full, unlimited output. See [Free plan limitations](#free-plan-limitations).

**Will companies without lead details be dropped?**
Never. **Enable lead details** adds contact information when it's available — it never filters. Every collected company is exported, so run time and cost stay predictable.

**How does the email finder work?**
It checks every way a company publicly publishes its address: the homepage, its contact/about pages, structured contact data on the page, protected/encoded addresses, and human-written forms like `hello [at] company [dot] com`. Emails without a usable public presence stay empty — nothing is invented.

**Why does a run take a little longer with lead details on?**
Lead details add a little extra time per company. Turn the toggle off if you only need company facts and jobs.

**Do keywords filter the feed?**
No — keywords are recorded on every row as context tags for your own filtering. The feed is the live startup feed.

**How many companies and jobs can one run collect?**
Both sections paginate rotating listing pages, so volume scales with your caps rather than stopping at a fixed pass size — hundreds to thousands of companies or jobs per run are normal. Set **Max companies this run** and **Max jobs this run** to bound runs and cost; scheduled runs combined with dedupe on your side build history over time.

**Some profile URLs produced no row — why?**
Rows are only exported for profiles that were successfully collected. Unavailable or invalid profiles are skipped and noted in the log with a fixed, generic message.

**How do webhooks work?**
Set `webhookUrl` (and `webhookFormat`) and each row is delivered to your URL as it is collected, in addition to being saved to the dataset. Failures log one warning and never stop the run.

**Can AI assistants use this Actor?**
Yes — connect the [Apify MCP server](https://docs.apify.com/platform/integrations/mcp) and run it with natural language. See [MCP usage](#mcp-usage-ai-assistants).

**How do I control cost?**
Set the per-section caps (**Max companies this run**, **Max jobs this run**) as your hard caps, and set a run spending limit in Apify — the Actor honors it and stops gracefully.

**What does `spendingLimitReached: true` mean?**
Your configured max charge for that run was reached. The run wrapped up cleanly with everything collected so far saved.

**What do the `paywall` fields mean?**
They make the free-plan gate visible in the run output: whether it was detected, whether the account is paying, the pricing tier, and whether results were capped or blocked.

**Is the data public?**
The Actor collects publicly displayed information only.

**How often can I run it?**
As often as you like, bounded by your Apify plan and spending limits. Scheduled runs are supported.

**What does "Webhook delivery failed" in the log mean?**
Your receiving URL rejected or timed out the delivery. Dataset writes are unaffected — fix the receiver or re-deliver from your own queue.

**How do I export data?**
Console export (JSON/CSV/Excel), the [Apify API](https://docs.apify.com/api/v2), native integrations, or your webhook.

***

### Support

Open an issue on this Actor's GitHub repository or contact the developer through the Apify Store page.

### License

ISC

# Actor input Schema

## `enableCompanies` (type: `boolean`):

Collect startup company profiles: name, description, markets, size, location, badges, and open-jobs count. Companies come from the paginated startup directory, so volume is not limited to the live feed's pass size. On by default.

## `companyKeywords` (type: `array`):

Optional keywords stored on every company row so you can tag and filter your exports (e.g. "AI", "fintech"). They label the data; the directory itself covers all startups.

## `feedMaxCompanies` (type: `integer`):

Cap the number of company rows for this run. This is the main cost lever for companies. 0 = no cap (keep paginating until the directory is exhausted).

## `enableJobs` (type: `boolean`):

Collect open job listings joined to their companies: title, salary, workplace type, location, and posting date. Job volume scales with this run's cap — pages of listings across roles and locations are collected until the cap is met. Off by default — turn on to include jobs.

## `jobKeywords` (type: `array`):

Optional keywords stored on every job row so you can tag and filter your exports (e.g. "engineering", "sales"). Keywords that clearly name a job family (engineering, sales, design, data, support…) also steer which role listings are collected first.

## `feedMaxJobs` (type: `integer`):

Cap the number of job rows for this run. This is the main cost lever for jobs. 0 = no cap (keep collecting until listings are exhausted).

## `feedRemoteOnly` (type: `boolean`):

Keep only job rows marked remote.

## `enablePeopleProfiles` (type: `boolean`):

Collect public member profiles from the profile URLs below: name, headline, avatar, and public social links when displayed. Off by default — turn on and add profile URLs.

## `peopleUrls` (type: `array`):

Public member profile URLs to collect (only used when People profiles is enabled).

## `enableLeadDetails` (type: `boolean`):

Add lead contact details to each company row — website, email, phone, social profiles — when available. Every company is still exported even when no lead details are found (nothing is dropped). Adds a little extra time per company. On by default.

## `webhookUrl` (type: `string`):

Optional. Every record is always saved to the run's dataset — this webhook is an ADDITIONAL real-time push. When set, each new record is also delivered to this URL (CRM, Slack incoming webhook, Zapier, Make, Google Sheets).

## `webhookFormat` (type: `string`):

json = full record object; slack = Slack-friendly message payload.

## `proxyConfiguration` (type: `object`):

Apify residential proxy (US) is enabled by default for reliable collection. Change here only if you need a different country or custom proxy URLs.

## Actor input object example

```json
{
  "enableCompanies": true,
  "companyKeywords": [],
  "feedMaxCompanies": 10,
  "enableJobs": false,
  "jobKeywords": [],
  "feedMaxJobs": 10,
  "feedRemoteOnly": false,
  "enablePeopleProfiles": false,
  "peopleUrls": [],
  "enableLeadDetails": true,
  "webhookUrl": "",
  "webhookFormat": "json",
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `allResults` (type: `string`):

Complete dataset with every field from all enabled features in this run.

## `companies` (type: `string`):

Startup company rows from the startup feed.

## `jobs` (type: `string`):

Job rows from the startup feed.

## `people` (type: `string`):

People profile rows from profile URLs.

## `runSummary` (type: `string`):

Per-run metadata: enabled features, exported record count, non-fatal errors, spending-limit status, and the paywall object (detected, isPaying, pricingTier, blocked, limited).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "enableCompanies": true,
    "companyKeywords": [],
    "feedMaxCompanies": 10,
    "enableJobs": false,
    "jobKeywords": [],
    "feedMaxJobs": 10,
    "enablePeopleProfiles": false,
    "enableLeadDetails": true,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("b2b_leads/wellfound-real-time-data-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "enableCompanies": True,
    "companyKeywords": [],
    "feedMaxCompanies": 10,
    "enableJobs": False,
    "jobKeywords": [],
    "feedMaxJobs": 10,
    "enablePeopleProfiles": False,
    "enableLeadDetails": True,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("b2b_leads/wellfound-real-time-data-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "enableCompanies": true,
  "companyKeywords": [],
  "feedMaxCompanies": 10,
  "enableJobs": false,
  "jobKeywords": [],
  "feedMaxJobs": 10,
  "enablePeopleProfiles": false,
  "enableLeadDetails": true,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call b2b_leads/wellfound-real-time-data-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,b2b_leads/wellfound-real-time-data-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3kwZz35ZfNaR1rbdU/builds/7wfUYZxqU9Ysk1FLH/openapi.json
