Wellfound Scraper — Startup & Job Data
Pricing
from $2.00 / 1,000 results
Wellfound Scraper — Startup & Job Data
Turn Wellfound (AngelList) into clean, structured startup data. Get company profiles with hiring signals and open jobs with salaries, plus websites, emails, and socials when available. Rows stream to your dataset in real time — no setup, pay per result.
Pricing
from $2.00 / 1,000 results
Rating
0.0
(0)
Developer
Emmanuel
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Wellfound Real-Time Data Scraper — Startup Companies, Jobs & People at Scale
Turn Wellfound (AngelList) into clean, structured startup data. One run collects the live startup feed — company profiles with descriptions, markets, size, and locations, plus their open job listings — and public people profiles, with optional lead details (website, email, phone, social profiles) added to each company when available. Rows stream into your Apify dataset as they are finalized, so even long runs stay light and nothing is lost mid-run.
No API keys. No spreadsheets. No cleanup. Just press Start.
⚡ Real-time streaming output • 🏢 Companies + jobs in one run • 👤 People profiles • 🎯 Lead details without filtering • 💸 Pay only for what you export
✅ Pricing & free-plan note (read first)
This Actor uses Apify's pay-per-result model — you're charged per exported row (see Pricing).
- Apify free plan: runs are limited to a 2-result sample per run, and the log tells you to upgrade. This is an Apify platform plan restriction, applied transparently — it is not an error.
- Any paid Apify plan (Bronze and above): full, uncapped output. You get exactly what your input asks for.
- Every run also honors your max run charge (spending limit): the Actor finishes gracefully the moment your configured limit is reached — no more charges than you approved.
Details: Free plan limitations · Pricing
Why this Actor
| Wellfound Real-Time Data Scraper | Typical scraper | |
|---|---|---|
| Output | Flat, stable JSON field names | Messy, inconsistent |
| Streaming | Rows saved live during the run | Results only at the end |
| Memory | 512 MB default | 2–4 GB+ |
| Lead details | Added when available — never filters | Filters and drops rows |
| Setup | Organized input UI — run immediately | Needs tuning |
| Freshness | Collected at run time ("real-time") | Cached or stale |
What you get
Every row includes featureType, recordType, and scrapedAt so you can filter, join, and pipe into any workflow.
Company rows (recordType: "company")
| Group | Fields |
|---|---|
| Identity | companyId, name, profileUrl, logoUrl |
| About | description (tagline, plus website description when found), location, markets[], companySize |
| Signals | badges[] (e.g. actively hiring), openJobsCount (open roles when shown) |
| Lead details | website, email, emails[], phone, phones[], linkedinUrl, twitterUrl, facebookUrl, instagramUrl, address |
| Structured view | leadDetails (object with all lead fields), searchQuery, source, scrapedAt |
Job rows (recordType: "job")
| Group | Fields |
|---|---|
| Identity | jobId, title, jobUrl, companyName, companyProfileUrl |
| Compensation & location | salary (e.g. $120k – $160k • 0.1–0.3% equity), location, workplaceType (ONSITE / REMOTE / HYBRID), remote |
| Timing | postedAt, searchQuery, source, scrapedAt |
People rows (recordType: "person")
| Group | Fields |
|---|---|
| Identity | name, title (headline), profileUrl, avatarUrl |
| Public links | linkedinUrl, twitterUrl, githubUrl, email (when displayed) |
| Structured view | leadDetails, source, scrapedAt |
Missing values are null — field names are stable across runs, so your transforms won't break.
Features
🏢 Companies (on by default)
Collect startup company profiles from the paginated startup directory: name, description, markets, size, location, badges, and the company's open-jobs count. Directory pages rotate, so a run is not limited to the live feed's pass size — collect 10 or 10,000 companies. Cap the number with Max companies this run (0 = keep going until the directory is exhausted).
💼 Jobs (off by default)
Collect open job listings joined to their company: title, salary, workplace type (onsite / remote), location, and posting date. Listings are collected across role families and locations in rotating pages, so job volume scales with your cap instead of stopping at a fixed pass size — the live feed pass is only used as a top-up. Cap the number with Max jobs this run, or keep only remote roles with Remote jobs only. Job keywords that clearly name a job family (engineering, sales, design, data, support…) steer which role listings are collected first.
🏷️ Keywords as context
Each section carries its own keywords, recorded on every exported row so you can tag and filter your own exports (e.g. AI for companies, engineering for jobs). For jobs, a keyword that names a job family also influences which listings are collected first; otherwise keywords label the data — they never filter it.
🎯 Enable lead details (on by default) — enrichment, never a filter
Adds lead contact details to each company row — website, email, phone, social profiles — when they are available. The email finder covers every way companies publish addresses: contact and about pages, structured contact data, protected/encoded addresses, and human-written forms like hello [at] company [dot] com.
Every company is still exported even when no lead details are found. Nothing is ever dropped for lacking contact info, so the time and cost of a run stay predictable (this is deliberate: you always know what a run of N companies costs). Adds a little extra time per company.
👤 People profiles
Paste public member profile URLs (wellfound.com/u/...) and collect people records as their own rows — name, headline, avatar, and the public social links the member displays. Ideal for founder and recruiting outreach lists.
⚙️ Output controls
Cap companies and jobs independently, keep only remote jobs, cap total items, and deliver rows to a webhook in real time (see Webhooks).
Who it's for
- Recruiters & talent teams — companies that are hiring right now, with roles, salaries, and direct contact channels.
- Sales & BD teams selling to startups — filter by market, size, and location; reach founders and hiring managers with verified contact details.
- Lead-generation agencies — a steady stream of fresh startup leads with website, email, and phone when available, streamed into your CRM.
- Growth & partnerships managers — track fast-growing companies by badge-worthy signals (actively hiring, growing fast) in your space.
- Investors & scouting analysts — market maps of active startups with markets, size, location, and descriptions for deal screening.
- Job seekers & career coaches — structured lists of open startup roles with salary ranges and workplace types.
- Market researchers & analysts — joinable JSON across runs for startup-ecosystem studies.
- AI & automation builders — clean rows for LLM prompts, scoring agents, and enrichment pipelines via the Apify API and MCP.
- CRM & data teams — stable flat JSON that drops straight into Postgres, BigQuery, Airtable, or Sheets.
Use cases
- Hiring-company prospecting — every run gives you companies with live job openings; combine with lead details for direct outreach.
- Recruiting-fee hunting — export roles with salary bands, find the hiring company's website and contact channels, pitch your services.
- SaaS lead gen by market — tag rows with your keywords (e.g.
AI,fintech,healthtech), route to your CRM by webhook, and let sales filter by market and size. - Foundering-target monitoring — schedule daily runs on your keywords; new companies appearing in the feed are fresh leads.
- Remote-job tracking — enable Remote jobs only and stream remote roles to Slack for a candidate community.
- Investor deal flow — collect startups in your thesis markets weekly;
markets[],companySize, anddescriptionfeed your screener. - Sales-trigger monitoring — companies that just posted roles are companies that are growing: feed them to your outbound sequences first.
- Talent-market mapping — compare salary ranges and workplace types across locations and markets over time.
- Enrichment pipeline input — take company rows into your own enrichment stack; website/socials are already filled when available.
- AI company research — feed descriptions, markets, and job data to an LLM to summarize and score startups automatically.
Quick start
- Open the Actor in Apify Console.
- Leave the defaults — the startup feed with lead details is on.
- Optionally add keywords to tag rows (e.g.
AI,fintech). - Set Max companies this run (Companies section) and Max jobs this run (Jobs section) — defaults: 10 each. Turn either section off if you only need the other.
- Press Start and watch rows stream into the dataset.
- Export as JSON / CSV / Excel, pull via the Apify API, wire up a webhook, or use MCP.
Example — startup feed with lead details (default)
{"enableCompanies": true,"companyKeywords": ["AI"],"feedMaxCompanies": 10,"enableJobs": true,"jobKeywords": ["engineering"],"feedMaxJobs": 10,"enableLeadDetails": true}
Example — remote jobs only, larger run
{"enableCompanies": false,"enableJobs": true,"feedRemoteOnly": true,"feedMaxJobs": 100}
Example — companies with lead details only
{"enableCompanies": true,"enableJobs": false,"feedMaxCompanies": 50,"enableLeadDetails": true}
Example — people profiles from URLs
{"enableCompanies": false,"enableJobs": false,"enablePeopleProfiles": true,"peopleUrls": ["https://wellfound.com/u/example-founder","https://wellfound.com/u/example-cto"]}
Example — results to a webhook (Slack-friendly)
{"enableCompanies": true,"feedMaxCompanies": 25,"webhookUrl": "https://hooks.slack.com/services/T000/B000/XXXX","webhookFormat": "slack"}
Input reference
| Field | Type | Default | Description |
|---|---|---|---|
enableCompanies | boolean | true | Companies section: collect startup company profiles from the paginated startup directory. |
companyKeywords | string[] | [] | Keywords recorded on every company row as context tags (e.g. AI, fintech). |
feedMaxCompanies | integer | 10 | Cap company rows this run — the main cost lever for companies. 0 = no cap (keep paginating until the directory is exhausted). |
enableJobs | boolean | false | Jobs section: collect open job listings joined to their companies. |
jobKeywords | string[] | [] | Keywords recorded on every job row as context tags (e.g. engineering, sales); job-family keywords also steer which listings are collected first. |
feedMaxJobs | integer | 10 | Cap job rows this run — the main cost lever for jobs. 0 = no cap (keep collecting until listings are exhausted). |
feedRemoteOnly | boolean | false | Keep only job rows marked remote (Jobs section). |
enablePeopleProfiles | boolean | false | Collect public member profiles from peopleUrls. |
peopleUrls | string[] | [] | Public member profile URLs (wellfound.com/u/...). |
enableLeadDetails | boolean | true | Add lead contact details to company rows when available. Never filters — every company is exported either way. Adds a little extra time per company. Free plan: capped at 2 results per run — see Free plan limitations. |
webhookUrl | string | "" | Optional. Each row is saved to the dataset and delivered to this URL in real time. |
webhookFormat | select | json | json (full record) or slack (message payload). |
proxyConfiguration | proxy | US residential | Apify proxy settings (US residential on by default). |
Output reference
Every exported record is a flat JSON object in the run's dataset, tagged by recordType:
recordType | Row shape |
|---|---|
company | Startup company from the feed (with lead details when available). |
job | Job listing from the feed, joined to its company. |
person | Public member profile from a profile URL. |
Company row — full field list
| Field | Type | Description |
|---|---|---|
featureType | string | Always company_profiles. |
recordType | string | Always company. |
companyId | string | null | Platform company ID. |
name | string | null | Company name. |
profileUrl | string | null | Company profile URL. |
description | string | null | Tagline / short description (longer website description when found). |
location | string | null | Primary location(s), e.g. New York City. |
markets | string[] | Market / industry tags, e.g. ["Artificial Intelligence", "SaaS"]. |
companySize | string | null | Size band, e.g. 11-50. |
badges | string[] | Public company signals, e.g. ACTIVELY_HIRING, YC, GROWING_FAST. |
openJobsCount | number | null | Open roles the company lists on the platform, when shown. |
logoUrl | string | null | Logo image URL. |
website | string | null | Company website (lead details). |
email | string | null | Public contact email (lead details). |
emails | string[] | Up to 3 public emails (lead details). |
phone | string | null | Public phone in (555) 555-5555 format (lead details). |
phones | string[] | Up to 5 public phones (lead details). |
linkedinUrl / twitterUrl / facebookUrl / instagramUrl | string | null | Social profiles (lead details). |
address | string | null | Public postal address (lead details). |
leadDetails | object | null | All lead fields grouped in one object (same values as the flat fields). |
searchQuery | string | null | Keywords this run tagged the row with. |
source | string | Which stage produced the row. |
scrapedAt | string | ISO 8601 UTC timestamp of collection. |
Job row — full field list
| Field | Type | Description |
|---|---|---|
featureType / recordType | string | jobs_feed / job (featureType is a stable data tag; the input sections are Companies and Jobs). |
jobId | string | null | Platform job ID. |
title | string | null | Job title. |
jobUrl | string | null | Job listing URL. |
companyName | string | null | Hiring company name. |
companyProfileUrl | string | null | Company profile URL. |
salary | string | null | Compensation string, e.g. $120k – $160k • 0.1–0.3% equity. |
location | string | null | Job location(s), e.g. San Francisco, New York City. |
workplaceType | string | null | ONSITE, REMOTE, or HYBRID. |
remote | boolean | null | True when the role is remote. |
postedAt | string | null | ISO 8601 UTC posting date. |
searchQuery / source / scrapedAt | — | Same as company rows. |
People row — full field list
| Field | Type | Description |
|---|---|---|
featureType / recordType | string | people_profiles / person. |
name | string | null | Member name. |
title | string | null | Headline. |
profileUrl | string | null | Profile URL you provided. |
avatarUrl | string | null | Avatar image URL. |
linkedinUrl / twitterUrl / githubUrl | string | null | Public social links displayed on the profile. |
email | string | null | Public email displayed on the profile. |
leadDetails / source / scrapedAt | — | Same as other rows. |
Run summary (OUTPUT key-value store)
Each run also writes a summary you can read via the Apify API:
{"totalPushed": 120,"enabledFeatures": ["companies", "jobs", "lead_details"],"errors": [],"spendingLimitReached": false,"paywall": {"detected": true,"isPaying": true,"pricingTier": "SILVER","blocked": false,"limited": false,"mode": "limit","freeTierMaxItems": null},"finishedAt": "2026-09-25T12:00:00.000Z"}
| Summary field | Meaning |
|---|---|
totalPushed | Rows exported this run. |
enabledFeatures | Features that ran (e.g. companies, jobs, people_profiles, lead_details). |
errors | Non-fatal failures, expressed as fixed user-facing sentences. |
spendingLimitReached | true when your max run charge was reached — the run then finishes gracefully. |
paywall | Free-plan gate transparency: detected, isPaying, pricingTier, blocked, limited, mode, freeTierMaxItems. |
finishedAt | When the run wrapped up. |
Dataset views
The dataset ships with ready-made views: Overview, Companies, Jobs, and People — switch between them in the Console's Dataset tab.
Webhooks — real-time delivery to your stack
Every row is always saved to the run's dataset. Optionally, set Webhook URL and each row is also delivered to your URL the moment it is collected — perfect for CRMs, Slack, Zapier, Make, n8n, or Google Sheets.
| Setting | Values |
|---|---|
webhookUrl | Your receiving URL (any service that accepts a POST). |
webhookFormat | json = the full record; slack = Slack message payload. |
Payload example (json format — company row):
{"featureType": "company_profiles","recordType": "company","name": "Example Startup","description": "AI for supply chains","location": "San Francisco","markets": ["Artificial Intelligence"],"companySize": "11-50","website": "https://examplestartup.com/","email": "hello@examplestartup.com","phone": "(555) 555-5555","linkedinUrl": "https://www.linkedin.com/company/examplestartup","profileUrl": "https://wellfound.com/company/example-startup","scrapedAt": "2026-09-25T12:00:00.000Z"}
Payload example (json format — job row):
{"featureType": "jobs_feed","recordType": "job","title": "Senior Software Engineer","companyName": "Example Startup","salary": "$150k – $200k • 0.1–0.3% equity","location": "New York City","workplaceType": "REMOTE","remote": true,"jobUrl": "https://wellfound.com/jobs/1234567","scrapedAt": "2026-09-25T12:00:00.000Z"}
Payload example (slack format):
{"text": ":office: *Example Startup*\nAI for supply chains\n*Location:* San Francisco\n*Size:* 11-50\n*Contact:* hello@examplestartup.com • (555) 555-5555\n*Website:* https://examplestartup.com/"}
Failed deliveries are logged as a single warning (Webhook delivery failed) and never interrupt the run or dataset writes.
Apify's own webhooks (run succeeded / failed / aborted) are configured separately under your run's Storage → Webhooks in the Console and work with any Actor.
Integrations — API, MCP, and automations
Apify API
Pull results programmatically as JSON, CSV, or Excel:
GET https://api.apify.com/v2/datasets/{DATASET_ID}/items— your rows.- Full reference: Apify API v2
Start and monitor runs with POST /v2/acts/{actorId}/runs — schedule them with Schedules for daily lead refreshes.
MCP usage (AI assistants)
Use the Apify MCP server so AI assistants (Claude, ChatGPT, Cursor, and others) can run this Actor and read its results directly in chat:
- Connect the Apify MCP server to your assistant (setup guide).
- Ask in natural language — the assistant calls this Actor with the right input.
- Results come back as dataset items the assistant can summarize, tabulate, or chain onward.
Example prompts:
"Collect 20 AI startups that are hiring and list each with its website and LinkedIn.""Run the Wellfound actor for fintech companies and turn the top 10 into a CSV with name, size, and location.""Get the latest remote startup jobs with salaries and summarize the market in 5 bullets."
Typical MCP flow:
User: "Find 10 healthtech startups hiring engineers and draft outreach to each"→ MCP runs Actor with companyKeywords=["healthtech"], feedMaxCompanies=10→ MCP reads dataset items (name, website, email, linkedinUrl, roles)→ Assistant drafts the outreach
LLM & RAG pipelines
Output is stable, flat JSON — ideal for ChatGPT, Claude, Gemini, LangChain, and LlamaIndex:
{"recordType": "company","name": "Example Startup","markets": ["Artificial Intelligence"],"companySize": "11-50","website": "https://examplestartup.com/","email": "hello@examplestartup.com","scrapedAt": "2026-09-25T12:00:00.000Z"}
Workflow: run the Actor → fetch dataset items via the API → pass records to your LLM or vector store.
Native integrations
Push rows straight from the dataset to Google Sheets, Airtable, Dropbox, Google Drive, Zapier, Make, and more via Apify Integrations — no code required.
Pricing
- Pay per exported result (pay-per-event). You are charged for rows actually delivered to your dataset — not for time, and not for work you don't receive.
- Spending limit respected: if you set a max charge for a run in Apify, the Actor stops gracefully the moment that limit is reached, with a clear status message — never an overrun. See Apify's pay-per-event docs.
- Cost levers: cap
feedMaxCompanies/feedMaxJobs, shorten thepeopleUrlslist, and turn off Enable lead details for lighter runs.
Free plan limitations
Apify free-plan accounts are limited to a 2-result sample per run. The run log shows an explicit notice to upgrade to a paid Apify plan for full, unlimited data, and the run's paywall summary object records the decision (detected, isPaying, pricingTier, blocked, limited). This is a transparent platform plan restriction — runs finish cleanly; nothing errors. Paid plans (Bronze and above) always get the full, normal output with no cap.
Performance & scale
- Streaming: rows are written to the dataset as they are finalized — memory stays light and a long run never loses completed work.
- 512 MB default memory (configurable up to 2 GB).
- 10,000-second run timeout — supports very long scheduled runs.
- Up to 50,000 rows per run (bounded by the per-section caps).
- US residential connectivity enabled by default for consistent results.
FAQ
Do I need a Wellfound account or API key? No. Just press Start.
What does "real-time" mean here? Data is collected when your run executes — not served from a stale cache. Schedule runs to keep it fresh.
I'm on the Apify free plan — why only 2 results? That's the platform free-plan restriction for this Actor: 2 results per run plus an upgrade notice in the log. Upgrade to any paid Apify plan for full, unlimited output. See Free plan limitations.
Will companies without lead details be dropped? Never. Enable lead details adds contact information when it's available — it never filters. Every collected company is exported, so run time and cost stay predictable.
How does the email finder work?
It checks every way a company publicly publishes its address: the homepage, its contact/about pages, structured contact data on the page, protected/encoded addresses, and human-written forms like hello [at] company [dot] com. Emails without a usable public presence stay empty — nothing is invented.
Why does a run take a little longer with lead details on? Lead details add a little extra time per company. Turn the toggle off if you only need company facts and jobs.
Do keywords filter the feed? No — keywords are recorded on every row as context tags for your own filtering. The feed is the live startup feed.
How many companies and jobs can one run collect? Both sections paginate rotating listing pages, so volume scales with your caps rather than stopping at a fixed pass size — hundreds to thousands of companies or jobs per run are normal. Set Max companies this run and Max jobs this run to bound runs and cost; scheduled runs combined with dedupe on your side build history over time.
Some profile URLs produced no row — why? Rows are only exported for profiles that were successfully collected. Unavailable or invalid profiles are skipped and noted in the log with a fixed, generic message.
How do webhooks work?
Set webhookUrl (and webhookFormat) and each row is delivered to your URL as it is collected, in addition to being saved to the dataset. Failures log one warning and never stop the run.
Can AI assistants use this Actor? Yes — connect the Apify MCP server and run it with natural language. See MCP usage.
How do I control cost? Set the per-section caps (Max companies this run, Max jobs this run) as your hard caps, and set a run spending limit in Apify — the Actor honors it and stops gracefully.
What does spendingLimitReached: true mean?
Your configured max charge for that run was reached. The run wrapped up cleanly with everything collected so far saved.
What do the paywall fields mean?
They make the free-plan gate visible in the run output: whether it was detected, whether the account is paying, the pricing tier, and whether results were capped or blocked.
Is the data public? The Actor collects publicly displayed information only.
How often can I run it? As often as you like, bounded by your Apify plan and spending limits. Scheduled runs are supported.
What does "Webhook delivery failed" in the log mean? Your receiving URL rejected or timed out the delivery. Dataset writes are unaffected — fix the receiver or re-deliver from your own queue.
How do I export data? Console export (JSON/CSV/Excel), the Apify API, native integrations, or your webhook.
Support
Open an issue on this Actor's GitHub repository or contact the developer through the Apify Store page.
License
ISC