Wellfound Scraper — Startup & Job Data avatar

Wellfound Scraper — Startup & Job Data

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Wellfound Scraper — Startup & Job Data

Wellfound Scraper — Startup & Job Data

Turn Wellfound (AngelList) into clean, structured startup data. Get company profiles with hiring signals and open jobs with salaries, plus websites, emails, and socials when available. Rows stream to your dataset in real time — no setup, pay per result.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Emmanuel

Emmanuel

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Wellfound Real-Time Data Scraper — Startup Companies, Jobs & People at Scale

Turn Wellfound (AngelList) into clean, structured startup data. One run collects the live startup feed — company profiles with descriptions, markets, size, and locations, plus their open job listings — and public people profiles, with optional lead details (website, email, phone, social profiles) added to each company when available. Rows stream into your Apify dataset as they are finalized, so even long runs stay light and nothing is lost mid-run.

No API keys. No spreadsheets. No cleanup. Just press Start.

⚡ Real-time streaming output • 🏢 Companies + jobs in one run • 👤 People profiles • 🎯 Lead details without filtering • 💸 Pay only for what you export


✅ Pricing & free-plan note (read first)

This Actor uses Apify's pay-per-result model — you're charged per exported row (see Pricing).

  • Apify free plan: runs are limited to a 2-result sample per run, and the log tells you to upgrade. This is an Apify platform plan restriction, applied transparently — it is not an error.
  • Any paid Apify plan (Bronze and above): full, uncapped output. You get exactly what your input asks for.
  • Every run also honors your max run charge (spending limit): the Actor finishes gracefully the moment your configured limit is reached — no more charges than you approved.

Details: Free plan limitations · Pricing


Why this Actor

Wellfound Real-Time Data ScraperTypical scraper
OutputFlat, stable JSON field namesMessy, inconsistent
StreamingRows saved live during the runResults only at the end
Memory512 MB default2–4 GB+
Lead detailsAdded when available — never filtersFilters and drops rows
SetupOrganized input UI — run immediatelyNeeds tuning
FreshnessCollected at run time ("real-time")Cached or stale

What you get

Every row includes featureType, recordType, and scrapedAt so you can filter, join, and pipe into any workflow.

Company rows (recordType: "company")

GroupFields
IdentitycompanyId, name, profileUrl, logoUrl
Aboutdescription (tagline, plus website description when found), location, markets[], companySize
Signalsbadges[] (e.g. actively hiring), openJobsCount (open roles when shown)
Lead detailswebsite, email, emails[], phone, phones[], linkedinUrl, twitterUrl, facebookUrl, instagramUrl, address
Structured viewleadDetails (object with all lead fields), searchQuery, source, scrapedAt

Job rows (recordType: "job")

GroupFields
IdentityjobId, title, jobUrl, companyName, companyProfileUrl
Compensation & locationsalary (e.g. $120k – $160k • 0.1–0.3% equity), location, workplaceType (ONSITE / REMOTE / HYBRID), remote
TimingpostedAt, searchQuery, source, scrapedAt

People rows (recordType: "person")

GroupFields
Identityname, title (headline), profileUrl, avatarUrl
Public linkslinkedinUrl, twitterUrl, githubUrl, email (when displayed)
Structured viewleadDetails, source, scrapedAt

Missing values are null — field names are stable across runs, so your transforms won't break.


Features

🏢 Companies (on by default)

Collect startup company profiles from the paginated startup directory: name, description, markets, size, location, badges, and the company's open-jobs count. Directory pages rotate, so a run is not limited to the live feed's pass size — collect 10 or 10,000 companies. Cap the number with Max companies this run (0 = keep going until the directory is exhausted).

💼 Jobs (off by default)

Collect open job listings joined to their company: title, salary, workplace type (onsite / remote), location, and posting date. Listings are collected across role families and locations in rotating pages, so job volume scales with your cap instead of stopping at a fixed pass size — the live feed pass is only used as a top-up. Cap the number with Max jobs this run, or keep only remote roles with Remote jobs only. Job keywords that clearly name a job family (engineering, sales, design, data, support…) steer which role listings are collected first.

🏷️ Keywords as context

Each section carries its own keywords, recorded on every exported row so you can tag and filter your own exports (e.g. AI for companies, engineering for jobs). For jobs, a keyword that names a job family also influences which listings are collected first; otherwise keywords label the data — they never filter it.

🎯 Enable lead details (on by default) — enrichment, never a filter

Adds lead contact details to each company row — website, email, phone, social profiles — when they are available. The email finder covers every way companies publish addresses: contact and about pages, structured contact data, protected/encoded addresses, and human-written forms like hello [at] company [dot] com.

Every company is still exported even when no lead details are found. Nothing is ever dropped for lacking contact info, so the time and cost of a run stay predictable (this is deliberate: you always know what a run of N companies costs). Adds a little extra time per company.

👤 People profiles

Paste public member profile URLs (wellfound.com/u/...) and collect people records as their own rows — name, headline, avatar, and the public social links the member displays. Ideal for founder and recruiting outreach lists.

⚙️ Output controls

Cap companies and jobs independently, keep only remote jobs, cap total items, and deliver rows to a webhook in real time (see Webhooks).


Who it's for

  • Recruiters & talent teams — companies that are hiring right now, with roles, salaries, and direct contact channels.
  • Sales & BD teams selling to startups — filter by market, size, and location; reach founders and hiring managers with verified contact details.
  • Lead-generation agencies — a steady stream of fresh startup leads with website, email, and phone when available, streamed into your CRM.
  • Growth & partnerships managers — track fast-growing companies by badge-worthy signals (actively hiring, growing fast) in your space.
  • Investors & scouting analysts — market maps of active startups with markets, size, location, and descriptions for deal screening.
  • Job seekers & career coaches — structured lists of open startup roles with salary ranges and workplace types.
  • Market researchers & analysts — joinable JSON across runs for startup-ecosystem studies.
  • AI & automation builders — clean rows for LLM prompts, scoring agents, and enrichment pipelines via the Apify API and MCP.
  • CRM & data teams — stable flat JSON that drops straight into Postgres, BigQuery, Airtable, or Sheets.

Use cases

  • Hiring-company prospecting — every run gives you companies with live job openings; combine with lead details for direct outreach.
  • Recruiting-fee hunting — export roles with salary bands, find the hiring company's website and contact channels, pitch your services.
  • SaaS lead gen by market — tag rows with your keywords (e.g. AI, fintech, healthtech), route to your CRM by webhook, and let sales filter by market and size.
  • Foundering-target monitoring — schedule daily runs on your keywords; new companies appearing in the feed are fresh leads.
  • Remote-job tracking — enable Remote jobs only and stream remote roles to Slack for a candidate community.
  • Investor deal flow — collect startups in your thesis markets weekly; markets[], companySize, and description feed your screener.
  • Sales-trigger monitoring — companies that just posted roles are companies that are growing: feed them to your outbound sequences first.
  • Talent-market mapping — compare salary ranges and workplace types across locations and markets over time.
  • Enrichment pipeline input — take company rows into your own enrichment stack; website/socials are already filled when available.
  • AI company research — feed descriptions, markets, and job data to an LLM to summarize and score startups automatically.

Quick start

  1. Open the Actor in Apify Console.
  2. Leave the defaults — the startup feed with lead details is on.
  3. Optionally add keywords to tag rows (e.g. AI, fintech).
  4. Set Max companies this run (Companies section) and Max jobs this run (Jobs section) — defaults: 10 each. Turn either section off if you only need the other.
  5. Press Start and watch rows stream into the dataset.
  6. Export as JSON / CSV / Excel, pull via the Apify API, wire up a webhook, or use MCP.

Example — startup feed with lead details (default)

{
"enableCompanies": true,
"companyKeywords": ["AI"],
"feedMaxCompanies": 10,
"enableJobs": true,
"jobKeywords": ["engineering"],
"feedMaxJobs": 10,
"enableLeadDetails": true
}

Example — remote jobs only, larger run

{
"enableCompanies": false,
"enableJobs": true,
"feedRemoteOnly": true,
"feedMaxJobs": 100
}

Example — companies with lead details only

{
"enableCompanies": true,
"enableJobs": false,
"feedMaxCompanies": 50,
"enableLeadDetails": true
}

Example — people profiles from URLs

{
"enableCompanies": false,
"enableJobs": false,
"enablePeopleProfiles": true,
"peopleUrls": [
"https://wellfound.com/u/example-founder",
"https://wellfound.com/u/example-cto"
]
}

Example — results to a webhook (Slack-friendly)

{
"enableCompanies": true,
"feedMaxCompanies": 25,
"webhookUrl": "https://hooks.slack.com/services/T000/B000/XXXX",
"webhookFormat": "slack"
}

Input reference

FieldTypeDefaultDescription
enableCompaniesbooleantrueCompanies section: collect startup company profiles from the paginated startup directory.
companyKeywordsstring[][]Keywords recorded on every company row as context tags (e.g. AI, fintech).
feedMaxCompaniesinteger10Cap company rows this run — the main cost lever for companies. 0 = no cap (keep paginating until the directory is exhausted).
enableJobsbooleanfalseJobs section: collect open job listings joined to their companies.
jobKeywordsstring[][]Keywords recorded on every job row as context tags (e.g. engineering, sales); job-family keywords also steer which listings are collected first.
feedMaxJobsinteger10Cap job rows this run — the main cost lever for jobs. 0 = no cap (keep collecting until listings are exhausted).
feedRemoteOnlybooleanfalseKeep only job rows marked remote (Jobs section).
enablePeopleProfilesbooleanfalseCollect public member profiles from peopleUrls.
peopleUrlsstring[][]Public member profile URLs (wellfound.com/u/...).
enableLeadDetailsbooleantrueAdd lead contact details to company rows when available. Never filters — every company is exported either way. Adds a little extra time per company. Free plan: capped at 2 results per run — see Free plan limitations.
webhookUrlstring""Optional. Each row is saved to the dataset and delivered to this URL in real time.
webhookFormatselectjsonjson (full record) or slack (message payload).
proxyConfigurationproxyUS residentialApify proxy settings (US residential on by default).

Output reference

Every exported record is a flat JSON object in the run's dataset, tagged by recordType:

recordTypeRow shape
companyStartup company from the feed (with lead details when available).
jobJob listing from the feed, joined to its company.
personPublic member profile from a profile URL.

Company row — full field list

FieldTypeDescription
featureTypestringAlways company_profiles.
recordTypestringAlways company.
companyIdstring | nullPlatform company ID.
namestring | nullCompany name.
profileUrlstring | nullCompany profile URL.
descriptionstring | nullTagline / short description (longer website description when found).
locationstring | nullPrimary location(s), e.g. New York City.
marketsstring[]Market / industry tags, e.g. ["Artificial Intelligence", "SaaS"].
companySizestring | nullSize band, e.g. 11-50.
badgesstring[]Public company signals, e.g. ACTIVELY_HIRING, YC, GROWING_FAST.
openJobsCountnumber | nullOpen roles the company lists on the platform, when shown.
logoUrlstring | nullLogo image URL.
websitestring | nullCompany website (lead details).
emailstring | nullPublic contact email (lead details).
emailsstring[]Up to 3 public emails (lead details).
phonestring | nullPublic phone in (555) 555-5555 format (lead details).
phonesstring[]Up to 5 public phones (lead details).
linkedinUrl / twitterUrl / facebookUrl / instagramUrlstring | nullSocial profiles (lead details).
addressstring | nullPublic postal address (lead details).
leadDetailsobject | nullAll lead fields grouped in one object (same values as the flat fields).
searchQuerystring | nullKeywords this run tagged the row with.
sourcestringWhich stage produced the row.
scrapedAtstringISO 8601 UTC timestamp of collection.

Job row — full field list

FieldTypeDescription
featureType / recordTypestringjobs_feed / job (featureType is a stable data tag; the input sections are Companies and Jobs).
jobIdstring | nullPlatform job ID.
titlestring | nullJob title.
jobUrlstring | nullJob listing URL.
companyNamestring | nullHiring company name.
companyProfileUrlstring | nullCompany profile URL.
salarystring | nullCompensation string, e.g. $120k – $160k • 0.1–0.3% equity.
locationstring | nullJob location(s), e.g. San Francisco, New York City.
workplaceTypestring | nullONSITE, REMOTE, or HYBRID.
remoteboolean | nullTrue when the role is remote.
postedAtstring | nullISO 8601 UTC posting date.
searchQuery / source / scrapedAt—Same as company rows.

People row — full field list

FieldTypeDescription
featureType / recordTypestringpeople_profiles / person.
namestring | nullMember name.
titlestring | nullHeadline.
profileUrlstring | nullProfile URL you provided.
avatarUrlstring | nullAvatar image URL.
linkedinUrl / twitterUrl / githubUrlstring | nullPublic social links displayed on the profile.
emailstring | nullPublic email displayed on the profile.
leadDetails / source / scrapedAt—Same as other rows.

Run summary (OUTPUT key-value store)

Each run also writes a summary you can read via the Apify API:

{
"totalPushed": 120,
"enabledFeatures": ["companies", "jobs", "lead_details"],
"errors": [],
"spendingLimitReached": false,
"paywall": {
"detected": true,
"isPaying": true,
"pricingTier": "SILVER",
"blocked": false,
"limited": false,
"mode": "limit",
"freeTierMaxItems": null
},
"finishedAt": "2026-09-25T12:00:00.000Z"
}
Summary fieldMeaning
totalPushedRows exported this run.
enabledFeaturesFeatures that ran (e.g. companies, jobs, people_profiles, lead_details).
errorsNon-fatal failures, expressed as fixed user-facing sentences.
spendingLimitReachedtrue when your max run charge was reached — the run then finishes gracefully.
paywallFree-plan gate transparency: detected, isPaying, pricingTier, blocked, limited, mode, freeTierMaxItems.
finishedAtWhen the run wrapped up.

Dataset views

The dataset ships with ready-made views: Overview, Companies, Jobs, and People — switch between them in the Console's Dataset tab.


Webhooks — real-time delivery to your stack

Every row is always saved to the run's dataset. Optionally, set Webhook URL and each row is also delivered to your URL the moment it is collected — perfect for CRMs, Slack, Zapier, Make, n8n, or Google Sheets.

SettingValues
webhookUrlYour receiving URL (any service that accepts a POST).
webhookFormatjson = the full record; slack = Slack message payload.

Payload example (json format — company row):

{
"featureType": "company_profiles",
"recordType": "company",
"name": "Example Startup",
"description": "AI for supply chains",
"location": "San Francisco",
"markets": ["Artificial Intelligence"],
"companySize": "11-50",
"website": "https://examplestartup.com/",
"email": "hello@examplestartup.com",
"phone": "(555) 555-5555",
"linkedinUrl": "https://www.linkedin.com/company/examplestartup",
"profileUrl": "https://wellfound.com/company/example-startup",
"scrapedAt": "2026-09-25T12:00:00.000Z"
}

Payload example (json format — job row):

{
"featureType": "jobs_feed",
"recordType": "job",
"title": "Senior Software Engineer",
"companyName": "Example Startup",
"salary": "$150k – $200k • 0.1–0.3% equity",
"location": "New York City",
"workplaceType": "REMOTE",
"remote": true,
"jobUrl": "https://wellfound.com/jobs/1234567",
"scrapedAt": "2026-09-25T12:00:00.000Z"
}

Payload example (slack format):

{
"text": ":office: *Example Startup*\nAI for supply chains\n*Location:* San Francisco\n*Size:* 11-50\n*Contact:* hello@examplestartup.com • (555) 555-5555\n*Website:* https://examplestartup.com/"
}

Failed deliveries are logged as a single warning (Webhook delivery failed) and never interrupt the run or dataset writes.

Apify's own webhooks (run succeeded / failed / aborted) are configured separately under your run's Storage → Webhooks in the Console and work with any Actor.


Integrations — API, MCP, and automations

Apify API

Pull results programmatically as JSON, CSV, or Excel:

  • GET https://api.apify.com/v2/datasets/{DATASET_ID}/items — your rows.
  • Full reference: Apify API v2

Start and monitor runs with POST /v2/acts/{actorId}/runs — schedule them with Schedules for daily lead refreshes.

MCP usage (AI assistants)

Use the Apify MCP server so AI assistants (Claude, ChatGPT, Cursor, and others) can run this Actor and read its results directly in chat:

  1. Connect the Apify MCP server to your assistant (setup guide).
  2. Ask in natural language — the assistant calls this Actor with the right input.
  3. Results come back as dataset items the assistant can summarize, tabulate, or chain onward.

Example prompts:

"Collect 20 AI startups that are hiring and list each with its website and LinkedIn."
"Run the Wellfound actor for fintech companies and turn the top 10 into a CSV with name, size, and location."
"Get the latest remote startup jobs with salaries and summarize the market in 5 bullets."

Typical MCP flow:

User: "Find 10 healthtech startups hiring engineers and draft outreach to each"
→ MCP runs Actor with companyKeywords=["healthtech"], feedMaxCompanies=10
→ MCP reads dataset items (name, website, email, linkedinUrl, roles)
→ Assistant drafts the outreach

LLM & RAG pipelines

Output is stable, flat JSON — ideal for ChatGPT, Claude, Gemini, LangChain, and LlamaIndex:

{
"recordType": "company",
"name": "Example Startup",
"markets": ["Artificial Intelligence"],
"companySize": "11-50",
"website": "https://examplestartup.com/",
"email": "hello@examplestartup.com",
"scrapedAt": "2026-09-25T12:00:00.000Z"
}

Workflow: run the Actor → fetch dataset items via the API → pass records to your LLM or vector store.

Native integrations

Push rows straight from the dataset to Google Sheets, Airtable, Dropbox, Google Drive, Zapier, Make, and more via Apify Integrations — no code required.


Pricing

  • Pay per exported result (pay-per-event). You are charged for rows actually delivered to your dataset — not for time, and not for work you don't receive.
  • Spending limit respected: if you set a max charge for a run in Apify, the Actor stops gracefully the moment that limit is reached, with a clear status message — never an overrun. See Apify's pay-per-event docs.
  • Cost levers: cap feedMaxCompanies / feedMaxJobs, shorten the peopleUrls list, and turn off Enable lead details for lighter runs.

Free plan limitations

Apify free-plan accounts are limited to a 2-result sample per run. The run log shows an explicit notice to upgrade to a paid Apify plan for full, unlimited data, and the run's paywall summary object records the decision (detected, isPaying, pricingTier, blocked, limited). This is a transparent platform plan restriction — runs finish cleanly; nothing errors. Paid plans (Bronze and above) always get the full, normal output with no cap.


Performance & scale

  • Streaming: rows are written to the dataset as they are finalized — memory stays light and a long run never loses completed work.
  • 512 MB default memory (configurable up to 2 GB).
  • 10,000-second run timeout — supports very long scheduled runs.
  • Up to 50,000 rows per run (bounded by the per-section caps).
  • US residential connectivity enabled by default for consistent results.

FAQ

Do I need a Wellfound account or API key? No. Just press Start.

What does "real-time" mean here? Data is collected when your run executes — not served from a stale cache. Schedule runs to keep it fresh.

I'm on the Apify free plan — why only 2 results? That's the platform free-plan restriction for this Actor: 2 results per run plus an upgrade notice in the log. Upgrade to any paid Apify plan for full, unlimited output. See Free plan limitations.

Will companies without lead details be dropped? Never. Enable lead details adds contact information when it's available — it never filters. Every collected company is exported, so run time and cost stay predictable.

How does the email finder work? It checks every way a company publicly publishes its address: the homepage, its contact/about pages, structured contact data on the page, protected/encoded addresses, and human-written forms like hello [at] company [dot] com. Emails without a usable public presence stay empty — nothing is invented.

Why does a run take a little longer with lead details on? Lead details add a little extra time per company. Turn the toggle off if you only need company facts and jobs.

Do keywords filter the feed? No — keywords are recorded on every row as context tags for your own filtering. The feed is the live startup feed.

How many companies and jobs can one run collect? Both sections paginate rotating listing pages, so volume scales with your caps rather than stopping at a fixed pass size — hundreds to thousands of companies or jobs per run are normal. Set Max companies this run and Max jobs this run to bound runs and cost; scheduled runs combined with dedupe on your side build history over time.

Some profile URLs produced no row — why? Rows are only exported for profiles that were successfully collected. Unavailable or invalid profiles are skipped and noted in the log with a fixed, generic message.

How do webhooks work? Set webhookUrl (and webhookFormat) and each row is delivered to your URL as it is collected, in addition to being saved to the dataset. Failures log one warning and never stop the run.

Can AI assistants use this Actor? Yes — connect the Apify MCP server and run it with natural language. See MCP usage.

How do I control cost? Set the per-section caps (Max companies this run, Max jobs this run) as your hard caps, and set a run spending limit in Apify — the Actor honors it and stops gracefully.

What does spendingLimitReached: true mean? Your configured max charge for that run was reached. The run wrapped up cleanly with everything collected so far saved.

What do the paywall fields mean? They make the free-plan gate visible in the run output: whether it was detected, whether the account is paying, the pricing tier, and whether results were capped or blocked.

Is the data public? The Actor collects publicly displayed information only.

How often can I run it? As often as you like, bounded by your Apify plan and spending limits. Scheduled runs are supported.

What does "Webhook delivery failed" in the log mean? Your receiving URL rejected or timed out the delivery. Dataset writes are unaffected — fix the receiver or re-deliver from your own queue.

How do I export data? Console export (JSON/CSV/Excel), the Apify API, native integrations, or your webhook.


Support

Open an issue on this Actor's GitHub repository or contact the developer through the Apify Store page.

License

ISC