# Wellfound Jobs Scraper (`scrapyx/wellfound-jobs-scraper`) Actor

Startup jobs from Wellfound (ex-AngelList Talent) by role and location: full description, compensation, remote flag, ATS source, plus the hiring company's size, funding badges and one-liner on every row. Refuses a search Wellfound silently widened instead of returning the wrong rows.

- **URL**: https://apify.com/scrapyx/wellfound-jobs-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Jobs
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.84 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Wellfound Jobs Scraper

Startup jobs from **Wellfound** (formerly AngelList Talent) — full job
descriptions, pay ranges, remote status, and the hiring company's size, stage
and funding badges, all in one row.

### Why use this actor

- **No account, no login, no API key.** Everything comes from public pages.
- **The whole job description** on every row, not a truncated teaser — and it
  arrives without a second request per job, so runs stay fast and cheap.
- **The company comes with the job.** Name, one-liner, headcount band, logo,
  and Wellfound's own badges (*Actively Hiring*, *YC Funded*, *Top Investors*,
  *Valuation $1B+*, Glassdoor ratings) are nested on every job row for free.
- **It tells you when a search did not run.** Wellfound answers `200 OK` to a
  misspelled role or city by quietly searching for something else — often
  thousands of unrelated jobs. This actor detects that and returns an error row
  instead of the wrong data. That is the main reason to pick it over a
  home-made scraper.
- **Honest counts.** Wellfound's own job total is larger than what its pages can
  actually hand back, and its "per page" number counts companies, not jobs.
  Every run returns a summary row that states both, so you never mistake one
  for the other.
- **Stable JSON** — the same envelope (`_input`, `_source`, `_scrapedAt`,
  `recordType`) as every other scraper in this collection, so one loader reads
  them all.

### How it works

1. You give it role slugs, location slugs, or Wellfound search URLs — the same
   words that appear in Wellfound's own address bar (`software-engineer`,
   `san-francisco`).
2. It walks the result pages for each search, page by page, stopping when
   Wellfound runs out of results, when your page cap is hit, or when a page
   returns nothing new.
3. Duplicate listings are removed — Wellfound repeats some rows across
   consecutive pages — and each job is checked to make sure it really belongs
   to the search you asked for.
4. Results land in your dataset as JSON, CSV or Excel, with one summary row per
   search telling you exactly what happened.

You do not manage browsers, sessions, retries or pagination.

### Input

```json
{
  "roles": ["software-engineer"],
  "locations": ["san-francisco"],
  "remote": false,
  "maxItemsPerSearch": 100,
  "maxPagesPerSearch": 5,
  "includeStartups": false,
  "includeJobDetails": false,
  "maxConcurrency": 3,
  "minRequestInterval": 0.8,
  "proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}
```

| Field | Type | Description |
|---|---|---|
| `roles` | array | Role slugs as they appear in Wellfound's URL — `software-engineer`, `product-manager`, `data-scientist`. Around 515 exist; browse them at [wellfound.com/browse/tech-jobs](https://wellfound.com/browse/tech-jobs). Combined with `locations` as a cross-product. |
| `locations` | array | Location slugs — `san-francisco`, `new-york`, `london`, `berlin`, `bangalore`, `united-states`. Coverage is limited and Wellfound publishes no list; an unsupported one comes back as an error row that says so. |
| `remote` | boolean | Search each role's remote listings instead of the general ones. Cannot be combined with `locations` — Wellfound has no "remote jobs in a city" page, so that combination is refused rather than quietly ignored. |
| `searchUrls` | array | Paste Wellfound search URLs directly. Supported: `/role/<role>`, `/role/r/<role>`, `/role/l/<role>/<location>`, `/location/<location>`. |
| `maxItemsPerSearch` | integer | Stop after this many unique jobs per search. `0` = no cap. Counted after duplicates are removed. Default `100`. |
| `maxPagesPerSearch` | integer | Ceiling on pages per search. Default `5`. Each page returns 20 companies and roughly 12–56 jobs. |
| `includeStartups` | boolean | Also emit one row per company as its own record. The same company data is already nested on every job row, so turn this on only if you want companies as a separate table. Default `false`. |
| `includeJobDetails` | boolean | Fetch each posting's own page to add the **employer's website** and geocoded locations. Costs one extra request per job (~45 per page). Default `false`. |
| `maxConcurrency` | integer | Requests in flight. Default `3`. |
| `minRequestInterval` | number | Seconds between request starts — the real speed control. Default `0.8`. |
| `proxyConfiguration` | object | Apify Proxy. Residential is the default — it was measured to be both more reliable and faster here than datacenter. |

### Output

Four record types share one dataset, told apart by `recordType`.

#### `JOB`

```json
{
  "_input": "role=software-engineer location=san-francisco",
  "_source": "S1-nextdata",
  "_scrapedAt": "2026-09-16T02:17:50Z",
  "recordType": "JOB",
  "id": "4639821",
  "title": "Software Engineer",
  "slug": "software-engineer",
  "jobUrl": "https://wellfound.com/jobs/4639821-software-engineer",
  "primaryRoleTitle": "Software Engineer",
  "compensation": "$150k – $176k",
  "jobType": "full-time",
  "remote": false,
  "remoteConfig": null,
  "locationNames": ["Denver", "San Francisco"],
  "acceptedRemoteLocationNames": [],
  "yearsExperienceMin": null,
  "yearsExperienceMax": null,
  "liveStartAt": 1787861184,
  "autoPosted": false,
  "atsSource": "AtsIntegration::Greenhouse::Listing",
  "isBookmarked": false,
  "description": "As a Software Engineer at Checkr, you will work on high-impact engineering projects that help build a fairer future for all. You will work in a collection of small services built on Ruby and Javascript, with both SQL and NoSQL databases, as well as message que… (truncated)",
  "startup": {
    "id": "395014",
    "name": "Checkr",
    "slug": "checkr",
    "companySize": "SIZE_501_1000",
    "highConcept": "The only background check company using artificial intelligence and machine learning",
    "logoUrl": "https://photos.wellfound.com/startups/i/395014-2a2164ef6a8cd5954eb970228e62ce15-medium_jpg.jpg?buster=1692326450",
    "companyUrl": "https://wellfound.com/company/checkr",
    "badges": [
      { "id": "ACTIVELY_HIRING", "label": "Actively Hiring", "tooltip": "Actively processing applications", "rating": null },
      { "id": "YC-395014", "label": "YC Funded", "tooltip": "Startup funded by Y Combinator", "rating": null },
      { "id": "HIGHLY_RATED-395014", "label": "Highly rated", "tooltip": "Checkr is highly rated on Glassdoor, with 4.1 out of 5 stars", "rating": "4.1" },
      "… 7 more"
    ]
  },
  "searchedRole": "software-engineer",
  "searchedLocation": "san-francisco",
  "searchedRemote": false,
  "pageNumber": 1
}
```

| Field | Type | Description |
|---|---|---|
| `id` | string | Wellfound's job id. |
| `title` | string | Job title as posted. |
| `jobUrl` | string | Direct link to the posting. |
| `primaryRoleTitle` | string | Wellfound's normalised role for the job. |
| `compensation` | string | Pay range exactly as Wellfound displays it. Empty when undisclosed. |
| `jobType` | string | `full-time`, `contract`, `internship`, … |
| `remote` | boolean | Whether the posting is remote. |
| `remoteConfig` | object | Remote policy detail when the employer set one. |
| `locationNames` | array | Cities the job is open in. |
| `acceptedRemoteLocationNames` | array | Regions accepted for remote applicants. |
| `yearsExperienceMin` / `Max` | integer | Experience band when stated. |
| `liveStartAt` | integer | Unix timestamp the posting went live. |
| `autoPosted` | boolean | Whether it was synced in automatically from the employer's system. |
| `atsSource` | string | Which applicant-tracking system it came from (Greenhouse, Lever, …). Useful for deduping against other job feeds. |
| `description` | string | The complete job description. |
| `startup` | object | The hiring company — name, slug, headcount band, one-liner, logo, profile URL and badges. |
| `searchedRole` / `searchedLocation` / `searchedRemote` | — | Which search produced this row. |
| `pageNumber` | integer | Which result page it came from. |

With `includeJobDetails: true`, each job also gains `detailCompanyWebsite`
(the employer's own site), `detailJobLocation` (geocoded), `detailDatePosted`,
`detailEmploymentType`, `detailIndustry`, `detailDirectApply`,
`detailBaseSalary` and `detailHiringOrganization`.

#### `SEARCH_SUMMARY`

One per search, so you can always tell what a run actually did.

```json
{
  "recordType": "SEARCH_SUMMARY",
  "searchUrl": "https://wellfound.com/role/l/software-engineer/san-francisco",
  "surface": "role_location",
  "totalJobCount": 763,
  "totalStartupCount": 290,
  "pageCount": 15,
  "perPage": 20,
  "pagesFetched": 2,
  "jobsReturned": 81,
  "startupsReturned": 38,
  "duplicateJobsDropped": 5,
  "duplicateStartupsDropped": 2,
  "stoppedBecause": "max_pages_reached",
  "filtersHeld": true,
  "routeRendered": "/seoLanding/roleLocationSearch",
  "effectiveFilters": { "role": "software-engineer", "location": "san-francisco", "page": 2 },
  "perPageCountsStartupsNotJobs": true,
  "totalJobCountExceedsReachable": true,
  "pagesOverlap": true
}
```

| Field | Description |
|---|---|
| `totalJobCount` / `totalStartupCount` | Wellfound's own totals, verbatim. |
| `perPage` | **Counts companies, not jobs** — `perPageCountsStartupsNotJobs` flags this. |
| `pageCount` | Wellfound's page count, derived from the company total. |
| `jobsReturned` | Unique jobs this run actually returned. |
| `duplicateJobsDropped` | Rows Wellfound repeated across pages and this actor removed. |
| `stoppedBecause` | `max_items_reached`, `max_pages_reached`, `page_count_reached`, `no_new_rows`, or `page_wrapped_to_N`. |
| `filtersHeld` | `true` means Wellfound ran the search you asked for. |
| `effectiveFilters` | The filters Wellfound actually applied. |
| `totalJobCountExceedsReachable` | Wellfound counts more jobs than its pages can return. |

#### `STARTUP`

Emitted only when `includeStartups: true` — one deduplicated row per company,
with the same fields as the nested `startup` block plus `highlightedJobCount`.

#### `ERROR`

One per search that failed or was refused, so every input maps to at least one
row.

| `_error` | Meaning |
|---|---|
| `filters_dropped` | Wellfound answered `200` but ran a different search — usually a role or city slug it does not recognise. The wrong rows were discarded. |
| `not_found` | Wellfound has no page for that slug. |
| `invalid_input` | Refused before any request — a malformed slug, or `remote` combined with a location. |
| `fetch_failed` | Network or server trouble after retries. |
| `path_blocked` | The page is not reachable by any automated client. |

### Known limits

- **Search is by role and location, not free text.** Wellfound's keyword search
  and its live job feed are not available to automated clients, so this actor
  uses the role and location pages instead.
- **Company profile pages are not scrapable** by any client, so company data
  comes from what the search results already carry — which is substantial, but
  not the full profile.
- **Not every job in a search is reachable.** Each page shows up to 3 jobs per
  company, so a company with 40 openings contributes 3. `totalJobCount` tells
  you how many exist; `jobsReturned` tells you how many you got.
- **Location coverage is limited** and Wellfound publishes no list of supported
  cities. Unsupported ones return an error row naming known-good slugs.

# Actor input Schema

## `roles` (type: `array`):

Job roles, as the lowercase kebab-case slug Wellfound uses in its own URL path — e.g. 'software-engineer' from wellfound.com/role/software-engineer. Roughly 515 role slugs exist; browse them at wellfound.com/browse/tech-jobs. Combined with 'locations' as a cross-product. A slug Wellfound does not recognise is refused with an ERROR row rather than silently widened (the site answers 200 and drops the unknown half instead of 404ing).

## `locations` (type: `array`):

Locations, as the lowercase kebab-case slug in Wellfound's URL path — e.g. 'san-francisco', 'new-york', 'london', 'berlin', 'bangalore', 'united-states'. Coverage is a closed set and Wellfound publishes no index, so an unsupported location returns an ERROR row explaining that. 'remote' is not a location — use the Remote only toggle instead.

## `remote` (type: `boolean`):

Search Wellfound's remote route for each role instead of the general one. Applies only when 'locations' is empty: Wellfound has no URL for 'remote jobs in a specific city', so combining this with a location is refused rather than quietly ignored.

## `searchUrls` (type: `array`):

Paste Wellfound search URLs directly instead of (or alongside) role and location slugs. Supported shapes: /role/<role>, /role/r/<role>, /role/l/<role>/<location>, /location/<location>. Company profile pages and the live /jobs feed are not supported — Cloudflare blocks both for every HTTP client.

## `maxItemsPerSearch` (type: `integer`):

Stop a search once this many unique jobs have been collected. 0 means no cap (still bounded by Max pages per search). Counted after deduplication — Wellfound's pages overlap, re-serving rows already returned on an earlier page.

## `maxPagesPerSearch` (type: `integer`):

Hard ceiling on page requests per search. Each page is a 500–700 KB document returning 20 companies and the up-to-3 jobs each of them highlights (~12–56 jobs). Paging also stops on Wellfound's own pageCount, and when a page returns nothing new.

## `includeStartups` (type: `boolean`):

Emit one deduplicated STARTUP record per company in addition to the job records. Company facts (name, slug, size, one-liner, logo, badges such as Actively Hiring / YC / funding stage) are ALREADY nested on every job row at no extra cost — turn this on only when you want companies as their own table.

## `includeJobDetails` (type: `boolean`):

Adds the employer's own website URL, geocoded job locations, structured salary and industry from each posting's detail page. Costs one extra request PER JOB (~45 per page), so it multiplies run time and proxy traffic roughly 45×. The full job description, compensation text, ATS source and posting date are already in the search response without it.

## `maxConcurrency` (type: `integer`):

Upper bound on requests in flight. Once the minimum request interval binds, raising this buys nothing — that interval, not this number, is the honest speed control.

## `minRequestInterval` (type: `number`):

The actual speed control: the shortest gap between two request starts, across all workers. No rate limiting was observed during testing, so this is routine pacing rather than a defensive measure — raise it if you see 429s.

## `proxyConfiguration` (type: `object`):

Apify Proxy settings. Residential is the default because it was measured to be both more reliable and faster here: on identical cloud runs, datacenter exits drew 7 challenges and lost one search outright (18.5s), while residential drew none and finished all three (10.3s). Datacenter still works often enough to be worth trying if you are optimising for proxy cost.

## Actor input object example

```json
{
  "roles": [
    "software-engineer"
  ],
  "locations": [
    "san-francisco"
  ],
  "remote": false,
  "maxItemsPerSearch": 100,
  "maxPagesPerSearch": 5,
  "includeStartups": false,
  "includeJobDetails": false,
  "maxConcurrency": 3,
  "minRequestInterval": 0.8,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "roles": [
        "software-engineer"
    ],
    "locations": [
        "san-francisco"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/wellfound-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "roles": ["software-engineer"],
    "locations": ["san-francisco"],
}

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/wellfound-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "roles": [
    "software-engineer"
  ],
  "locations": [
    "san-francisco"
  ]
}' |
apify call scrapyx/wellfound-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/wellfound-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/8WNdAvTrulaAep9Di/builds/0fBp5MnuwDodCokIZ/openapi.json
