# Built In Jobs Scraper — Salaries, Skills & Job Monitoring (`corvuslab/builtin-scraper`) Actor

\[💰 $0.90 / 1K] Scrape jobs from builtin.com — salary min/max, required skills, benefits, company & industry profile and full descriptions. Filter by keyword, category, location & remote, or monitor a search for new postings with Telegram/Slack/webhook alerts. No code; export JSON, CSV, Excel, API.

- **URL**: https://apify.com/corvuslab/builtin-scraper.md
- **Developed by:** [Corvuslab](https://apify.com/corvuslab) (community)
- **Categories:** Jobs, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.90 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Built In Jobs Scraper — Salaries, Skills & Job Monitoring

### What does the Built In Jobs Scraper do?

> **Turn every live builtin.com tech & startup job into a clean, structured record — salary, skills, benefits, company profile and full description — in seconds, at scale.**

Scrape tech and startup job listings from **builtin.com** with the richest per-job detail available: **salary ranges** (min/max, currency, period), **required tech skills**, the full **company & industry profile**, the complete **benefits list**, and the entire **job description** as text, HTML and Markdown. Filter by keyword, job category, location, remote and posted-within date; schedule it in **incremental monitoring** mode to capture only new and changed postings; and get pinged on **Telegram, Slack, Discord or any webhook**. No code, no login, no API key — export to JSON, CSV, Excel or the API.

**Why this scraper**

- ⚡ **Fast & low-cost** — a direct, lightweight extraction path keeps large runs cheap, so scanning thousands of jobs stays affordable.
- 🧾 **Rich, typed records** — dozens of structured fields per job (salary min/max, skills, benefits, company & industry), not raw HTML.
- ♻️ **Cheap to monitor** — incremental mode re-scrapes only what changed (see below).
- 🔔 **Notifications built in** — Telegram, Slack, Discord or any webhook (n8n / Make / Zapier).
- 🤖 **AI- & API-ready** — compact output, MCP-friendly, one-click integrations.

***

### ✨ Key features

- 🔎 **Search or URL scraping** — run a keyword search with filters, or paste builtin.com search and job URLs directly.
- 💰 **Structured salary** — parsed `salaryMin` / `salaryMax` plus currency and period, not just a raw pay string — sort and filter on real numbers.
- 🧠 **Required skills & benefits** — every job's `skills[]` and its full `benefits[]` list, extracted and ready to analyze.
- 🏢 **Company & industry profiles** — company name, URL, logo and industry tags on each record, so a job is also a company signal.
- 🎚️ **Rich filters** — keyword, job category (16 categories), city/metro, workplace type (remote / hybrid / on-site), experience level (entry → expert) and posted-within (24h → 30 days).
- 🧾 **Full descriptions, three ways** — the complete job description as plain **text**, **HTML** and **Markdown**, plus a short summary.
- 📇 **Contacts & lead-gen** — emails, phones, URLs and social profiles pulled from each posting; keep only jobs that carry a contact with `requireContact`.
- ♻️ **Incremental monitoring** — schedule it and get only what changed, tagged **NEW / UPDATED / UNCHANGED / EXPIRED**, with repost detection; unchanged items are skipped before their page is even fetched.
- 🔔 **Notifications** — Telegram, Slack, Discord or any webhook (n8n / Make / Zapier), fanned out on every run.
- 🤖 **AI-ready** — compact mode and drop-empty output keep payloads small for LLMs and MCP clients.

***

### 📤 Example output

```json
{
  "id": "9132993",
  "title": "Engineer, Applied AI",
  "url": "https://builtin.com/job/sr-applied-ai-engineer/9132993",
  "company": "Zapier",
  "companyUrl": "https://builtin.com/company/zapier",
  "companyLogo": "https://builtin.com/sites/www.builtin.com/files/2024-10/150@x2.png",
  "companyIndustry": "Artificial Intelligence, Productivity, Software, Automation",
  "industries": ["Artificial Intelligence", "Productivity", "Software", "Automation"],
  "location": "29 Locations",
  "workplaceType": "Remote",
  "isRemote": true,
  "experienceLevel": "Senior level",
  "salaryText": "192K-287K Annually",
  "salaryMin": 192000,
  "salaryMax": 287000,
  "salaryCurrency": "USD",
  "salaryPeriod": "YEARLY",
  "employmentType": "FULL_TIME",
  "postedDate": "2026-08-14",
  "validThrough": "2026-09-13T00:20:50+00:00",
  "easyApply": false,
  "skills": ["Cloud Infrastructure", "Llm Ops", "Ml Ops", "Python", "Typescript"],
  "benefits": ["Offers 401(K)", "Provides 401(K) matching", "Offers company equity", "Offers generous PTO", "Offers a remote work program", "…"],
  "summary": "As a Sr. Applied AI Engineer at Zapier, you will build and enhance AI platform capabilities, focusing on LLM Ops and ML Ops to support scalable AI development across teams.",
  "description": "AI at Zapier — Are you excited about building the platform that makes AI and machine learning development faster, safer, and more reliable across an entire company? …",
  "source": "builtin.com",
  "searchKeyword": "engineer",
  "scrapedAt": "2026-08-14T08:25:17Z",
  "detailFetched": true
}
```

The full record also carries `descriptionHtml`, `descriptionMarkdown`, the complete `benefits[]` list, `city` / `state` / `country`, and (in incremental mode) a `changeType` tag.

***

### 📥 Example input

A few ready-to-run configurations — set these in the visual editor or pass them as JSON via the API:

Broad keyword search across all of Built In:

```json
{ "query": "software engineer", "maxResults": 100 }
```

Filtered — remote data & analytics roles posted in the last 7 days, with full details:

```json
{
  "category": "data-analytics",
  "location": "austin",
  "remoteMode": "remote",
  "postedWithinDays": 7,
  "includeDetails": true,
  "maxResults": 200
}
```

Incremental monitoring — poll new product-manager postings on a schedule and push to Telegram:

```json
{
  "query": "product manager",
  "incrementalMode": true,
  "telegramToken": "123:ABC",
  "telegramChatId": "@myjobsfeed"
}
```

***

### 📚 What data can you extract from builtin.com?

- **Core** — `id`, `title`, `url`, `company`, `location`, `salaryText`, `postedDate`, `source`, `searchKeyword`, `scrapedAt`.
- **Salary** — `salaryMin`, `salaryMax`, `salaryCurrency`, `salaryPeriod` (parsed to real numbers).
- **Role details** — `experienceLevel`, `employmentType`, `workplaceType`, `isRemote`, `easyApply`, `validThrough`, `skills[]`, `benefits[]`, `summary`.
- **Descriptions** (with `includeDetails`) — `description` (text), `descriptionHtml`, `descriptionMarkdown`.
- **Company** — `company`, `companyUrl`, `companyLogo`, `companyIndustry`, `industries[]`.
- **Location** — `location`, `city`, `state`, `country`.
- **Contacts & signals** — `extractedEmails`, `extractedPhones`, `extractedUrls`, `socialProfiles`, `applyEmail`, `contactName`, `phoneNumber`.
- **Monitoring** — `changeType` (NEW / UPDATED / UNCHANGED / EXPIRED), `isRepost`, `repostOfId`, `contentHash`.

Every field is present in standard mode (missing values are `null`); **compact mode** returns the core fields only, for lean AI/MCP payloads.

***

### ⚙️ Input

Configure it in the visual editor — no code needed — or pass JSON via the API.

| Field | What it does |
|---|---|
| `query` | Keyword search, e.g. `python`, `product manager` (comma-separate for multiple searches). |
| `startUrls` | Scrape specific builtin.com search or job URLs directly. |
| `category` | Restrict to one of 16 job categories (Developer + Engineer, Data + Analytics, Product, Design + UX, …). |
| `location` | City or metro, e.g. `austin`, `new york`, `boston`. |
| `remoteMode` | `any`, `remote`, `hybrid` or `onsite` — restrict to one workplace type. |
| `experienceLevel` | Keep only chosen seniority levels — Entry / Junior / Senior / Expert + Leader. |
| `postedWithinDays` | Only jobs updated within 1 / 3 / 7 / 14 / 30 days. |
| `includeDetails` | Fetch each job's page for descriptions, benefits and the full company profile. |
| `requireContact` | Keep only jobs that carry a contact (email / phone / either / both). |
| `incrementalMode` | Emit only what changed since the last run (NEW / UPDATED / EXPIRED). |
| `compact` | Return core fields only — ideal for AI agents and MCP. |
| `maxResults` | Cap the number of records (0 = unlimited). |
| `proxyConfiguration` | Optional — works without a proxy at typical volumes. |

…and **27 inputs** in total — the table shows the essentials; the rest cover description format (text/HTML/Markdown), drop-empty output, notification channels (Telegram, Slack, Discord, webhook) and advanced tuning, all in the visual editor.

***

### 💡 What can you do with builtin.com data?

- **Job-board & aggregator feeds** — populate your own tech-jobs site with clean, structured, salary-rich listings.
- **Compensation & market research** — analyze salary ranges, in-demand skills and benefits across companies, categories and cities.
- **Recruiting & sourcing** — track new tech roles by category and location; pull contacts for lead generation.
- **Company intelligence** — build a picture of who's hiring, in which industries, and for what skills.
- **New-job monitoring** — schedule incremental mode + notifications for a live change feed of fresh postings.
- **AI agents & pipelines** — compact output plugs straight into LLM / MCP workflows.

***

### ♻️ Incremental monitoring — pay for changes, not repeats

Schedule the actor and turn on **incremental mode**: each run compares against the last and emits only **NEW / UPDATED / EXPIRED** records — unchanged items are skipped *before* their detail page is fetched, so a daily watch costs a fraction of a full re-scrape.

| Daily churn | of 1,000 tracked | billable records | you save |
|---|---|---|---|
| 5 % | 1,000 | 50 | **95 %** |
| 15 % | 1,000 | 150 | **85 %** |
| 30 % | 1,000 | 300 | **70 %** |

The first run seeds the baseline and bills in full; every run after that bills only the delta.

***

### 🚀 How to scrape builtin.com

1. Open the actor and enter a **search keyword** (e.g. `software engineer`) and/or pick filters — or paste a builtin.com URL.
2. Set **Max results** and choose whether to **fetch full details** (descriptions, benefits, company profile).
3. (Optional) Turn on **incremental mode** and a **notification** channel, then **Schedule** it.
4. Click **Start**.
5. Download the data as **JSON, CSV or Excel**, or pull it from the **API**.

New to Apify? Create a free account — it comes with monthly credit, no credit card required.

***

### 🔌 Integrations & export

Export to **JSON, CSV, Excel** or an HTML table, or pull from the **REST API** and the **JavaScript / Python** clients. Runs on a **schedule**, connects to **Google Sheets, Slack, Make, Zapier and n8n**, and works as an **MCP tool** for AI agents — compact mode keeps token usage small.

***

### ❓ FAQ

**Do I need a proxy or login?** No — it runs out of the box with no login or API key; Apify Proxy is available under Advanced for high-volume runs.

**Can I get only new jobs on a schedule?** Yes — turn on incremental mode and schedule it; each run emits only what changed (NEW / UPDATED / EXPIRED) and can notify your channel.

**Does it include salary data?** Yes — where a posting lists pay, you get the raw `salaryText` plus parsed `salaryMin`, `salaryMax`, `salaryCurrency` and `salaryPeriod`.

**Can I get the full job description?** Yes — enable full details for the complete description as plain text, HTML and Markdown, plus benefits and the company profile.

**What formats can I export?** JSON, CSV, Excel, HTML table, or via the API.

**Is it good for AI agents?** Yes — enable compact mode; the output is MCP-friendly and payloads stay small.

**How many records can I get?** Set `maxResults` (0 = unlimited, bounded by how many the search returns).

***

### ⚖️ Is it legal to scrape builtin.com?

This actor accesses only publicly available data on builtin.com. You are responsible for how you use the extracted data — in particular any personal information — and for complying with the site's terms and applicable law (including the GDPR where it applies). Not affiliated with, endorsed by, or sponsored by Built In.

***

**Keywords:** builtin scraper · builtin.com scraper · built in jobs scraper · built in jobs api · tech jobs scraper · startup jobs scraper · job listings scraper · salary data scraper · tech skills scraper · remote jobs scraper · job monitoring · new job alerts · recruiting data · job board api · export jobs to CSV · MCP tool for AI agents

# Actor input Schema

## `query` (type: `string`):

Keywords to search for (e.g. "python", "product manager"). Separate multiple searches with commas — each runs as its own search and results are merged and de-duplicated.

## `startUrls` (type: `array`):

Paste site search or listing URLs to scrape directly. Detail-page URLs work here too.

## `maxResults` (type: `integer`):

Maximum number of records to return. Set 0 for unlimited (bounded by how many the search has).

## `ignoreUrlFailures` (type: `boolean`):

Skip URLs that cannot be interpreted instead of failing the whole run.

## `category` (type: `string`):

Restrict to one Built In job category. Leave blank for all categories.

## `location` (type: `string`):

City or metro to restrict to, e.g. "austin", "new york", "boston". Leave blank for nationwide. Ignored when Remote only is selected.

## `remoteMode` (type: `string`):

Restrict to one workplace type, or include all. Remote uses Built In's remote landing page; Hybrid and On-site are screened on each listing (jobs that don't state a workplace type are excluded).

## `experienceLevel` (type: `array`):

Keep only jobs at one or more experience levels. Screened on each listing against the level Built In shows; jobs with no stated level are excluded when this is set.

## `postedWithinDays` (type: `string`):

Only jobs updated within this window. Leave as Any time for no date filter.

## `requireContact` (type: `string`):

Keep only records that include a contact. off = keep everything; email / phone = require that channel; either = at least one; both = email and phone. Screened at emit time, so it also trims what you're billed for.

## `includeDetails` (type: `boolean`):

Fetch each item's detail page for the richer fields. Turn off for the fastest, cheapest runs.

## `descriptionFormat` (type: `string`):

Which representation(s) of any long-text field to include.

## `compact` (type: `boolean`):

Emit only the core fields. Ideal for AI agents and MCP clients.

## `excludeEmptyFields` (type: `boolean`):

Remove null, empty-string and empty-array fields from each record.

## `incrementalMode` (type: `boolean`):

Track state between runs and tag every record with a changeType (NEW / UPDATED / UNCHANGED / EXPIRED).

## `stateKey` (type: `string`):

Stable name for the tracked search. Leave empty to derive one automatically from your search settings.

## `emitUnchanged` (type: `boolean`):

Also emit records that have not changed since the previous run.

## `emitExpired` (type: `boolean`):

Emit records for items present last run but gone now.

## `telegramToken` (type: `string`):

Bot token from @BotFather.

## `telegramChatId` (type: `string`):

Chat or channel ID, e.g. "-100123456789" or "@yourchannel".

## `slackWebhookUrl` (type: `string`):

Slack incoming-webhook URL.

## `discordWebhookUrl` (type: `string`):

Discord incoming-webhook URL.

## `webhookUrl` (type: `string`):

Any HTTPS endpoint. Receives a JSON POST with the matched records — works with n8n, Make and Zapier.

## `webhookHeaders` (type: `object`):

Extra headers for the webhook request, e.g. {"Authorization": "Bearer xyz"}.

## `notificationLimit` (type: `integer`):

How many records to include in each notification message.

## `proxyConfiguration` (type: `object`):

Optional. Built In runs fine with no proxy (the default), which is the fastest and cheapest option. Only enable Apify Proxy if you scrape at high volume and start seeing blocks — datacenter first, residential as a last resort.

## `maxRequestRetries` (type: `integer`):

How many times to retry a failed request before giving up on it.

## Actor input object example

```json
{
  "query": "software engineer",
  "maxResults": 25,
  "ignoreUrlFailures": true,
  "category": "",
  "remoteMode": "any",
  "postedWithinDays": "0",
  "requireContact": "off",
  "includeDetails": true,
  "descriptionFormat": "all",
  "compact": false,
  "excludeEmptyFields": false,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "notificationLimit": 5,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "maxRequestRetries": 3
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `allItems` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "",
    "location": ""
};

// Run the Actor and wait for it to finish
const run = await client.actor("corvuslab/builtin-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "",
    "location": "",
}

# Run the Actor and wait for it to finish
run = client.actor("corvuslab/builtin-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "",
  "location": ""
}' |
apify call corvuslab/builtin-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,corvuslab/builtin-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/89dMQFKtZ0a07zzbd/builds/hTBqB2lOFsNSSNueq/openapi.json
