# Built In Jobs Scraper (`scrapyx/builtin-jobs-scraper`) Actor

Tech jobs from Built In: title, company, location, salary range, employment type, industries, benefits and the full description. Browse by category, city or remote, or search by keyword. A category Built In does not recognise is reported, not quietly replaced.

- **URL**: https://apify.com/scrapyx/builtin-jobs-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Jobs
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.84 / 1,000 results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Built In Jobs Scraper

Tech jobs from **Built In** — title, company, location, salary range,
employment type, industries, benefits and the full description.

### Why use this actor

- **Complete rows, not just links.** Built In's listing pages publish only four
  things per job: title, link, position and a one-line summary. No company, no
  location, no salary, no date. This actor reads each job's own page to fill
  all of that in, and does it by default.
- **Salary ranges as numbers**, with currency and period — not a string you
  have to parse.
- **Per-job benefits lists.** 401(k) matching, parental leave, fertility
  benefits, mental-health cover and the rest, as a structured list. Very few
  job boards publish this.
- **A typo can't hand you the wrong jobs.** Ask Built In for a category it does
  not recognise and it returns ten perfectly real, completely unrelated jobs
  with a `200 OK` — no error, no redirect. This actor spots that and returns an
  error row instead of the wrong data.
- **Nothing depends on page markup.** Every field comes from the structured
  data Built In publishes for search engines, so a site redesign does not
  silently empty your dataset.
- **No account, no login, no API key.**

### How it works

1. You pick categories (`dev-engineering`, `remote`, `san-francisco`), keywords,
   or nothing at all for the full listing.
2. The actor reads the listing pages, collecting job links and deduplicating.
3. It then reads each job's own page for the employer, salary, location, date,
   benefits and full description.
4. Results land in your dataset as JSON, CSV or Excel, with one summary row per
   query.

### Input

```json
{
  "categories": ["dev-engineering"],
  "searchTerms": [],
  "includeJobDetails": true,
  "maxItemsPerQuery": 100,
  "maxPagesPerQuery": 5,
  "maxConcurrency": 5,
  "minRequestInterval": 0.5,
  "proxyConfiguration": { "useApifyProxy": true }
}
```

| Field | Type | Description |
|---|---|---|
| `categories` | array | The path segment from Built In's own URL — `dev-engineering`, `remote`, `san-francisco`. An unrecognised one is reported as an error rather than quietly replaced with generic jobs. |
| `searchTerms` | array | Keyword search. Combine with categories and every category is searched with every term. |
| `includeJobDetails` | boolean | **On by default** — without it you get titles and links and nothing else, because that is all the listing publishes. Costs one request per job. |
| `maxItemsPerQuery` | integer | Stop after this many unique jobs per query. `0` = no cap. Default `100`. |
| `maxPagesPerQuery` | integer | Listing-page ceiling. Page size varies between about 10 and 27 jobs, so use the item cap for a precise budget. Default `5`. |
| `maxConcurrency` | integer | Requests in flight, which matters most during the detail pass. Default `5`. |
| `minRequestInterval` | number | Seconds between request starts. Default `0.5`. |
| `proxyConfiguration` | object | Apify Proxy. Datacenter is the default and is sufficient. |

### Output

Three record types share one dataset, told apart by `recordType`.

#### `JOB`

```json
{
  "_input": "category=dev-engineering",
  "_source": "S1-jsonld-list+S2-jsonld-detail",
  "_scrapedAt": "2026-09-16T04:13:58Z",
  "recordType": "JOB",
  "jobId": "11198462",
  "title": "Software Engineer, Behavior Validation",
  "jobUrl": "https://builtin.com/job/software-engineer-behavior-validation/11198462",
  "summary": "Develop behavior validation products for autonomous vehicle software…",
  "detailCompanyName": "General Motors",
  "detailCompanyUrl": "https://builtin.com/company/general-motors",
  "detailCompanyLogo": "https://builtin.com/sites/www.builtin.com/files/2025-06/gm_avatar_white_blue%201.png",
  "detailSalaryMin": 123200,
  "detailSalaryMax": 189100,
  "detailSalaryCurrency": "USD",
  "detailSalaryPeriod": "YEAR",
  "employmentType": "FULL_TIME",
  "datePosted": "2026-09-16",
  "validThrough": "2026-10-16T03:22:34+00:00",
  "directApply": false,
  "industry": ["Automotive", "Big Data", "Information Technology", "Robotics", "… 3 more"],
  "jobBenefits": ["Offers 401(K)", "Provides 401(K) matching", "Offers generous parental leave", "… 17 more"],
  "applicantLocationRequirements": { "@type": "Country", "name": "USA" },
  "hiringOrganization": { "@type": "Organization", "name": "General Motors", "sameAs": "https://builtin.com/company/general-motors" },
  "baseSalary": { "@type": "MonetaryAmount", "currency": "USD", "value": { "minValue": 123200, "maxValue": 189100, "unitText": "YEAR" } },
  "description": "<b>Description</b><br><strong>Role</strong>… (full HTML)",
  "listPosition": 1,
  "pageNumber": 1
}
```

| Field | Description |
|---|---|
| `jobId` / `jobUrl` | Built In's job id and a direct link. |
| `title` | Job title. |
| `summary` | The one-line summary from the listing page. |
| `description` | The full job description (HTML) from the job's own page. |
| `detailCompanyName` / `detailCompanyUrl` / `detailCompanyLogo` | The employer. |
| `detailSalaryMin` / `detailSalaryMax` / `detailSalaryCurrency` / `detailSalaryPeriod` | Salary range, already split into numbers. |
| `detailCity` / `detailRegion` / `detailCountry` / `detailPostalCode` | Location, flattened from the posting's address. |
| `employmentType` | `FULL_TIME`, `PART_TIME`, `CONTRACTOR`, … |
| `datePosted` / `validThrough` | When it was posted and when it expires. |
| `industry` | The employer's industry tags. |
| `jobBenefits` | Structured benefits list — often 20+ entries. |
| `applicantLocationRequirements` | Where applicants must be based. |
| `directApply` | Whether Built In accepts the application directly. |
| `baseSalary` / `hiringOrganization` / `jobLocation` / `identifier` | The original structured objects, kept whole. |
| `_detailError` | Present only when a job's own page could not be read. |

#### `SEARCH_SUMMARY`

| Field | Description |
|---|---|
| `category` / `searchTerm` | What was queried. |
| `pagesFetched` / `jobsReturned` / `duplicatesDropped` | What the run did. |
| `collectionPageTitle` | The page title Built In returned — this is what proves the category applied. |
| `unfilteredPageTitle` | The generic title, for comparison. |
| `jobDetailsRequested` / `jobDetailFailures` | Whether the detail pass ran, and how many pages failed. |
| `stoppedBecause` | `max_items_reached`, `max_pages_reached`, `no_new_jobs`, `no_more_results`. |
| `upstreamTotalNotPublished` | Always `true` — Built In publishes no result total anywhere, so none is invented. |
| `pageSizeVaries` | Always `true` — pages hold roughly 10–27 jobs. |

#### `ERROR`

| `_error` | Meaning |
|---|---|
| `category_not_applied` | Built In ignored the category and served generic jobs. The rows were discarded. |
| `invalid_input` | A malformed category slug, refused before any request. |
| `blocked` | Every attempt was refused; try the residential proxy group. |
| `unexpected_shape` | A `200` without the structured data this actor depends on. |
| `fetch_failed` | Network or server trouble after retries. |

### Known limits

- **No result total.** Built In does not publish one, so runs are bounded by
  your page and item caps rather than by a known total.
- **Page size is not fixed** (roughly 10–27 jobs), so a page count is not a
  precise job budget.
- **The detail pass is one request per job.** A 500-job run is ~500 requests.
- **Company profiles are not scraped** — each job carries its employer's name,
  Built In page and logo, but not the full company profile.

# Actor input Schema

## `categories` (type: `array`):

The path segment Built In uses in its own URL — 'dev-engineering' from builtin.com/jobs/dev-engineering, or 'remote', or a city like 'san-francisco'. Built In does NOT return a 404 for a category it does not recognise; it quietly serves a generic listing of unrelated jobs instead. This actor detects that and returns an error row rather than the wrong jobs.

## `searchTerms` (type: `array`):

Keyword search, e.g. 'python', 'machine learning'. Can be combined with categories — every category is then searched with every term.

## `includeJobDetails` (type: `boolean`):

ON by default, and you almost certainly want it. Built In's listing pages publish only four things per job: title, link, position and a one-line summary — no company, no location, no salary, no date. All of that lives on the job's own page. Turning this off costs one request per page instead of one per job, and gives you a list of links.

## `maxItemsPerQuery` (type: `integer`):

Stop a query once this many unique jobs have been collected. 0 means no cap (still bounded by max pages).

## `maxPagesPerQuery` (type: `integer`):

Ceiling on listing requests per query. Built In's page size varies between roughly 10 and 27 jobs, so this is not a precise job budget — use the item cap for that.

## `maxConcurrency` (type: `integer`):

Requests in flight. This matters most during the detail pass, which is one request per job. Once the interval below binds, raising this buys nothing.

## `minRequestInterval` (type: `number`):

Shortest gap between two request starts, across all workers. The honest speed control. No rate limiting was observed during testing, so this is routine pacing.

## `proxyConfiguration` (type: `object`):

Apify Proxy settings. Datacenter is the default and is sufficient — both the listing and detail pages answered normally to a plain unproxied request during testing.

## Actor input object example

```json
{
  "categories": [
    "dev-engineering"
  ],
  "includeJobDetails": true,
  "maxItemsPerQuery": 100,
  "maxPagesPerQuery": 5,
  "maxConcurrency": 5,
  "minRequestInterval": 0.5,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "categories": [
        "dev-engineering"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/builtin-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "categories": ["dev-engineering"] }

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/builtin-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "categories": [
    "dev-engineering"
  ]
}' |
apify call scrapyx/builtin-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/builtin-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/jRIFnSzcdMKbxgWWQ/builds/70JGLI1qiti0HZf7A/openapi.json
