# BuiltIn Scraper – Tech Company Profiles & Jobs (`bovi/builtin-scraper`) Actor

Scrape **tech company profiles** from BuiltIn.com: name, industries, **employee count**, founded year, all office locations, **hiring status**, open **job counts by department**, benefits, and website. Filter by industry, size, and remote-hiring flag. No API key or proxy. **Pay per result.**

- **URL**: https://apify.com/bovi/builtin-scraper.md
- **Developed by:** [Vitalii Bondarev](https://apify.com/bovi) (community)
- **Categories:** Lead generation, Jobs, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.75 / 1,000 builtin scraper – tech company profiles & jobs

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## BuiltIn Scraper – Tech Company Profiles & Jobs

Scrape **tech company profiles** from [BuiltIn.com](https://builtin.com/companies) — the leading tech career site covering 10,000+ companies. Returns hiring status, **open job counts by department**, employee count, founded year, all office locations, industry categories, benefits count, and company description. Filter by industry, size band, and remote-hiring flag. No API key, no proxy, no authentication. **Pay per result.**

If you're building B2B lead lists, hiring intelligence dashboards, or tech ecosystem maps, BuiltIn is one of the richest free sources — and this actor delivers the structured data at **$1.50 per 1,000 companies**.

***

### What you get per company

| Field | Description |
|---|---|
| `company_id` | BuiltIn internal company ID |
| `name` | Company name |
| `builtin_url` | Full URL to company profile on builtin.com |
| `website_url` | Company's own website |
| `description` | Company tagline / mission text |
| `industries` | List of industry categories (e.g. `["Software", "Fintech"]`) |
| `employee_count` | Total employee count (integer) |
| `founded_year` | Year founded (integer) |
| `headquarters` | Primary HQ location ("City, State, Country") |
| `all_locations` | All office locations (list of formatted strings) |
| `offices_count` | Total number of offices |
| `total_jobs` | Number of open positions right now |
| `job_categories` | Department breakdown: `[{"name": "Engineering", "count": 42}, ...]` |
| `hiring_now` | Boolean — company is actively hiring |
| `benefits_count` | Number of listed benefits |
| `logo_url` | URL to company logo image |
| `parse_confidence` | Data quality score (1.0 = perfect; <0.5 = drift detected) |
| `warnings` | List of quality warning codes |
| `scraped_at` | ISO-8601 timestamp |

***

### Who uses BuiltIn company data?

- **B2B sales teams** — filter for companies actively hiring in Engineering or Sales as a buying-intent signal
- **Recruiters** — identify companies hiring now in a specific city or industry
- **Investors / analysts** — map the tech ecosystem by industry, headcount, and growth
- **Lead generation** — export enriched company lists with HQ, website, and industry for outreach tools
- **Market research** — benchmark employee counts and job category distribution across sectors

***

### How to use

#### Scrape all AI/Software companies hiring now

```json
{
  "maxResults": 200,
  "industries": "artificial-intelligence,software",
  "hasOpenRoles": true,
  "enrichApi": true,
  "enrichJobs": true
}
```

#### Filter by company size

```json
{
  "maxResults": 500,
  "companySizes": "51-200,201-500",
  "enrichApi": true
}
```

#### Remote-hiring companies in fintech

```json
{
  "maxResults": 100,
  "industries": "fintech",
  "remoteHiring": true,
  "enrichJobs": true
}
```

***

### Input parameters

| Parameter | Type | Default | Description |
|---|---|---|---|
| `maxResults` | integer | 100 | Maximum companies to return |
| `industries` | string | (all) | Comma-separated industry aliases (see list below) |
| `companySizes` | string | (all) | Comma-separated size filters (e.g. `"51-200,201-500"`) |
| `remoteHiring` | boolean | false | Only companies hiring remote |
| `hasOpenRoles` | boolean | false | Only companies with open positions |
| `enrichApi` | boolean | true | Fetch website URL, founded year, office details via API |
| `enrichJobs` | boolean | true | Fetch open job count + department breakdown |

### Common industry aliases

`software` · `fintech` · `healthtech` · `edtech` · `marketing-tech` · `artificial-intelligence` · `cybersecurity` · `data-analytics` · `consumer-web` · `ecommerce` · `cloud` · `blockchain` · `payments`

***

### Example output

```json
{
  "company_id": "84805",
  "name": "Stripe",
  "builtin_url": "https://builtin.com/company/stripe",
  "website_url": "https://stripe.com",
  "description": "Stripe is a technology company that builds economic infrastructure for the internet...",
  "industries": ["Software", "Payments"],
  "employee_count": 5360,
  "founded_year": 2010,
  "headquarters": "Dublin, Dublin, IE",
  "all_locations": ["San Francisco, CA", "New York, NY", "London, GB", "Dublin, Dublin"],
  "offices_count": 9,
  "total_jobs": 407,
  "job_categories": [
    {"name": "Engineering", "count": 84},
    {"name": "Sales", "count": 43}
  ],
  "hiring_now": true,
  "benefits_count": 12,
  "logo_url": "https://builtin.com/sites/stripe.jpeg",
  "parse_confidence": 1.0,
  "warnings": [],
  "scraped_at": "2026-06-04T12:00:00+00:00"
}
```

***

### How it works (three complementary sources)

1. **Listing pages** (`/companies?handler=SearchResults`) — HTML with 20 companies per page. Provides name, industries, employee count, office count, hiring status, benefits count, and description.
2. **Company API** (`api.builtin.com/companyapi/byIds`) — JSON endpoint for website URL, founding year, structured industry list, and all office locations.
3. **Job sections** (`/companies?handler=JobSectionsEncoded`) — Batched HTML endpoint for total open positions count and department-level job categories.

***

### Why this beats competitors

Other BuiltIn scrapers on the Store either miss the job department breakdown, skip the API enrichment step, or break on minor CSS changes. This actor:

- Returns `job_categories` (department breakdown) — which companies are hiring Engineers vs Sales vs Marketing vs Design
- `parse_confidence` on every record — no competitor has it
- Uses **data-company-id structural anchors** — not fragile CSS class selectors
- Works without proxy — builtin.com serves 200 from datacenter IPs
- Combines three complementary data sources for maximum field coverage

***

### FAQ

**Does it require login or authentication?**
No. All three data sources (listing pages, company API, job sections) are publicly accessible without credentials.

**Does it scrape individual company pages?**
For the base listing, no — the listing API is faster and more reliable. With `enrichApi: true`, it calls the dedicated `api.builtin.com` endpoint for additional fields (website, offices, founded year).

**Can it filter by US city (e.g. Chicago, NYC)?**
The filtering API supports national/global queries. City-level filtering on BuiltIn requires scraping their city sub-sites — not yet supported in this version.

**What is `parse_confidence`?**
A 0.0–1.0 per-record data quality score. 1.0 = all fields populated. <0.5 = structural drift detected — some fields may be stale or missing.

***

### Pricing

**$1.50 per 1,000 results** (Pay Per Event). Charged per company record pushed. Failed or filtered-out items are not charged.

### Integrations

Built for B2B sales teams and hiring-intel analysts building tech-company lead lists with job counts and firmographics — the JSON/dataset output drops into the tools you already run, no glue code:

- **n8n / Make / Zapier** — trigger a run or pipe every new dataset item into 500+ apps (Google Sheets, Airtable, Slack, HubSpot, your database) with no code: [n8n](https://docs.apify.com/platform/integrations/n8n), [Make](https://docs.apify.com/platform/integrations/make), [Zapier](https://docs.apify.com/platform/integrations/zapier).
- **Webhooks** — fire your own endpoint the moment a run finishes, to push results straight into your pipeline ([docs](https://docs.apify.com/platform/integrations/webhooks)).
- **MCP server** — expose this actor as a tool to Claude, Cursor, or any [MCP client](https://mcp.apify.com) so an AI agent can pull this data mid-conversation ([guide](https://blog.apify.com/how-to-use-mcp/)).
- **API & SDKs** — fetch the dataset as JSON, CSV, or Excel through the Apify REST API or the Python / JS SDKs.

See all [Apify integrations](https://apify.com/integrations).

### Legal

This actor scrapes publicly available data from builtin.com. No authentication is required. Use responsibly and in accordance with BuiltIn's Terms of Service.

# Actor input Schema

## `maxResults` (type: `integer`):

Maximum number of company records to return. Each page contains 20 companies. Default: 100.

## `industries` (type: `string`):

Comma-separated industry aliases to filter by (e.g. 'software,fintech,healthtech'). Leave blank for all industries. Common values: software, fintech, healthtech, edtech, marketing-tech, artificial-intelligence, cybersecurity, data-analytics, consumer-web, ecommerce.

## `companySizes` (type: `string`):

Comma-separated company size aliases to filter by (e.g. '51-200,201-500'). Leave blank for all sizes. Values: 1-50, 51-200, 201-500, 501-1000, 1001-5000, 5000+.

## `remoteHiring` (type: `boolean`):

Set to true to filter to companies currently hiring remotely.

## `hasOpenRoles` (type: `boolean`):

Set to true to filter to companies with at least one open position.

## `enrichApi` (type: `boolean`):

Set to true (default) to enrich each company with website URL, founding year, structured industries list, and all office locations from the BuiltIn internal API. Adds one API call per company.

## `enrichJobs` (type: `boolean`):

Set to true (default) to fetch open job counts and job category breakdown per company (e.g. 'Engineering: 42, Sales: 15'). Uses a batch endpoint — 5 companies per request.

## Actor input object example

```json
{
  "maxResults": 100,
  "industries": "",
  "companySizes": "",
  "remoteHiring": false,
  "hasOpenRoles": false,
  "enrichApi": true,
  "enrichJobs": true
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset containing Builtin Scraper records (company_id, name, builtin_url, website_url, industries, employee_count, founded_year, headquarters, hiring_now, total_jobs, benefits_count, parse_confidence).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxResults": 100,
    "remoteHiring": false,
    "hasOpenRoles": false,
    "enrichApi": true,
    "enrichJobs": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("bovi/builtin-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxResults": 100,
    "remoteHiring": False,
    "hasOpenRoles": False,
    "enrichApi": True,
    "enrichJobs": True,
}

# Run the Actor and wait for it to finish
run = client.actor("bovi/builtin-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxResults": 100,
  "remoteHiring": false,
  "hasOpenRoles": false,
  "enrichApi": true,
  "enrichJobs": true
}' |
apify call bovi/builtin-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bovi/builtin-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/JrFKsc7fbIiWujpRa/builds/FWbQNJxnsrDwsyHbi/openapi.json
