# Ashby Jobs Scraper (`scrapyx/ashby-jobs-scraper`) Actor

Jobs from any company's Ashby job board (jobs.ashbyhq.com) via the public Job Posting API: title, department, team, location, remote/hybrid, employment type, salary range and currency, publish date, apply URL and full description. Filter by keyword, department, location, remote.

- **URL**: https://apify.com/scrapyx/ashby-jobs-scraper.md
- **Developed by:** [Ibnu Adzim](https://apify.com/scrapyx) (community)
- **Categories:** Jobs
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.84 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Ashby Jobs Scraper

Every open job on any company's **Ashby** job board — Ramp, Linear, Notion
and the thousands of others hosted at `jobs.ashbyhq.com/<org>` — through
Ashby's public Job Posting API: title, department, team, location and
address, remote / hybrid / on-site, employment type, **salary range with
currency and interval**, publish date, apply URL and the full description
as text and HTML.

HTTP only, no login, no key, no browser. **One request per organization**
returns the whole board.

### What it is for

- **Salary datasets** — Ashby is the ATS where compensation is a structured
  field; Ramp publishes a range on 141 of 148 postings.
- **Company watchlists** — a list of orgs on a schedule, diff by `jobId`.
- **Multi-company search** — many orgs, one keyword, one dataset.
- **Remote-only slices** — `remoteOnly` uses the board's own flags.

### Input

| field | what it does |
| --- | --- |
| `organizations` | Org slugs (`ramp`) or URLs (`jobs.ashbyhq.com/ramp`, a job URL under the org). |
| `searchTerms`, `departments`, `locations`, `remoteOnly` | Local filters on the full list (the API has no search). |
| `includeDescription` | On by default; same response, bigger rows. |
| `maxItems`, `maxConcurrency`, `minRequestInterval`, `proxyConfiguration` | Limits. |

Each org gets a `BOARD_SUMMARY` with its total, how many rows carry a
salary, what the filters dropped, and the board's own department and
location vocabularies.

### Three things about this API worth knowing before you trust a run

#### 1. There is no pagination, and page parameters are silently ignored

`?page=2&limit=5` answers with all 148 Ramp jobs, same as no parameters. A
client that walks pages gets the whole board again on every "page". This
Actor makes one request per org and applies `maxItems` locally.

#### 2. Compensation is off unless asked for, and nested three deep when on

Without `includeCompensation=true` there is no `compensation` key at all —
a scraper that forgets the flag reports "no salaries on Ashby". With it, the
numbers live in `compensation.compensationTiers[].components[]` beside
equity and bonus components. The Actor always sends the flag and lifts the
first tier's Salary component to `salaryMin`, `salaryMax`, `salaryCurrency`,
`salaryInterval`; the human string (`"$211.4K – $290.6K • Offers Equity"`)
is `compensationSummary`, `offersEquity` is derived, and the raw object is
kept in `compensation` for boards with several tiers.

#### 3. Titles can carry leading whitespace

`" Security Engineer, Cloud"` — sorted, grouped or matched raw it becomes a
different title. Stripped.

### Other things measured

- An unknown org answers a bare `404 Not Found` (text, not JSON). Reported
  as `organization_not_found`; the other orgs in the run are unaffected.
- `publishedAt` is normalised to UTC (`Z`).
- `address` is `{locality, region, country}` from the posting's postal
  address, or null; `secondaryLocations` is the list of extra offices.
- No throttling or anti-bot layer; responses up to ~3 MB per board.

### Output

- **`JOB`** — `jobId`, `title`, `url`, `applyUrl`, `department`, `team`,
  `employmentType`, `workplaceType`, `isRemote`, `location`, `address`,
  `secondaryLocations`, `publishedAt`, `isListed`, `salaryMin`, `salaryMax`,
  `salaryCurrency`, `salaryInterval`, `compensationSummary`, `offersEquity`,
  `compensation`, `description`, `descriptionHtml`, `organization`,
  `resultPosition`.
- **`BOARD_SUMMARY`** — `organization`, `boardUrl`, `totalJobsOnBoard`,
  `jobsReturned`, `rowsWithSalary`, `stoppedReason`, `filteredOut`,
  `departmentsOnBoard`, `locationsOnBoard`.
- **`ERROR`** — `invalid_input`, `organization_not_found`,
  `payload_shape_changed`, `fetch_failed`, with detail.

### Known limits

- Salary is what the company chose to publish; boards outside pay-
  transparency jurisdictions often have none (`rowsWithSalary` says how
  many).
- Filters are substring matches on the board's own words
  (`departmentsOnBoard` in the summary is the vocabulary).
- Unlisted postings are not served by the API and cannot be fetched.

# Actor input Schema

## `organizations` (type: `array`):

Ashby org slugs or job-board URLs: 'ramp', 'linear', 'notion', https://jobs.ashbyhq.com/ramp (a job URL under the org works too). One request per org returns the whole board with compensation. An unknown slug is reported as organization\_not\_found.

## `searchTerms` (type: `array`):

Keep jobs whose title, department, team or description contains any of these (case-insensitive). Applied locally - the API has no search. Empty = all jobs.

## `departments` (type: `array`):

Keep jobs whose department or team name contains any of these. departmentsOnBoard in the summary row lists the vocabulary.

## `locations` (type: `array`):

Keep jobs whose location, secondary locations or address (city, region, country) contains any of these ('new york', 'remote', 'US').

## `remoteOnly` (type: `boolean`):

Keep only jobs flagged isRemote or with workplaceType Remote.

## `includeDescription` (type: `boolean`):

On by default: the job's full description as text (description) and HTML (descriptionHtml). Same response either way; rows are much smaller without it.

## `maxItems` (type: `integer`):

Overall cap on JOB rows across every organization. A board arrives whole in one response (Ramp: ~150 jobs); the cap is applied locally.

## `maxConcurrency` (type: `integer`):

Organizations fetched in parallel. Responses can be a few MB each.

## `minRequestInterval` (type: `integer`):

Politeness delay between request starts. One request per organization; rarely needed.

## `proxyConfiguration` (type: `object`):

Optional. A public API published for embedding; no anti-bot layer. Enable Apify's free datacenter proxy only if a cloud run reports fetch\_failed.

## Actor input object example

```json
{
  "organizations": [
    "ramp"
  ],
  "remoteOnly": false,
  "includeDescription": true,
  "maxItems": 1000,
  "maxConcurrency": 3,
  "minRequestInterval": 0,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per scraped record. See the dataset's default view for field definitions.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "organizations": [
        "ramp"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapyx/ashby-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "organizations": ["ramp"] }

# Run the Actor and wait for it to finish
run = client.actor("scrapyx/ashby-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "organizations": [
    "ramp"
  ]
}' |
apify call scrapyx/ashby-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapyx/ashby-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zvbkUnNN0wJ5UnbHY/builds/S8QC7K9eX73k69UCe/openapi.json
