# Hiring Signal Check & ATS Detector — Greenhouse, Lever, Ashby (`brenton8907/hiring-signal-check`) Actor

Hiring signals and ATS detection from company domains: is the company hiring, on Greenhouse, Lever or Ashby, and for what. Open role count, departments, seniority split and the technologies its job ads name. Each match is checked against the board's own links or published website. JSON/CSV/API/MCP.

- **URL**: https://apify.com/brenton8907/hiring-signal-check.md
- **Developed by:** [Brenton Keller](https://apify.com/brenton8907) (community)
- **Categories:** Lead generation, SEO tools, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.60 / 1,000 hiring check (board found and corroborated)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hiring Signal Check

**Prototype, not published.** Working against 12 graded companies; no store listing, no
icon, and pricing is provisional. See "What is not done".

Is this company hiring, on which applicant tracking system, and for what?

Input company domains. Output one row each: whether they have a live job board on
Greenhouse, Lever or Ashby, how many roles are open, which departments, the seniority
split, and the technologies their job ads name.

### Why the match is trustworthy

Nothing in `stripe.com` tells you the company's board is called `stripe`, so resolving a
domain to a board is a **guess**. This actor makes the guess and then tries to corroborate
it, and every row says which evidence it had:

| `match_confidence` | What it means |
|---|---|
| `verified_url` | The board points at the company's own domain: a job link on it (`stripe.com`, `careers.airbnb.com`), or, on Ashby, the company website the board publishes (`linear` → `linear.app`). Strongest. |
| `verified_name` | The board publishes a company name that matches the domain (`GitLab` / `gitlab.com`). Used when job links point at the ATS vendor instead. |
| `verified_content` | The board's own postings name the company. The best Lever can reach, and Ashby's fallback when a board publishes no website — but it **cannot tell a same-named different company apart**, so it is the weakest of the three. |
| `slug_only` | The slug exists but nothing tied it to the domain. **Treat with suspicion** — it may be a different company with a similar name. |

Measured on the 12-company set: 3 `verified_url`, 4 `verified_name`, 3 `verified_content`,
1 `slug_only`, 1 correct no-match. A `slug_only` row is returned with the caveat stated and
billed at half price, never silently upgraded or quietly dropped.

On the live 30-company sweep: 3 `verified_url` on Greenhouse plus 4 on Ashby, 6
`verified_name`, 1 `verified_content`, 0 wrong companies.

### "No board" and "empty board" are different facts

A board that exists with **zero open roles** answers differently from a domain with no
board at all, and the actor keeps them apart:

- `ats_platform: "lever"`, `is_hiring: false`, `open_roles_count: 0` — they use Lever and
  are not hiring right now. For a prospect list, knowing their ATS is still worth having.
- `ats_platform: null`, `is_hiring: false` — nothing found on these three platforms.

> **A null platform does not mean "not hiring".** Plenty of companies use Workday,
> SmartRecruiters, Recruitee, Teamtailor or a hand-built careers page, none of which this
> actor reads. It means "no board found on Greenhouse, Lever or Ashby". If you know a
> company's board name, pass it in `extraSlugs` and it is tried first.

### Output

| Field | Example |
|---|---|
| `is_hiring` | `true` |
| `ats_platform` / `ats_slug` | `greenhouse` / `stripe` |
| `match_confidence` | `verified_url` |
| `open_roles_count` | `721` |
| `departments` | `["Global Operations", "Engineering", …]` |
| `tech_stack_mentioned` | `["Python", "Spark", "Databricks", "LLM"]` |
| `seniority_distribution` | `{"exec": 1, "vp": 14, "director": 1, "manager": 227, "senior_ic": 63, "junior_ic": 50, "ic": 364}` |
| `locations` | `["Dublin", "New York", …]` (first 25) |
| `latest_role_posted` | `2026-10-06T17:01:49-04:00`: when the newest open role was first published, ISO 8601 on every platform |
| `company_name_on_board` / `company_website_on_board` | what the board says about whose it is (`Linear` / `https://linear.app/`) |
| `query` / `query_domain` | the string you submitted / the domain it resolved to |
| `hint` | set whenever an answer is provisional — read it |

`job-dump` mode returns one row per open role instead: title, department, location,
seniority, apply URL and posting date.

### Cost and speed

These are public, unauthenticated JSON APIs with no bot protection, so this actor is
dramatically cheaper and faster than a scraper: **4.4 requests and 1.0 s per company**
across the 12-company graded set, no proxy required.

A *match* usually costs 2 requests, because the first slug guess is normally right. A
*no-match* is the expensive case: it has to rule out every slug candidate on every
platform before it can answer, so at the default 6 candidates across 3 platforms it can
reach 18. That is a deliberate trade — a short domain often has a longer board name
(`scale.com` is `scaleai`, with 188 open roles; `box.com` is `boxinc`, with 129), and
answering "no board" for those is far worse than a slow negative. Narrow `platforms` or
lower `maxSlugCandidates` if you would rather have cheap negatives than complete ones.
`maxRequestsPerDomain` (24 by default) is a hard ceiling, and a row stopped by it says so
in `hint` rather than reporting a clean "no board".

### Pricing

| Event | Price | When |
|---|---|---|
| `hiring-check` | $0.008 | a board was found **and corroborated** |
| `hiring-unverified` | $0.004 | a board matched on the slug alone, nothing corroborating |
| `hiring-no-match` | $0.003 | no board on any selected platform |
| `job-listing` | $0.001 | one open role, in `job-dump` mode |

Prices follow how much the answer is worth, not how much it cost to produce. An
uncorroborated match carries the same hiring data as a verified one but might be a
different company, so it is half price rather than sharing an event with an answer we can
stand behind. A no-match costs less to buy and more to produce, deliberately: it is the
weaker answer, so charging match price for it would be selling less for the same money. Rows that failed,
were rate limited, or could not be parsed are written to the dataset so the failure is
auditable and carry **no charge**; so do public-suffix and non-domain inputs.

### Known limitations, measured

**Three platforms, not all of them.** Greenhouse, Lever and Ashby only. This is the main
source of false "no board found".

**Seniority comes from job titles, so sales titles skew it.** "Account Manager" and
"Engagement Manager" are individual-contributor roles that land in the `manager` bucket,
which is why a sales-heavy board like Stripe's shows 227 managers in 720 roles. Read the
split as a shape, not a headcount. A missing title is reported as `unknown`, never guessed
as mid-level.

The `exec` bucket deserves a specific note, because it was wrong until 2026-10-07. It
matched the word "partner", so every "Administrative Business Partner" and "HR Business
Partner" was counted as an executive — Stripe's reported 40 executives were all Business
Partner roles, with zero genuine C-level. A real partner title (`Partner`,
`Managing Partner`) still counts as exec; a compound one does not. Stripe now reports 1.

**`tech_stack_mentioned` is what the ads *say*, not what the company runs.** It is a
keyword scan over a sample of postings. A company can run Postgres and never name it, and
an agency-written listing can name Kubernetes for a company that has none. The row reports
`descriptions_sampled` so you can see how much it read; raise `minTechMentions` to drop
technologies named only once.

**Technologies come from a sample.** Scanning all 721 of Stripe's postings to discover
"Python" is waste, so `techSampleSize` (40 by default) bounds it. The row says how many
postings were read, and the hint says when the sample was smaller than the board.

**Department names are reported verbatim and are often messy.** Stripe's Greenhouse board
has 203 departments with names like `1185 Account Executives (EMEA)` and
`5112 General University`. Normalising those into a tidy `["Engineering", "Sales"]`
taxonomy would mean inventing a mapping and presenting a guess as a fact, so the raw names
are returned, busiest first.

**Only Greenhouse exposes departments directly.** For Lever and Ashby the names come from
each posting's own department field, which is usually tidier but can be team-level
(`Delta`, `Echo`) rather than functional.

**Slug resolution is a guess, with limits.** The actor tries the brand, then common
qualifiers both stripped (`ashbyhq.com` → `ashby`) and added (`scale.com` → `scaleai`,
`box.com` → `boxinc`). A board named nothing like its domain is still missed —
`extraSlugs` is the reliable fix, and it is tried before every guess.

**A different company with the same name is the main false-positive risk.** Two measured
cases, both on generic brand words, both Ashby boards:

- `kernel.org` (the Linux kernel) resolves to `kernel` on Ashby, which belongs to an
  unrelated London startup at `kernel.ai`.
- `atlas.co` (an AI GIS company) resolves to `atlas` on Ashby, which belongs to a credit
  card company at `atlascard.com`.

Their postings name *themselves*, so content corroboration confirmed the wrong answer and
both rows used to read `verified_content`, billed at the full rate.

**Ashby boards are now checked against the website they publish.** Ashby's posting API
names no company and hosts every job link itself, but its hosted board page publishes the
company's website. On an Ashby hit the actor reads it:

- **Website is this domain** → `verified_url`. 21 of 22 Ashby customers tested.
- **Website differs, but one redirects to the other** → `verified_url`, and the hint says
  so. `notion.so` → `notion.com`, `temporal.com` → `temporal.io`, `zapier.org` →
  `zapier.com`. A brand comparison cannot separate these from `kernel.org`/`kernel.ai`; the
  redirect can.
- **Website is another company's** → the board is **rejected** and the search continues
  with the remaining slugs and platforms. If nothing else turns up the row is a no-match
  whose hint names the rejected board and its website. `kernel.org` and `atlas.co` now come
  back this way.
- **Website differs and the domain cannot be fetched to check** (a timeout, for example) →
  `slug_only` at half price, with the conflicting website in the hint.
- **No website published** (the company turned the hosted page off, as PostHog has) → the
  old ladder applies, so `verified_content` is still possible there.

Measured on 20 adversarial domains whose brand word is a real Ashby customer's slug
(`linear.com`, `loom.io`, `vanta.io`, `plaid.org`…): before, 18 were sold as corroborated
matches and 12 of those were the wrong company. Now 13 are rejected, 6 genuine aliases
verify, 1 is `slug_only`, and none is a wrong company at full rate. No legitimate board
was rejected across the 62 other live domains tested.

Where each tier is reachable:

| Platform | `verified_url` | `verified_name` | `verified_content` |
|---|---|---|---|
| Greenhouse | reachable (job links are often on the company's own domain) | reachable (publishes `company_name`) | reachable |
| Lever | **never** — all job links are on `jobs.lever.co` | **never** — publishes no company name | only option |
| Ashby | reachable via the published website | not used — the website is stronger | only if no website is published |

**Lever is now the weak spot.** Its hosted page links a logo to whatever the company chose
(Spotify's goes to `spotifyjobs.com`, Palantir's back to its own board), so it cannot
reject a board the way Ashby's website can.

**How to protect yourself:** filter on `match_confidence == "verified_url"` when you need
certainty; treat `verified_content` on a short or generic brand word as a lead rather than
a fact; and pass a known board name in `extraSlugs` to bypass guessing entirely.

### What is not done

- Not published. The store listing, icon and SEO copy are ready in `.actor/`; publishing is
  held for lead.
- Pricing is provisional — set from measured request counts, not from observed demand.
- No canary. The APIs are stable JSON rather than scraped HTML, so a silent shape change
  is less likely than on the Play actor, but it is not impossible and nothing watches for
  it yet.
- No rate-limit evidence at scale. 28 requests across 10 companies hit no throttling; that
  says nothing about 500 companies in one run.
- Lever publishes no company name or website, so a Lever match is `verified_content` at
  best and a same-named Lever board cannot be rejected.

# Actor input Schema

## `mode` (type: `string`):

hiring-check: one enrichment row per domain with counts and signals. job-dump: every open role on the company's board.

## `domains` (type: `array`):

One company domain per line. The board slug is resolved from the domain and then corroborated against the board's own job links, so a domain is answerable and a company name is not.

## `platforms` (type: `array`):

Narrowing this makes a no-match cheaper and faster, at the cost of missing boards on the platforms you drop.

## `extraSlugs` (type: `array`):

If you already know a company's board name, put it here and it is tried before the guesses. The most reliable way to fix a false 'no board found'.

## `maxSlugCandidates` (type: `integer`):

Nothing in a domain says what a company called its board, so the actor guesses: the brand, then common qualifiers stripped (ashbyhq.com -> ashby) and added (scale.com -> scaleai, box.com -> boxinc). Six covers the variants that produced real hits. Lowering it makes a no-match cheaper and risks reporting 'no board' for a company that is hiring; raising it costs more requests on negatives only.

## `maxRequestsPerDomain` (type: `integer`):

Absolute cap, whatever the other limits say. Six slug guesses across three platforms is 18 requests for a true no-match, so 24 leaves headroom. A domain stopped by this ceiling is reported with a hint saying the answer is provisional rather than as a clean 'no board'.

## `includeTechStack` (type: `boolean`):

Scans postings for named technologies. Requires the fuller Greenhouse payload (5 MB on a 700-role board), so turning it off makes large boards noticeably lighter.

## `techSampleSize` (type: `integer`):

Reading all 721 of Stripe's postings to discover 'Python' is waste; the row reports how many were sampled.

## `minTechMentions` (type: `integer`):

Raise to 2 or 3 to drop technologies named in a single passing reference.

## `maxRolesPerCompany` (type: `integer`):

Each role is one row and one job-listing charge ($0.001), so this caps the cost per company: at most $0.20 at the default. Only job-dump mode uses it.

## `requestDelaySeconds` (type: `number`):

These are public APIs that did not rate-limit at this actor's pace, so the default is no delay. Raise it if you see HTTP 429.

## `proxyConfiguration` (type: `object`):

Not required: all three APIs are public and unauthenticated, with no bot protection observed. Provided for accounts that must route egress through a proxy.

## Actor input object example

```json
{
  "mode": "hiring-check",
  "domains": [
    "stripe.com",
    "gitlab.com",
    "spotify.com",
    "berkshirehathaway.com"
  ],
  "platforms": [
    "greenhouse",
    "lever",
    "ashby"
  ],
  "extraSlugs": [],
  "maxSlugCandidates": 6,
  "maxRequestsPerDomain": 24,
  "includeTechStack": true,
  "techSampleSize": 40,
  "minTechMentions": 1,
  "maxRolesPerCompany": 200,
  "requestDelaySeconds": 0,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

One row per company domain (or per open role in job-dump mode).

## `csv` (type: `string`):

The same rows as CSV, for spreadsheets.

## `run` (type: `string`):

The run page with the Overview table, status message and log.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "stripe.com",
        "gitlab.com",
        "spotify.com",
        "berkshirehathaway.com"
    ],
    "extraSlugs": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("brenton8907/hiring-signal-check").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "stripe.com",
        "gitlab.com",
        "spotify.com",
        "berkshirehathaway.com",
    ],
    "extraSlugs": [],
}

# Run the Actor and wait for it to finish
run = client.actor("brenton8907/hiring-signal-check").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "stripe.com",
    "gitlab.com",
    "spotify.com",
    "berkshirehathaway.com"
  ],
  "extraSlugs": []
}' |
apify call brenton8907/hiring-signal-check --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,brenton8907/hiring-signal-check"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vzv9UvSEe2Ndnt2S4/builds/AGvYN7507GN2gks1O/openapi.json
