# Jobs Scraper - Verified Employer & ATS Job Postings (`solutionssmart/jobs-scraper`) Actor

Scrape verified job postings from direct employer career pages and ATS boards (Greenhouse, Lever, Ashby, Personio) with LinkedIn, Indeed and Stepstone lead discovery. Returns normalized jobs with title, company, location, posting URL, status and confidence.

- **URL**: https://apify.com/solutionssmart/jobs-scraper.md
- **Developed by:** [Solutions Smart](https://apify.com/solutionssmart) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 jobs

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Jobs Scraper — Verified Direct-Employer Job Postings

Jobs Scraper finds verified job postings directly from employer career pages and applicant tracking systems (ATS). Instead of scraping job aggregators, it discovers and extracts jobs from the source—company career pages and public ATS boards like Greenhouse, Lever, Ashby, and Personio.

Give Jobs Scraper a hiring goal ("Find robotics software jobs in Germany") or a list of career page URLs, and it returns normalized, structured job data ready for recruiting, market research, or lead generation.

### What you get

Each run produces a dataset of verified employer jobs. Every record includes:

- **Job title and description** (when available)
- **Company name and location**
- **Job URL** (canonical, deduplicated)
- **ATS platform** detected (Greenhouse, Lever, Ashby, Personio, Workday, SmartRecruiters, Teamtailor, Recruitee, or generic)
- **Status** (active/inactive when detectable)
- **Confidence score** based on extraction method
- **Tags** (employment type, remote, salary presence, custom tags)
- **Salary data** (when published by the employer)

All dataset items have `recordType: "verified-job"` to clearly identify what each row represents.

### How to run

Jobs Scraper offers two modes:

1. **Discover mode** — Provide a natural-language task like "Find direct-employer robotics software jobs in Germany." Jobs Scraper searches for relevant career pages, identifies ATS platforms, and extracts verified postings.

2. **Crawl mode** — Provide a list of known employer career page or ATS URLs. Jobs Scraper extracts jobs directly from those sources.

The default `auto` mode picks the right workflow based on your input.

#### Quick start examples

Runnable example inputs live in the [`examples/`](examples/) folder of this repo, not in this file:

- [`examples/germany-robotics-software.json`](examples/germany-robotics-software.json) — discover direct-employer robotics software jobs in Germany.
- [`examples/us-software-engineering.json`](examples/us-software-engineering.json) — discover direct-employer software engineering jobs in the United States.
- [`examples/uk-software-engineering.json`](examples/uk-software-engineering.json) — discover direct-employer software engineering jobs in the United Kingdom.
- [`examples/crawl-career-pages.json`](examples/crawl-career-pages.json) — crawl known employer career or ATS URLs (replace the placeholder URLs first).

Named, reusable tasks cannot be shipped inside the Actor code — they live on the platform. To create them, open [Jobs Scraper in the Apify Store](https://apify.com/solutionssmart/jobs-scraper), go to the **Tasks** tab, and create one task per example above by pasting the matching JSON file as the task input. (The old-style `console.apify.com/actors/solutionssmart~jobs-scraper` link only works while logged in and only after the Actor is published, which is why a bare console link can appear broken — always share the Store URL above.)

### Filters and options

#### Keywords

Keywords filter jobs by checking the **job title** only. At least one keyword must appear in the title:

```json
{
  "keywords": ["robotics", "software"],
  "excludeKeywords": ["intern", "student"]
}
```

Excluded keywords match against the title, location, or description.

#### Countries and locations

Filter jobs by country or specific city/region:

```json
{
  "countries": ["Germany", "Austria", "Switzerland"],
  "locations": ["Berlin", "Munich"]
}
```

#### Posted date filter

Filter jobs by posting date. When `postedWithinDays` is set, jobs that publish a date are kept only if that date is inside the window. Jobs with no published date are still kept:

```json
{
  "postedWithinDays": 30
}
```

#### Salary filter

Exclude jobs below a minimum salary:

```json
{
  "minSalary": 100000,
  "salaryCurrency": "USD"
}
```

Jobs without salary data are excluded when this filter is enabled.

#### Direct employers only

Exclude staffing agencies and recruiters:

```json
{
  "directEmployersOnly": true
}
```

#### Custom tags

Add your own tags to every job in the dataset:

```json
{
  "tags": ["q1-2026", "high-priority"]
}
```

Jobs Scraper also automatically adds tags like `remote`, `ats-greenhouse`, `full-time`, `has-salary`, and region tags when relevant.

### Browser rendering for JavaScript-heavy sites

Some employer career pages load jobs via JavaScript. Jobs Scraper includes Chromium and can render these pages when needed.

Set `enableBrowser: true` and configure a browser page budget:

```json
{
  "enableBrowser": true,
  "maxBrowserPages": 30
}
```

Jobs Scraper uses the browser only when a career page requires JavaScript rendering and is not a supported ATS board. Public ATS APIs (Greenhouse, Lever, Ashby, Personio) are fetched via HTTP without a browser.

### Optional discovery hints from regional job boards

Jobs Scraper optionally supports curated regional job board presets as discovery seeds. These provide starting points for web searches but do not guarantee specific coverage or job counts.

Set `jobSourcesPreset` to `usa`, `eu`, `remote`, or `serbia` to include regional board URLs in the discovery process.

**Note:** These presets are discovery hints only. Jobs Scraper resolves leads back to employer career pages and ATS boards before saving a job to the dataset. Aggregator results are not included as final output.

The presets are derived from [cursustrace](https://github.com/riccione/cursustrace), licensed under Apache-2.0.

### Google Sheets export (optional)

Export results directly to Google Sheets after each run. This feature is **disabled by default**.

To enable:

```json
{
  "googleSheetsEnabled": true,
  "googleSheetsSpreadsheetId": "your-spreadsheet-id",
  "googleSheetsServiceAccountKey": "{ ... service account JSON ... }",
  "googleSheetsOperation": "APPEND"
}
```

The export runs in the background after Jobs Scraper completes. Check the run logs for the export actor run ID and status. Use Apify's secret input feature to protect your service account key.

### Apify Proxy (optional)

Jobs Scraper can route requests through Apify Proxy when career pages block direct access or when you need geo-specific results.

**Proxy is disabled by default** and billed separately by Apify.

Set `proxyMode` to:

- `unblocker` — Route HTTP and browser requests through Apify Unblocker
- `google-serp` — Route web discovery searches through Google SERP Proxy
- `both` — Enable both proxy services
- `none` — No proxy (default)

Example with Unblocker:

```json
{
  "proxyMode": "unblocker",
  "proxyCountryCode": "DE"
}
```

Example with Google SERP Proxy for discovery:

```json
{
  "proxyMode": "google-serp",
  "googleDomain": "google.de",
  "proxyCountryCode": "DE"
}
```

See [Apify Proxy Unblocker](https://docs.apify.com/proxy/unblocker) and [Google SERP Proxy](https://docs.apify.com/proxy/google-serp-proxy) for pricing details.

### Optional Jev controller

Jobs Scraper's default controller uses deterministic rules to plan searches and inspect candidate pages. For faster, more adaptive discovery, you can optionally enable **Jev** (TypeSafe System One), an intelligent routing controller.

When you configure a TypeSafe API key as an Apify Actor secret (`TYPESAFE_API_KEY`), Jev becomes available. Set `agentController: "jev"` to use it.

If the TypeSafe API is unavailable or times out, Jobs Scraper automatically falls back to the rule-based controller. ATS extraction, canonicalization, deduplication, and validation remain deterministic regardless of controller choice.

Obtain an API key from [TypeSafe](https://typesafe.ai).

### Output example

Here's what a single dataset item looks like:

```json
{
  "recordType": "verified-job",
  "title": "Senior Robotics Software Engineer",
  "company": "Acme Robotics GmbH",
  "location": "Berlin, Germany",
  "country": "DE",
  "jobUrl": "https://boards.greenhouse.io/acme/jobs/5678901",
  "canonicalUrl": "https://boards.greenhouse.io/acme/jobs/5678901",
  "sourceDomain": "boards.greenhouse.io",
  "ats": "greenhouse",
  "status": "active",
  "directEmployer": true,
  "extractionMethod": "ats-api",
  "confidence": 1.0,
  "postedAt": "2026-09-15T10:00:00Z",
  "tags": ["ats-greenhouse", "full-time", "has-salary", "salary-eur"],
  "salary": {
    "min": 80000,
    "max": 120000,
    "currency": "EUR",
    "period": "year"
  },
  "description": "We are looking for an experienced robotics engineer..."
}
```

### Pricing

Jobs Scraper uses a **pay-per-result** model:

- **$0.01 USD per job** added to the dataset
- Apify platform compute (CPU, memory) is billed separately as standard Apify usage
- Optional proxy usage (Unblocker, Google SERP) is billed separately by Apify

You only pay for verified employer jobs that Jobs Scraper saves to your dataset. There is no upfront fee.

### Limitations

- Jobs Scraper reads **public pages only**. It does not log in, solve CAPTCHAs, or bypass access controls.
- Not all ATS platforms have public listing APIs. Workday, SmartRecruiters, Teamtailor, and Recruitee rely on generic extraction and may return fewer structured fields.
- JavaScript-heavy career pages depend on the configured `maxBrowserPages` budget. When the budget is exhausted, Jobs Scraper falls back to HTTP-only extraction.
- When `postedWithinDays` is set, jobs that publish a date outside the window are excluded; jobs without a date are kept.
- LinkedIn, Indeed, and Stepstone are used only as discovery hints during web search. Jobs Scraper resolves these leads back to employer career pages or ATS boards. Unresolved aggregator leads are not saved to the dataset.

### Attribution

The optional regional job board presets (`jobSourcesPreset`) are derived from [cursustrace](https://github.com/riccione/cursustrace), licensed under Apache License 2.0.\
Copyright 2026 cursustrace contributors.

# Actor input Schema

## `mode` (type: `string`):

Choose automatic, discovery, or crawl workflow.

## `task` (type: `string`):

Natural-language hiring goal.

## `careerPages` (type: `array`):

Known employer career or ATS board URLs.

## `maxResults` (type: `integer`):

Maximum normalized jobs to save.

## `maxRequests` (type: `integer`):

Maximum HTTP requests for this run.

## `maxCareerPages` (type: `integer`):

Maximum pages discovered by search.

## `maxPagesPerCareerSite` (type: `integer`):

Maximum candidate job pages followed per source.

## `maxAgentSteps` (type: `integer`):

Maximum controller actions.

## `maxBrowserPages` (type: `integer`):

Maximum browser-rendered pages per run when enableBrowser is true.

## `enableBrowser` (type: `boolean`):

Render public JavaScript-heavy non-ATS employer career pages with the bundled browser. Browser support is included in the deployed Actor image.

## `includeDescriptions` (type: `boolean`):

Include job descriptions where available.

## `validateJobs` (type: `boolean`):

Check pages for active or inactive signals.

## `directEmployersOnly` (type: `boolean`):

Keep only confidently direct-employer postings.

## `countries` (type: `array`):

Optional country filters.

## `locations` (type: `array`):

Optional location filters.

## `keywords` (type: `array`):

Required text keywords.

## `excludeKeywords` (type: `array`):

Excluded text keywords.

## `postedWithinDays` (type: `integer`):

When set, jobs that publish a posting date are kept only if that date is inside the window. Jobs with no published date are still kept.

## `agentController` (type: `string`):

Jev (TypeSafe System One) is the default when TYPESAFE_API_KEY is set; rules is the stable fallback; FunctionGemma is optional.

## `debug` (type: `boolean`):

Enable additional run diagnostics.

## `discoverySources` (type: `array`):

Discovery hints: web searches and LinkedIn leads help find employer career pages. JobCook verifies jobs on the employer's own ATS or career page before saving them to the dataset.

## `proxyMode` (type: `string`):

Use Unblocker for career/job pages, Google SERP Proxy for web discovery, or both. Proxy usage is billed separately by Apify.

## `proxyCountryCode` (type: `string`):

Optional two-letter country code for Unblocker, for example US or DE. Leave empty to let Unblocker choose automatically.

## `googleDomain` (type: `string`):

Google domain used by Google SERP Proxy, for example google.com or google.de.

## `googleSheetsEnabled` (type: `boolean`):

Send results to Google Sheets after run completes.

## `googleSheetsSpreadsheetId` (type: `string`):

The spreadsheet ID from the Google Sheets URL (required when export enabled).

## `googleSheetsOperation` (type: `string`):

APPEND adds new rows; UPSERT updates existing rows based on upsert key.

## `googleSheetsUpsertKey` (type: `string`):

Column name to use as unique key for UPSERT operation (required when operation is UPSERT).

## `googleSheetsServiceAccountKey` (type: `string`):

Google service account JSON key with Sheets API access. Use secret input to protect this credential.

## `googleSheetsFirstRowHeaders` (type: `boolean`):

Treat first row of spreadsheet as column headers.

## `jobSourcesPreset` (type: `string`):

Optional curated regional job board preset (usa, eu, remote, serbia). Adds board URLs as discovery seeds. Derived from cursustrace under Apache-2.0.

## `tags` (type: `array`):

Optional tags to add to all emitted jobs. Auto-tags (e.g., 'remote', region) are added automatically when using presets.

## `minSalary` (type: `integer`):

Optional minimum salary filter. Jobs with no salary data or salary below this minimum are excluded. Requires salaryCurrency.

## `salaryCurrency` (type: `string`):

Three-letter ISO currency code for salary filtering (e.g., USD, EUR, GBP). Required when minSalary is set.

## Actor input object example

```json
{
  "mode": "auto",
  "task": "Find direct-employer robotics software jobs in Germany",
  "careerPages": [],
  "maxResults": 100,
  "maxRequests": 500,
  "maxCareerPages": 50,
  "maxPagesPerCareerSite": 50,
  "maxAgentSteps": 30,
  "maxBrowserPages": 20,
  "enableBrowser": false,
  "includeDescriptions": true,
  "validateJobs": true,
  "directEmployersOnly": false,
  "countries": [],
  "locations": [],
  "keywords": [],
  "excludeKeywords": [],
  "agentController": "rules",
  "debug": false,
  "discoverySources": [
    "web",
    "linkedin"
  ],
  "proxyMode": "none",
  "googleDomain": "google.com",
  "googleSheetsEnabled": false,
  "googleSheetsOperation": "APPEND",
  "googleSheetsFirstRowHeaders": true,
  "tags": []
}
```

# Actor output Schema

## `jobs` (type: `string`):

Verified job listings extracted from employer career pages and ATS boards. Each record includes title, company, location, job URL, ATS type, status, confidence score, and optional tags and salary information.

## `summary` (type: `string`):

Statistics about the actor run including pages visited, jobs found, quality metrics, and discovery provider performance.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "task": "Find direct-employer robotics software jobs in Germany"
};

// Run the Actor and wait for it to finish
const run = await client.actor("solutionssmart/jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "task": "Find direct-employer robotics software jobs in Germany" }

# Run the Actor and wait for it to finish
run = client.actor("solutionssmart/jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "task": "Find direct-employer robotics software jobs in Germany"
}' |
apify call solutionssmart/jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,solutionssmart/jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Naduk7Uy5qpm21mQ8/builds/eQxjjCQFQKQOwHKjw/openapi.json
