# Remote Jobs Aggregator & Job Alerts - RemoteOK, Himalayas, WWR (`lukehunter/remote-jobs-aggregator`) Actor

Scrape remote job listings from RemoteOK, Himalayas, Jobicy and We Work Remotely in one run. Deduplicated, filterable by keyword, salary, location and date. Turn on job alerts and run on a schedule to get only new listings each time, billed once each.

- **URL**: https://apify.com/lukehunter/remote-jobs-aggregator.md
- **Developed by:** [Luke Hunter](https://apify.com/lukehunter) (community)
- **Categories:** Jobs, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.50 / 1,000 jobs

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Remote Jobs Aggregator & Job Alerts — RemoteOK, Himalayas, Jobicy & WWR

**For job-alert bots, remote-job newsletter writers, recruiters and job-board sites who need one clean feed instead of four.** One run pulls current remote listings from RemoteOK, Himalayas, Jobicy and We Work Remotely, normalises them into a single flat schema, merges cross-posted duplicates, and filters by keyword, recency, location and salary. Turn on **job alerts** mode (`onlyNewSinceLastRun`) and run it on a schedule to get billed only for genuinely new listings each time.
Pay-per-result: **$0.0015 per job — 1,000 jobs = $1.50.**
Try it free with Apify's monthly platform credit.

This Actor exists because a "remote python jobs" search touches four different sites with four different formats. This Actor is the merge step: one request, one schema, no duplicates.

### Quick start (1 minute)

Open the **Input** tab and use this prefill (swap in your own keywords):

```json
{
  "keywords": ["python"],
  "postedWithinDays": 7,
  "maxItems": 20
}
```

Click **Start**. It finishes in a few seconds. Export the dataset to CSV/JSON, or pull it through the API shown below.

### Use cases

- **Job-alert bots** polling on a schedule for new listings matching a keyword set, then posting to Slack/Discord/email.
- **Remote-job newsletter writers** pulling a fresh, de-duplicated shortlist every week without visiting four sites.
- **Recruiters and sourcers** scanning what's currently open across boards for a given skill or location.
- **Job-board and aggregator sites** backfilling or supplementing their own listings with a normalised external feed.

### Job alerts: get only new jobs

Turn on `onlyNewSinceLastRun` and put this Actor on a **schedule** for a real job alert: every run after the first delivers (and charges for) only jobs it hasn't delivered before for that exact search — cross-posted duplicates are still recognised even if a job moves between boards between runs, using the same company+title identity as the dedupe above.

1. Set your keywords/filters as usual, plus:

```json
{
  "keywords": ["python"],
  "onlyNewSinceLastRun": true,
  "maxItems": 50
}
```

2. Click **Schedule** on the run page (or create one under **Schedules** in the Apify Console) — daily is typical for an active search.
3. **The first scheduled run is a baseline**: it delivers every job currently matching (it's all new to you) and remembers it — you're charged normally for that run. Every run after that only delivers jobs it hasn't seen before for this exact search.
4. Add a **Webhook** under the schedule's **Integrations** for `ACTOR.RUN.SUCCEEDED`, pointing at:
   - **Slack**: Apify's own Slack integration (or a Zapier/Make webhook step) posting each run's new dataset items to a channel.
   - **Email**: a Zapier/Make "on webhook, send email" step, or Apify's own email integration, summarising that run's new jobs.
   - **Google Sheets**: the Apify-to-Google-Sheets integration, appending each run's rows to a sheet you watch.

The run's status message tells you what happened: `"First run: 12 job(s) delivered and remembered; next runs return only new ones."` on the baseline, then `"2 new job(s) since last run (10 already seen)."` on later runs. A run that finds nothing new still **succeeds** with an empty dataset — that's the point of the mode, not a failure.

Each distinct combination of `keywords`/`excludeKeywords`/`sources`/`postedWithinDays`/`locationContains`/`worldwideOnly`/`minSalary`/`includeUnknownSalary` is tracked as its own watch, so several alert schedules (e.g. one per keyword set) don't mix up each other's history. `stateStoreName` only needs changing if you want to reset a watch's memory or explicitly isolate it.

### Input

```json
{
  "keywords": ["python"],
  "excludeKeywords": [],
  "sources": ["remoteok", "himalayas", "jobicy", "wwr"],
  "postedWithinDays": 7,
  "locationContains": "",
  "worldwideOnly": false,
  "minSalary": null,
  "includeUnknownSalary": true,
  "maxItems": 20,
  "includeFullDescription": false,
  "onlyNewSinceLastRun": false,
  "stateStoreName": "remote-jobs-aggregator-seen-jobs"
}
```

| Field | Type | Default | Description |
|---|---|---:|---|
| `keywords` | string\[] | `[]` (all) | Keep jobs whose title, tags or description match ANY of these words (case-insensitive) |
| `excludeKeywords` | string\[] | `[]` | Drop jobs matching ANY of these words |
| `sources` | string\[] | all 4 | `remoteok`, `himalayas`, `jobicy`, `wwr` — any subset |
| `postedWithinDays` | integer | 7 | 1–90. Jobs with no parseable posting date are kept regardless |
| `locationContains` | string | — | Case-insensitive substring match on `location`. Ignored if `worldwideOnly` is on |
| `worldwideOnly` | boolean | false | Only jobs open to applicants anywhere (`isWorldwide: true`) |
| `minSalary` | integer | — | Keep only jobs whose stated salary (min or max, whichever is higher) meets this. Currency-naive — see Limitations |
| `includeUnknownSalary` | boolean | true | With `minSalary` set, whether to still keep jobs that state no salary |
| `maxItems` | integer | 50 | 1–1000. Hard cap on jobs delivered — this is your cost cap |
| `includeFullDescription` | boolean | false | Off truncates `descriptionText` to 2000 characters |
| `onlyNewSinceLastRun` | boolean | false | Job alerts mode — see above. Delivers and charges only jobs not delivered by a previous run of the same search |
| `stateStoreName` | string | `remote-jobs-aggregator-seen-jobs` | Name of the store that remembers what a job-alerts search has already delivered. Change it only to isolate or reset a watch |

For most users, only **keywords** and **maximum jobs** matter — the defaults handle the rest.

### Output fields

| Category | Fields |
|---|---|
| Identity | `id`, `source`, `duplicateOf`, `alsoPostedOn` |
| Job | `title`, `company`, `companyUrl`, `companyLogo`, `employmentType`, `tags` |
| Links | `url`, `applyUrl` |
| Timing | `postedAt` |
| Location | `location`, `isWorldwide` |
| Salary | `salaryMin`, `salaryMax`, `salaryCurrency` |
| Description | `descriptionText` |

Missing values are returned as `null` rather than guessed.

### Output example

```json
{
  "id": "remoteok:1137421",
  "title": "Senior Backend Engineer (Python)",
  "company": "Acme Robotics",
  "companyUrl": null,
  "companyLogo": "https://remoteok.com/assets/logo.png",
  "source": "remoteok",
  "url": "https://remoteok.com/remote-jobs/senior-backend-engineer-python-acme-robotics-1137421",
  "applyUrl": "https://remoteok.com/remote-jobs/senior-backend-engineer-python-acme-robotics-1137421",
  "postedAt": "2026-09-24T16:00:06.000Z",
  "location": "Germany",
  "isWorldwide": false,
  "employmentType": null,
  "salaryMin": 70000,
  "salaryMax": 80000,
  "salaryCurrency": "USD",
  "tags": ["python", "backend", "senior"],
  "descriptionText": "We are looking for a Senior Backend Engineer...",
  "duplicateOf": null,
  "alsoPostedOn": [
    { "source": "himalayas", "url": "https://himalayas.app/companies/acme-robotics/jobs/senior-backend-engineer" }
  ]
}
```

### Cross-posting dedupe

Many roles are posted to more than one board word-for-word. Jobs are grouped by normalised **company + title**; only the first one seen is delivered (and charged for) — the rest are merged into its `alsoPostedOn` list, so a cross-posted job is never billed twice. `duplicateOf` is always `null` on delivered rows: it is reserved for a possible future partial-dedupe mode and does nothing yet.

### Use it as an API

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("lukehunter/remote-jobs-aggregator").call(
    run_input={"keywords": ["python"], "postedWithinDays": 7, "maxItems": 50}
)

for job in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(job["title"], job["company"], job["url"])
```

Use Apify schedules to poll on a cadence for a job-alert bot, and webhooks to push new results into Slack, Discord or email.

### Pricing and cost control

This Actor uses **pay per delivered job**.

**Current configured rate: $0.0015 per job delivered.** Check the Apify **Pricing** tab for the latest published rate.

| Jobs delivered | Cost at $0.0015/job |
|---:|---:|
| 20 | $0.03 |
| 100 | $0.15 |
| 1,000 | $1.50 |

There is no charge for merely starting a run, and no charge for a job dropped by filtering or dedupe. `maxItems` gives you a clear upper bound on the number of billable jobs.

### Reliability

- A failure fetching one source (site down, format changed) is logged and the run continues with the other three — it never fails the whole run over one board. Which sources failed is recorded in the run's key-value store under `OUTPUT`.
- The run only fails outright if **every** requested source failed to fetch — a narrow keyword/location/salary filter that matches nothing is a normal, successful, empty result.
- `descriptionText` is plain text (HTML stripped), truncated to 2000 characters by default.

### Important limitations

- **RemoteOK**: its public API returns one fixed page of its current ~100 most recent listings — there is no further pagination to request.
- **Himalayas**: paginated up to 10 pages (up to ~1,000 jobs) per run to bound runtime; increasing `maxItems` beyond that will not pull more from Himalayas alone.
- **Jobicy**: requested in a single page, capped at 50 jobs per run from this source (Jobicy's own documented ceiling for one request).
- **We Work Remotely**: only the single combined feed (`weworkremotely.com/remote-jobs.rss`) is read, not the per-category RSS feeds.
- **Salary comparisons are currency-naive.** `minSalary` compares raw numbers regardless of `salaryCurrency` — no FX conversion is performed.
- `postedWithinDays` cannot filter a job whose source did not publish a usable date; such jobs are kept rather than guessed away.
- Descriptions, salaries and locations are exactly what each source publishes — this Actor does not verify a listing is still live or accurate.
- **RemoteOK attribution**: per RemoteOK's API Terms of Service, every RemoteOK-sourced row keeps its `url` linking back to the RemoteOK job page and `source: "remoteok"` naming RemoteOK — please preserve these if you republish RemoteOK-sourced listings.
- This is an independent aggregator and is **not affiliated with, endorsed by, or connected to** RemoteOK, Himalayas, Jobicy or We Work Remotely. Each is a trademark of its respective owner.

### FAQ

#### Does this bypass any login, CAPTCHA or paywall?

No. All four sources are public, unauthenticated feeds intended for programmatic access (three JSON APIs, one RSS feed). No personal data is collected — only public job listings.

#### Why is a job I saw on RemoteOK also on Himalayas in the results?

It's the same listing, merged. See "Cross-posting dedupe" above — check `alsoPostedOn` on the delivered row.

#### Can I get only fully remote/worldwide jobs?

Yes — set `worldwideOnly: true`.

#### Can I run this on a schedule for job alerts?

Yes — set `onlyNewSinceLastRun: true`, create an Apify Schedule with your keyword input, and use a webhook or integration to forward new dataset items to Slack/Discord/email. See "Job alerts" above. The first scheduled run delivers everything (and remembers it); later runs only deliver, and charge for, what's new.

### Related Actors

Other data tools from the same developer, built to the same standard: official or public sources, hard cost caps, and honest documentation of limits.

- **[Wellfound (AngelList) Jobs Scraper](https://apify.com/lukehunter/wellfound-jobs-scraper)**: startup jobs from Wellfound with salary ranges, equity and company stage.
- **[Google Play App Scraper](https://apify.com/lukehunter/google-play-scraper)**: ratings, installs, developer contact info and pricing for any Google Play app.
- **[Apple App Store Reviews Scraper](https://apify.com/lukehunter/app-store-reviews-scraper)**: Apple App Store reviews for any iOS app, across countries, with rating, version and date.
- **[Spotify Scraper](https://apify.com/lukehunter/spotify-scraper)**: play counts, monthly listeners and playlist track lists for any public Spotify artist, playlist, album or track.
- **[Walmart Category Scraper](https://apify.com/lukehunter/walmart-category-scraper)**: product names, prices, was-prices and ratings from Walmart category pages.
- **[Shopify Store Products Scraper](https://apify.com/lukehunter/shopify-store-products-scraper)**: full product catalogues from any Shopify store, with prices, sale prices, variants and stock.
- **[Vinted Scraper](https://apify.com/lukehunter/vinted-scraper)**: Vinted search results with prices, brands, sizes and favourites, across any Vinted country.
- **[AliExpress Search Scraper](https://apify.com/lukehunter/aliexpress-scraper)**: AliExpress search results with prices, discounts, ratings and sold counts, by keyword.

# Actor input Schema

## `keywords` (type: `array`):

Match jobs whose title, tags or description contain ANY of these words (case-insensitive). Leave empty to keep every job from the selected sources.

## `excludeKeywords` (type: `array`):

Drop jobs whose title, tags or description contain ANY of these words (case-insensitive). Leave empty to exclude nothing.

## `sources` (type: `array`):

Which job boards to pull from. Default is all four.

## `postedWithinDays` (type: `integer`):

Only keep jobs posted within this many days, 1-90. A job whose posting date could not be determined is kept regardless.

## `locationContains` (type: `string`):

Only keep jobs whose location text contains this (case-insensitive), e.g. "Europe" or "United States". Ignored if "Worldwide only" is on. Leave empty for no location filter.

## `worldwideOnly` (type: `boolean`):

Only keep jobs open to applicants anywhere in the world (no country/region restriction stated by the source).

## `minSalary` (type: `integer`):

Only keep jobs whose stated salary (min or max, whichever is higher) meets this figure. Currency-naive: compared as a raw number regardless of the job's currency. Leave empty for no salary filter.

## `includeUnknownSalary` (type: `boolean`):

When a minimum salary is set, whether to still keep jobs that state no salary at all (most listings don't). Ignored if no minimum salary is set.

## `maxItems` (type: `integer`):

Hard cap on the number of jobs delivered across all sources, 1-1000, default 50. You are charged per delivered job, so this is your cost cap.

## `includeFullDescription` (type: `boolean`):

By default, descriptionText is truncated to 2000 characters. Turn this on to include the full plain-text description (larger dataset items).

## `onlyNewSinceLastRun` (type: `boolean`):

Turn this on and run the Actor on a schedule to get only jobs you haven't seen before: jobs already delivered on a previous run for the same keywords/sources/filters are skipped and never charged again (cross-posted duplicates are matched by company+title, so a job that moves between boards is still recognised). The first run for a given search delivers everything (it's all new) and remembers it; later runs only deliver what's new since then. Off by default so a one-off run always gets the full current results.

## `stateStoreName` (type: `string`):

Only used when "Job alerts" is on. Name of the key-value store that remembers which jobs this search has already delivered, so it persists across scheduled runs. Leave as the default unless you're running several different alert searches that must not share history — give each its own name in that case. Letters, numbers and hyphens only.

## Actor input object example

```json
{
  "keywords": [
    "engineer",
    "developer"
  ],
  "excludeKeywords": [],
  "sources": [
    "remoteok",
    "himalayas",
    "jobicy",
    "wwr"
  ],
  "postedWithinDays": 14,
  "worldwideOnly": false,
  "includeUnknownSalary": true,
  "maxItems": 20,
  "includeFullDescription": false,
  "onlyNewSinceLastRun": false,
  "stateStoreName": "remote-jobs-aggregator-seen-jobs"
}
```

# Actor output Schema

## `jobs` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "engineer",
        "developer"
    ],
    "excludeKeywords": [],
    "postedWithinDays": 14,
    "locationContains": "",
    "maxItems": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("lukehunter/remote-jobs-aggregator").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": [
        "engineer",
        "developer",
    ],
    "excludeKeywords": [],
    "postedWithinDays": 14,
    "locationContains": "",
    "maxItems": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("lukehunter/remote-jobs-aggregator").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "engineer",
    "developer"
  ],
  "excludeKeywords": [],
  "postedWithinDays": 14,
  "locationContains": "",
  "maxItems": 20
}' |
apify call lukehunter/remote-jobs-aggregator --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,lukehunter/remote-jobs-aggregator"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/2qt8ZDcmXo3Uu6wM9/builds/hOJCdght6MU4PLmBc/openapi.json
