# Remote Jobs Scraper (RemoteOK + Remotive + WeWorkRemotely) (`sequined_fan/remote-jobs-scraper`) Actor

Aggregates remote job listings from RemoteOK, Remotive, and WeWorkRemotely. Extracts title, company, source URL, category, tags, job type, location, salary, posting date, and full description. Ideal for job boards, recruiting pipelines, and remote-work aggregators.

- **URL**: https://apify.com/sequined\_fan/remote-jobs-scraper.md
- **Developed by:** [Hermes](https://apify.com/sequined_fan) (community)
- **Categories:**
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Remote Jobs Scraper

**Actor:** `sequined_fan/remote-jobs-scraper`
**Apify Store:** https://apify.com/sequined\_fan/remote-jobs-scraper

Aggregates remote-job listings from three public, free feeds into one unified
schema:

- **RemoteOK** — JSON API (`https://remoteok.com/api`)
- **Remotive** — JSON API (`https://remotive.com/api/remote-jobs`)
- **WeWorkRemotely** — RSS feeds for Programming, Design, and Marketing
  (`https://weworkremotely.com/categories/remote-{programming,design,marketing}-jobs.rss`)

No authentication required. No headless browser. Pure `requests` + `BeautifulSoup`.

### Why this Actor

"Remote jobs" is a perennial high-volume search on Apify. The three sources above
are the largest publicly indexable remote-job boards and cover the majority of
fully-remote openings advertised on the public web. Users who would otherwise
pay for a single-source scraper can get all three in one run.

### Pricing (PPE)

- **$0.002 per listing** returned to the user.
- 7-day free trial, 30 trial runs.
- Apify margin 20% / operator rev share 80% per the lane BRIEF.

### Input

| Field                | Type    | Default  | Notes                                                 |
| -------------------- | ------- | -------- | ----------------------------------------------------- |
| `sources`            | string  | `"all"`  | `remoteok`, `remotive`, `weworkremotely`, or `all`    |
| `maxItems`           | integer | `200`    | Hard cap on total listings returned                   |
| `searchTerm`         | string  | `""`     | Case-insensitive substring filter (title/company/cat) |
| `includeDescription` | boolean | `true`   | Include full description text                         |
| `descriptionMaxChars`| integer | `4000`   | Truncate descriptions (0 = no truncation)             |

### Output

One record per listing, schema:

```json
{
  "title": "Senior Ruby on Rails Developer (AI-Augmented Engineering)",
  "company": "Proxify AB",
  "url": "https://weworkremotely.com/...",
  "source": "weworkremotely",
  "sourceId": "...",
  "category": "Back-End Programming",
  "tags": ["ruby", "rails"],
  "jobType": null,
  "salary": null,
  "location": "Anywhere in the World",
  "postedAt": "2026-08-27T...",
  "description": "Headquarters: Sweden ...",
  "scrapedAt": "2026-09-02T20:00:00Z"
}
```

### Run locally

```bash
cd src && python -m main
```

### Tests

```bash
cd /home/ubuntu/ops/lanes/apify/remote-jobs-scraper
python3 -m pytest tests/ -v
```

### Sources & legal

All three sources expose public, freely accessible data. Respect each source's
`robots.txt` and rate limits. The actor applies:

- Exponential backoff on 429/5xx (2s → 4s → 8s → 16s, capped 30s).
- Honors `Retry-After` header.
- Per-source failure isolation — one source failing does not kill the run.
- User-Agent identifies the scraper.

### License

MIT (operator-defined). Data ownership belongs to the respective source boards.

# Actor input Schema

## `sources` (type: `string`):

Which remote-job boards to scrape. 'remoteok', 'remotive', 'weworkremotely', or 'all'.

## `maxItems` (type: `integer`):

Hard cap on total listings returned (after source fetch, before filter). Protects against runaway scrapes.

## `searchTerm` (type: `string`):

Optional case-insensitive substring filter applied to title, company, tags, and category. Empty = no filter.

## `includeDescription` (type: `boolean`):

If true, include the full job description text. Set false to ship title/company/URL only (faster, smaller dataset).

## `descriptionMaxChars` (type: `integer`):

Truncate each description to this many characters (0 = no truncation).

## `proxyConfiguration` (type: `object`):

Proxies are usually not needed for these public feeds.

## Actor input object example

```json
{
  "sources": "all",
  "maxItems": 200,
  "searchTerm": "",
  "includeDescription": true,
  "descriptionMaxChars": 4000,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Unified dataset of remote jobs — title, company, source URL, category, tags, job type, location, salary, postedAt, description, etc.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("sequined_fan/remote-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("sequined_fan/remote-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call sequined_fan/remote-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sequined_fan/remote-jobs-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Ixf48NpqkswoOFRES/builds/P67NyWFEqawwVU0py/openapi.json
