# AI Training Jobs Scraper: Mercor, Alignerr, micro1, xAI & More (`wulfcare/ai-training-jobs-scraper`) Actor

AI training, AI tutor and data-annotation jobs from Mercor, Alignerr, micro1, xAI, Handshake AI, Appen, Remotasks and HumanSignal in one list: hourly pay as numbers, hours, location, posted date and apply link. Filter by keyword and pay, or get only new jobs.

- **URL**: https://apify.com/wulfcare/ai-training-jobs-scraper.md
- **Developed by:** [Wulfcare Data](https://apify.com/wulfcare) (community)
- **Categories:** Jobs, AI, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## AI Training Jobs Scraper

Every open **AI training, AI tutor and data annotation job** from the platforms that hire people to teach AI models, in one table:

| Platform | What it lists | Open jobs (Oct 2026) |
|---|---|---|
| **Alignerr** (Labelbox) | Expert AI trainers in every field and language | ~550 roles (5,600 city listings) |
| **micro1** | AI trainers and experts for frontier labs, plus micro1's own roles | ~375 |
| **Mercor** | Domain experts (law, medicine, finance, coding...) evaluating AI, often $60-200/hr | ~330 |
| **HumanSignal** (Label Studio) | AI trainers and annotators | ~35 |
| **Appen** | AI training and data collection projects | ~35 |
| **xAI** | AI Tutors for Grok (languages, STEM, coding, specialist fields) | ~35 |
| **Handshake AI** | AI trainers and fellows for top labs | a few |
| **Remotasks** (Scale AI) | Contributor and AI trainer roles | a few |

Each job has the **pay range as numbers** (min, max, currency and per hour / year / task), the date it was posted, the full description and the **apply link**, plus **hours per week**, commitment, remote or onsite, location, **eligible countries**, skills and category wherever the platform publishes them. Every platform is cut down to the same columns, so the results sort and filter as one list.

**Filter the way a job seeker or recruiter would**: keywords (`law`, `medicine`, `python`, `spanish`), excluded words, **minimum hourly pay**, remote only, and *posted in the last 7 days*. Turn on **Only new jobs** and a daily schedule returns only the jobs it hasn't sent you before.

### What people use it for

- **Job alerts**: a daily schedule with *Only new jobs* that sends new $50+/hr legal or medical AI-training gigs to Slack, email or a Google Sheet
- **Job boards and newsletters**: feed an AI-jobs site, Discord, Telegram channel or weekly newsletter with fresh listings from every platform
- **Pay research**: what AI labs pay for each skill, language and country, and how that changes over time
- **Recruiting and market tracking**: which fields and languages the AI labs are hiring for right now, and on which platforms
- **Freelancers and agencies**: find every open project that fits your expertise without checking eight sites

### How to use it

1. Pick the **platforms** (all of them by default).
2. Optionally add **keywords**, a **minimum hourly pay**, **Remote only** or **Posted within**.
3. Run it, then download JSON, CSV or Excel, or call the API.

A run with every platform and no filters takes under a minute. All the platforms together have about **1,350 open jobs**. With filters it only charges for the jobs that match.

### What you get

One row per job, newest first:

```json
{
  "platform": "Mercor",
  "jobId": "list_AAABoKOQUds_uf4uJMtGK64s",
  "title": "Legal Expert",
  "company": null,
  "category": "Law",
  "skills": null,
  "payMin": 60,
  "payMax": 150,
  "payCurrency": "USD",
  "payPeriod": "hour",
  "payText": "$60-150/hr",
  "hoursPerWeek": 40,
  "commitment": "part-time",
  "workArrangement": "remote",
  "location": "Remote",
  "eligibleCountries": null,
  "postedAt": "2026-09-15T05:35:39Z",
  "url": "https://work.mercor.com/jobs/list_AAABoKOQUds_uf4uJMtGK64s/legal-expert",
  "description": "The research teams behind the best-known AI models come to Mercor for legal judgment...",
  "referralBonus": 600,
  "openings": null,
  "isNew": null,
  "scrapedAt": "2026-10-01T09:01:29Z",
  "error": null
}
```

- **payPeriod** is `hour`, `year`, `month`, `week`, `task` or `one-time` (`null` if the platform doesn't say). `payText` is the pay exactly as the platform shows it.
- Every column is on every row (`null` when a platform doesn't publish that value), so the table never changes shape.
- **isNew** is `true` on every row of an *Only new jobs* run.
- `company` is the client or employer when the platform names it (most platforms don't say which AI lab a project is for).
- **openings** is the number of positions when the platform says, and for Alignerr the number of city listings of the role.

### Pricing

**Free during launch**: no start fee, no rental, no per-job charge. You only pay Apify's own platform usage, which is tiny: a full run of every platform (about 1,350 jobs) uses roughly $0.001 of compute.

If the Actor moves to per-job pricing later, you'll get Apify's 14-day notice first.

### Examples

**Law and legal AI-training jobs paying $50/hr or more, from the last 30 days**

```json
{ "keywords": ["law*", "legal", "attorney", "lawyer"], "minHourlyRate": 50, "postedWithin": "30 days" }
```

**Daily alert: new remote coding jobs (for a schedule)**

```json
{ "keywords": ["python", "software", "coding", "javascript"], "remoteOnly": true, "onlyNewJobs": true }
```

**Every AI tutor job for Spanish speakers**

```json
{ "keywords": ["spanish", "español"] }
```

**Every Mercor and micro1 job, without descriptions (small rows)**

```json
{ "platforms": ["mercor", "micro1"], "maxJobs": 0, "includeDescriptions": false }
```

**Everything from every platform**

```json
{ "maxJobs": 0 }
```

### Good to know

- **Keywords** match whole words in the title, category, skills and client name: `law` doesn't match `lawn`. End a word with `*` to match longer forms too (`law*` = law, laws, lawyer). Turn on *Also search job descriptions* to look in the full text as well; that finds more but looser matches.
- **Only new jobs** remembers which jobs earlier runs returned (for 180 days), in a key-value store called `ai-training-jobs-monitor` in your account. Runs with the same platforms and filters share one list. Give runs a **Monitor name** to share a list across different filters, or to keep two alerts apart.
- **Alignerr lists most roles once per city**: one *Software Engineer (AI Training)* role has 300+ copies, all remote, each with a reworded summary. You get **one row per role** (about 550), with the number of copies in `openings`, so you don't pay for duplicates. Turn on *Every Alignerr city listing* to get all 5,600 copies as separate rows.
- Alignerr's list has a short summary and no date, so the Actor opens the page of each Alignerr role it returns to get the full description, the date it was first listed and the pay. Alignerr re-posts old roles every day, so the date is when the role was **first** listed, not the latest re-post. *Posted within* opens up to 2,000 Alignerr pages to check dates; add keywords to narrow big searches. Listings whose page is gone (closed roles) are skipped.
- **micro1**'s list has no description, so the Actor also opens each micro1 job it returns for the full description, the number of openings and the referral reward. Turn off *Full job details* for a faster run without these page visits.
- **Minimum hourly pay** compares the top of each hourly range, in the job's own currency (nearly all of these jobs pay in USD). Yearly salaries, per-task pay and jobs without a stated rate are left out when it's set.
- **Everything comes from the platforms' public job listings**, as a logged-out visitor sees them. No account, no login and no personal data: job listings only.
- Platform missing? Open an issue with its jobs page and it may be added.

### Other Actors by Wulfcare Data

- [Google Hotels Prices Scraper](https://apify.com/wulfcare/google-hotels-scraper): every booking site's price for any hotel and dates, and hotel searches
- [Free YouTube Scraper](https://apify.com/wulfcare/youtube-scraper): YouTube videos, Shorts, channels and searches with views, likes, comments and channel stats, free
- [Google Jobs Scraper](https://apify.com/wulfcare/google-jobs-scraper): job listings from Google Jobs, for any search and location
- [ChatGPT Search Scraper & AI Brand Visibility Tracker](https://apify.com/wulfcare/chatgpt-search-scraper): what ChatGPT answers and cites for any prompt, with a brand share-of-voice report
- [Google Trends Scraper](https://apify.com/wulfcare/google-trends-scraper): interest over time, rising queries and Trending Now
- [YouTube Transcript Scraper](https://apify.com/wulfcare/youtube-transcript-scraper): transcripts of videos, channels, playlists and searches
- [Google Play Scraper](https://apify.com/wulfcare/google-play-scraper) and [App Store Scraper](https://apify.com/wulfcare/app-store-scraper): reviews, app details and charts

Found a problem or need a field? Open an issue on the Issues tab.

# Actor input Schema

## `platforms` (type: `array`):

Which AI-training platforms to read. Leave empty for all of them.

## `keywords` (type: `array`):

Return only jobs that mention at least one of these in the title, category, skills or client name (e.g. 'law', 'medicine', 'python', 'spanish'). Whole words: 'law' does not match 'lawn'. End a word with \* to allow endings ('math\*' matches math and mathematics). Leave empty for every job.

## `excludeKeywords` (type: `array`):

Drop jobs whose title, category or skills mention any of these (e.g. 'audio', 'transcription').

## `searchDescriptions` (type: `boolean`):

Match keywords inside the full job description too. Finds more jobs but also loose matches (a physics job that mentions 'laws of motion' would match 'law').

## `minHourlyRate` (type: `integer`):

Keep only hourly jobs whose top rate is at least this much. Jobs paid per task, one-off or without a stated rate are dropped when this is set.

## `remoteOnly` (type: `boolean`):

Keep only jobs listed as remote.

## `postedWithin` (type: `string`):

Only jobs first posted in this period, e.g. '24 hours', '7 days', '2 weeks', or since a date like 2026-09-01. Alignerr dates come from each job page (needs 'Full job details').

## `onlyNewJobs` (type: `boolean`):

Return, and charge for, only jobs that earlier runs with the same filters have not returned yet. Perfect for a daily schedule that feeds Slack, email or a sheet. The list of seen jobs is kept in a key-value store named 'ai-training-jobs-monitor' in your own account.

## `monitorName` (type: `string`):

Optional. Runs with the same monitor name share one 'already seen' list. By default the list is tied to your platform and filter choices.

## `maxJobs` (type: `integer`):

Stop after this many jobs (newest first). 0 = no limit. All platforms together have about 1,350 distinct jobs (about 6,400 with every Alignerr city listing).

## `includeDescriptions` (type: `boolean`):

Full job description text in each row. Turn off for smaller rows.

## `fullDetails` (type: `boolean`):

Open the page of each returned Alignerr and micro1 job. Alignerr's list has only a short summary and no date; micro1's list has no description or number of openings. Turn off for a faster run with less detail.

## `alignerrAllLocations` (type: `boolean`):

Alignerr lists most remote roles once per city (one role can have 300+ identical copies). By default you get one row per role, with the number of copies in 'openings'. Turn on to get every copy as its own row (about 5,600).

## `proxyConfiguration` (type: `object`):

Not needed: these platforms publish their job lists openly. Only set a proxy if a platform starts blocking your runs.

## Actor input object example

```json
{
  "platforms": [
    "mercor",
    "alignerr",
    "micro1",
    "xai",
    "handshake",
    "appen",
    "remotasks",
    "humansignal"
  ],
  "keywords": [],
  "excludeKeywords": [],
  "searchDescriptions": false,
  "remoteOnly": false,
  "onlyNewJobs": false,
  "maxJobs": 200,
  "includeDescriptions": true,
  "fullDetails": true,
  "alignerrAllLocations": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "platforms": [
        "mercor",
        "alignerr",
        "micro1",
        "xai",
        "handshake",
        "appen",
        "remotasks",
        "humansignal"
    ],
    "keywords": [],
    "excludeKeywords": [],
    "maxJobs": 200,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("wulfcare/ai-training-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "platforms": [
        "mercor",
        "alignerr",
        "micro1",
        "xai",
        "handshake",
        "appen",
        "remotasks",
        "humansignal",
    ],
    "keywords": [],
    "excludeKeywords": [],
    "maxJobs": 200,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("wulfcare/ai-training-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "platforms": [
    "mercor",
    "alignerr",
    "micro1",
    "xai",
    "handshake",
    "appen",
    "remotasks",
    "humansignal"
  ],
  "keywords": [],
  "excludeKeywords": [],
  "maxJobs": 200,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call wulfcare/ai-training-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,wulfcare/ai-training-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aUlLAbWvuWcowc7zV/builds/XfsudvTBhtDDG1XJb/openapi.json
