# AI Training Jobs Scraper — Mercor, micro1, Turing (`tacps126/ai-training-jobs-scraper`) Actor

Scrape jobs from Mercor, micro1 and Turing — expert and AI-training gigs from Mercor, micro1 and Turing in one list. Pay, hours, skills and seniority in one clean format. Keyword filter and new-job monitoring.

- **URL**: https://apify.com/tacps126/ai-training-jobs-scraper.md
- **Developed by:** [Tapaswai Ashok Choudhary](https://apify.com/tacps126) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 1 total users, 0 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.80 / 1,000 jobs

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## AI Training Jobs Scraper — Mercor, micro1, Turing

Get expert and AI-training gigs from Mercor, micro1 and Turing in one list — as a clean spreadsheet-ready dataset in seconds. Export to JSON, CSV or Excel, send to Google Sheets or your database, or call it as an API from your own app or AI agent.

**Built for** recruiters and sourcers, job boards and aggregators, sales teams tracking who's hiring, analysts studying the job market, and job seekers who want every new opening first.

### Why this scraper

- **Fast.** The full list comes back in seconds. No browser is started — the Actor reads the site's own data directly.
- **Costs almost nothing to run.** It's a small native program that runs in 256 MB of memory, so the platform usage of a typical run is a fraction of a cent. You mostly pay for results, not for machine time.
- **Starts instantly.** A tiny container means runs begin in seconds, and the always-on API mode (below) answers without any start-up wait.
- **Handles rate limits for you.** When a site says “slow down” (HTTP 429), the Actor waits exactly as long as the site asks, then continues. Server hiccups, timeouts and dropped downloads are retried automatically, and request speed adapts to how the site responds.
- **Never loses or double-bills work.** Results are saved as each search finishes. If the platform moves your run to another server, it continues where it stopped without charging twice. Set a maximum charge and it stops cleanly at your limit, keeping everything collected so far.
- **Clean, consistent data.** Every job has the same fields — structured city/region/country with ISO code, remote/hybrid/on-site (only when the posting says so), employment type, seniority, years of experience, job function, skills, and salary converted to a yearly figure for easy comparison. The same format is used by all of our job scrapers, so you can mix sources in one table.
- **Simple to set up.** Press Start — no input is required. Add keywords to narrow the list. No accounts or keys needed.

### How to use it

1. Optionally enter **Search keywords** to narrow the list.
2. Optionally narrow results with the **Filters** (title words, location, posted within N days, remote only).
3. Press **Start**. Download the results from the **Output** tab, or connect an integration.

#### Example input

```json
{
  "query": "python",
  "maxItems": 100
}
```

### What you get

Board-specific extras for Mercor, micro1 and Turing: hourly rate, hours per week, domain, work arrangement.

| Field | Description |
|-------|-------------|
| `title` | Job title |
| `company` / `companySlug` | Company name and its board handle |
| `location` / `locations` | Location as listed, and every listed location |
| `city` / `region` / `country` / `countryCode` | Structured location (ISO country code) |
| `workMode` / `remote` | `remote`, `hybrid` or `onsite` — only when the posting says so |
| `employmentType` | `full_time`, `part_time`, `contract`, `internship`, `temporary`, `freelance` |
| `seniority` | `intern`, `junior`, `mid`, `senior`, `lead`, `principal`, `director`, `executive` |
| `yearsExperienceMin` / `yearsExperienceMax` | Required experience, when stated |
| `jobFunction` | Engineering, data, product, design, sales, marketing, AI training, … |
| `skills` | Skills named in the posting |
| `salary` | `{min, max, currency, period, annualMin, annualMax}` when pay is published |
| `educationLevel` / `visaSponsorship` | Degree asked for, and visa sponsorship when stated |
| `department` / `team` | Organization details, when available |
| `summary` / `description` | Short summary and full description |
| `url` / `applyUrl` | Posting and apply links |
| `postedAtIso` / `postedDaysAgo` | Posting date (ISO 8601) and its age in days |
| `source` / `scrapedAt` | Where the job came from and when it was collected |

Board labels are translated into one vocabulary (for example `FullTime`, `fulltime_permanent` and `Full-time` all become `full_time`), and the original label is kept alongside so nothing is lost. Each run's summary also shows **field coverage** — what share of jobs carry each field.

### Monitor new jobs on a schedule

Turn on **Only new jobs** and schedule the Actor (hourly, daily, weekly). Each run returns only postings earlier runs haven't delivered — you're charged only for new ones. Connect Slack, email, Google Sheets, Zapier, Make or a webhook to get them pushed to you automatically.

### Use it as an instant API

The Actor also runs in **Standby mode** — an always-ready endpoint that returns jobs straight in the response, ideal for apps, spreadsheets and AI agents. Find the URL on the **Standby** tab:

```bash
curl "https://<standby-url>/?query=python&maxItems=20" \
  -H "Authorization: Bearer <YOUR_APIFY_TOKEN>"
```

Any input field works as a query parameter, or `POST` the full input as JSON. The response is `{"success": true, "count": 20, "items": [ … ]}`.

### Pricing

Pay per result: about **$1 per 1,000 jobs** (lower on higher Apify plans) plus a tiny start fee. Filtered-out, duplicate and (with *Only new jobs*) already-seen jobs are free. Because the Actor needs so little memory and time, platform usage stays close to zero. See the **Pricing** tab for exact rates, and set a maximum charge per run to cap spend.

On Apify's **free plan** each run returns up to 500 jobs — plenty to try everything. Any paid Apify plan removes the limit.

### More job sources

Need several boards at once? **[Job Board Harvester — Multi-Platform](https://apify.com/tacps126/jobs-harvester)** combines all of them (and career-page links from any supported board) in one run, in the same format.

Single-board scrapers: [Greenhouse](https://apify.com/tacps126/greenhouse-jobs-scraper), [Lever](https://apify.com/tacps126/lever-jobs-scraper), [Ashby](https://apify.com/tacps126/ashby-jobs-scraper), [Workday](https://apify.com/tacps126/workday-jobs-scraper), [SmartRecruiters](https://apify.com/tacps126/smartrecruiters-jobs-scraper), [Recruitee](https://apify.com/tacps126/recruitee-jobs-scraper), [Personio](https://apify.com/tacps126/personio-jobs-scraper), [Teamtailor](https://apify.com/tacps126/teamtailor-jobs-scraper), [LinkedIn](https://apify.com/tacps126/linkedin-jobs-scraper), [Himalayas](https://apify.com/tacps126/himalayas-remote-jobs-scraper).

### FAQ

**Is it legal?** The Actor collects publicly listed job postings. You're responsible for how you use the data — for personal data in descriptions, follow the laws that apply to you (e.g. GDPR).

**Can I get the data in Excel or Google Sheets?** Yes — export from the Output tab, or add the Google Sheets integration to fill a sheet automatically after every run.

### Support

Found a problem or need a field that isn't there yet? Open an issue on the **Issues** tab with your input and what you expected — requests genuinely decide what gets built next.

# Actor input Schema

## `query` (type: `string`):

Optional — narrow the list to these keywords, e.g. `python`, `medicine`, `finance`. Leave empty for everything.

## `maxPages` (type: `integer`):

Optional cap on result pages to read. Leave empty to read until “Max jobs”.

## `maxItems` (type: `integer`):

Maximum number of jobs to return.

## `maxItemsPerSource` (type: `integer`):

Optional cap per marketplace, so every one gets its share of “Max jobs”.

## `titleIncludes` (type: `array`):

Keep a job only if its title contains any of these words (case-insensitive), e.g. `engineer`, `developer`.

## `titleExcludes` (type: `array`):

Drop a job whose title contains any of these words (case-insensitive), e.g. `intern`, `manager`.

## `locationIncludes` (type: `array`):

Keep a job only if its location contains any of these (case-insensitive), e.g. `Berlin`, `Remote`.

## `postedWithinDays` (type: `integer`):

Keep only jobs posted in the last N days. Jobs whose board doesn't publish a date are kept.

## `remoteOnly` (type: `boolean`):

Keep only jobs the posting itself marks as remote.

## `onlyNew` (type: `boolean`):

Return only jobs that no earlier run with the same input has delivered. Schedule the Actor and you pay only for new jobs.

## `stateKey` (type: `string`):

Optional name for the “only new” history, so several monitors can run side by side. Leave empty to derive it from the input.

## `proxyConfiguration` (type: `object`):

Not needed for Mercor, micro1 and Turing in most runs. Turn on if you schedule very large or very frequent runs.

## `enrich` (type: `boolean`):

Normalize every job into one schema: structured location + country code, work mode, employment type, seniority, years of experience, job function, skills, annualized salary, education and visa signals.

## `dedupe` (type: `boolean`):

Drop duplicate postings (by url + title + company).

## `tenantId` (type: `string`):

Optional label used to keep separate histories for different clients or teams.

## `debug` (type: `boolean`):

Save extra diagnostics with the run (for support requests).

## Actor input object example

```json
{
  "query": "python",
  "maxItems": 200,
  "remoteOnly": false,
  "onlyNew": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "enrich": true,
  "dedupe": true,
  "debug": false
}
```

# Actor output Schema

## `jobs` (type: `string`):

Every job posting found, in one consistent format.

## `summary` (type: `string`):

Counts, warnings, and whether the run stopped early (e.g. at your spending limit).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "python"
};

// Run the Actor and wait for it to finish
const run = await client.actor("tacps126/ai-training-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "python" }

# Run the Actor and wait for it to finish
run = client.actor("tacps126/ai-training-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "python"
}' |
apify call tacps126/ai-training-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tacps126/ai-training-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/M9fyLtI9sK4Ln1goi/builds/bxzqCrFGxQrnZOZmb/openapi.json
