# TalentLyft Jobs Monitor & Scraper (`cliqtomedia/talentlyft-jobs-monitor-scraper`) Actor

Collect public TalentLyft jobs from known company boards. Get job text, dates and links. Compare full public sitemap snapshots to find new, updated and closed roles.

- **URL**: https://apify.com/cliqtomedia/talentlyft-jobs-monitor-scraper.md
- **Developed by:** [Cliqto Media](https://apify.com/cliqtomedia) (community)
- **Categories:** Jobs, Automation
- **Stats:** 2 total users, 1 monthly users, 50.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.25 / 1,000 saved job or changes

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Collect job listings from known TalentLyft companies. Get titles, job descriptions, dates and links. Compare full public sitemap snapshots to find new, edited and closed roles.

![Read jobs, save a snapshot, compare changes](https://api.apify.com/v2/key-value-stores/Sj6FdnVkhvfq0Vizo/records/hero.png)

One run reads up to 3 company boards. Tested live: 11 jobs from 3 boards. The technical cap is 50 scanned jobs. Larger live boards have not been proven. The price is $0.04 per start plus $0.00225 per saved Dataset row. Empty and unchanged runs still have the start fee.

### Contents

[What it does](#what-it-does) · [Data](#what-data-you-get) · [Quick start](#quick-start) · [Inputs](#all-inputs) · [Examples](#input-examples) · [Output](#output-example) · [Fields](#all-output-fields) · [Price](#price-and-examples) · [Monitoring](#schedule-and-monitoring) · [API](#api-and-export) · [Limits](#limits-and-partial-results) · [Help](#troubleshooting) · [FAQ](#faq) · [Support](#support)

### What it does

1. Read each company's public `sitemap.xml`.
2. Open all public `/jobs/` pages in that list.
3. Save job data and a full `SNAPSHOT` in the Key-value store.
4. If you provide a previous full snapshot, compare jobs by `board` and `jobId`.

A full snapshot covers all job URLs in that public sitemap. A job that the company leaves out of its sitemap cannot be found. Source pages may change during the scan.

### What data you get

Dataset: one job per row, or one new, updated or closed job per row in `changes` mode. Each row has public job text, source dates, a link and a stable content hash. Missing source metadata is `null`. We do not guess salary or remote work.

Key-value store: `SNAPSHOT` for the next comparison, `CHANGES` for all differences, and `RUN_SUMMARY` plus `OUTPUT` for counts and source status. KVS copies are included in the start fee. Only Dataset rows count as row events.

### Quick start

1. Open Input and add `decathlon-philippines-inc` to `boards`.
2. Keep `mode` as `snapshot` and `maxResults` as `50`.
3. Set run memory to 256 MB and timeout to 120 seconds. Set a charge cap of $0.20.
4. Click Start. Open Dataset and check a job's `title`, `descriptionText` and `url`.
5. Open `RUN_SUMMARY`. Check `status` and `snapshotComplete` before reusing `SNAPSHOT`.

This input returned 8 jobs in our live check on 4 October 2026. The company's current list can change.

### All inputs

| Key | Type | Default / required | Range | Effect and advice |
| --- | --- | --- | --- | --- |
| `boards` | array of strings | Required; prefill is `decathlon-philippines-inc` | 1–3 entries; each 1–63 lowercase letters, digits or hyphens | Known subdomains only. Do not add `.talentlyft.com` or URLs. Duplicates are merged. |
| `mode` | string | `snapshot` | `snapshot`, `changes` | Use snapshot first. Changes needs a previous full snapshot. |
| `maxResults` | integer | `50` | 1–50 | Cap Dataset rows after the full scan. Does not reduce requests or the snapshot. Use 1 to inspect one row. |
| `keywords` | array of strings | `[]` | Up to 10 strings, 1–100 characters each | Case-insensitive title filter. A title needs any one word. Applied after the full scan, including closed rows. Use an empty list for all titles. |
| `previousSnapshot` | object | Omitted | Complete version 1 TalentLyft snapshot; same boards; ≤50 jobs; ≤6 MiB | Paste the entire saved SNAPSHOT object. Required for changes. Invalid, edited or partial snapshots are rejected. |

Extra input keys are rejected. A filter and an output cap do not reduce full scan work. They do not remove jobs from the saved snapshot.

### Input examples

#### Read one company

```json
{"boards":["decathlon-philippines-inc"],"mode":"snapshot","maxResults":50,"keywords":[]}
```

#### Read more boards

```json
{"boards":["decathlon-philippines-inc","colliers-international","careers"],"mode":"snapshot","maxResults":50}
```

#### Filter titles

```json
{"boards":["decathlon-philippines-inc"],"keywords":["Manager"],"maxResults":10}
```

#### Compare two full scans

Run snapshot mode first. Copy the **whole** `SNAPSHOT` object from its Key-value store. In the next input use the same `boards`, set `mode` to `changes`, and paste that object as `previousSnapshot`. Do not paste just the jobs array or use a filtered Dataset as the previous snapshot.

### Output example

This is one trimmed row from the one-company snapshot input. Description text is shortened here; real rows hold the full available job description within the safety cap. Hash and timestamp come from the original full content.

```json
{
  "board": "decathlon-philippines-inc",
  "jobId": "v71",
  "title": "Department Manager",
  "company": "Decathlon Philippines Inc.",
  "location": "Metro Manila, Cebu, Iloilo, Clark, Philippines",
  "country": "PH",
  "employmentType": "FULL_TIME",
  "datePosted": "2026-06-25",
  "url": "https://decathlon-philippines-inc.talentlyft.com/jobs/department-manager-v71",
  "applyUrl": "https://decathlon-philippines-inc.talentlyft.com/jobs/department-manager-v71",
  "descriptionHtml": "<p>We are looking for <strong>Future Store Managers</strong>!</p>",
  "descriptionText": "We are looking for Future Store Managers!",
  "contentHash": "be6c978faf457eac95e888085abd201537f1f90e3718314aa307369b63b00544",
  "changeType": null,
  "scrapedAt": "2026-10-04T16:25:11.953Z"
}
```

### All output fields

| Field | Type | Can be null? | Example | Meaning |
| --- | --- | --- | --- | --- |
| `board` | string | No | `decathlon-philippines-inc` | Known company subdomain. |
| `jobId` | string | No | `v71` | Source URL suffix. Use board + jobId as the stable key. |
| `title` | string | No | `Department Manager` | Public job title. |
| `company` | string | Yes | `Decathlon Philippines Inc.` | Source employer name. |
| `location` | string | Yes | `Metro Manila, Cebu, Iloilo, Clark, Philippines` | First source job location text. |
| `country` | string | Yes | `PH` | First source location country. |
| `employmentType` | string | Yes | `FULL_TIME` | Source employment type. |
| `datePosted` | string | Yes | `2026-06-25` | Source posting date. Old dates stay old. |
| `url` | string | No | `https://decathlon-philippines-inc.talentlyft.com/jobs/department-manager-v71` | Public job page. |
| `applyUrl` | string | No | `https://decathlon-philippines-inc.talentlyft.com/jobs/department-manager-v71` | Same job page, which has the application form. |
| `descriptionHtml` | string | No | `<p>Role text...</p>` | Decoded source description. Executable tags and attributes are removed. |
| `descriptionText` | string | No | `Role text...` | Description as plain text with normalized spaces. |
| `contentHash` | string | No | `64 hex characters` | SHA-256 of stable source content; fetch time is excluded. |
| `changeType` | string | Yes | `updated` | new, updated or closed in changes mode; null in snapshot mode. |
| `scrapedAt` | string | No | `2026-10-04T16:25:11.953Z` | UTC fetch time in ISO format. |

#### Reports and snapshot fields

| Record | Fields | Meaning |
| --- | --- | --- |
| `OUTPUT`, `RUN_SUMMARY` | `schemaVersion` integer; `status` string; `complete`, `snapshotComplete` booleans | Overall output status and full scan state. |
| Same reports | `counts` object: `scanned`, `selected`, `delivered`, `new`, `updated`, `closed` integers | Scanned current jobs; selected before output cap; saved Dataset rows; total differences before filters. |
| Same reports | `boards` array: `board`, `status`, `sourceCount`, `scanned`, `error` | Per-board result. Source count and error may be null. |
| Same reports | `errors` array of strings; `datasetId` string or null; `snapshotKey`, `changesKey`, `sourceCommit`, `createdAt` strings | Diagnostics, output location and scan time. |
| Same reports | `limits` object; `metrics` object: `requests`, `retries`, `bytes`, `durationMs` integers | Safety caps and measured HTTP work, including retries. |
| `SNAPSHOT` | `schemaVersion` integer; `source` string; `complete` boolean; `boards` array; `jobs` array; `createdAt` string | Unfiltered scan. Source is `talentlyft-public-sitemap`. Reuse only when complete is true. |
| `CHANGES` | `complete` boolean; `counts` object with `new`, `updated`, `closed`; `rows` array | All computed differences before title filter and Dataset cap. These rows use the same job fields. |

The Dataset is capped separately. `snapshotComplete: true` with `status: partial` can mean the full scan was saved but Dataset output reached a cap. `CHANGES` can have more rows than Dataset. No extra row fee applies to it.

### Price and examples

You pay **$0.04 per Actor start + $0.00225 per saved Dataset row**. A closed job in changes mode is also one row. The start fee pays for the full scan attempt, including valid empty results, filters with no match, source errors and unchanged monitoring runs. Do not start with a cap below the start fee. With 256 MB, there is one start event.

| Scenario | Saved rows | Event charge |
| --- | ---: | ---: |
| Empty board, filter with no match or unchanged scan | 0 | $0.04000 |
| Inspect one saved row | 1 | $0.04225 |
| One Decathlon board in our live check | 8 | $0.05800 |
| Ten saved rows | 10 | $0.06250 |
| Tested three-board scan | 11 | $0.06475 |

100 or 1,000 rows in one run are outside the cap and have not been tested. We do not quote those batch sizes. For paid users, platform usage is included in PPE and paid by the developer. Free-plan runs follow Apify's platform rules. Private owner runs measure platform usage; they do not prove a paying customer's charge.

A partial run charges the start plus the Dataset rows saved before the stop. A failed or cancelled run may still have the start fee and saved rows. Unsaved rows are not billed. Storage writes and charging are not one transaction; do not resurrect a run to get an exactly-once guarantee. Check saved rows and the charge log first.

### Use cases

| Task | Use these fields | Next action |
| --- | --- | --- |
| Fill a small job board | `title`, `company`, `location`, `descriptionText`, `url` | Export Dataset to your job system. |
| Watch a target company | `board`, `jobId`, `changeType`, `contentHash` | Compare full snapshots and send relevant changes to your own workflow. |
| Review hiring signals | `employmentType`, `datePosted`, `country` | Group roles in Sheets or a database. Null means the source did not give a value. |

### Schedule and monitoring

Repeat snapshot scans on the same boards. To compare runs, your workflow must pass the last complete `SNAPSHOT` as `previousSnapshot` and use `changes` mode. A saved Task with a fixed old snapshot always compares against that old scan; it does not update itself. A schedule alone does not pass the last snapshot automatically. Store the last complete snapshot in your workflow and update it after a successful scan. Unchanged monitoring has 0 Dataset rows and the $0.04 start fee.

A source failure prevents closed-job output. Retry with the last complete snapshot after the source is available. Do not replace a trusted snapshot with one whose complete flag is false.

### API and export

Call this Actor with the same JSON input through the [Apify API](https://docs.apify.com/api/v2). Use `POST /v2/acts/cliqtomedia~talentlyft-jobs-monitor-scraper/runs?build=latest&memory=256&timeout=120&maxTotalChargeUsd=0.20`. Authenticate with your existing Apify token. Keep it out of shared inputs and logs.

Open the run's Dataset to export JSON, CSV or Excel. Read `SNAPSHOT` from that run's Key-value store for the next comparison. SDK and no-code clients use the same inputs and output records. An MCP client can pass that JSON through the [Apify MCP server](https://docs.apify.com/platform/integrations/mcp). You choose and configure integrations; this Actor does not create them.

### Limits and partial results

- Only known `*.talentlyft.com` company subdomains. No company discovery, custom domains, protected customer API or candidate data.
- One public sitemap file per board. Sitemap indexes, missing/blocked maps, redirects and unexpected formats are errors. Jobs absent from the map are outside this snapshot.
- Technical limits: 3 boards, 50 scanned jobs, 50 Dataset rows, 65 HTTP attempts, 2 MiB per response, 16 MiB combined response bytes, 100 kB sanitized description, 90-second source deadline, 256 MB; recommended run timeout 120 seconds. The 300-second Store test and larger finite positive timeouts are supported without extending the 90-second source deadline.
- Live capacity proved: 11 jobs across 3 boards. A 50-job fixture checks the code cap; it is not a live-source capacity promise.
- Only the first location is returned. Missing fields stay null. A generic open-application posting is kept if the company publishes it as a job.
- `complete`: scan and selected output saved. `partial`: some data exists but a source, output or charge limit stopped completion. `not_found`: all board sitemaps returned 404. `error`: invalid input or all sources failed.
- A run can have Apify status SUCCEEDED with summary partial or not_found. Always check `RUN_SUMMARY`. A forced platform kill or storage failure may prevent final reports.

### Troubleshooting

#### No Dataset rows

Check `RUN_SUMMARY.status`. Complete with scanned 0 means the public sitemap has no job URLs. Complete with scanned greater than 0 may mean no title matched or no job changed. Error or not_found is a source/input problem, not an empty company list.

#### Partial result

Check `errors` and each board status. `OUTPUT_LIMIT` means increase `maxResults` within 50. `CHARGE_LIMIT` means the charge cap stopped saved rows. `SCAN_LIMIT` means split the company list; a single board above 50 jobs is not supported. Never use an incomplete snapshot as the next baseline.

#### Invalid previous snapshot

Copy the whole complete `SNAPSHOT` from a prior run, keep its fields intact and use exactly the same set of boards. Do not use `CHANGES` or Dataset rows.

#### Source error or timeout

Open the public board in your browser. Check the slug, then retry a small input later. Redirects, inaccessible pages and changed formats need a new valid source or a parser fix; they are not reported as successful zero jobs.

### FAQ

#### Do I need a TalentLyft API key?

No. This Actor reads public sitemap and job pages. It does not read customer API endpoints or applicant records.

#### Can it find every TalentLyft company?

No. Add company subdomains you already know.

#### Does the output cap make the snapshot smaller?

No. The scan happens first. Title filters and maxResults affect Dataset only. Check snapshotComplete to know whether the scan was full.

#### Why is an old job still present?

The source sitemap still lists it. Source dates and old public jobs are preserved. Check the original job link before making a business decision.

#### Are descriptions safe to display as HTML?

Executable tags and attributes are removed. Use your application's own HTML sanitizer when displaying any external text.

### Support

Use the Issues tab on this Actor's Apify page. Send the run ID, the board slug and the error code from RUN_SUMMARY. Do not send tokens, private applicant data or account credentials.

# Actor input Schema

## `boards` (type: `array`):

1–3 known TalentLyft subdomains, without .talentlyft.com. Duplicates are merged. Custom domains and full URLs are not supported.

## `mode` (type: `string`):

snapshot returns current jobs. changes requires a full previousSnapshot and returns new, updated and closed jobs.

## `maxResults` (type: `integer`):

1–50, default 50. Caps Dataset output after the full scan. Does not reduce source requests. Snapshot stays full when output is capped.

## `keywords` (type: `array`):

Up to 10 title words, OR, case-insensitive. Default empty: all titles. Filtering occurs after the full scan and does not change the saved snapshot.

## `previousSnapshot` (type: `object`):

Paste the whole SNAPSHOT object from a previous run with complete=true and the same boards. Required for changes. Maximum 50 jobs and 6 MiB. A partial snapshot is rejected.

## Actor input object example

```json
{
  "boards": [
    "decathlon-philippines-inc"
  ],
  "mode": "snapshot",
  "maxResults": 50,
  "keywords": []
}
```

# Actor output Schema

## `jobs` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `snapshot` (type: `string`):

No description

## `changes` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "boards": [
        "decathlon-philippines-inc"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("cliqtomedia/talentlyft-jobs-monitor-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "boards": ["decathlon-philippines-inc"] }

# Run the Actor and wait for it to finish
run = client.actor("cliqtomedia/talentlyft-jobs-monitor-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "boards": [
    "decathlon-philippines-inc"
  ]
}' |
apify call cliqtomedia/talentlyft-jobs-monitor-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cliqtomedia/talentlyft-jobs-monitor-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CsL9Wm2di5YUsndQ0/builds/3HOwGW6UHDWOzZU3Z/openapi.json
