# Job Feed Change Detector (`ceddl/job-feed-change-detector`) Actor

Compare baseline and current job-feed exports to find new, changed, removed, and duplicate postings with stable keys and field-level evidence.

- **URL**: https://apify.com/ceddl/job-feed-change-detector.md
- **Developed by:** [Cedric Günther](https://apify.com/ceddl) (community)
- **Categories:** Jobs, Automation, Open source
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $30.00 / 1,000 job snapshots compareds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Job Feed Change Detector

Compares two supplied job-feed snapshots and returns deterministic added, removed, changed, and duplicate-posting evidence without contacting a job board. It is built for recruiting-data teams, job-board operators, ETL and monitoring engineers. The main result is structured, deterministic evidence that can be consumed from the default dataset or an Apify automation.

### When to use this Actor

- Detect newly listed and removed jobs between scheduled exports.
- Review field-level changes to titles, locations, descriptions, or other selected fields.
- Find duplicate stable identifiers or fallback keys before loading a feed downstream.

### How it works

- Validate both bounded inline snapshots and normalize scalar source fields.
- Choose a stable key from sourceId, canonical HTTP(S) URL, or normalized company-title-location in that order.
- Compare selected fields, sort evidence deterministically, and persist dataset plus complete REPORTS output before charging.

The Actor validates only the declared product contract. It does not infer facts outside the supplied data or claim outcomes that the source material cannot prove.

### Quick start

1. Open the Actor's **Input** tab or create a Task from one of the public examples.
2. Paste or adapt this bounded example.
3. Click **Start** and inspect the default dataset plus the output links shown on the run page.

```json
{
  "comparisonId": "weekly-engineering-feed",
  "baseline": [
    {
      "sourceId": "J-1",
      "title": "Data Analyst",
      "company": "Northwind",
      "location": "Berlin",
      "url": "https://jobs.example/J-1"
    }
  ],
  "current": [
    {
      "sourceId": "J-1",
      "title": "Senior Data Analyst",
      "company": "Northwind",
      "location": "Berlin",
      "url": "https://jobs.example/J-1"
    },
    {
      "sourceId": "J-2",
      "title": "Data Engineer",
      "company": "Northwind",
      "location": "Remote",
      "url": "https://jobs.example/J-2"
    }
  ],
  "trackedFields": [
    "title",
    "location"
  ]
}
```

Expected result: one changed J-1 record, one added J-2 record, and a summary for the supplied comparison.

### Input

The quick-start example is intentionally small. These are the material controls; the Input tab remains authoritative for the complete current schema.

| Field | Purpose and format | Default | Important bounds or interaction |
|---|---|---|---|
| `comparisonId` | Stable caller-defined label copied to every output record so repeated feed comparisons can be routed and reconciled. | No implicit default | minimum length 1; maximum length 120 |
| `baseline` | Earlier bounded posting snapshot; each record should carry a stable source ID or canonical URL whenever available. | No implicit default | maximum items 1000 |
| `current` | Later bounded posting snapshot compared with baseline under the same identity and field rules. | No implicit default | maximum items 1000 |
| `trackedFields` | Scalar posting fields evaluated for CHANGED evidence; identity fields are still used even when not listed here. | No implicit default | minimum items 1; maximum items 30 |
| `maxRecordsPerSnapshot` | Optional lower safety cap applied independently to baseline and current without raising the product maximum. | No implicit default | minimum 1; maximum 1000 |

Unknown top-level fields and invalid field combinations fail validation rather than being guessed.

### Output

The default dataset contains typed records. The run's Output tab links the dataset and any key-value-store reports declared by the current output schema.

| Field | Meaning |
|---|---|
| `recordType` | Discriminates job-change, duplicate-group, and job-summary records. |
| `comparisonId` | Current dataset field. |
| `stableKey` | Current dataset field. |
| `changeType` | Current dataset field. |
| `changedFields` | Deterministic before/after evidence for each selected field that changed. |
| `snapshot` | Current dataset field. |
| `indexes` | Source indexes participating in a duplicate stable-key group. |
| `count` | Current dataset field. |
| `engineVersion` | Current dataset field. |

Representative current-schema dataset item:

```json
{
  "recordType": "job-change",
  "comparisonId": "weekly-engineering-feed",
  "stableKey": "source:J-2",
  "changeType": "ADDED",
  "engineVersion": "1.0.0"
}
```

When the snapshots are equivalent, the run succeeds with a job-summary record and zero job-change records; zero change is not treated as failure.

### Pricing and billing

This Actor uses `PAY_PER_EVENT`; platform usage is included in event prices. A charge is eligible only after the billable unit described below is durably completed. Validation failures and the non-billable failure classes in the product contract do not emit the custom completion event. The current live policy uses the same event price at every Store tier; no tier discount is active. The Apify **Pricing** tab is authoritative if a later approved pricing change takes effect.

| Event | What triggers it | FREE | BRONZE | SILVER | GOLD | PLATINUM | DIAMOND |
|---|---|---:|---:|---:|---:|---:|---:|
| `job-snapshots-compared` | One baseline/current job snapshot pair converted into durable change evidence. | $0.03000000 | $0.03000000 | $0.03000000 | $0.03000000 | $0.03000000 | $0.03000000 |
| `apify-actor-start` | Platform-managed Actor start event. | $0.00005000 | $0.00005000 | $0.00005000 | $0.00005000 | $0.00005000 | $0.00005000 |

The Actor does not have Task-specific prices: public Tasks use this same live Actor pricing. Third-party costs are not implied; see the data and security section for external services actually contacted.

### Limits and bounds

- Each baseline and current snapshot is capped at 1,000 postings; a lower maxRecordsPerSnapshot may be selected.
- Only scalar source fields are preserved and compared; nested arbitrary documents are outside the contract.
- The Actor compares supplied snapshots only and never crawls a job board or infers when a posting actually changed.

These are product-facing limits, not targets. Use smaller inputs when you need faster feedback or simpler evidence.

### Failure and edge-case behavior

- Invalid top-level input, missing required posting fields, unsupported values, or an over-limit snapshot fails closed before the custom event.
- An empty snapshot is valid when supplied explicitly and can truthfully produce additions or removals.
- Duplicate keys are represented as duplicate-group output rather than silently merged.

Operationally:

- Use the same comparison context and stable source identifiers across scheduled exports.
- Removed means absent from the supplied current snapshot, not confirmed closed by the source site.

### Use with Tasks and automation

Public Tasks provide reusable saved inputs for distinct supported workflows. Start with the closest Example Task, review its visible fields and scope caveat, then save your own Task for schedules or repeated runs. Do not treat an Example Task as evidence that unsupported behavior exists.

- [Compare two job feed snapshots](https://apify.com/ceddl/job-feed-change-detector/examples/compare-job-feed-snapshots): Identify new, changed, removed, and duplicate job postings between two safe inline exports.
- [Detect new and removed job postings](https://apify.com/ceddl/job-feed-change-detector/examples/detect-new-and-removed-job-postings): Identify newly posted and removed jobs between two exported feed snapshots.
- [Find duplicate job postings](https://apify.com/ceddl/job-feed-change-detector/examples/find-duplicate-job-postings): Find duplicate listings inside a current job-feed export using stable posting keys.

### Integration and API usage

Every saved Task can be started manually, through the Apify API, or from an Apify schedule. Run-completion webhooks can notify a downstream system after output is durable. Actor-to-Actor calls should consume the typed dataset/output links instead of scraping the Store page.

- Schedule a saved Task after each feed export and send completion webhooks to an ingestion or alerting workflow.
- Use stableKey and recordType for idempotent downstream routing.

No third-party integration is claimed unless it is named above and supported by the current product contract.

### Data, privacy, and security

- Input postings and output evidence are written only to the run's Apify input, dataset, and key-value store.
- No external service is contacted and no credentials are required.

Set Apify storage retention and access according to the sensitivity of your inputs and outputs. This documentation does not create legal, privacy, compliance, or security certification.

### Support and known limitations

- Fallback company-title-location matching cannot prove two independently authored postings are the same job.
- The Actor does not fetch feeds, classify job quality, or interpret hiring intent.

For support, use the [Actor Issues page](https://apify.com/ceddl/job-feed-change-detector/issues). Include the run ID, a minimal reproducible input with sensitive values removed, the failing record or error code, and what you expected. Do not post credentials, private source files, customer data, or full confidential payloads.

# Actor input Schema

## `comparisonId` (type: `string`):

Input field comparisonId.

## `baseline` (type: `array`):

Input field baseline.

## `current` (type: `array`):

Input field current.

## `trackedFields` (type: `array`):

Input field trackedFields.

## `maxRecordsPerSnapshot` (type: `integer`):

Input field maxRecordsPerSnapshot.

## Actor input object example

```json
{
  "comparisonId": "sample-jobs",
  "baseline": [
    {
      "sourceId": "J-1",
      "title": "Data Analyst",
      "company": "Northwind",
      "location": "Berlin",
      "url": "https://jobs.example/J-1",
      "description": "SQL and dashboards"
    }
  ],
  "current": [
    {
      "sourceId": "J-1",
      "title": "Senior Data Analyst",
      "company": "Northwind",
      "location": "Berlin",
      "url": "https://jobs.example/J-1",
      "description": "SQL, dashboards, and Python"
    },
    {
      "sourceId": "J-2",
      "title": "Data Engineer",
      "company": "Northwind",
      "location": "Remote",
      "url": "https://jobs.example/J-2",
      "description": "Build data pipelines"
    }
  ]
}
```

# Actor output Schema

## `evidence` (type: `string`):

No description

## `summary` (type: `string`):

No description

## `reports` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "comparisonId": "sample-jobs",
    "baseline": [
        {
            "sourceId": "J-1",
            "title": "Data Analyst",
            "company": "Northwind",
            "location": "Berlin",
            "url": "https://jobs.example/J-1",
            "description": "SQL and dashboards"
        }
    ],
    "current": [
        {
            "sourceId": "J-1",
            "title": "Senior Data Analyst",
            "company": "Northwind",
            "location": "Berlin",
            "url": "https://jobs.example/J-1",
            "description": "SQL, dashboards, and Python"
        },
        {
            "sourceId": "J-2",
            "title": "Data Engineer",
            "company": "Northwind",
            "location": "Remote",
            "url": "https://jobs.example/J-2",
            "description": "Build data pipelines"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ceddl/job-feed-change-detector").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "comparisonId": "sample-jobs",
    "baseline": [{
            "sourceId": "J-1",
            "title": "Data Analyst",
            "company": "Northwind",
            "location": "Berlin",
            "url": "https://jobs.example/J-1",
            "description": "SQL and dashboards",
        }],
    "current": [
        {
            "sourceId": "J-1",
            "title": "Senior Data Analyst",
            "company": "Northwind",
            "location": "Berlin",
            "url": "https://jobs.example/J-1",
            "description": "SQL, dashboards, and Python",
        },
        {
            "sourceId": "J-2",
            "title": "Data Engineer",
            "company": "Northwind",
            "location": "Remote",
            "url": "https://jobs.example/J-2",
            "description": "Build data pipelines",
        },
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("ceddl/job-feed-change-detector").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "comparisonId": "sample-jobs",
  "baseline": [
    {
      "sourceId": "J-1",
      "title": "Data Analyst",
      "company": "Northwind",
      "location": "Berlin",
      "url": "https://jobs.example/J-1",
      "description": "SQL and dashboards"
    }
  ],
  "current": [
    {
      "sourceId": "J-1",
      "title": "Senior Data Analyst",
      "company": "Northwind",
      "location": "Berlin",
      "url": "https://jobs.example/J-1",
      "description": "SQL, dashboards, and Python"
    },
    {
      "sourceId": "J-2",
      "title": "Data Engineer",
      "company": "Northwind",
      "location": "Remote",
      "url": "https://jobs.example/J-2",
      "description": "Build data pipelines"
    }
  ]
}' |
apify call ceddl/job-feed-change-detector --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ceddl/job-feed-change-detector"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/zg9Hi5FyUx4wZdQmG/builds/FUlLBQCDOr7RjTODl/openapi.json
