# EU Pay Transparency Job Ad Checker (`gravelly_caladium/eu-pay-transparency-checker`) Actor

Checks job ad URLs or raw text against the EU Pay Transparency Directive (2023/970) and a bonus US salary-disclosure ruleset: pay range disclosure, pay-history questions, gender-neutral titles, transparency statements. Returns a verdict, rule refs, salary, and a fix suggestion.

- **URL**: https://apify.com/gravelly\_caladium/eu-pay-transparency-checker.md
- **Developed by:** [Relay Data Tools](https://apify.com/gravelly_caladium) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## EU Pay Transparency Job Ad Checker

Checks job ad URLs or raw job ad text against the **EU Pay Transparency Directive
((EU) 2023/970)** and its national implementations -- plus a bonus ruleset covering
existing US state/city salary-disclosure laws (NYC, CA, WA, CO, IL, and more). Returns a
per-ad compliance verdict (compliant / non-compliant / cannot-determine), the specific
rule findings behind it (with a reference and source URL), the extracted salary, and a
short plain-language fix suggestion.

**This tool is informational only and is not legal advice.** See "Legal" below.

### Why this exists

Directive (EU) 2023/970 required all EU member states to transpose it into national law
by **7 June 2026**. That deadline has now passed. Some countries transposed on time, some
are late, some still have a bill pending -- and several national implementations go
further than the Directive's own minimum (e.g. requiring the pay range directly *in* the
job ad, not just before an interview). For anyone posting or auditing job ads across
several EU countries -- HR teams, recruiters, job boards, compliance/legal ops -- knowing
which rule applies to which country's ad, and whether a given ad actually meets it, is
genuinely hard to track by hand. This Actor encodes that per-country research into a
versioned, sourced rules file and checks ads against it automatically.

### Who it's for

- **HR / talent acquisition teams** publishing job ads across multiple EU countries who
  want a pre-publish compliance check.
- **Recruiters and staffing agencies** auditing client job ads before they go live.
- **Job boards and ATS platforms** wanting a bulk compliance pass across listings.
- **Compliance / legal ops teams** building a paper trail of what was checked, when, and
  against which rule version.

### How it works

1. **Fetch**: for each job ad URL, a plain HTTP GET is sent (no headless browser).
2. **Parse**: the page's embedded JSON-LD `JobPosting` data (schema.org) is extracted
   when present -- Greenhouse, Lever, Workday, Personio, Teamtailor, SmartRecruiters and
   many other ATS platforms include this for Google for Jobs SEO, and it's a far more
   reliable source for salary/title/location than scraping visible text. When it's
   absent, readable page text is extracted instead (scripts/styles/nav stripped).
3. **Resolve country**: the job's country is taken from JSON-LD `jobLocation`, then the
   URL (ccTLD / locale path segment), then a country/city name mention in the ad text,
   then your input's `country` field as a final fallback -- see "Input".
4. **Check**: the resolved country's rule (from `rules/eu_pay_transparency.json`, or
   `rules/us_state_laws.json` for a US jurisdiction) is loaded, and the ad is checked for:
   a numeric pay range/starting figure, a pay-history question, a gender-neutral job
   title (heuristic), and a pay-transparency statement (informational).
5. **Push**: one dataset item per ad, plus a run-level summary written to the key-value
   store (`SUMMARY`) with the overall compliance rate and a breakdown by country/company.

Career pages can also be crawled: give a `careersPageUrls` entry and the Actor follows
same-domain links that look like individual job postings (up to `maxJobsPerCareersPage`
each), then checks each one the same way.

#### How compliance status is decided

For each dimension (salary range in ad, pay-history question, gender-neutral title, pay-
transparency statement), the resolved country's rule says either **required**, **not
required**, or **unclear**/**only required before interview, not in the ad**. Combining
all dimensions:

- **`non-compliant`**: at least one dimension has a confirmed violation of a rule that is
  **currently in force** (e.g. the rule requires a range in the ad and none was found).
  One clear violation is enough, regardless of ambiguity elsewhere.
- **`cannot-determine`**: no ruleset exists for the resolved country; OR a ruleset exists
  but at least one dimension is genuinely ambiguous (e.g. the rule only requires
  disclosure before the interview, not in the ad itself, and the ad text has no pay info
  either way -- it may still be disclosed later, off-ad) with zero confirmed violations
  elsewhere; OR the resolved country's national transposition of the Directive is **not
  yet in force** (still a draft bill, enacted but not yet effective, etc.) -- a would-be
  violation there can't be graded `non-compliant` against a law that isn't binding yet, so
  it's downgraded to an informational finding (evaluated against the EU Directive's own
  baseline / the draft's likely shape instead) and the ad comes back `cannot-determine`.
  Check `transpositionStatus` on the result and `rules/eu_pay_transparency.json` for that
  country's actual status and expected effective date.
- **`compliant`**: a ruleset exists and is in force, zero violations, zero ambiguous
  dimensions.

A range that's found but implausibly wide (more than 2x between min and max) is a
`warning`-level finding, not a violation -- it doesn't flip the status, but is worth a
look. US bonus jurisdictions are always treated as current law (state/city statutes
already in force), so this "not yet in force" downgrade only applies to EU countries.

### Input

```json
{
  "jobAdUrls": ["https://boards.greenhouse.io/example/jobs/1234567"],
  "jobAdTexts": ["Software Engineer (m/w/d), Berlin. We offer EUR 60,000-70,000/year..."],
  "careersPageUrls": [],
  "maxJobsPerCareersPage": 20,
  "country": "",
  "autoDetectCountry": true,
  "language": "",
  "includeUsBonusRules": true,
  "maxConcurrency": 3,
  "proxyConfiguration": { "useApifyProxy": true }
}
```

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `jobAdUrls` | array of strings | -- | Public job ad page URLs. |
| `jobAdTexts` | array of strings | -- | Raw job ad text, one full ad per entry. |
| `careersPageUrls` | array of strings | -- | Company careers pages to crawl for job links. |
| `maxJobsPerCareersPage` | integer | 20 | Cap on job links followed per careers page. |
| `country` | string | -- | ISO 3166-1 alpha-2 (e.g. `DE`), or a US bonus code like `US-NYC`/`US-CA`. Used when auto-detect is off, or as the final fallback. |
| `autoDetectCountry` | boolean | `true` | Try JSON-LD -> URL -> ad text before falling back to `country`. |
| `language` | string | -- | ISO 639-1 override for which phrase list is prioritized. All supported languages are scanned regardless -- see "Limitations". |
| `includeUsBonusRules` | boolean | `true` | Apply the bonus US ruleset when a job resolves to a US jurisdiction. |
| `maxConcurrency` | integer | 3 | Concurrent fetches. |
| `proxyConfiguration` | object | Apify Proxy | Usually not needed for ATS-hosted pages. |

### Output (one dataset item per ad)

```json
{
  "sourceType": "url",
  "sourceUrl": "https://boards.greenhouse.io/example/jobs/1234567",
  "companyName": "Acme GmbH",
  "jobTitle": "Senior Backend Engineer (m/f/d)",
  "country": "DE",
  "countryDetectionMethod": "jsonld",
  "language": "en",
  "ruleSource": "eu",
  "status": "compliant",
  "findings": [
    { "ruleId": "pay_transparency_statement_found", "severity": "info",
      "message": "Ad includes pay-transparency/equal-pay statement language ('pay transparency').",
      "ruleReference": "Germany's implementing law", "sourceUrl": "https://..." }
  ],
  "extractedSalary": { "min": 70000, "max": 85000, "currency": "EUR", "period": "year", "raw": "70000-85000 EUR" },
  "salaryMin": 70000, "salaryMax": 85000, "salaryCurrency": "EUR", "salaryPeriod": "year",
  "payHistoryQuestionFound": false,
  "genderNeutralTitle": true,
  "transparencyStatementFound": true,
  "fixSuggestion": null,
  "lawName": "...", "lawUrl": "...", "transpositionStatus": "in_force",
  "jsonLdFound": true,
  "robotsNoindex": null,
  "checkedAt": "2026-09-28T12:00:00+00:00",
  "error": null
}
```

If a URL/text item can't be fetched or processed at all, one item is pushed with every
field `null` except `sourceType`/`sourceUrl` and `error` (a description) -- the run does
not stop. A ready-to-use **Overview** table view is available in the dataset UI/API. A
run-level **summary** (compliance rate, breakdown by country and company) is written to
the default key-value store under the key `SUMMARY` -- see the Actor's Output tab, or
`GET .../key-value-stores/{id}/records/SUMMARY`.

See `samples/sample_output.json` for a real local run's output.

### Rules coverage

Rules live in two versioned, sourced JSON files, each entry carrying a `last_verified`
date and source URLs -- refresh them periodically, laws change:

- `rules/eu_pay_transparency.json` -- all 27 EU member states' transposition status of
  Directive (EU) 2023/970, as of `last_verified: 2026-09-28`: **5 in force** (IT, LT, MT,
  SK, and partially EE), **1 partially in force** (PL), **1 enacted but not yet effective**
  (EL, effective 2026-11-01), **18 still a draft bill / not yet started**, and **1 unclear**
  (SE, which has publicly said it isn't pursuing transposition and wants the Directive
  renegotiated). See "How compliance status is decided" above for how a not-yet-in-force
  country is graded (`cannot-determine`, not `non-compliant`).
- `rules/us_state_laws.json` -- a bonus set of 17 US state/city salary-transparency laws,
  all currently in force (NYC, NY State, CA, WA, CO, IL, CT, MD, NJ, RI, HI, MN, MA, NV,
  Cincinnati/Toledo OH, DC).
- `rules/eu_research_2026-09-28.json` -- the raw per-country research notes (status, law,
  effective date, requirement flags, source, and caveats) that `eu_pay_transparency.json`
  was converted from; kept for provenance/audit trail.

See the coverage table in the build report for the full per-country breakdown (status,
law name, effective date, source). A country/jurisdiction with no entry, or a US ad with
no specific jurisdiction resolved, comes back `cannot-determine` rather than a guess.

### Testing

- **Unit + integration tests** (pytest, offline, no network by default): salary/range
  extraction across currencies and languages, pay-history and vague-salary phrase
  detection, the gender-neutral-title heuristic, JSON-LD `JobPosting` parsing, country/
  language auto-detection, the rules engine (including a structural check that all 27 EU
  countries and all US bonus jurisdictions resolve to a well-formed rule with a source),
  the compliance status algorithm (isolated from the rules file's actual content via
  synthetic fixtures), and a >=30-ad realistic fixture set across German, French, Spanish,
  Italian, Dutch, Polish, Portuguese, Swedish, Danish, Finnish, Czech, Latvian,
  Lithuanian, and English ads, plus JSON-LD HTML fixtures. Run with `pytest`.
- **Live network smoke test** (excluded by default): fetches a real Greenhouse-hosted
  careers page and one real job posting. Run with `pytest -m network`.
- **Local end-to-end run**: see the build report for actual pytest output and a real run
  against fixtures + live public job ad URLs.

### Limitations

- **This is a rules engine over hand-researched legal data, not a law firm.** Every
  rule entry has a `last_verified` date and source URLs -- always check those before
  relying on a result. Transposition status changes; some countries' rules were still
  moving as of this build (see `notes` per country in the rules file).
- **Salary extraction is regex-based pattern matching**, not an NLP/financial model. It
  requires a currency symbol/code or a "k" suffix (e.g. "60k-80k") to accept a range --
  a bare number range with neither (which is at least as likely to be an age, date, or
  hours-per-week range) is deliberately rejected. Numbers spelled out in words, and
  highly unusual formats, are not handled. See `src/salary.py`'s docstring.
- **Pay-history and vague-salary phrase lists are hand-curated, not a translation API.**
  They cover 15 languages with a deliberately narrow, high-confidence phrase set to keep
  false positives low; absence of a match means "not detected", not "confirmed absent".
- **The gender-neutral-title check is an explicitly-labelled heuristic** covering
  English (legacy gendered nouns) and six grammatically-gendered languages (German,
  French, Spanish, Italian, Polish, Portuguese) with a curated term list. Other
  languages return "not evaluated", not a false "neutral". This is a signal worth a
  human look, not a linguistic authority -- see `src/gender.py`'s docstring.
- **Country/language auto-detection is heuristic** (JSON-LD > URL ccTLD/locale segment >
  country/city name mention in text > your input's `country`). Ambiguous text (e.g.
  multiple countries mentioned) deliberately returns "not detected" rather than guessing.
- **`robotsNoindex` only records the page's own `<meta name="robots">` signal** for
  reference; it is not itself a compliance finding and `/robots.txt` is not checked.
- **Careers-page crawling is a same-domain, path-keyword heuristic** (`/jobs/`,
  `/careers/`, etc.), not a sitemap parser -- it can miss JS-rendered link lists (no
  headless browser is used) or pick up a non-job page that happens to match the pattern.
- **US coverage is a bonus set of 17 jurisdictions**, not all 50 states -- an ad resolved
  to a US state/city not in `rules/us_state_laws.json` (or to plain "US" with no more
  specific jurisdiction determinable) returns `cannot-determine`.

### Legal

This Actor is an **informational compliance-checking tool, not legal advice**. It
reflects a snapshot of publicly available legal research (see each rule's source URLs
and `last_verified` date in `rules/eu_pay_transparency.json` and
`rules/us_state_laws.json`) and a set of heuristics for text patterns -- it can miss
violations, flag false positives, and go stale as laws change. Before making any hiring,
publishing, or compliance decision, verify against the official legal text and/or consult
a qualified employment lawyer in the relevant jurisdiction. This Actor only fetches
public job ad pages, the same way a browser would; it does not access any account or
bypass any paywall or access control.

### FAQ

**Why did an ad come back `cannot-determine`?**
Either its country couldn't be resolved (no ruleset to check against), or the resolved
country's rule for some dimension is itself unclear/ambiguous (see `findings`) with no
other confirmed violation. Check `countryDetectionMethod` and the `findings` list.

**Why didn't it find the salary I can clearly see in the ad?**
See "Limitations" -- extraction requires a currency symbol/code or a "k" suffix. A range
with neither (or spelled-out numbers) won't be picked up. `extractedSalary.raw` shows
exactly what was matched, if anything.

**Is a flagged gendered title actually illegal?**
Not necessarily -- it's a heuristic signal (`genderTitleHeuristicConfidence: "heuristic"`
in the underlying check), not a certified legal determination. Always verify manually.

**How is this priced?**
See [PRICING.md](PRICING.md) for the proposed pay-per-event plan.

# Actor input Schema

## `jobAdUrls` (type: `array`):

Public job ad page URLs to check, from any site. JSON-LD JobPosting data is used when the page has it (Greenhouse, Lever, Workday, Personio, Teamtailor, SmartRecruiters, and many others); otherwise readable page text is extracted and checked.

## `jobAdTexts` (type: `array`):

Raw job ad text, one full ad per array entry, as an alternative or addition to Job ad URLs. Use this for ads you already have the text of (e.g. from a PDF or an internal system).

## `careersPageUrls` (type: `array`):

Public company careers/jobs listing page URLs. Each is fetched and scanned for links that look like individual job postings (up to 'Max jobs per careers page' each), which are then checked the same as Job ad URLs.

## `maxJobsPerCareersPage` (type: `integer`):

Upper limit on how many job links to follow from each careers page in 'Careers page URLs to crawl'.

## `country` (type: `string`):

e.g. 'DE', 'FR', 'US-NYC', 'US-CA'. Used when auto-detection is off, or as a fallback when auto-detection can't determine a job's country. See README for supported US bonus jurisdiction codes.

## `autoDetectCountry` (type: `boolean`):

Try to detect each job's country from JSON-LD data, the URL (domain/locale), and the ad text, before falling back to the 'Country' field above. Turn off to force every ad to use 'Country'.

## `language` (type: `string`):

ISO 639-1 code (e.g. 'de', 'fr') to force which language's phrase list is prioritized for pay-history/gender-title checks. Leave empty to auto-detect per ad -- all supported languages are always scanned regardless, so this mainly affects which match is reported first.

## `includeUsBonusRules` (type: `boolean`):

When a job's country resolves to the US (or a US jurisdiction code like 'US-NYC'), also apply the bonus US state/city salary-transparency ruleset (rules/us\_state\_laws.json). Turn off to only ever evaluate against the EU ruleset.

## `maxConcurrency` (type: `integer`):

How many URLs to fetch at the same time. Kept low by default to be polite to the sites being checked.

## `proxyConfiguration` (type: `object`):

Optional. Only needed if a target site blocks direct requests. The Actor works without a proxy for most ATS-hosted job pages.

## Actor input object example

```json
{
  "jobAdUrls": [],
  "jobAdTexts": [
    "Sviluppatore/Sviluppatrice Backend (Milano). Retribuzione: 38.000 - 45.000 EUR lordi annui. Contratto a tempo indeterminato.",
    "Programuotojas (-a), Vilnius. Atlyginimas: 2500-3200 EUR/mėn. neatskaičius mokesčių. Prašome nurodyti dabartinį atlyginimą."
  ],
  "careersPageUrls": [],
  "maxJobsPerCareersPage": 20,
  "country": "DE",
  "autoDetectCountry": true,
  "includeUsBonusRules": true,
  "maxConcurrency": 3,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `resultsAll` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "jobAdUrls": [],
    "jobAdTexts": [
        "Sviluppatore/Sviluppatrice Backend (Milano). Retribuzione: 38.000 - 45.000 EUR lordi annui. Contratto a tempo indeterminato.",
        "Programuotojas (-a), Vilnius. Atlyginimas: 2500-3200 EUR/mėn. neatskaičius mokesčių. Prašome nurodyti dabartinį atlyginimą."
    ],
    "careersPageUrls": [],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("gravelly_caladium/eu-pay-transparency-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "jobAdUrls": [],
    "jobAdTexts": [
        "Sviluppatore/Sviluppatrice Backend (Milano). Retribuzione: 38.000 - 45.000 EUR lordi annui. Contratto a tempo indeterminato.",
        "Programuotojas (-a), Vilnius. Atlyginimas: 2500-3200 EUR/mėn. neatskaičius mokesčių. Prašome nurodyti dabartinį atlyginimą.",
    ],
    "careersPageUrls": [],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("gravelly_caladium/eu-pay-transparency-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "jobAdUrls": [],
  "jobAdTexts": [
    "Sviluppatore/Sviluppatrice Backend (Milano). Retribuzione: 38.000 - 45.000 EUR lordi annui. Contratto a tempo indeterminato.",
    "Programuotojas (-a), Vilnius. Atlyginimas: 2500-3200 EUR/mėn. neatskaičius mokesčių. Prašome nurodyti dabartinį atlyginimą."
  ],
  "careersPageUrls": [],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call gravelly_caladium/eu-pay-transparency-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gravelly_caladium/eu-pay-transparency-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/fqAJinB6dcfTL5Ye2/builds/TW7gtxNGiaRo03mVD/openapi.json
