# EURAXESS Jobs Scraper — Research & Academic Jobs (`nomad-agent/euraxess-enrich-translate-normalize-scraper`) Actor

Search EURAXESS PhD, postdoc, fellowship, research, and faculty vacancies. Get structured records with complete descriptions, requirements, funding, deadlines, contacts, and locations when published.

- **URL**: https://apify.com/nomad-agent/euraxess-enrich-translate-normalize-scraper.md
- **Developed by:** [Nomad Dev](https://apify.com/nomad-agent) (community)
- **Categories:** Jobs
- **Stats:** 1 total users, 0 monthly users, 60.7% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.90 / 1,000 euraxess job results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## EURAXESS Jobs Scraper — Research & Academic Jobs

Search public EURAXESS PhD, postdoctoral, fellowship, research, and faculty
vacancies and receive consistent, source-linked records with complete job
descriptions, requirements, funding, deadlines, contacts, and locations when
EURAXESS publishes them.

Optional features can expand a keyword across major European languages, fill
missing description-backed facts, or translate selected output fields to
English. Source facts always win over optional enrichment.

> This is an independent Actor. It is not affiliated with or endorsed by
> EURAXESS or the European Commission.

**Ready-made resources:** [Agent skill](https://github.com/Exdenta/nomad-agent-job-scrapers/blob/main/.agents/skills/euraxess-enrich-translate-normalize-scraper/SKILL.md) · [output contract](https://github.com/Exdenta/nomad-agent-job-scrapers/blob/main/.agents/skills/euraxess-enrich-translate-normalize-scraper/references/output-contract.md) · [Python parser](https://github.com/Exdenta/nomad-agent-job-scrapers/blob/main/.agents/skills/euraxess-enrich-translate-normalize-scraper/scripts/parse_output.py) · [n8n, Make, Airtable, MCP, API, and webhook integrations](https://github.com/Exdenta/nomad-agent-job-scrapers)

### Why use this Actor

- **Research-focused coverage:** PhD, postdoc, fellowship, research, and
  academic vacancies from the EURAXESS portal.
- **One stable schema:** every job uses `nomad-agent-job-v1`, the same envelope
  used by other normalized Nomad Agent job sources.
- **Rich research-job details:** research fields, seniority, requirements,
  education, experience, funding, benefits, deadlines, application channels,
  contacts, and all explicit locations when available.
- **Source-faithful data:** EURAXESS taxonomy stays distinct from applicant
  requirements, missing facts stay unknown, and optional enrichment never
  overwrites source facts.
- **Multilingual search and output:** optionally expand an exact keyword across
  major EURAXESS languages or translate selected display fields to English.
- **Alert-friendly delivery:** strict normalized filters and built-in cross-run
  deduplication support scheduled feeds without title-based guessing.
- **Integration-ready:** use the same n8n, Make, Airtable, MCP, REST API,
  webhook, parser, and Claude/Codex skill surfaces documented for LinkedIn.

### What each result contains

Every item has exactly six top-level fields:

| Field | Meaning |
|---|---|
| `schemaVersion` | Always `nomad-agent-job-v1` |
| `identity` | EURAXESS source, posting ID, and canonical job URL |
| `data` | Normalized job, organisation, research, location, employment, application, requirement, benefit, funding, compensation, and constraint fields |
| `custom` | Versioned EURAXESS-only facts, including source academic-level taxonomy |
| `llm` | Status and provenance for optional description-backed enrichment |
| `raw` | Complete description text and source HTML when `includeRaw` is true; otherwise `null` |

All declared fields are present. `null` means unknown or unavailable. `[]`
means EURAXESS explicitly established an empty collection.

Important accuracy rules:

- EURAXESS `Positions` or academic-level labels stay in
  `custom.data.academicLevelRaw`; they are not converted into applicant
  education requirements;
- a city, country, facility, or address does not by itself prove `onsite`;
- only named people are returned as hiring contacts;
- the EURAXESS posting URL and a separate application URL or email remain
  distinct;
- optional AI fills only allowlisted fields that remain `null` after source
  parsing;
- raw HTML is untrusted source content and must be sanitized before rendering.

Dataset rows remain one `nomad-agent-job-v1` record per job. Separately, the
Actor writes a `nomad-agent-run-summary-v4` record to the default key-value
store under `RUN-SUMMARY`. Use it to read the run outcome, delivered count,
whether results were limited, and any bounded retry recommendation.

### Input

| Field | Default | Purpose |
|---|---:|---|
| `schemaVersion` | required | Must be `nomad-agent-job-search-input-v1` |
| `keyword` | `""` | Job title, discipline, skill, or research term |
| `location` | `""` | Text matched against the location and country published by EURAXESS |
| `euraxessSearch` | omitted | Optional translations of the exact keyword across major EURAXESS languages; never similar-title expansion |
| `postedWithin` | `"24h"` | `24h`, `7d`, `30d`, or `any` |
| `workArrangements` | omitted | Any combination of explicitly stated `remote`, `hybrid`, and `onsite` |
| `filters` | omitted | Versioned filters over normalized job fields |
| `maxItems` | `100` | Maximum returned jobs; `0` requests the bounded 200-item window |
| `dedupe` | enabled | Suppress jobs already delivered in the same scope |
| `aiEnrichment` | disabled | Fill selected missing facts from the complete description |
| `translateToEnglish` | `false` | Translate selected normalized display fields |
| `includeRaw` | `true` | Include complete description text and HTML |
| `analyticsEnabled` | `false` | Share privacy-preserving aggregate operational analytics |

Unknown input fields are rejected. This prevents misspelled or retired options
from silently changing the meaning of a run.

#### Example search

```json
{
  "schemaVersion": "nomad-agent-job-search-input-v1",
  "keyword": "postdoctoral machine learning",
  "location": "Germany",
  "postedWithin": "30d",
  "maxItems": 25,
  "dedupe": {"enabled": false, "key": ""},
  "aiEnrichment": {"enabled": false, "accuracy": "silver"},
  "translateToEnglish": false,
  "includeRaw": false,
  "analyticsEnabled": false
}
```

EURAXESS publishes posting dates as calendar dates rather than exact
timestamps. `24h` therefore includes the current and previous UTC calendar
date. It is not an exact rolling 24-hour window.

Disable cross-run deduplication for repeatable one-off searches. Leave it
enabled for scheduled alerts, using a stable public `key` only when multiple
searches should intentionally share delivery history. Use a new key when an
intentional fresh delivery scope is needed.

### Optional features

#### Multilingual keyword expansion

The EURAXESS search extension keeps the original keyword and adds faithful
translations used across major portal languages. It is an explicit, off-by-
default input. It does not broaden the search to related titles, roles, or
disciplines, and failure falls back to the original term. Similar-job search is
not implemented because it would need a separate matching contract and
per-result match evidence to remain explainable.

```json
{
  "euraxessSearch": {
    "schemaVersion": "nomad-agent-euraxess-search-v1",
    "translateKeywords": true
  }
}
```

#### Description-backed AI enrichment

Enable `aiEnrichment` with a `silver` or `gold` accuracy profile. The Actor
reads the complete plain-text posting description and fills only supported
fields that source parsing left `null`. Source facts and source-established
empty arrays always win. No customer model key is required, and an enrichment
failure leaves the base job unchanged with a failed status in `llm`.

```json
{"aiEnrichment": {"enabled": true, "accuracy": "silver"}}
```

#### English translation

`translateToEnglish: true` translates selected normalized display fields:
title, domains, applicant-requirement prose, benefits, eligibility and
selection text, work authorization, security clearance, and location
preference. Organisation names, locations, identifiers, URLs, source-raw
labels, skills, qualifications, certifications, programme names, descriptions,
raw HTML, and provenance remain unchanged. No customer translation key is
required.

#### Normalized filters

Use `filters` for versioned AND/OR expressions over supported normalized job
fields. Filters run on source-language values before optional output
translation. Unknown facts do not become guessed matches.

```json
{
  "filters": {
    "schemaVersion": "nomad-agent-job-filter-v1",
    "expression": {
      "all": [
        {"field": "data.locations[].countryCode", "operator": "eq", "value": "DE"},
        {"field": "data.title", "operator": "not_contains", "value": "internship"}
      ]
    }
  }
}
```

### Run with Python

```python
from decimal import Decimal

from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor(
    "nomad-agent/euraxess-enrich-translate-normalize-scraper"
).call(run_input={
    "schemaVersion": "nomad-agent-job-search-input-v1",
    "keyword": "postdoctoral machine learning",
    "location": "Germany",
    "postedWithin": "30d",
    "maxItems": 25,
}, build="1.0.20", max_items=25, max_total_charge_usd=Decimal("0.10"))

if run["buildNumber"] != "1.0.20":
    raise RuntimeError("unexpected Actor build")

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["data"]["title"], item["data"]["company"]["name"])
```

### Pricing

The current pay-per-event prices are:

| Event | Price | Charged when |
|---|---:|---|
| Job result | **$0.0009** | A confirmed normalized job is delivered |
| English translation | **$0.006** | A delivered row contains translated selected fields |
| Silver enrichment | **$0.006** | A delivered row contains successful Silver enrichment |
| Gold enrichment | **$0.010** | A delivered row contains successful Gold enrichment |

Optional events are not charged when the feature is unnecessary or fails.
Check the Apify **Pricing** tab before a paid run in case prices have changed,
and use Apify's maximum-cost-per-run setting when you need a hard budget.

### Limits

- Public EURAXESS pages can change, block requests, or omit fields.
- A run returns at most 200 jobs, even when `maxItems` is `0`.
- Location search matches source-published location text; it is not a geocoder.
- Work-arrangement filters match only explicit source or opted-in enrichment
  evidence.
- Optional enrichment improves coverage but does not make unknown source facts
  certain or guarantee accuracy on every future posting.

For lossless validation, table flattening, and integration examples, use the
[public integration repository](https://github.com/Exdenta/nomad-agent-job-scrapers).

# Actor input Schema

## `schemaVersion` (type: `string`):

Version of the shared normalized job-search Actor input contract.

## `keyword` (type: `string`):

Exact job title, discipline, skill, or research term matched against job titles, domains, classifications, and full descriptions. Related roles are not added. Leave empty to search the bounded current job window.

## `location` (type: `string`):

City, region, postal code, facility, address, or country text matched case-insensitively against source-published locations. A missing or ambiguous location does not count as a match.

## `euraxessSearch` (type: `object`):

Optionally add faithful translations of the exact keyword across major EURAXESS languages. The original term is retained and related titles, roles, and disciplines are never added.

## `postedWithin` (type: `string`):

Return jobs within the selected publication-date window. EURAXESS publishes calendar dates, so 24h includes the current and previous UTC date; exact hour-level filtering is unavailable.

## `workArrangements` (type: `array`):

Optional union of remote, hybrid, and on-site arrangements. EURAXESS rows match only when the source explicitly establishes an arrangement; a city, country, office address, or missing workplace label is never interpreted as on-site evidence.

## `maxItems` (type: `integer`):

Maximum number of job postings to return. Set 0 to request the Actor's complete bounded delivery window of 200 items; it never means unlimited.

## `dedupe` (type: `object`):

Suppress jobs already delivered in the same search or alert scope. Leave key empty for a scope derived from this Apify user and the search, or provide a public opaque alert/profile key to share history intentionally. Disable it for repeatable one-off searches.

## `filters` (type: `object`):

Use a nomad-agent-job-filter-v1 expression for country or location requirements, title or organisation exclusions, work arrangements, and other supported normalized fields. Unknown facts do not become guessed matches.

## `aiEnrichment` (type: `object`):

Fill supported fields that remain null using the complete public plain-text job description. Source facts always win. Choose the Silver or Gold accuracy profile; failed enrichment leaves the base job unchanged. No customer model key is required.

## `translateToEnglish` (type: `boolean`):

Translate selected normalized display fields to English: title, domains, applicant requirement prose, benefits, eligibility and selection text, work authorization, security clearance, and location preference. Descriptions, organisation names, locations, identifiers, URLs, source-raw labels, skills, qualifications, certifications, programme names, and provenance remain unchanged. No customer translation key is required.

## `includeRaw` (type: `boolean`):

When enabled, each result includes the complete EURAXESS plain-text detail description and source HTML in raw. Disable it to return raw: null; requested position enrichment still runs before raw output is removed.

## `analyticsEnabled` (type: `boolean`):

Share one privacy-preserving aggregate operational event. It can include version, outcome category, coarse duration, result count, enabled features, and aggregate health counters, but never caller IDs, searches, records, URLs, source text, translations, extracted values, errors, or secrets.

## Actor input object example

```json
{
  "schemaVersion": "nomad-agent-job-search-input-v1",
  "keyword": "postdoctoral machine learning",
  "location": "Germany",
  "euraxessSearch": {
    "schemaVersion": "nomad-agent-euraxess-search-v1",
    "translateKeywords": true
  },
  "postedWithin": "30d",
  "maxItems": 5,
  "dedupe": {
    "enabled": false,
    "key": ""
  },
  "filters": {
    "schemaVersion": "nomad-agent-job-filter-v1",
    "expression": {
      "all": [
        {
          "field": "data.locations[].countryCode",
          "operator": "eq",
          "value": "DE"
        },
        {
          "field": "data.title",
          "operator": "not_contains",
          "value": "internship"
        }
      ]
    }
  },
  "aiEnrichment": {
    "enabled": true,
    "accuracy": "silver"
  },
  "translateToEnglish": false,
  "includeRaw": true,
  "analyticsEnabled": false
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `runSummary` (type: `string`):

nomad-agent-run-summary-v4 outcome with the delivered count, results-limited status, and one optional bounded retry recommendation.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "postedWithin": "30d",
    "maxItems": 5,
    "dedupe": {
        "enabled": false,
        "key": ""
    },
    "includeRaw": false
};

// Run the Actor and wait for it to finish
const run = await client.actor("nomad-agent/euraxess-enrich-translate-normalize-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "postedWithin": "30d",
    "maxItems": 5,
    "dedupe": {
        "enabled": False,
        "key": "",
    },
    "includeRaw": False,
}

# Run the Actor and wait for it to finish
run = client.actor("nomad-agent/euraxess-enrich-translate-normalize-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "postedWithin": "30d",
  "maxItems": 5,
  "dedupe": {
    "enabled": false,
    "key": ""
  },
  "includeRaw": false
}' |
apify call nomad-agent/euraxess-enrich-translate-normalize-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nomad-agent/euraxess-enrich-translate-normalize-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/2JPXQf3gbDCUAi57w/builds/5n7huEXQYFUBdndQP/openapi.json
