Personio Jobs Scraper avatar

Personio Jobs Scraper

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Personio Jobs Scraper

Personio Jobs Scraper

Scrape any company's Personio job board and get clean, de-duplicated job JSON: title, department, office, employment type and canonical apply URLs. One click, no required fields, no LLM.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Turgay NANTA

Turgay NANTA

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

a day ago

Last modified

Share

Scrape any company's Personio job board and get clean, de-duplicated job JSON: title, department, office, employment type and canonical apply URLs. One click, no required fields, no LLM.

What it does

Personio is Europe's leading HR platform, powering hiring for 12,000+ companies (mostly EU/DACH-region). This actor reads any company's public Personio XML career feed and returns every open role as clean JSON: title, department, office, recruiting category, employment type, seniority and the canonical apply URL. Give it one company's Personio subdomain (4401) or several, and it handles the .de/.com domain variation automatically — de-duplicated and ready for your pipeline.

Why this one:

  • Official public XML career feed — no HTML parsing, no proxies
  • Automatic .de/.com domain fallback (Personio accounts use either)
  • Multi-company runs + monitor mode for scheduled diffing
  • No LLM anywhere — deterministic output, predictable costs, no hallucinated fields
  • Clean by default — canonical URLs (tracking parameters stripped), parsed numbers, merged duplicates

Quick start (no code)

  1. Click Try for free / Start — every field has a working default, nothing is required.
  2. (Optional) change query to what you need.
  3. Open the Dataset tab when the run finishes → export as JSON, CSV or Excel.

Input

FieldRequiredDefaultDescription
queryno4401Personio career-site subdomain, comma-separated for multiple. Find it in the careers URL: <subdomain>.jobs.personio.de (or .com).
maxResultsno20Maximum clean results (capped at 500)
enrichnofalseDeterministic enrichment per record — see below
monitornofalseCompare with the previous run, flag NEW records only

Example input:

{
"query": "4401",
"maxResults": 50
}

Output

Real example record (from a live run):

{
"id": "2726014",
"title": "Operations Superintendent",
"url": "https://4401.jobs.personio.de/job/2726014",
"location": "UAE",
"seller": "4401",
"team": "Contract Hire",
"department": "Engineering - Operations",
"commitment": "full-time",
"created_at": "2026-07-23T15:54:10+00:00"
}

The final _summary row carries run totals (total_clean, deduped, enriched); in monitor mode a _changes row lists keys new since the last run.

Field reference

FieldMeaning
idPersonio position ID (stable, deduplication key)
titleJob title
urlCanonical apply URL, built as `
sellerCompany's Personio subdomain
locationOffice, as published
team / departmentRecruiting category and department
commitmentSchedule / employment type (full-time, fixed-term, etc.)
created_atPosition creation timestamp
price / price_textParsed numeric value + original text, when the source publishes one
completeness0–1 filled-fields score (with enrich)

Use cases

  • EU/DACH hiring watch — Personio's base skews strongly European; track hiring where most other ATS scrapers don't reach.
  • Competitor hiring watch — schedule with monitor mode; change-alert fires only on NEW roles.
  • Talent mappingdepartment/team show exactly where a company is growing.
  • Market research — measure hiring velocity by office/region across an industry.
  • AI agents — live 'is company X hiring in Europe?' answers via MCP.

Enrichment (optional, charged only when it produces something)

Set enrich: true and every record additionally gets: e-mail addresses extracted from the description (when present), the canonical domain of the record's URL, and a completeness score (0–1, how many core fields are filled). Deterministic — the same input always yields the same output — and you are only charged for records that actually got enriched. Records where enrichment adds nothing are free.

Monitor mode — change alerts on a schedule

Set monitor: true and the actor compares the current run with the previous one (per-actor named storage) and flags only NEW records. Combine with Apify Schedules for a daily/hourly watch: the _changes summary row lists what appeared since the last run, and the change-alert event is charged per new record only — an unchanged run costs you almost nothing.

Use it from your code

Python

from apify_client import ApifyClient
client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("EnezLi/personio-scraper").call(run_input={ "query": "4401", "maxResults": 50 })
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_API_TOKEN>' });
const { defaultDatasetId } = await client.actor('EnezLi/personio-scraper').call({ "query": "4401", "maxResults": 50 });
const { items } = await client.dataset(defaultDatasetId).listItems();
console.log(items);

curl

curl -X POST "https://api.apify.com/v2/acts/EnezLi~personio-scraper/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \
-H "Content-Type: application/json" \
-d '{ "query": "4401", "maxResults": 50 }'

Use it with AI agents (MCP)

This actor is agent-ready: it appears in Apify's AGENTS / MCP servers catalog, so any MCP-capable assistant (Claude, custom agents, LangGraph tools) can discover and call it with a one-line tool call — zero required fields means an agent can run it safely with defaults. Connect your agent to the Apify MCP server and ask for live data in natural language.

Pricing — Pay-Per-Event, start is free

EventWhen charged
Actor startFree ($0) — try it with one click
resultPer clean result returned
enrichmentOnly per record that actually got enriched
change-alertMonitor mode: per NEW record since the previous run

No subscription, no minimum. Volume discounts apply automatically through Apify account tiers (up to −44% on GOLD). Typical run cost example: 20 results ≈ a few cents total — you can predict your bill from the numbers above before you run.

This actor collects publicly available data only — the same information any visitor sees in a browser, via public endpoints. It does not bypass logins, collect private personal data, or store credentials. You are responsible for using the output in compliance with the source site's terms and the laws that apply to you (e.g. GDPR when the output contains personal data).

Support & feedback

Found a bug, need another field, or want a variant for a related platform? Open an issue on the Issues tab — issues are monitored and answered, and frequently-requested fields get added to the standard output. The actor is maintained as part of a scraper family built on one shared, tested core (bugs fixed once are fixed everywhere).

Changelog

  • 0.1 (2026-07) — initial public release: search, dedup, optional enrichment, monitor mode, PPE pricing.

Limitations (honest ones)

Reads Personio's official public XML career feed; companies that left Personio or unlisted their board return zero rows. Full HTML job description is not included yet (plain-text excerpt only).

FAQ

How do I find a company's Personio subdomain?

Open the company's careers page; if it's on Personio the URL looks like 4401.jobs.personio.de — the first part is the subdomain to use.

Why XML instead of JSON?

Personio's official public career feed is XML by design; the actor parses it for you and returns clean JSON either way — no difference in what you receive.

A company returned zero jobs — why?

Either the subdomain is misspelled, the company has no open roles, or they use the other domain (.de vs .com — the actor tries both automatically). Nothing is charged for a zero-result run.

Do I need an API key or account on the source platform?

No. The actor uses public endpoints — you only need your Apify account.

Does it use AI / an LLM?

No. The core is fully deterministic: same input, same output, no hallucinations, no per-token costs.

Can I run it on a schedule?

Yes — use Apify Schedules; combine with monitor mode to pay only for what's new.

What's the maximum number of results?

500 per run (memory-safe cap). Run multiple queries or schedule runs for more.

How is my bill calculated?

Only from the events in the Pricing table — start is free, and there is no subscription.