# Jobindex.dk Jobs Scraper — Danish Job Ads & Employers (`scrapersdelight/jobindex-jobs-scraper`) Actor

From $0.50 per 1,000 ads. Scrape Denmark's largest job board by keyword, one of 81 job categories, any of 110 locations, address radius, employer, contract type or date posted. Employer name + website, address & coordinates, deadline, rating. Optional full ad text and contact e-mail.

- **URL**: https://apify.com/scrapersdelight/jobindex-jobs-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Jobs, Lead generation, Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.50 / 1,000 per job ad returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 🇩🇰 Jobindex.dk Jobs Scraper — Danish Job Ads, Employers & Hiring Signals

**Scrape Jobindex.dk — Denmark's largest job board (18,970 live ads measured 2026-09-05) — by keyword, any of 81 job categories, any of 110 locations, address + radius, employer, contract type, working hours, remote mode or date posted. Every ad returns the employer's name *and their own website*, the workplace address with coordinates, posted date, application deadline, home-working flag and employer rating — plus the ad teaser, free. $0.50 per 1,000 ads, with no per-run fee.**

### Why this one?

| | **This actor** | Typical Jobindex scraper |
|---|---|---|
| **Price per 1,000 ads** | **$0.50** | $0.99 – $2.00 |
| **Per-run start fee** | **none** | $0.05 – $10.00 per 1,000 runs (13 of the 14 live rivals charge one) |
| Crawl order | **`sort=date` by default — 0.00 % duplicate rows measured** | relevance order, **9.64 % duplicates measured on the identical search** |
| Employer's own website (`companyWebsite`) | ✅ 96 % / 89 % filled (two measured samples) | ❌ |
| Workplace coordinates (lat/lng) | ✅ 67 % / 74 % filled | ❌ |
| Employer rating + review count | ✅ 56 % / 52 % filled | ❌ |
| Ad description | ✅ included free (no second request) | often a paid "detail" tier |
| Full advert text + contact e-mail/phone | ✅ optional, **included in the same per-ad price** | separate $5.00/1k detail event |
| Beats the site's 1,000-ad ceiling | ✅ built-in category × region partitioning | ❌ |
| Server-side filters | ✅ 13, every one verified against live hit counts | keyword + location |
| Incremental / monitor mode | ✅ `postedWithin` + `onlyNewSince` | ❌ |

No login, no API key, no browser. The actor reads the structured payload Jobindex's own search page is rendered from — 20 ads per request, 512 MB, plain HTTP.

***

### What does Jobindex.dk Jobs Scraper do?

It extracts **Danish job postings** from [jobindex.dk](https://www.jobindex.dk) and returns clean rows you can export to **JSON, CSV, Excel** or pull via API:

- 🔎 **Search the way the site does** — free-text keywords, 81 job subcategories across 10 families, 110 locations (country / region / area / all 98 municipalities), address + radius in km, employer id, 12 contract types, full-time vs part-time, remote mode, and date posted. All server-side, all verified against live hit counts.
- 🏢 **Employer intelligence, not just a job title** — company name, Jobindex company id (joins every ad by the same employer), **the employer's own website**, their Jobindex profile URL, follower count, and their public rating + review count.
- 📍 **Where the work actually is** — street address, postcode, city, region, and **latitude/longitude** — measured on 67 % and 74 % of ads in the two samples below. Ads with no structured address keep `area`.
- 📅 **Time-aware** — first published date, listing expiry, and the real application deadline (or the "as soon as possible" flag).
- 🧾 **The ad text, free** — every search result ships its own rendered teaser, so the description costs **zero extra requests**. Measured across 2,300 ads: 30–399 characters, median 286–300. It is a teaser; `fetchFullDescription` gets the whole advert.
- 📖 **Optional full advert + contact details** — turn on `fetchFullDescription` and the actor opens each ad's own page (which for two thirds of ads redirects to the employer's own careers/ATS site) for the complete advert body, plus any contact e-mail and phone. **Included in the per-ad price — there is no second event.**
- 🎯 **ATS technographics** — `externalApplyUrl` is the employer's own off-site apply link, and its host names the applicant-tracking system they run. Measured, not estimated: **35.6 %** of the 1,000 newest ads and **25.5 %** of an all-81-categories sample carry one, and every single vendor seen across those 2,300 ads was **hr-manager** (267 / 237), **emply** (82 / 85) or **elvium** (7 / 9). Jobindex's own in-house quick-apply endpoint is deliberately **not** reported here — it is not an external ATS and carries no vendor signal (see "What this actor does NOT return").
- 🔁 **Monitor mode** — `postedWithin: "7"` plus `onlyNewSince` returns only ads new since your last run. Measured 2026-09-05: **7,480 ads posted in the last 7 days**, so a weekly pull is a recurring feed, not a one-off dump.

***

### What data does it extract?

Every job ad is one dataset row.

**Fill rates are measured, on two real runs on 2026-09-05 — no estimates, and no single flattering sample.** Jobindex mixes two kinds of ad and they fill different fields, so one number would mislead you. Both runs are reproducible from the inputs given:

| Sample | Input | Rows |
|---|---|---|
| **A — the default path at scale** | `{ "maxItems": 1000 }` (whole board, `sort=date`) | 1,000 of the 18,970 live ads |
| **B — a spread across the whole taxonomy** | all 81 `subcategoryIds`, `partitionBy: "subcategory"`, `maxPagesPerQuery: 1`, `maxItems: 0` | 1,300 unique (1,596 raw, 296 cross-category duplicates collapsed) |

**Always 100 % in both samples**

- 🆔 `tid` — the stable ad id (`h…` Jobindex-hosted, `r…` aggregated partner feed, a handful of `o…`). Dedupe key.
- 🏷️ `title` · 🔗 `jobUrl` (the canonical `vis-job` link, never the tracker)
- 📅 `publishedDate` · `lastDate` · `applyDeadlineAsap`
- 🚩 `homeWorkplace` · `isArchived` · `isSponsored` · `hasVideo` · `addressCount`
- 🕒 `scrapedAt` · `sourceSearchUrl`

**Everything else, measured (Sample A / Sample B)**

| Field | A (1,000 newest) | B (all 81 categories) |
|---|---|---|
| 📝 `description` (free teaser) | 100.0 % | 99.7 % |
| `geoAreaIds` | 98.9 % | 98.7 % |
| 🏢 `companyName` · `companyId` · `companyProfileUrl` · 👥 `companyFollowers` | 96.6 % | 90.0 % |
| 🌐 `companyWebsite` — the employer's own domain, the lead hook | 95.9 % | 88.7 % |
| `companyText` | 92.0 % | 93.7 % |
| `area` | 91.3 % | 90.0 % |
| 📍 `address` · `latitude` · `longitude` | 66.5 % | 73.8 % |
| 📍 `zipcode` | 66.4 % | 73.7 % |
| 📍 `city` | 66.2 % | 73.7 % |
| ⭐ `ratingScore` · `ratingCount` | 55.6 % | 51.7 % |
| ⏳ `applyDeadline` | 40.7 % | 50.9 % |
| 🖱️ `externalApplyUrl` — the employer's off-site ATS | 35.6 % | 25.5 % |
| 🔀 `source` — the partner feed an aggregated ad came from | 10.9 % | 24.2 % |

`source` is **not** a dead column and it is **not** random: it is filled only on aggregated partner-feed ads (`tid` starting `r`), never on Jobindex's own ads. Measured that way it is 24.0 % of the `r` ads in Sample A and 56.0 % in Sample B. Values seen: Jobcenter, Powermatch, SOSU.nu, frivilligjob.dk, The Hub, Studerende Online, Akademikernes Jobbank, Skolejobs.dk, Sundhedsjobs.dk. If your search is all Jobindex-hosted ads the column will be empty — that is the ad mix, not a parse failure.

**Optional blocks**

- `addresses[]` + `geojson` (`addressFormat: "nested"`) · `rawHtml` (`includeRawHtml`)
- `fullDescription`, `fullDescriptionChars`, `contactEmails[]`, `contactPhones[]`, `adPageUrl`, `adPageHost`, `adPageIsEmployerSite`, and `ldEmploymentType` / `ldDatePosted` / `ldValidThrough` / `ldHiringOrganization` (`fetchFullDescription`)

#### ⚠️ What this actor does NOT return

Being straight about it, because Jobindex simply does not publish these:

- ❌ **No salary or compensation** — there is no salary field anywhere in the payload.
- ❌ **No seniority level.**
- ❌ **Contact e-mail/phone only via `fetchFullDescription`**, and only on the ads whose text contains one — measured **25 of 50** (e-mail) and **22 of 50** (phone) on run `aQj5NkYyWihxUc0TG`. It is never guessed or pattern-generated.
- ❌ **`externalApplyUrl` does not mean "every ad has an ATS".** Jobindex serves its own in-house quick-apply endpoint (`/api/app/user/qa-partial-login/…`) under the same payload key as the employers' real ATS links. On the 1,000 newest ads, 435 ads carried that key and **79 of them (18 % of the filled values) were that internal endpoint** — a jobindex.dk URL with zero vendor signal, on a path jobindex.dk's own `robots.txt` disallows. Those are **not** reported as an external apply URL, which is why the measured figure is 35.6 % and not 43.5 %. Nothing is lost: those ads are still reachable through `jobUrl`, which is 100 % filled.
- ❌ **A PDF advert is not turned into text.** A small minority of ad pages redirect to a PDF file. The actor does not shovel the raw PDF stream into `fullDescription` (measured once in 50 ads); such a row keeps its free teaser and the run log counts it.

***

### Who is it for?

- 🤝 **Staffing & recruitment agencies** — every Danish employer hiring right now, with their own website (96 % / 89 % filled) and, where they run one, their external ATS. 7,480 new ads in the last measured week means a standing weekly feed.
- 🎯 **B2B lead-gen & GTM teams** — hiring is a buying signal. Filter to "companies hiring warehouse staff within 25 km of Odense" and you have a named, qualified, geocoded list.
- 🧰 **ATS & HR-tech vendors** — where an employer runs an external ATS, `externalApplyUrl` names it (hr-manager, emply or elvium in every case measured). Expect it on a quarter to a third of rows, not all of them.
- 📊 **Labour-market analysts** — 81 categories × 110 regions × date, with coordinates for mapping.
- 🏗️ **Job-board & aggregator builders** — a clean Danish feed with descriptions and deadlines.

***

### Quick start

```json
{ "maxItems": 100 }
```

That returns the 100 most recently posted ads on the whole board.

#### Companies hiring developers around Copenhagen

```json
{
  "subcategoryIds": ["1"],
  "geoAreaIds": ["15182"],
  "maxItems": 200
}
```

#### Weekly lead feed with full ad text and contacts

```json
{
  "subcategoryIds": ["80"],
  "postedWithin": "7",
  "fetchFullDescription": true,
  "maxItems": 500
}
```

#### Paste a search you already built in the browser

```json
{
  "searchUrls": ["https://www.jobindex.dk/jobsoegning?q=sygeplejerske&geoareaid=7"],
  "maxItems": 300
}
```

The actor parses the querystring, keeps every filter (including ones Jobindex adds in future), and takes over paging and sort order itself.

#### Sweeping more than 1,000 ads — the ceiling and how to beat it

Jobindex serves 20 ads per page and stops at page 50, so **any single search tops out at 1,000 ads** (page 51 answers 404 — verified). The whole corpus is 18,970 ads (jobindex.dk's own hit count, measured 2026-09-05), so a full sweep needs partitioning:

```json
{
  "subcategoryIds": ["27", "47", "70"],
  "geoAreaIds": ["2", "3", "4", "15182", "15179", "15180"],
  "partitionBy": "subcategoryAndGeoArea",
  "maxItems": 0
}
```

`partitionBy` turns one over-ceiling search into a grid of small ones and de-duplicates across them. Measured 2026-09-05 by reading the live hit count of all 81 categories in one run, only 3 of them exceed 1,000 ads on their own (Pædagog 1,841 · Pleje og omsorg 1,508 · Detailhandel 1,275); splitting those by region puts every cell comfortably under the cap. The run log warns you whenever a search hit the ceiling.

***

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `searchUrls` | array | `[]` | Paste any jobindex.dk search URL; its filters are parsed out. Overrides the fields below. |
| `queries` | array | `[]` | Free-text keywords — one search per keyword. |
| `subcategoryIds` | array (81-value picker) | `[]` | Job categories, grouped by their 10 parent families. |
| `geoAreaIds` | array (110-value picker) | `[]` | Country / region / area / municipality. |
| `address` + `radiusKm` | string + int | – | Distance search around a Danish address. Both are required together. |
| `companyIds` | array | `[]` | Every open ad for named employers — the "watch my target accounts" mode. |
| `jobTitleIds` | array | `[]` | Jobindex's normalised job-title taxonomy ids. |
| `employmentTypes` | array (12) | `[]` | Permanent, fixed-term, student job, apprentice, graduate, internship, freelance… |
| `workingHoursTypes` | array (3) | `[]` | Full-time / part-time / unspecified. |
| `remoteMode` | select (5) | any | Remote offered / not offered / 100 % home working. |
| `postedWithin` | select (5) | all current | Today · last 7 days · last 30 days · all currently online · no filter. |
| `sortBy` | select | `date` | **Leave on `date`** — see "Correctness" below. |
| `partitionBy` | select (4) | `none` | Split a search by category, location, or both, to beat the 1,000-ad ceiling. |
| `onlyNewSince` | date | – | Only ads first published on/after this date. |
| `maxItems` | int | `100` | Row cap. `0` = no cap. |
| `maxPagesPerQuery` | int | `50` | Page cap per search (site maximum is 50). |
| `excludeSponsored` | bool | `false` | Drop sponsored/branded placements. |
| `dedupeByTid` | bool | `true` | Collapse repeated ads before delivery **and before billing**. |
| `fetchFullDescription` | bool | `false` | Full advert body + contact e-mail/phone. No extra charge. |
| `includeDescription` / `includeGeo` / `includeCompanyRatings` | bool | `true` | Trim blocks you do not need. |
| `includeRawHtml` | bool | `false` | Keep each ad's raw HTML blob (10–20 KB/row). |
| `addressFormat` | select | `flat` | `flat` inlines the first address; `nested` keeps `addresses[]` + `geojson`. |
| `proxyConfiguration` | object | Apify datacenter | Datacenter is measured-sufficient; residential/BYO is available. |
| `maxConcurrency` | int | `4` | Parallel page fetches. |
| `requestRetries` | int | `3` | Retries per request, each on a fresh proxy session. |

***

### Correctness — why `sort=date` is the default

Jobindex paginates by offset over a **live, re-ranked index**. In relevance order the ranking shifts while you page, so ads repeat across adjacent pages and others are never served at all. Re-measured 2026-09-05 on the identical search (`subid=1`, hit count 432), contiguous pages 1–14 through the Apify proxy, with de-duplication switched **off** so the raw truth shows:

| Sort order | Raw rows | Unique ads | Duplicates | Run |
|---|---|---|---|---|
| **`date` (default)** | 280 | **280** | **0.00 %** | `8h4zXIkfHLsiXSjmh` |
| `score` (relevance) | 280 | 253 | **9.64 %** | `czaTeJMVWob4cXYyc` |

A full-depth date-order crawl of the whole board went 50 pages deep and returned **1,000 raw / 1,000 unique — 0 duplicates** (run `DQRWh1jhWtGL9gpoW`).

So relevance order would have made you **pay for ~10 % duplicate rows while silently missing part of the corpus**. Date order is the default, and rows are de-duplicated on `tid` **before** they are delivered or billed either way.

***

### Reliability — measured, not claimed

Everything below ran on the Apify platform through the **Apify datacenter proxy**, never a home IP, on 2026-09-05:

- **Search surface: 185 result pages fetched, 0 failed (100 %).** That is the full 50-page depth of the whole board twice (`DQRWh1jhWtGL9gpoW`, `rfJ1F9RuIfpU8HYf5`), page 1 of all 81 categories twice (`UiAezXGVmGlve4MmR`, `Zgg2eAiPuuf2JaoHk` — 81 pages each), and 14 contiguous pages in each sort order. Zero 403s, zero CAPTCHAs, zero rate limiting, no login wall, no Cloudflare.
- **Ad-page enrichment: 43 of 50 ad pages returned a usable advert** (`aQj5NkYyWihxUc0TG`) — 6 third-party employer sites did not answer and 1 served a PDF instead of a page. Those rows keep their free teaser; enrichment reaches off jobindex.dk onto employer domains, so it depends on those sites, not on our transport.
- **Billing matched delivery on every run**: `chargedEventCounts` == dataset item count on all of them (1,000 / 1,300 / 100 / 100 / 50 / 280 / 280 / 1).
- A page that fails after all retries is logged as a **hole and skipped** — never treated as end-of-list, because that is how a crawl silently truncates.

***

### Output sample

One real row, verbatim from run `DQRWh1jhWtGL9gpoW` (description truncated here for length):

```json
{
  "tid": "h1688024",
  "title": "Fagligt stærk kollega til økonomistyring og skriftlige forelæggelser",
  "jobUrl": "https://www.jobindex.dk/vis-job/h1688024",
  "companyName": "Motorstyrelsen",
  "companyId": 51699,
  "companyWebsite": "https://motorst.dk/",
  "companyProfileUrl": "https://www.jobindex.dk/virksomhed/51699/motorstyrelsen#om-virksomhed",
  "area": "Aalborg",
  "publishedDate": "2026-09-05",
  "lastDate": "2026-09-06",
  "applyDeadline": "2026-09-06T21:59:59Z",
  "applyDeadlineAsap": false,
  "externalApplyUrl": "https://candidate.hr-manager.net/ApplicationForm/SinglePageApplicationForm.aspx?cid=3010&ProjectId=166642&DepartmentId=19142&MediaId=5130&apply_with=jobdk",
  "homeWorkplace": false,
  "isArchived": false,
  "isSponsored": true,
  "hasVideo": false,
  "source": null,
  "companyText": "Motorstyrelsen",
  "description": "Er din talforståelse i top, og kan du omsætte tallene til en organisatorisk kontekst…",
  "companyFollowers": 234,
  "ratingScore": 4,
  "ratingCount": 91,
  "city": "Aalborg",
  "zipcode": "9000",
  "address": "Lauritzens Plads 1",
  "latitude": 57.04899389,
  "longitude": 9.94566904,
  "addressCount": 1,
  "scrapedAt": "2026-09-05T21:45:03.454Z"
}
```

***

### How much does it cost?

**$0.0005 per job ad delivered — $0.50 per 1,000.** One event, `job-scraped`. That is it.

- **No per-run start fee.** 13 of the 14 live Jobindex actors charge one on top of their row price.
- **No detail-fetch upsell.** Full advert text and contact details are included in the same per-ad price.
- **You are never billed for a duplicate.** Rows are de-duplicated on `tid` before delivery.
- **Rows dropped by your filters or your `maxItems` cap are never billed**, because delivery and billing are the same atomic call.

| Run | Ads | Cost |
|---|---|---|
| Quick test | 100 | **$0.05** |
| A category in one region | 500 | **$0.25** |
| One search to the site's ceiling | 1,000 | **$0.50** |
| The whole live board | 18,970 | **~$9.49** |
| A weekly feed for a year | ~389,000 | **~$194** |

Plus Apify platform compute, which is small: the actor runs at 512 MB with no browser.

***

### Is it legal to scrape Jobindex.dk?

You are responsible for your own compliance; this section states the facts rather than giving legal advice.

- The data is **public, employer-side business information** — company names, company websites, workplace addresses and job adverts. It is not personal data about candidates, and no login or paywall is bypassed at any point.
- **`robots.txt` does disallow the search parameters this actor uses.** Quoted verbatim from `https://www.jobindex.dk/robots.txt` on 2026-09-04:

  ```
  Disallow: /c?
  Disallow: /api/
  Disallow: /jobsoegning/*/*/*?
  Disallow: /jobsoegning/*?*&
  Disallow: /jobsoegning?*&*&
  Disallow: /jobsoegning*page=
  Disallow: /jobsoegning*sort=
  Disallow: /jobsoegning*jobage=
  Disallow: /jobsoegning*archive=
  Disallow: /jobsoegning*_supid=
  Disallow: /jobsoegning*companyid=
  Disallow: /jobsoegning*radius=
  Disallow: /jobsoegning*subid=
  Disallow: /jobsoegning*flag=
  Disallow: /jobsoegning*geoareaid=
  ```

  `robots.txt` is a crawler directive, not a contract or a technical access control. We publish the lines so the decision is yours to make with the facts in front of you.
- The actor **never emits a `/c?` click-tracker URL nor a jobindex.dk `/api/` URL** — both disallowed paths. It uses the canonical `vis-job` link. Verified by scanning every field of all 2,300 rows in the two samples above: **0 occurrences of either**. (An earlier build did ship Jobindex's `/api/app/user/qa-partial-login/…` quick-apply link inside `externalApplyUrl`; that is fixed, see "What this actor does NOT return".)
- Requests are polite: concurrency 4 by default, one ad per row, no login, no CAPTCHA solving, no signature forging.
- If you export contact details via `fetchFullDescription`, **GDPR compliance for that data is your responsibility** as the data controller.

***

### FAQ

**Why is my run capped at 1,000 ads?**
That is Jobindex's own ceiling — 20 ads/page × 50 pages, and page 51 returns 404. Use `partitionBy` to split the search into smaller cells. The log tells you whenever a search hit the cap.

**Does it return salaries?**
No. Jobindex's payload contains no salary field of any kind. Any actor claiming Danish salaries from this source is inferring them from ad text.

**How do I get contact e-mails?**
Set `fetchFullDescription: true`. Measured on 50 consecutive ads (run `aQj5NkYyWihxUc0TG`): 43 returned the full advert body (median 7,467 characters, every one longer than the teaser), 25 yielded a contact e-mail and 22 a phone number; 36 of the 50 ad pages were on the employer's own domain. Ads that do not publish a contact return an empty array — nothing is guessed.

**Will it find remote jobs?**
Yes, via `remoteMode`, but be aware "100 % home working" is a genuinely tiny slice of the Danish market — measured at a few dozen ads board-wide. An almost-empty run there is the market, not a bug.

**Can I monitor for new ads?**
Yes. Schedule it with `postedWithin: "7"` and `onlyNewSince` set to your last run date. Measured 2026-09-05, jobindex.dk's own hit count for the last 7 days was **7,480 ads**.

**Can I search expired ads?**
No. Every `postedWithin` option scopes *currently live* ads: today, the last 7 or 30 days, everything online, or no filter at all. There is no archive value — sending one is rejected by the input schema before the run starts.

**Why do some rows have no city or coordinates?**
Because a substantial minority of ads publish no structured address at all — nationwide roles, agencies, aggregated partner-feed ads. Measured, not guessed: coordinates were present on **66.5 %** of the 1,000 newest ads and **73.8 %** of an all-categories sample, so roughly a quarter to a third of rows have none. Those rows keep `area` (90–91 % filled), which is usually the town name.

**Do I need residential proxies?**
No. Datacenter proxies were byte-identical to direct and residential requests in testing, with zero blocks. Residential and BYO proxies are available in `proxyConfiguration` if you want a Danish exit IP.

**What is `tid`?**
The stable Jobindex ad id, 100 % filled. It is the dedupe key and it round-trips: `https://www.jobindex.dk/vis-job/<tid>`. Its first letter also tells you which kind of ad you have: `h…` is hosted by Jobindex, `r…` came in from a partner feed (those are the ones that fill `source` and usually lack a structured address), plus a handful of `o…`.

**Why is `externalApplyUrl` empty on most rows?**
Because most Danish employers apply through Jobindex itself rather than an external system, and the field only reports a genuine off-site ATS link. Measured: 35.6 % of the 1,000 newest ads and 25.5 % of an all-categories sample — and every vendor seen across those 2,300 rows was hr-manager, emply or elvium. Jobindex's own internal quick-apply endpoint is filtered out rather than counted as an ATS, so the number is lower and honest.

**Does the actor fail if my search matches nothing?**
No. An honestly empty search exits cleanly with a status message and charges nothing. It fails loudly only when the page loads but the data payload cannot be parsed — which means the site changed and the parser needs fixing, not a retry.

***

### Feedback

Found a field that should be here, or an ad that parsed wrong? Open an issue on the actor page — include the `tid` and the input you used and it can be reproduced exactly.

# Actor input Schema

## `searchUrls` (type: `array`):

Paste any jobindex.dk/jobsoegning URL straight out of your browser after building a search in their UI — the query string is parsed and used as-is. This is the fastest way to reproduce a filter you already made on the site. When you supply URLs, the individual filter fields below are IGNORED (each URL is its own search); "Sort order", "Max pages per search" and all the output toggles still apply.

## `queries` (type: `array`):

Free-text keywords, one search per entry. Danish terms match best. Measured hit counts on 2026-09-05: "udvikler" 8,322 · "sygeplejerske" 1,609 · "elektriker" 385. Leave empty to search the whole corpus (18,970 live ads measured 2026-09-05).

## `subcategoryIds` (type: `array`):

Jobindex's own 81-value subcategory taxonomy, grouped by its 10 parent families. Several categories are OR-ed together (measured: subid 1 = 431 ads, subid 17 = 755 ads, both together = 1,186 = the exact sum). Combined with a region or keyword they are AND-ed (subid 1 + København = 132).

## `geoAreaIds` (type: `array`):

The live 110-value location vocabulary in four tiers — Land (Danmark, Grønland, Færøerne, Udlandet), Region, Område (Storkøbenhavn, Nordsjælland, Fyn, Sydjylland, Skåne) and all 98 kommuner. OR-ed together (measured: København 2,567 + Aarhus 1,419 → both 3,940, the 46-ad difference being ads listed in both).

## `address` (type: `string`):

A town, postcode or street ("Aarhus", "8000 Aarhus C", "Odense"). Only has an effect together with "Radius" below — on its own it is ignored by the site and you get the whole corpus back. Measured from Aarhus: 5 km 812 ads · 10 km 1,235 · 25 km 1,705 · 50 km 2,698 · 100 km 6,989.

## `radiusKm` (type: `integer`):

Distance in kilometres from "Address" above. Must be 1 or more — jobindex.dk answers radius=0 with a 404, so the actor drops a 0 rather than sending it. Ignored unless an address is set.

## `companyIds` (type: `array`):

Watch specific employers: every open ad for one company. The id is the number in a jobindex.dk/virksomhed/<ID>/<slug> profile URL, and the actor also returns it as companyId on every row, so you can pull a list once and then monitor those accounts forever. Measured: companyid 24096 = 39 open ads.

## `jobTitleIds` (type: `array`):

Jobindex's normalised job-title taxonomy. There is no public list of these ids; take them from a jobtitleid= parameter in a search URL you built on the site. Verified as a real filter (jobtitleid=100 → 3 ads, jobtitleid=2000 → 0). Most users should use Keywords or Job categories instead.

## `employmentTypes` (type: `array`):

OR-ed together (measured: Studiejob 1,869 → Studiejob + Praktikplads 1,988).

## `workingHoursTypes` (type: `array`):

Full-time / part-time. Measured: Deltid (part-time) = 4,374 ads.

## `remoteMode` (type: `string`):

WARNING — "100 % hjemmearbejde" is a genuinely tiny slice of the Danish market: measured 9 live ads today, against 630 for "Muligt" (remote possible). If you pick it and get ~9 rows, the run is not broken.

## `postedWithin` (type: `string`):

The monitoring lever. Jobindex's own hit counts, measured 2026-09-05: "Sidste 7 dage" 7,480 · "Sidste 30 dage" 21,176 · "Online" 23,706 · no filter 18,970. "Fra i dag" tracks the current day and so swings from a handful late at night to the low hundreds by late afternoon. Run this daily or weekly with "Sidste 7 dage" and you get a live feed of newly hiring Danish employers instead of one static dump.

## `sortBy` (type: `string`):

CORRECTNESS SETTING, not cosmetics. Jobindex paginates by offset over a live re-ranked index. Measured on contiguous page bands: relevance order returned 240 raw / 198 unique rows over pages 1-12 (17.5% duplicates) and silently lost about a sixth of the corpus, while date order returned 431 raw / 431 unique over the full 23-page depth of a 431-ad search — exactly the hit count, zero duplicates. Keep "Newest first" unless you specifically want relevance ranking.

## `maxItems` (type: `integer`):

Hard cap on delivered — and therefore billed — rows across the whole run. 0 means no cap. Duplicates and rows dropped by your filters never count against it and are never charged.

## `maxPagesPerQuery` (type: `integer`):

20 ads per page. The site itself stops at page 50, so anything above 50 is clamped to 50.

## `partitionBy` (type: `string`):

By default all the categories/locations you picked are OR-ed into ONE search, which the site caps at 1,000 ads. Splitting runs a SEPARATE search per category and/or per location, so each cell stays under the cap and you can pull far more of the corpus. Measured 2026-09-05: only 3 of the 81 categories exceed 1,000 on their own (Pædagog 1,841 · Pleje og omsorg 1,508 · Detailhandel 1,275) and splitting those by region drops every cell well below it. Rows are still de-duplicated on the ad id across every partition, so an ad filed under two categories is delivered and billed once.

## `onlyNewSince` (type: `string`):

ISO date, e.g. 2026-09-01. Compared against the ad's own first-published date. Older ads are dropped before delivery, so they are never billed — this is the cheap way to run the same saved search on a schedule and pay only for what is new.

## `excludeSponsored` (type: `boolean`):

Jobindex re-injects paid, branded placements into result pages. They are real ads, but if you are building a clean corpus you may not want them. Dropped rows are never billed.

## `dedupeByTid` (type: `boolean`):

Collapse repeats on the stable ad id (tid) BEFORE delivery and before billing. Leave this on — it is what stops you paying twice for the same ad when it appears in two categories or on two pages.

## `fetchFullDescription` (type: `boolean`):

OFF by default. The search payload already carries a free teaser (measured 30-399 characters, median ~300). Turn this on and the actor also opens each ad's own page at jobindex.dk/jobannonce/<id>, which for most ads redirects to the employer's own careers/ATS page — and mines the full advert text, any contact e-mail and any phone number from it. Measured on 50 consecutive live ads 2026-09-05 (run aQj5NkYyWihxUc0TG): 43/50 returned a usable advert, all 43 longer than the teaser (median 7,467 characters), 25/50 yielded a contact e-mail, 22/50 a phone number, 36/50 landed on the employer's own domain; 6 third-party sites did not answer and 1 served a PDF instead of a page, and those rows keep the free teaser. It costs one extra HTTP request per ad — the 50-ad run took 73 seconds end to end at the default concurrency of 4 — and it is NOT charged as a separate event: it is included in the per-ad price.

## `includeDescription` (type: `boolean`):

The ad teaser rendered from the markup that ships with the search results. Measured across 2,300 ads on 2026-09-05: 30–399 characters, median 286–300. Free — no extra request. Verified present on 100% of ads.

## `includeGeo` (type: `boolean`):

Street address, postcode, city, latitude and longitude. Measured fill: ~84% of ads carry a mapped address.

## `includeCompanyRatings` (type: `boolean`):

Jobindex's own employer rating (1-5), the number of reviews behind it, and how many people follow the company. Measured fill: ~60% of ads have a rating.

## `includeRawHtml` (type: `boolean`):

Adds the ad's raw markup blob (roughly 9-20 KB per row). Useful if you want to re-parse yourself; it will bloat your dataset badly on large runs.

## `addressFormat` (type: `string`):

Flat inlines the first address into the row (best for CSV/Sheets). Nested keeps jobindex's full addresses\[] array and its GeoJSON feature collection intact (best for ads with several worksites).

## `proxyConfiguration` (type: `object`):

Apify DATACENTER proxy is measured-sufficient here: 195+ live calls, zero 403s, zero CAPTCHAs, no rate limiting. Residential is NOT needed and only costs you more; switch only if you want Danish exit IPs or you are bringing your own proxies.

## `maxConcurrency` (type: `integer`):

How many pages/ad pages to fetch at once. 3-5 is polite and was measured clean; the site showed no rate limiting at these levels.

## `requestRetries` (type: `integer`):

Keep at 3 or more. The single failure across 195 measured live calls was a transient "HTTP/2 stream early terminated" that succeeded on retry — it was not a block.

## Actor input object example

```json
{
  "searchUrls": [],
  "queries": [],
  "subcategoryIds": [],
  "geoAreaIds": [],
  "companyIds": [],
  "jobTitleIds": [],
  "employmentTypes": [],
  "workingHoursTypes": [],
  "remoteMode": "",
  "postedWithin": "",
  "sortBy": "date",
  "maxItems": 100,
  "maxPagesPerQuery": 50,
  "partitionBy": "none",
  "excludeSponsored": false,
  "dedupeByTid": true,
  "fetchFullDescription": false,
  "includeDescription": true,
  "includeGeo": true,
  "includeCompanyRatings": true,
  "includeRawHtml": false,
  "addressFormat": "flat",
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "maxConcurrency": 4,
  "requestRetries": 3
}
```

# Actor output Schema

## `items` (type: `string`):

The dataset of scraped jobindex.dk job ads (one ad per row, de-duplicated on the ad id).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchUrls": [],
    "queries": [],
    "subcategoryIds": [],
    "geoAreaIds": [],
    "companyIds": [],
    "jobTitleIds": [],
    "employmentTypes": [],
    "workingHoursTypes": [],
    "sortBy": "date",
    "maxItems": 100,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/jobindex-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchUrls": [],
    "queries": [],
    "subcategoryIds": [],
    "geoAreaIds": [],
    "companyIds": [],
    "jobTitleIds": [],
    "employmentTypes": [],
    "workingHoursTypes": [],
    "sortBy": "date",
    "maxItems": 100,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/jobindex-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchUrls": [],
  "queries": [],
  "subcategoryIds": [],
  "geoAreaIds": [],
  "companyIds": [],
  "jobTitleIds": [],
  "employmentTypes": [],
  "workingHoursTypes": [],
  "sortBy": "date",
  "maxItems": 100,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call scrapersdelight/jobindex-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,scrapersdelight/jobindex-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yi7MiJFjI74Ssx7tv/builds/BgjpBXRp8yfg6qNAR/openapi.json
