Duunitori Jobs Scraper - Finland Job Ads, Apply URL & Y-tunnus avatar

Duunitori Jobs Scraper - Finland Job Ads, Apply URL & Y-tunnus

Pricing

from $1.00 / 1,000 per job returneds

Go to Apify Store
Duunitori Jobs Scraper - Finland Job Ads, Apply URL & Y-tunnus

Duunitori Jobs Scraper - Finland Job Ads, Apply URL & Y-tunnus

Scrape duunitori.fi, Finland's largest job board (~17,500 live ads), by keyword, city, region, industry, occupation, contract type, language, remote and salary-published filters. Optional detail fetch adds the external apply URL + ATS domain; optional employer fetch adds the Y-tunnus business ID.

Pricing

from $1.00 / 1,000 per job returneds

Rating

0.0

(0)

Developer

Scrapers Delight

Scrapers Delight

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

3 days ago

Last modified

Share

🇫🇮 Duunitori Jobs Scraper — Finland's whole job board, as structured rows

Turn duunitori.fi, Finland's largest job board, into a clean dataset: ~17,500 live job ads, searchable by keyword, city, region, industry, occupation, contract type, ad language, remote, published-salary and more — with optional enrichment that adds the employer's external apply URL and the ATS vendor behind it, and the employer's Y-tunnus (Finnish Business ID).

No login, no API key, no CAPTCHA. Pay per row you actually receive.


🎯 What does Duunitori Jobs Scraper do?

It drives Duunitori's own search, reads every result page, deduplicates by job id, and gives you one row per job ad. Three levels, each priced separately so you only pay for the depth you use:

LevelExtra request?What you get
Search results (always)nojob id, title, company, location, industry, occupation bucket, ad tier, badge, posted date, logo, ad URL
Job details (optional)1 per jobfull description, structured salary (min/max/currency/period), employment type, exact posted timestamp, application deadline, the ad's own location string (usually longer than the card's) — plus the external apply URL and the ATS domain
Employer profile (optional)1 per unique employerY-tunnus (Finnish Business ID), officially registered company name, official toimiala, company website, social profiles, company description

💡 Why the apply URL and the Y-tunnus matter

Every other Duunitori scraper on this Store stops at title / company / location / salary text. We pulled all four rivals' output schemas live on 2026-09-04: not one returns the external apply URL, and not one returns a business ID.

  • applyDomain names the recruiting software each Finnish employer runs. From one 60-ad sample: ats.talentadore.com, jobilla.talentadore.com, emp.jobylon.com, recright.com, korsisaari.solaforce.com, steadyenergy.careers.haileyhr.app, corehr.hrcloud.hr, tamk.rekrytointi.com, georgfischer.wd103.myworkdayjobs.com, plus employers' own careers domains. That is a technographic signal you would otherwise buy from a data vendor.
  • employerBusinessId (Y-tunnus) is the join key into PRH / YTJ and every Finnish B2B database — turnover, employee count, directors, credit rating. It turns a job ad into a company record.

Together they answer questions a job feed alone cannot: which companies are hiring, which HR software do they already pay for, and who exactly are they on the business register?


📊 Field fill rates — measured, not claimed

One unfiltered sweep of the newest 60 ads on 2026-09-04, details + employer profiles on:

FieldFillFieldFill
jobId, title, company, location100%descriptionText100%
industry, adTier, badge, url, slug100%datePostedpostedDate100%
companyLogoUrl, allLocations100%validThroughexpiresInDays100%
employmentType100%occupation98.3%
primaryLocation91.7%applyUrl / applyDomain65.0%
companyWebsite63.3%salaryCurrency / salaryUnit30.0%
salaryMin26.7%salaryMax30.0%
employerBusinessId (Y-tunnus)31.7%employerIndustry, employerRegisteredName31.7%
employerSocials30.0%employerProfileUrl46.7%
directApply20.0%hiringOrganization100%

An independent re-measurement on a separate 60-row unfiltered sweep landed within 3.3 points on every one of these (applyUrl 63.3%, companyWebsite 63.3%, employerBusinessId 33.3%, occupation 95.0%, primaryLocation 90.0%, salaryCurrency 28.3%), and put employerProfileUrl at 60–68% across its larger sample.

Honest reading of the low numbers (these are properties of Duunitori, not of the scraper):

  • applyUrl is 65% because 21 of 60 employers take applications inside Duunitori rather than linking out. Those rows carry applyIsExternal: false, which is what tells "no external ATS" apart from "we failed to find one". A separate 28-ad sample measured 71.4%.
  • directApply is 20% and is never false. It is a straight passthrough of the ad's own schema.org directApply flag, and Duunitori only publishes that flag when it is true — measured over 172 rows: true on 20%, absent on 80%, literally false on none. So null means "the ad does not state it", not "false"; use applyIsExternal (derived from the page's own apply anchors) when you need a yes/no.
  • hiringOrganization is the ad's JSON-LD organisation name, which is normally the same employer as company, just lower-cased ("wsp finland" vs "WSP Finland") — 21 distinct values against 21 on the same 25 rows. It is kept for structured-data fidelity, not as extra information: use company for the display name and employerRegisteredName for the official registry name.
  • allLocations is the ad's own location string, which is usually longer than the card's ("Tampere, Oulu, Rovaniemi" where the card says "Tampere ja 4 muuta") — but it is what the employer published, not a guaranteed expansion of "ja N muuta". Measured counter-example: an ad whose card read "Forssa ja 60 muuta" published allLocations: "Forssa".
  • Salary is ~30% because only ~2,450 of the ~17,500 live ads publish pay at all. Switch on "Only ads that publish a salary" and every delivered row carries structured salary — measured on a 15-row run: salaryCurrency 15/15, salaryMax 15/15, salaryMin 13/15 (87%), because a handful of ads publish a single figure rather than a range (e.g. 3,600–4,300 EUR/MONTH, or a flat 20.80 EUR/HOUR that lands in salaryMax only).
  • Y-tunnus is 32% because only Duunitori's modern employer profiles carry the registry block, and only ~47% of ads link to an employer profile at all. Of the profiles that exist, 12 of 21 unique employers in that run carried usable data. You are never charged for a profile that returned nothing.
  • descriptionText runs 1,748–11,851 characters (median 4,220).

🔎 Every filter, verified against live result counts

Each filter below was run against the live site on 2026-09-04 and changed the result count — none of them is decorative. Board size that day: 17,596 live ads.

InputLive countInputLive count
(no filter)17,596remoteOnly926
searchQueries: ["developer"]122withSalaryOnly2,452
+ searchDescriptions190greatPlaceToWorkOnly79
municipalities: ["helsinki"]3,429diversityPromiseOnly1,697
municipalities: ["oulu"]903goodSummerJobOnly5
regions: ["uusimaa"]6,143ageGroup: students1,074
regions: ["pirkanmaa"]2,201ageGroup: below_1840
regions: ["lappi"]1,314employmentType: full_time15,469
industries: ["talonrakennus"]1,444employmentType: part_time2,956
industries: ["lakiala"]84contractType: permanent13,569
occupations: ["kuljettaja"]744contractType: fixed_term4,658
adLanguage: eng_lang1,778contractType: summer_job89
adLanguage: swe_lang326adLanguage: fi_lang13,560

Filters combine: industries:["lakiala"] + municipalities:["helsinki"] → 53 ·

industries:["lakiala"] + searchQueries:["juristi"]
→ 23 · regions:["pirkanmaa"] + employmentType:["part_time"] → 320.

Two traps this Actor handles for you, both measured:

  1. Duunitori keeps only the LAST value of a repeated filter. Sending filter_work_type=full_time&filter_work_type=part_time returns 2,956 — the part-time count, not the union — and a comma-joined value is ignored entirely (returns the full 17,596). So every multi-select here is expanded into separate searches and merged by job id. You are never billed twice for a job that two of your searches both matched.
  2. The site's ?category= parameter does not filter anything. category=lakiala, category=rakennusala and category=tietotekniikka all returned the identical 17,537 total. Only the /tyopaikat/ala/<slug> browse path filters — which is what this Actor uses.

A region/industry/occupation slug Duunitori doesn't know returns HTTP 200 with zero rows, which looks exactly like a real empty search. So slugs are checked against the site's own taxonomy (19 regions, 191 industries, 10 occupations, enumerated live) before a request is spent, the closest matches are named in the log, and if none of your slugs are valid the run stops clean rather than quietly scraping the whole board.


📅 Freshness, deltas and the recurring-run shape

5,786 of 17,537 ads were posted in the last 7 days — about 33% weekly churn. That is what makes a schedule worth setting up.

  • postedWithinDays — verified: with 2, 53 of 80 cards were dropped and every delivered row came back at 0, 1 or 2 days old.
  • expiringWithinDays — the renewal/re-post trigger a staffing firm calls on. Verified: with 5, 9 of 10 rows were filtered and the survivor expired in 3 days.
  • onlyNewSince — incremental mode. Verified across three runs: run 1 returned 7 jobs, run 2 on the same key returned 0 ("already seen"), run 3 on a different key returned 7 again. State lives in a named key-value store, so it survives between runs.

ℹ️ Duunitori has no date sort and no date filter of its own — order_by=date, -date, published and search_rank all return byte-identical pages — so these three are applied by this Actor after reading each ad's date. Rows they remove are never charged.


🧾 Deduplication — and why it changes your bill

Duunitori's index churns ~33% a week, so offset pagination re-shows some ads at depth. Measured on contiguous page bands: sivu=1..12 → 240 fetched / 238 unique (0.8%); sivu=600..619 → 400 fetched / 362 unique (9.5%).

The Actor deduplicates by jobId across the whole run, before charging. Proven offline: three identical searches for kehittäjä read 21 cards, dropped 14 duplicates, and delivered and billed exactly 7 unique jobs. Turn deduplicateByJobId off and the same input delivers 21.

For the same reason this page does not promise "all 17,537 jobs". It promises every job matching your search, deduplicated.


🚦 Reliability

duunitori.fi sits behind Cloudflare. Measured escalation ladder on 2026-09-04:

TransportResult
Plain HTTP, no proxy403 — 6,023-byte "Just a moment" challenge
HTTP + datacenter proxy403 — challenge
HTTP + Apify RESIDENTIAL, country FI200 — 209,303 bytes of real job cards
HTTP + RESIDENTIAL, country US / GB403 — challenge

No browser is needed and none is used (512 MB, plain HTTP + HTML parsing — which is why this Actor is cheap to run). FI, SE, NO, DK, DE and EE exits all clear; GB and US are challenged, so the Actor pins the residential exit to FI unless you choose a country yourself. Leave the proxy on Apify Proxy → RESIDENTIAL.

Across 323 requests in 36 validation runs: 0 Cloudflare challenges, 1 transient proxy hiccup (590 UPSTREAM504), retried on a fresh session and recovered — 99.7% first-attempt, 100% eventual. An earlier 158-call load test measured 98.7%.

Three behaviours that keep runs honest:

  • A 404 past the last result page is the END of results, not an error. The last page is exactly ceil(total/20); the Actor computes it and also follows the site's own <link rel="next">, so it doesn't even fetch pages that cannot exist. Verified: a 122-result search read exactly 7 pages in 7 requests with 0 duplicates.
  • A retry always mints a new proxy session — the residential pool's transient failure throws before any HTTP response exists, so a normal HTTP retry cannot see it.
  • The run FAILS LOUDLY when the data is wrong, and exits clean when a search is legitimately empty. Run it with no proxy and it fails in seconds with an actionable message instead of returning an empty dataset that looks like "no jobs found".

💰 Pricing

EventPriceWhen it fires
Per job returned$1.00 / 1,000 rows ($0.001)every deduplicated job delivered to your dataset
Per job detail fetched$1.50 / 1,000 ($0.0015)only with Fetch job details on, and only for rows you receive
Per employer profile fetched$2.00 / 1,000 ($0.002)only with Fetch employer profile on, and only for rows that came back with data

There is no start fee and no platform-usage surcharge. Rows are billed with a budget-aware push, so if you set a max charge you keep exactly the rows you paid for — never more delivered than billed, never more billed than delivered.

What things actually cost (from the measured corpus of 17,537 live ads):

JobRowsCost
One-off snapshot of the whole board17,537$17.54
Same, with full detail on every ad17,537$43.84
Weekly delta (~5,786 new ads)5,786$5.79/week ≈ $301/yr
Every English-language ad, enriched1,778$4.45
Every ad that publishes a salary, enriched2,452$6.13
All 84 legal-sector ads, fully enriched84$0.38

The cheapest existing Duunitori scraper on this Store charges $1.20 per 1,000 and offers no detail or employer tier at all.


🧑‍💼 Who buys this

  • Finnish HR-tech and recruitment-marketing vendors (Jobilla, Talented, Talentadore resellers) — prospect on applyDomain: you can see exactly which ATS every hiring employer already runs.
  • Staffing and RPO firmsexpiringWithinDays surfaces ads about to lapse; that is the call.
  • Sales-intelligence and lead-gen agenciesemployerBusinessId joins each hiring company straight into PRH/YTJ, Vainu or Fonecta.
  • Compensation and labour-market analystswithSalaryOnly plus structured min/max/period.
  • Job aggregators and AI job agents — a full, deduplicated, incremental Finnish feed.

📤 Output sample

{
"jobId": "20479402",
"title": "WSP:llä avoinna useita rooleja rakennuttamisessa (Tampere, Oulu, Rovaniemi, Vaasa tai Seinäjoki)",
"company": "WSP Finland",
"hiringOrganization": "wsp finland",
"location": "Tampere ja 4 muuta",
"allLocations": "Tampere, Oulu, Rovaniemi",
"primaryLocation": "Tampere",
"countryCode": "FI",
"industry": "rakennusala",
"occupation": "projektijohtaja",
"employmentType": "FULL_TIME",
"postedRaw": "Julkaistu 12.8.",
"postedDate": "2026-08-12",
"postedDaysAgo": 23,
"validThrough": "2026-09-09T20:59:00+00:00",
"expiresInDays": 6,
"applyUrl": "https://rekry.wsp.com/jobs?split_view=true&department=Rakennuttaminen",
"applyDomain": "rekry.wsp.com",
"applyIsExternal": true,
"directApply": null,
"url": "https://duunitori.fi/tyopaikat/tyo/wsp-finland-wsplla-avoinna-useita-rooleja-rakennuttamisessa-tampere-oulu-tai-rovaniemi-sdsuu-20479402",
"companyLogoUrl": "https://duunitori.imgix.net/media/images/logos/WSP_logo.png?auto=format&w=59",
"companyWebsite": "https://www.wsp.com/fi-FI",
"adTier": "Ultrakampanja",
"badge": "Katso",
"salaryMin": null, "salaryMax": null, "salaryCurrency": null, "salaryUnit": null,
"employerBusinessId": "0875416-5",
"employerRegisteredName": "WSP Finland",
"employerIndustry": "Yhdyskuntasuunnittelu",
"employerSocials": [
"https://www.facebook.com/WSPglobal/",
"https://www.instagram.com/lifeatwspfinland/",
"https://www.linkedin.com/company/1483604/"
],
"employerProfileUrl": "https://duunitori.fi/yritys/wsp-finland",
"descriptionText": "Rakennuttamisen palvelumme kasvavat – tule mukaan rakentamaan …",
"detailFetched": true,
"employerProfileFetched": true,
"scrapedAt": "2026-09-04T06:21:00.690Z"
}

Three ready-made dataset views ship with it: Overview, Employer leads (business ID, website, ATS vendor, apply link) and Published salaries.

Row shape: outputFormat

ValueShapeKeys per row
flat (default)salaryMin, salaryMax, salaryCurrency, salaryUnit are top-level columns — what a CSV/spreadsheet export wants42
nestedthose four move into a single salary: { min, max, currency, unit } object39

nested groups only the salary fields. Everything else stays flat — including the employer block (employerBusinessId, employerRegisteredName, employerIndustry, employerSocials, employerDescription, employerProfileUrl) and the apply block. That asymmetry is deliberate (the salary quartet is the only group a JSON consumer usually wants as one object), but it is worth knowing before you write a parser against it. The Published salaries dataset view reads both shapes; whichever you run, the other shape's four salary columns come back empty.

Rows are billed identically in both shapesoutputFormat changes presentation only.

Running with no input at all

The Console form prefills Fetch job details on. An API or scheduled call that sends an empty input {} gets the schema default, which is off — you receive list-level rows only, and 24 of the 42 columns (description, salary, apply URL/ATS domain, deadline, employment type, occupation, company website and the whole employer block) come back null on every row, still billed at $0.001/row. The Actor logs that list in a warning at the start of any such run. Send "fetchJobDetails": true if you want the enriched row.


❓ FAQ

Do I need a Duunitori account or an API key? No. Only public job-ad pages are read.

Why must I use a residential proxy? Cloudflare fronts the site and only serves Nordic residential exits. Measured: no proxy and datacenter proxies both get a challenge; RESIDENTIAL pinned to FI returns real data. Leave the proxy field on its default and it is handled for you.

Can I paste a search URL straight from my browser? Yes — Start URLs. Every filter in the URL is kept exactly as you set it, including facets this form doesn't expose. Any ?sivu= page number is stripped; the Actor paginates the whole result set itself.

Can I get all 17,500 jobs in one run? Set Max jobs to 0 and leave the filters empty. Note the honest caveat: Duunitori's index churns ~33% a week and pages shift under offset pagination, so what you get is every job your search matched, deduplicated — not a guaranteed frozen census of the board at one instant.

Does it scrape salaries? It returns the ad's own structured salary when the employer published one — about 30% of ads. Switch on Only ads that publish a salary and every delivered row carries structured salary (measured 15/15 for currency and max; salaryMin was 13/15, because some ads publish one figure, not a range).

What is a Y-tunnus? The Finnish Business ID (e.g. 2854570-7) — the national company registry number. It is the join key into PRH/YTJ and Finnish B2B data providers.

Why is the apply URL missing on some rows? Because those employers take applications on Duunitori itself. applyIsExternal: false tells you so explicitly, rather than leaving you guessing.

Can I run it on a schedule and only get new jobs? Yes — switch on Only jobs not seen in a previous run. Use a different state key per saved search.

Will it search in English? Yes. Duunitori's search works in both languages (developer 122 hits, kehittäjä 7), and adLanguage: eng_lang restricts to the 1,778 ads written in English.

How fast is it? The site answers in ~0.9–1.2 s. A 60-job run with full detail and employer profiles took 84 requests. Concurrency defaults to 2 — there is nothing to gain from hammering a site that answers this quickly.

What happens if my filters match nothing? The run ends successfully with a status message explaining it, and you are charged nothing. Only a genuinely broken run (no page readable, or a page that reports results but yields no parseable cards) is failed.


  • robots.txt. https://duunitori.fi/robots.txt reads, verbatim:

    User-agent: *
    Disallow: /

    with named search-engine crawlers allowlisted below it. This Actor reads only public job-ad pages that Duunitori serves to any visitor. It does not log in, does not create an account, does not solve or bypass a CAPTCHA or any other anti-abuse control, and does not access anything behind authentication. Whether your use is consistent with Duunitori's Terms of Service is your call as the operator — read them before you run this at scale.

  • Personal data (GDPR). Job ads are published in Finland, in the EU. Contact details, recruiter names and any personal data in an ad are personal data under the GDPR, and you are the controller for whatever you do with them. Have a lawful basis, honour erasure requests, and don't use this to build a database of individuals.

  • Rate. Requests are sequential by default and concurrency is capped at 10. Please don't raise it beyond what your use actually needs.

  • Trademarks and content belong to Duunitori Oy and the advertising employers. This Actor is not affiliated with, endorsed by, or connected to Duunitori.


All numbers on this page were measured live on 2026-09-04 (board size 17,596 ads; 60-ad enrichment sample; 323 requests across 36 validation runs). None of them are estimates.