JobKorea Search Scraper
Pricing
from $2.99 / 1,000 jobkorea job records
JobKorea Search Scraper
Scrape job listings from JobKorea.co.kr, South Korea's leading job portal. Extract job titles, companies, locations, salary ranges, and descriptions for Korean recruitment.
Pricing
from $2.99 / 1,000 jobkorea job records
Rating
0.0
(0)
Developer
Jobs API
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
11 days ago
Last modified
Categories
Share
JobKorea public jobs scraper
This Apify Actor extracts publicly available JobKorea (jobkorea.co.kr) job postings. It reads official search pages, verifies each posting on its official detail page, and writes only complete, identity-checked job records to the default dataset.
The implementation uses bounded native HTTPS requests and Cheerio. It does not use a proxy, browser automation, fingerprinting, or challenge bypass. This keeps runs fast and makes target blocking visible in RUN_DIAGNOSTICS instead of producing guessed records.
Modes
mode selects one of five workflows:
search— search one query and enrich up tomaxItemspostings.searchMultiple— run up to five queries, taking up tomaxItemsunique postings per query.single— fetch one official/Recruit/GI_Read/<id>detail URL.multiple— fetch a bounded list of official detail URLs.startUrls— accept official search and/or detail URLs.
Search pagination is bounded by maxPages (1–3). Detail requests run sequentially. maxItems is limited to 10 per query/direct-input mode, and each request has a 5–45 second timeout.
Input
The complete input definition is .actor/input_schema.json. A normal search input is:
{"mode": "search","query": "developer","location": "Seoul","maxItems": 3,"maxPages": 2,"requestTimeoutSecs": 25}
For multiple and startUrls, URL entries are objects:
{"mode": "multiple","urls": [{ "url": "https://www.jobkorea.co.kr/Recruit/GI_Read/49844042" },{ "url": "https://www.jobkorea.co.kr/Recruit/GI_Read/49856065" }],"maxItems": 2}
Only HTTPS URLs on the official JobKorea host are accepted. Detail identifiers, canonical URLs, titles, employers, and search-query matches are checked before a record is emitted.
Dataset
The strict output definition is .actor/dataset_schema.json. Records include, when publicly present:
- stable posting ID, title, employer, company URL/logo, canonical detail URL, and source page;
- advertised location, structured address, remote indicator, employment type, experience, education, schedule, salary, categories, skills, qualifications, preferred qualifications, and benefits;
- normalized posting/start/deadline dates and original deadline text;
- JSON-LD summary, cleaned public detail-section HTML, a readable source-backed description, company information, map URL, and explicit application URL provenance;
- source receipts, identity/canonical verification, search rank/query, HTTP status, and quality counts.
JobKorea currently exposes the concise JSON-LD summary and selected server-rendered public sections to this local client. Some full advertisement-body content is client-only or access-controlled, so fullDescriptionAvailable is explicitly false when that body is not publicly available; the Actor never fabricates it or invents an application URL. Optional fields are omitted rather than emitted as null, blank, placeholder, or empty values.
Run state is kept outside the dataset in the default key-value store:
RUN_SUMMARY— mode, counts, timing, requests, environment, and cloud build/run provenance when available;RUN_DIAGNOSTICS— target, transport, or parsing errors;RUN_SKIPS— candidates rejected after verification (for example, a query mismatch);RUN_HEALTH— dataset quality receipts;REQUEST_RECEIPTS— request URLs, purposes, HTTP statuses, timestamps, and response-body hashes.
Local development
From this directory:
npm install --ignore-scriptsnpm testnpm run lintnpx --yes apify-cli validate-schema .actor/input_schema.jsonnpm run validate
Before a local Actor run, follow the root AGENTS.md storage checks: preserve any existing apify_storage, confirm a new relative APIFY_LOCAL_STORAGE_DIR does not exist, and use --resurrect without --purge. HTTP 401/403/429/451 or a source-authored challenge stops the run, discards buffered jobs, and records resultStatus: SKIPPED with request receipts. Redirects, HTTP 407, and unverified request failures are non-block failures, not source blocks. npm run validate checks required fields, official URL/ID consistency, duplicate IDs/URLs, recursive null/blank/placeholder/empty values, application provenance, and key-value-store consistency.
RUN_SUMMARY.status is the Apify platform run state (SUCCEEDED or FAILED); resultStatus separately reports COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED. A generic transport, redirect, or non-barrier HTTP failure with no rows is DEFERRED, not a source block. Successfully stored rows are preserved as SUCCEEDED/LIMITED if a later non-block error occurs. A zero-row fatal failure is summarized best-effort and uses Actor.fail(). A handled source skip reports platform SUCCEEDED only when its run evidence and summary persist. Diagnostics and request receipts stay in the key-value store, never in the normal dataset.
Cloud deployment
apify pushapify call jobsapi/jobkorea-jobs-search-scraper --memory 512 --timeout 300 --input-file INPUT.json
Keep cloud validation bounded with the provided input. The default dataset contains only persisted complete jobs; diagnostics remain in the default key-value store.
Responsible use
Collect only public job information, respect JobKorea’s terms and robots directives, and use modest bounded request limits. Do not use this Actor for authenticated pages, paywall or challenge bypassing, or collection of candidate or recruiter personal data.