# JobKorea Search Scraper (`jobsapi/jobkorea-jobs-search-scraper`) Actor

Scrape job listings from JobKorea.co.kr, South Korea's leading job portal. Extract job titles, companies, locations, salary ranges, and descriptions for Korean recruitment.

- **URL**: https://apify.com/jobsapi/jobkorea-jobs-search-scraper.md
- **Developed by:** [Jobs API](https://apify.com/jobsapi) (community)
- **Categories:** Jobs, Automation, Developer tools
- **Stats:** 1 total users, 0 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.99 / 1,000 job details

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## JobKorea public jobs scraper

This Apify Actor extracts publicly available JobKorea (`jobkorea.co.kr`) job postings. It reads official search pages, verifies each posting on its official detail page, and writes only complete, identity-checked job records to the default dataset.

The implementation uses bounded native HTTPS requests and Cheerio. It does not use a proxy, browser automation, fingerprinting, or challenge bypass. This keeps runs fast and makes target blocking visible in `RUN_DIAGNOSTICS` instead of producing guessed records.

### Modes

`mode` selects one of five workflows:

- `search` — search one query and enrich up to `maxItems` postings.
- `searchMultiple` — run up to five queries, taking up to `maxItems` unique postings per query.
- `single` — fetch one official `/Recruit/GI_Read/<id>` detail URL.
- `multiple` — fetch a bounded list of official detail URLs.
- `startUrls` — accept official search and/or detail URLs.

Search pagination is bounded by `maxPages` (1–3). Detail concurrency is bounded by `concurrency` (1–5). `maxItems` is limited to 10 per query/direct-input mode, and each request has a 5–45 second timeout.

### Input

The complete input definition is [`.actor/input_schema.json`](.actor/input_schema.json). A normal search input is:

```json
{
  "mode": "search",
  "query": "developer",
  "location": "Seoul",
  "maxItems": 3,
  "maxPages": 2,
  "concurrency": 4,
  "requestTimeoutSecs": 25
}
```

For `multiple` and `startUrls`, URL entries are objects:

```json
{
  "mode": "multiple",
  "urls": [
    { "url": "https://www.jobkorea.co.kr/Recruit/GI_Read/49844042" },
    { "url": "https://www.jobkorea.co.kr/Recruit/GI_Read/49856065" }
  ],
  "maxItems": 2
}
```

Only HTTPS URLs on the official JobKorea host are accepted. Detail identifiers, canonical URLs, titles, employers, and search-query matches are checked before a record is emitted.

### Dataset

The strict output definition is [`.actor/dataset_schema.json`](.actor/dataset_schema.json). Records include, when publicly present:

- stable posting ID, title, employer, company URL/logo, canonical detail URL, and source page;
- advertised location, structured address, remote indicator, employment type, experience, education, schedule, salary, categories, skills, qualifications, preferred qualifications, and benefits;
- normalized posting/start/deadline dates and original deadline text;
- JSON-LD summary, cleaned public detail-section HTML, a readable source-backed description, company information, map URL, and explicit application URL provenance;
- source receipts, identity/canonical verification, search rank/query, HTTP status, and quality counts.

JobKorea currently exposes the concise JSON-LD summary and selected server-rendered public sections to this local client. Some full advertisement-body content is client-only or access-controlled, so `fullDescriptionAvailable` is explicitly `false` when that body is not publicly available; the Actor never fabricates it or invents an application URL. Optional fields are omitted rather than emitted as null, blank, placeholder, or empty values.

Run state is kept outside the dataset in the default key-value store:

- `RUN_SUMMARY` — mode, counts, timing, requests, environment, and cloud build/run provenance when available;
- `RUN_DIAGNOSTICS` — target, transport, or parsing errors;
- `RUN_SKIPS` — candidates rejected after verification (for example, a query mismatch);
- `RUN_HEALTH` — dataset quality receipts.

### Local development

From this directory:

```powershell
npm install
npm test
npm run lint
npx --yes apify-cli validate-schema .actor/input_schema.json
npx --yes apify-cli run --purge --input-file INPUT.json
npm run validate
```

The local run writes to `storage/`. `npm run validate` checks required fields, official URL/ID consistency, duplicate IDs/URLs, recursive null/blank/placeholder/empty values, application provenance, and key-value-store consistency.

### Cloud deployment

```powershell
apify push
apify call jobsapi/jobkorea-jobs-search-scraper --memory 512 --timeout 300 --input-file INPUT.json
```

Keep cloud validation bounded with the provided input. The default dataset contains only persisted complete jobs; diagnostics remain in the default key-value store.

### Responsible use

Collect only public job information, respect JobKorea’s terms and robots directives, and use modest bounded request limits. Do not use this Actor for authenticated pages, paywall or challenge bypassing, or collection of candidate or recruiter personal data.

# Actor input Schema

## `mode` (type: `string`):

search, searchMultiple, single, multiple, or startUrls.

## `query` (type: `string`):

Keyword or job title for search mode.

## `location` (type: `string`):

JobKorea location filter used in search modes.

## `queries` (type: `array`):

Queries for searchMultiple mode; each query is bounded to maxItems records.

## `url` (type: `string`):

Official public https://www.jobkorea.co.kr/Recruit/GI\_Read/<id> URL for single mode.

## `urls` (type: `array`):

Official public JobKorea posting URLs for multiple mode.

## `startUrls` (type: `array`):

Official JobKorea search or posting URLs for startUrls mode.

## `maxItems` (type: `integer`):

Maximum records per search query or direct-input mode.

## `maxPages` (type: `integer`):

Bounded public result pages per query.

## `concurrency` (type: `integer`):

Concurrent official detail requests.

## `requestTimeoutSecs` (type: `integer`):

Maximum time for one official request.

## Actor input object example

```json
{
  "mode": "search",
  "query": "developer",
  "location": "Seoul",
  "maxItems": 3,
  "maxPages": 2,
  "concurrency": 4,
  "requestTimeoutSecs": 25
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "developer",
    "location": "Seoul"
};

// Run the Actor and wait for it to finish
const run = await client.actor("jobsapi/jobkorea-jobs-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "developer",
    "location": "Seoul",
}

# Run the Actor and wait for it to finish
run = client.actor("jobsapi/jobkorea-jobs-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "developer",
  "location": "Seoul"
}' |
apify call jobsapi/jobkorea-jobs-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jobsapi/jobkorea-jobs-search-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vzyyYbzzRqhA5D71o/builds/JfghPMwPRGiuhMZWR/openapi.json
