# 104.com.tw Taiwan Jobs Scraper (`corvuslab/taiwan104-jobs-scraper`) Actor

Scrape jobs from Taiwan's 104.com.tw — salary, location, company and HR contact details. Provisional listing; final SEO copy set post-verification.

- **URL**: https://apify.com/corvuslab/taiwan104-jobs-scraper.md
- **Developed by:** [Corvuslab](https://apify.com/corvuslab) (community)
- **Categories:** Jobs, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.25 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 104.com.tw Jobs Scraper — Salary, Contacts & Job Monitoring for Taiwan

### What does the 104.com.tw Jobs Scraper do?

> **Turn every live 104.com.tw job listing into a clean, structured record — salary (TWD), location with GPS, company profile, HR contact email/phone, applicant demand signals and full description — in seconds, at scale.**

Scrape job listings from **104.com.tw**, Taiwan's largest job board, with the richest per-job detail available: **salary ranges** (min/max in TWD, period), **HR contact email and phone**, **applicant count and recruiter response score**, the full **company & industry profile** (size, website, logo), **welfare/benefits**, **required skills, languages and certificates**, and the complete **job description** as text, HTML and Markdown. Filter by keyword, region, remote work, experience level, posted-within date and sort order; schedule it in **incremental monitoring** mode to capture only new and changed postings; and get pinged on **Telegram, Slack, Discord or any webhook**. No code, no login, no API key — export to JSON, CSV, Excel or the API.

**Why this scraper**

- ⚡ **Fast & low-cost** — browser-backed API extraction, no proxy needed. Clears Cloudflare automatically.
- 🧾 **Rich, typed records** — 60+ structured fields per job (salary min/max, HR contact, applicant demand, welfare tags, skills, GPS coordinates), not raw HTML.
- ♻️ **Cheap to monitor** — incremental mode re-scrapes only what changed (see below).
- 🔔 **Notifications built in** — Telegram, Slack, Discord or any webhook (n8n / Make / Zapier).
- 🤖 **AI- & API-ready** — compact output, MCP-friendly, one-click integrations.

***

### ✨ Key features

- 🔎 **Search or URL scraping** — run a keyword search (Chinese or English) with filters, or paste 104.com.tw search and job URLs directly.
- 💰 **Structured salary** — parsed `salaryMin` / `salaryMax` in TWD plus `salaryCurrency`, `salaryPeriod` and `salaryNegotiable` — sort and filter on real numbers.
- 📇 **HR contact details** — `applyEmail`, `phoneNumber`, `contactName` plus `extractedEmails[]` and `extractedPhones[]` pulled from the posting — the lead-gen edge competitors bury.
- 📊 **Applicant demand signals** — `applicantCount` and `hrResponseScore` on every listing — spot low-competition, fast-replying employers.
- 🏢 **Company intelligence** — company name, industry, employee count, website, logo and verified-advertiser flag on every record.
- 📍 **Location with GPS** — `location`, full `address`, `latitude` / `longitude` and `nearestMRT` station.
- 🎚️ **Rich filters** — keyword, 20 Taiwan regions, remote-only, posted within (3–30 days), minimum experience (1–10+ years) and sort by relevance / date / salary.
- 🧾 **Full descriptions, three ways** — the complete job description as plain **text**, **HTML** and **Markdown**, with optional truncation.
- 🎁 **Welfare & benefits** — `welfareTags[]` and `welfareText` with the full benefits breakdown.
- 📋 **Requirements** — `experienceRequired`, `educationRequired`, `languages[]`, `skills[]`, `specialties[]`, `certificates[]` and `acceptRoles[]`.
- ♻️ **Incremental monitoring** — schedule it and get only what changed, tagged **NEW / UPDATED / UNCHANGED / EXPIRED**, with repost detection; unchanged items are skipped before their detail page is even fetched.
- 🔔 **Notifications** — Telegram, Slack, Discord or any webhook (n8n / Make / Zapier), fanned out on every run.
- 🤖 **AI-ready** — compact mode and drop-empty output keep payloads small for LLMs and MCP clients.

***

### 📤 Example output

```json
{
  "id": "77wla",
  "jobNo": "12126142",
  "title": "Principal Scientist – Generative AI & NLP",
  "url": "https://www.104.com.tw/job/77wla",
  "source": "104.com.tw",
  "searchKeyword": "python",
  "scrapedAt": "2026-08-17T20:08:13.158479+00:00",
  "detailFetched": true,
  "company": "易捷系統股份有限公司",
  "companyNo": "130000000072555",
  "companyUrl": "https://www.104.com.tw/company/1a2x6bjjaj",
  "companyLogo": "https://static.104.com.tw/b_profile/cust_picture/2555/130000000072555/logo.png",
  "companyIndustry": "人力派遣服務",
  "employeeCount": 31000,
  "companySize": "31000人",
  "appearDate": "2026-08-11",
  "closeDate": "2026-09-13",
  "applicantCount": 15,
  "hrResponseScore": 0.935,
  "salaryMin": 4000000,
  "salaryMax": null,
  "salaryText": "年薪4,000,000元以上",
  "salaryCurrency": "TWD",
  "salaryPeriod": "yearly",
  "salaryNegotiable": false,
  "location": "印度",
  "address": "印度矽谷 - Bengaluru",
  "latitude": 12.9715987,
  "longitude": 77.5945627,
  "remoteWork": false,
  "employmentType": "full-time",
  "headcount": 1,
  "experienceRequired": "10年以上",
  "educationRequired": "碩士以上",
  "languages": ["英文"],
  "jobCategories": ["資料科學家", "資料工程師"],
  "welfareTags": ["退職金提撥", "電信費補助", "週休二日", "勞保", "健保", "特別休假"],
  "contactName": "HR",
  "applyEmail": "kevin.pi@nityo.com",
  "phoneNumber": "02-1234-5678",
  "extractedEmails": ["kevin.pi@nityo.com", "janet.han@nityo.com", "charlyn.lu@nityo.com"],
  "extractedPhones": ["02-1234-5678"]
}
```

The full record also carries `descriptionHtml`, `descriptionMarkdown`, `welfareText`, `skills[]`, `specialties[]`, `certificates[]`, `acceptRoles[]`, `socialProfiles`, and (in incremental mode) `changeType`, `isRepost` and `repostOfId`.

***

### 📥 Example input

A few ready-to-run configurations — set these in the visual editor or pass them as JSON via the API:

Broad keyword search across all of 104.com.tw:

```json
{ "query": "python", "maxResults": 100 }
```

Filtered — Taipei remote AI/data roles posted in the last 7 days, sorted by salary:

```json
{
  "query": "資料分析師",
  "area": ["6001001000"],
  "remoteWork": true,
  "datePosted": "7",
  "sortBy": "salary",
  "includeDetails": true,
  "maxResults": 200
}
```

Lead-gen — only jobs with HR email, in Hsinchu:

```json
{
  "query": "engineer",
  "area": ["6001006000"],
  "requireContact": "email",
  "maxResults": 50
}
```

Incremental monitoring — poll new postings on a schedule and push to Telegram:

```json
{
  "query": "軟體工程師",
  "incrementalMode": true,
  "telegramToken": "123:ABC",
  "telegramChatId": "@myjobsfeed"
}
```

***

### 📚 What data can you extract from 104.com.tw?

- **Core** — `id`, `jobNo`, `title`, `url`, `source`, `searchKeyword`, `scrapedAt`, `detailFetched`.
- **Salary** — `salaryMin`, `salaryMax`, `salaryText`, `salaryCurrency`, `salaryPeriod`, `salaryNegotiable` (parsed to real numbers in TWD).
- **Demand signals** — `applicantCount`, `applicantVolumeLabel`, `hrResponseScore` — how many people applied and how fast the recruiter replies.
- **Company** — `company`, `companyNo`, `companyUrl`, `companyLogo`, `companyIndustry`, `companySize`, `employeeCount`, `companyWebsite`, `companyDescription`, `isVerifiedAdvertiser`.
- **Location** — `location`, `address`, `latitude`, `longitude`, `nearestMRT`.
- **Employment** — `employmentType`, `remoteWork`, `headcount`, `manageResponsibility`, `businessTrip`, `vacationPolicy`, `startWorkingDay`, `workHours`, `jobCategories[]`, `majors[]`.
- **Requirements** — `experienceRequired`, `educationRequired`, `requiredMajors[]`, `languages[]`, `skills[]`, `specialties[]`, `certificates[]`, `otherRequirements`, `acceptRoles[]`.
- **Descriptions** (with `includeDetails`) — `description` (text), `descriptionHtml`, `descriptionMarkdown`.
- **Welfare & benefits** — `welfareTags[]`, `welfareText`.
- **Contacts & lead-gen** — `contactName`, `applyEmail`, `phoneNumber`, `applicationNote`, `extractedEmails[]`, `extractedPhones[]`, `extractedUrls[]`, `socialProfiles`.
- **Dates** — `appearDate`, `closeDate`.
- **Monitoring** — `changeType` (NEW / UPDATED / UNCHANGED / EXPIRED), `isRepost`, `repostOfId`, `repostDetectedAt`, `contentHash`.

Every field is present in standard mode (missing values are `null`); **compact mode** returns the core fields only, for lean AI/MCP payloads.

***

### ⚙️ Input

Configure it in the visual editor — no code needed — or pass JSON via the API.

| Field | What it does |
|---|---|
| `query` | Keyword search (Chinese or English), e.g. `python`, `資料分析師`. Comma-separate for multiple searches. |
| `startUrls` | Scrape specific 104.com.tw search or job-detail URLs directly. |
| `area` | Restrict to one or more of 20 Taiwan regions (Taipei, New Taipei, Taoyuan, Taichung, Tainan, Kaohsiung, Hsinchu, …). |
| `remoteWork` | Limit results to remote-friendly positions. |
| `datePosted` | Only jobs posted within 3 / 7 / 14 / 30 days. |
| `minExperience` | Keep only jobs requiring 1+ / 3+ / 5+ / 10+ years of experience. |
| `sortBy` | Sort by relevance, newest first or highest salary first. |
| `requireContact` | Keep only jobs that carry HR contact details (email / phone / either / both). |
| `includeApplicantInsights` | Add applicant count and recruiter response score to every record. |
| `includeDetails` | Fetch each job's page for the full description, requirements, welfare and HR contact email/phone. |
| `descriptionFormat` | Output the description as text, HTML, Markdown or all three. |
| `compact` | Return core fields only — ideal for AI agents and MCP. |
| `incrementalMode` | Emit only what changed since the last run (NEW / UPDATED / EXPIRED). |
| `maxResults` | Cap the number of records (0 = unlimited). |
| `proxyConfiguration` | Optional — works without a proxy at typical volumes. |

…plus notification channels (Telegram, Slack, Discord, webhook), incremental tuning (`stateKey`, `emitUnchanged`, `emitExpired`), description truncation, drop-empty output and retry settings — all in the visual editor.

***

### 💡 What can you do with 104.com.tw data?

- **Taiwan job-market intelligence** — track salary trends, demand by region and industry, and competitive hiring across Taiwan's largest job board.
- **Recruiting & sourcing** — surface HR email and phone for direct outreach; filter to only jobs with contacts using `requireContact`.
- **Compensation benchmarking** — analyze salary ranges by role, experience level and region with real TWD numbers.
- **New-job monitoring** — schedule incremental mode + notifications for a live change feed of fresh postings, paying only for the delta.
- **Company intelligence** — build a picture of who's hiring, their size, industry and welfare packages.
- **AI agents & pipelines** — compact output plugs straight into LLM / MCP workflows for automated analysis.
- **Job-board & aggregator feeds** — populate your own Taiwan jobs site with clean, structured, salary-rich listings.

***

### ♻️ Incremental monitoring — pay for changes, not repeats

Schedule the actor and turn on **incremental mode**: each run compares against the last and emits only **NEW / UPDATED / EXPIRED** records — unchanged items are skipped *before* their detail page is fetched, so a daily watch costs a fraction of a full re-scrape.

| Daily churn | of 1,000 tracked | billable records | you save |
|---|---|---|---|
| 5 % | 1,000 | 50 | **95 %** |
| 15 % | 1,000 | 150 | **85 %** |
| 30 % | 1,000 | 300 | **70 %** |

The first run seeds the baseline and bills in full; every run after that bills only the delta.

***

### 🚀 How to scrape 104.com.tw

1. Open the actor and enter a **search keyword** (e.g. `python` or `工程師`) and/or pick filters — or paste a 104.com.tw URL.
2. Set **Max results** and choose whether to **fetch full details** (descriptions, welfare, HR contact).
3. (Optional) Turn on **incremental mode** and a **notification** channel, then **Schedule** it.
4. Click **Start**.
5. Download the data as **JSON, CSV or Excel**, or pull it from the **API**.

New to Apify? Create a free account — it comes with monthly credit, no credit card required.

***

### 🔌 Integrations & export

Export to **JSON, CSV, Excel** or an HTML table, or pull from the **REST API** and the **JavaScript / Python** clients. Runs on a **schedule**, connects to **Google Sheets, Slack, Make, Zapier and n8n**, and works as an **MCP tool** for AI agents — compact mode keeps token usage small.

***

### ❓ FAQ

**Do I need a proxy or login?** No — it runs out of the box with no login or API key; Apify Proxy is available under Advanced for high-volume runs.

**Can I get only new jobs on a schedule?** Yes — turn on incremental mode and schedule it; each run emits only what changed (NEW / UPDATED / EXPIRED) and can notify your channel.

**Does it include salary data?** Yes — where a posting lists pay, you get the raw `salaryText` plus parsed `salaryMin`, `salaryMax` in TWD, `salaryCurrency` and `salaryPeriod`.

**Can I get HR contact details?** Yes — enable full details for the recruiter's `applyEmail`, `phoneNumber` and `contactName`, plus all `extractedEmails[]` and `extractedPhones[]` found in the posting.

**What about applicant competition?** Every listing includes `applicantCount` and `hrResponseScore` — spot high-demand roles and fast-replying recruiters.

**Can I get the full job description?** Yes — enable full details for the complete description as plain text, HTML and Markdown, plus welfare benefits and requirements.

**What formats can I export?** JSON, CSV, Excel, HTML table, or via the API.

**Is it good for AI agents?** Yes — enable compact mode; the output is MCP-friendly and payloads stay small.

**How many records can I get?** Set `maxResults` (0 = unlimited, bounded by how many the search returns).

***

### ⚖️ Is it legal to scrape 104.com.tw?

This actor accesses only publicly available data on 104.com.tw. You are responsible for how you use the extracted data — in particular any personal information — and for complying with the site's terms and applicable law (including the GDPR and Taiwan's PDPA where they apply). Not affiliated with, endorsed by, or sponsored by 104 Corporation.

***

**Keywords:** 104 scraper · 104.com.tw scraper · taiwan jobs scraper · taiwan job board scraper · 104 jobs api · taiwan salary data · taiwan recruitment · HR contact scraper · taiwan job listings · job monitoring taiwan · new job alerts · recruiting data · job board api · export jobs to CSV · MCP tool for AI agents · 台灣工作 · 104人力銀行

# Actor input Schema

## `query` (type: `string`):

Keywords to search for (Chinese or English), e.g. 「工程師」 or "python". Separate multiple searches with commas — each runs as its own search and results are merged and de-duplicated. Leave empty to scrape the broadest listing.

## `startUrls` (type: `array`):

Paste 104.com.tw search URLs (…/jobs/search/…) or job-detail URLs (…/job/<id>) to scrape directly.

## `maxResults` (type: `integer`):

Maximum number of jobs to return across all searches. Set 0 for unlimited (bounded by how many the search has).

## `ignoreUrlFailures` (type: `boolean`):

Skip URLs that cannot be interpreted instead of failing the whole run.

## `area` (type: `array`):

Limit results to one or more Taiwan regions. Leave empty for nationwide.

## `remoteWork` (type: `boolean`):

Limit results to remote-friendly positions.

## `datePosted` (type: `string`):

Keep only jobs published within this many days.

## `minExperience` (type: `string`):

Keep only jobs requiring at least this many years of experience.

## `sortBy` (type: `string`):

Order in which 104 returns results.

## `requireContact` (type: `string`):

Keep only jobs that include an HR contact. off = keep everything; email / phone = require that channel; either = at least one; both = email and phone. Screened at emit time, so it also trims what you're billed for. Requires detail fetching.

## `includeApplicantInsights` (type: `boolean`):

Add how many people have applied and the recruiter's response score — handy for spotting low-competition, fast-replying employers.

## `includeDetails` (type: `boolean`):

Fetch each job's detail page for the full description, requirements, welfare benefits, close date and HR contact email/phone. Turn off for the fastest, cheapest runs.

## `descriptionFormat` (type: `string`):

Which representation(s) of the job description to include.

## `descriptionMaxLength` (type: `integer`):

Truncate each description to this many characters. Leave empty for full text.

## `compact` (type: `boolean`):

Emit only the core fields (title, company, salary, location, applicant count). Ideal for AI agents and MCP clients.

## `excludeEmptyFields` (type: `boolean`):

Remove null, empty-string and empty-array fields from each record.

## `incrementalMode` (type: `boolean`):

Track state between runs and tag every job with a changeType (NEW / UPDATED / UNCHANGED / EXPIRED).

## `stateKey` (type: `string`):

Stable name for the tracked search. Leave empty to derive one automatically from your search settings.

## `emitUnchanged` (type: `boolean`):

Also emit jobs that have not changed since the previous run.

## `emitExpired` (type: `boolean`):

Emit jobs that were present last run but are gone now.

## `telegramToken` (type: `string`):

Bot token from @BotFather.

## `telegramChatId` (type: `string`):

Chat or channel ID, e.g. "-100123456789" or "@yourchannel".

## `slackWebhookUrl` (type: `string`):

Slack incoming-webhook URL.

## `discordWebhookUrl` (type: `string`):

Discord incoming-webhook URL.

## `webhookUrl` (type: `string`):

Any HTTPS endpoint. Receives a JSON POST with the matched jobs — works with n8n, Make and Zapier.

## `webhookHeaders` (type: `object`):

Extra headers for the webhook request, e.g. {"Authorization": "Bearer xyz"}.

## `notificationLimit` (type: `integer`):

How many jobs to include in each notification message.

## `proxyConfiguration` (type: `object`):

Optional. 104.com.tw is scraped over direct HTTP with no proxy — leave this off for the cheapest runs. Enable a proxy only if you hit rate limits at very high volume.

## `maxRequestRetries` (type: `integer`):

How many times to retry a failed request before giving up on it.

## Actor input object example

```json
{
  "query": "python, 資料分析師",
  "maxResults": 25,
  "ignoreUrlFailures": true,
  "remoteWork": false,
  "datePosted": "0",
  "minExperience": "0",
  "sortBy": "relevance",
  "requireContact": "off",
  "includeApplicantInsights": true,
  "includeDetails": true,
  "descriptionFormat": "all",
  "compact": false,
  "excludeEmptyFields": false,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "notificationLimit": 5,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "maxRequestRetries": 3
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `allItems` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "python"
};

// Run the Actor and wait for it to finish
const run = await client.actor("corvuslab/taiwan104-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "python" }

# Run the Actor and wait for it to finish
run = client.actor("corvuslab/taiwan104-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "python"
}' |
apify call corvuslab/taiwan104-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,corvuslab/taiwan104-jobs-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/amsYJGGGdCpD8cJag/builds/JvClNT2dpQJEKTi6u/openapi.json
