# Daum News (다음뉴스) Scraper (`ardent_fork/daum-news`) Actor

Korean news articles from Daum News sections: title, publisher, reporter, publish time, full body text, images and Daum's auto-summary. No login, no API key.

- **URL**: https://apify.com/ardent\_fork/daum-news.md
- **Developed by:** [KF P](https://apify.com/ardent_fork) (community)
- **Categories:** News
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 articles

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Daum News (다음뉴스) Scraper

Collects Korean news articles from [news.daum.net](https://news.daum.net), Korea's second-largest news portal, which aggregates every major Korean outlet (연합뉴스, 뉴스1, 조선일보, 한겨레, KBS, …). Everything scraped is public: no login, no API key, no proxy needed. Pay-per-event: **one `article` event per article saved**.

### What you get

One record per article:

| field | example |
|---|---|
| `id` | `20260901122450927` |
| `url` | `https://v.daum.net/v/20260901122450927` |
| `title` | 전기차 보급 '속도전'·충전기는 '고도화'… |
| `section` | `economy` — the section page it was found on; an article listed in several requested sections is saved once, under whichever was crawled first |
| `publisher`, `publisherUrl` | `뉴스1`, `https://v.daum.net/channel/396/home` |
| `reporter` | `황덕현 기후환경전문기자` |
| `publishedAt` | ISO 8601 in KST (`2026-09-01T12:24:50+09:00`) |
| `modifiedAt` | usually `null`; set only when Daum prints a second timestamp |
| `description` | Daum's meta description (first ~200 chars) |
| `body` | full article text: paragraphs and sub-headings, one per line (`\n`) |
| `autoSummary` | Daum's own 3–4 sentence auto summary, when present |
| `image`, `images`, `captions`, `thumbnail` | og:image, in-body images, their captions, list thumbnail |

With `fetchArticles: false` you get only `id`, `url`, `title`, `thumbnail`, `publisher`, `section` — one request per section, seconds instead of minutes.

### Input

Example: economy and tech, full articles —

```json
{ "sections": ["economy", "tech"], "maxItems": 200 }
```

- **sections** — slugs after `news.daum.net/`. Main: `home`, `politics`, `society`, `economy`, `world`, `culture`, `tech`, `life`, `people`, `climate`. Economy sub-sections: `policy`, `industry`, `stock`, `finance`, `estate`, `coin`, `pension`, `employ`, `consumer`, `startup`, `worldeconomy`, `autos`. Any other slug that exists on the site works too.
- **maxItems** — total cap across sections (default 100). Each section page currently lists roughly 20–80 articles.
- **fetchArticles** — open each article (default `true`).
- **maxConcurrency**, **proxyConfiguration** — optional.

### Pricing

Pay-per-event: one `article` event is charged per article pushed to the dataset (in list-only mode, `fetchArticles: false`, one per list entry). Nothing else is charged. If a run's budget (`maxTotalChargeUsd`) is exhausted before `maxItems`, the actor stops cleanly with the articles it could pay for and says so in the status message.

### Notes

- Section pages are curated front pages, not archives: each shows the current 20–80 articles. Run the actor on a schedule to build a continuous feed; articles are de-duplicated within a run by id (and across resumes of the same run).
- Times are Korea Standard Time (UTC+9) as printed by Daum.
- Articles removed between listing and fetch (Daum answers HTTP 500) are skipped with a warning. A section that fails is named in the run's status message; the run only fails if no article at all could be saved.

### Use cases

Media monitoring, Korean-language NLP corpora, brand/keyword tracking across all Korean outlets from a single source, comparing outlet coverage of an event.

# Actor input Schema

## `sections` (type: `array`):

Daum News section slugs to scrape (the path after news.daum.net/). Main sections: home, politics, society, economy, world, culture, tech, life, people, climate. Economy sub-sections: policy, industry, stock, finance, estate, coin, pension, employ, consumer, startup, worldeconomy, autos.

## `maxItems` (type: `integer`):

Stop after this many articles have been saved (across all sections). Each section page lists roughly 20–80 articles.

## `fetchArticles` (type: `boolean`):

Open every article page for the body text, reporter, publisher and exact publish time. Turn off to get only the list data (id, url, title, thumbnail) — much faster.

## `maxConcurrency` (type: `integer`):

Parallel page fetches.

## `proxyConfiguration` (type: `object`):

Optional. Daum serves these pages without a proxy; use one only if your IP is blocked.

## Actor input object example

```json
{
  "sections": [
    "economy",
    "tech"
  ],
  "maxItems": 200,
  "fetchArticles": true,
  "maxConcurrency": 5
}
```

# Actor output Schema

## `articles` (type: `string`):

Dataset items, one per Daum News article (title, publisher, reporter, publishedAt, body, images, summary).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sections": [
        "home"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ardent_fork/daum-news").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "sections": ["home"] }

# Run the Actor and wait for it to finish
run = client.actor("ardent_fork/daum-news").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sections": [
    "home"
  ]
}' |
apify call ardent_fork/daum-news --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ardent_fork/daum-news"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Nb4IPtZTzc2RtuH7S/builds/pkvTtbXGhE3zpWBwa/openapi.json
