# Dzen.ru Scraper - Articles, Videos, Channels & Comments (`abotapi/dzen-ru-scraper`) Actor

Scrape Dzen.ru (ex Yandex.Zen): articles, videos and channels. Full article text with likes and comment counts, channel subscriber counts with their on-page publications, video metadata. Paste any Dzen link; incremental mode reports only new and changed content.

- **URL**: https://apify.com/abotapi/dzen-ru-scraper.md
- **Developed by:** [Abot API](https://apify.com/abotapi) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.80 / 1,000 content records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Dzen.ru Scraper: Articles, Videos & Channels

Dzen.ru Scraper turns Dzen (ex Yandex.Zen), Russia's largest content platform, into a clean JSON API. Walk the discovery feed page by page, or paste any Dzen link - an article, a video, or a channel - and get full metadata: article text with like and comment counts, channel subscriber counts with all their articles, video views and dates. Export to JSON, CSV or Excel, or pull results straight into your app through the API.

### Why This Scraper?

- **Two lanes in one actor.** Feed mode walks Dzen's discovery feed page by page; URL mode deep-dives pasted article, video and channel links.
- **Channel monitoring.** Paste channel links and get the channel record (subscriber count included) plus its articles, newest first, as far back as Max records allows - incremental mode reports only the new ones on scheduled runs.
- **Full article records.** Title, description, body text, images, author channel, like and comment counts.
- **Comments included.** Turn on Include comments to get each article's and video's comments (text, author, date, likes, dislikes, emoji reactions and reply count) inside its record, plus a separate Comments output with one row per comment. Sort by popularity, newest or oldest.
- **Built for schedules.** Incremental mode returns only new and changed content on recurring runs, and a refused run fails loudly instead of returning an empty dataset.
- **Reliable RU connection.** Runs over a residential RU connection by default, the setup Dzen was tested on.

### Use Cases

- **Media analytics:** track subscriber counts, like counts and comment counts across Dzen channels on a schedule.
- **Content monitoring:** walk the discovery feed daily and get only the new publications.
- **Journalism and research:** pull full article text and metadata for RU-media datasets.
- **Competitive tracking:** monitor specific channels' publication cadence and engagement.

### Data You Get

> Sample shape: values are illustrative placeholders, not from a live record.

| Field | Example |
| --- | --- |
| `id` | `"article:arjYU4pA7HGwblbw"` (prefixed with the entity type) |
| `dzenType` | `"article"` (also `"video"`, `"channel"`, `"native"` for feed rows) |
| `title` | `"Article headline"` |
| `url` | `https://dzen.ru/a/...` canonical link |
| `publicationDate` | `"2026-09-29"` or an ISO timestamp |
| `textPreview` | `"Short description from the page"` |
| `textContent` | full article body text (pasted article links, or channel articles with Full article text on); channel articles otherwise carry the lead paragraph in `textPreview` |
| `images` | `["https://avatars.dzeninfra.ru/..."]` |
| `authorName` / `authorUrl` | channel name and link |
| `subscribers` | `869053` (channels) |
| `likes` / `commentsCount` | `124` / `44` where the page exposes them |
| `views` | `15234` (feed rows and videos) |
| `isPremium` | `false` |
| `timeToReadSeconds` | `42` (feed rows) |
| `collectionContext` | `"Channel Name"` (publications walked from a channel) |
| `comments` | `[{"commentText": "...", "commentAuthor": "...", "commentDate": "2026-09-29T14:42:40Z", "commentLikes": 3, "commentReplies": 1, ...}]` (with Include comments on) |
| `detailFetched` | `true` when the record's own page was read |
| `changeType` / `changedFields` / `firstSeenAt` / `lastSeenAt` | incremental mode only |

### How to Use

1. Pick a **mode**: `feed` (walk the discovery feed), `url` (paste links) or `search` (search by keyword).
2. In feed mode set **Max feed pages**; in URL mode paste article, video or channel links (for channels, choose the order and whether you want articles, videos or shorts); in search mode enter queries and a result type.
3. Set **Max records** to control run size and cost, then click **Start**.
4. Download the dataset as JSON, CSV or Excel, or read it through the API.

**Search Dzen for videos:**

```json
{
  "mode": "search",
  "searchQueries": ["грибы", "рецепты выпечки"],
  "searchType": "videos",
  "maxItems": 50
}
```

**A channel's most popular videos:**

```json
{
  "mode": "url",
  "urls": [{ "url": "https://dzen.ru/tass" }],
  "channelSort": "popular",
  "channelContent": "videos",
  "maxItems": 30
}
```

**Walk the discovery feed:**

```json
{
  "mode": "feed",
  "maxPages": 5,
  "maxItems": 100
}
```

**Deep-read specific articles:**

```json
{
  "mode": "url",
  "urls": [
    "https://dzen.ru/a/arjYU4pA7HGwblbw",
    "https://dzen.ru/a/arnnyamrzEdmCVFV"
  ]
}
```

#### Run it from your code

Python:

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("abotapi/dzen-ru-scraper").call(run_input={
    "mode": "feed",
    "maxItems": 100,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item["dzenType"], item["title"])
```

JavaScript:

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('abotapi/dzen-ru-scraper').call({
    mode: 'url',
    urls: ['https://dzen.ru/tass'],
    maxItems: 50,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
```

Or connect it to Make, Zapier, n8n, Google Sheets or webhooks from the **Integrations** tab.

#### Resume and recurring updates

- **Resume** (`resumeFromRunId`) continues one interrupted run: paste its run or dataset ID and the actor skips everything already collected there, so you don't pay twice.
- **Incremental mode** (`incrementalMode`) is for scheduled runs over the same links. Each record is classified `NEW`, `UPDATED` (with `changedFields`), `UNCHANGED` (suppressed and not billed unless `emitUnchanged` is on), `REAPPEARED` or `EXPIRED` (only after a run that scanned every link, and only with `emitExpired`). `stateKey` names or shares the stored state. With incremental mode off, output is exactly as before.
- **What counts as a change.** `views` never makes a record `UPDATED`, because view counts rise on almost every run; the field is still returned. Likes, comment counts and subscriber counts do count, so a record is returned (and billed) again when they move. That is how a like-count jump or subscriber growth is spotted. To hear only about new content, add those fields to `ignoreFieldsForChanges`. Links in the output are always clean canonical dzen.ru URLs, so they never cause a false change.

### Input Parameters

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `mode` | string | `feed` | `feed` (walk the discovery feed), `url` (paste links) or `search` (search by keyword). |
| `searchQueries` | array | (none) | Search mode: keywords or phrases, one per row. |
| `searchType` | string | `all` | Search mode: `all`, `articles`, `videos` or `channels`. |
| `channelSort` | string | `newest` | Channel links: `newest`, `popular` or `oldest` first. |
| `channelContent` | string | `articles` | Channel links: the channel's `articles`, `videos` or `shorts`. |
| `fetchArticleText` | boolean | `false` | Channel links: read each article in full (complete text and images), charged as a full page read. Off returns the listing data (title, lead paragraph, counts, date, cover). |
| `urls` | array | sample | Dzen links (url mode): articles, videos, channels. Mixed sets are fine. |
| `maxItems` | integer | `20` | Stop after this many records (`0` = no limit). |
| `maxPages` | integer | `0` | Max result pages per feed or search query (`0` = no limit). |
| `includeComments` | boolean | `false` | Also collect each article's and video's comments (nested in the record, and one row per comment in the Comments output). |
| `maxCommentsPerPost` | integer | `20` | Max comments per post. Does not count toward `maxItems`. |
| `commentsSort` | string | `top` | `top` (most popular, at most 20 per post), `newest` or `oldest`. |
| `resumeFromRunId` | string | (none) | Continue one interrupted run. |
| `incrementalMode` | boolean | `false` | Return only new and changed records on scheduled runs. |
| `stateKey` | string | (none) | Name or share an incremental-mode state. |
| `emitUnchanged` | boolean | `false` | Also return (and bill) unchanged records. |
| `emitExpired` | boolean | `false` | Also return (and bill) expired records. |
| `ignoreFieldsForChanges` | array | (none) | Extra output fields that should not make a record `UPDATED` (`views` is always ignored). |
| `proxy` | object | residential RU | Connection settings. |
| `mcpConnectors` | array | (none) | Optional: send a summary of each record to apps you authorized under Integrations. |
| `notionParentPageUrl` | string | (none) | Notion connector only: page under which records are created. |
| `maxNotifyListings` | integer | `50` | Cap on records written to each connector per run. |

### Output Example

> Sample shape: values are illustrative placeholders, not from a live record.

```json
{
  "id": "article:arjYU4pA7HGwblbw",
  "dzenType": "article",
  "title": "Article headline",
  "url": "https://dzen.ru/a/arjYU4pA7HGwblbw",
  "publicationDate": "2026-09-29T10:00:00Z",
  "textPreview": "Short description from the page",
  "textContent": "Full article body text...",
  "images": ["https://avatars.dzeninfra.ru/get-zen-pub/..."],
  "authorName": "Channel Name",
  "subscribers": null,
  "likes": 124,
  "commentsCount": 44,
  "comments": [],
  "views": null,
  "collectionContext": null,
  "detailFetched": true,
  "changeType": null,
  "changedFields": [],
  "firstSeenAt": null,
  "lastSeenAt": null
}
```

### Plan Requirement

The default connection uses a residential RU proxy; a residential plan gives the most reliable results. Expect 15-30 seconds before the first records while the connection is set up.

### FAQ

#### How much does it cost?

You pay per record returned: articles, videos and channels are charged as content records, and each collected comment (only with Include comments on) is charged as a comment, at a lower rate. Records whose full page is read, which means every pasted article or video link and channel articles with **Full article text** on, are also charged a full page read. Channel articles without it, channel videos and shorts, search results and feed rows never are. The **Pricing** tab shows the current rates. Use **Max records** and **Max comments per post** to cap the cost of any run; a run also stops cleanly at the spending limit you set.

#### How do comments work?

Comments sit inside their article or video record, in the `comments` list, so each post stays one record. The **Comments** output (next to the default output on the run page, or `?view=comments` through the API) shows one row per comment next to the post's title and link, ready for CSV or Excel. Only top-level comments are returned; `commentReplies` tells you how many replies each has. With **Most popular first**, Dzen shows at most 20 comments per post; choose **Newest first** or **Oldest first** to get more. In incremental mode a post's comments are returned (and charged) again only when the post itself is `NEW`, `UPDATED` or `REAPPEARED`. A new comment changes `commentsCount`, which counts as a change.

#### Is it legal to scrape Dzen?

This actor collects only publicly available content metadata and article text. You are responsible for how you use the data: follow Dzen's terms and the laws that apply to you, and get legal advice if you plan commercial redistribution. Article text and media can be subject to third-party rights.

#### What is the difference between feed, URL and search mode?

Feed mode walks Dzen's discovery feed and returns a broad, paginated stream of publications. URL mode deep-dives specific links: full article text, channel subscriber counts and their articles, videos or shorts in newest, most popular or oldest order, and video metadata. Search mode returns Dzen's search results for your queries, optionally only articles, only videos or only channels. Channels found by search carry a rounded subscriber count (for example 15.3K); paste the channel link in URL mode for the exact figure.

#### Can I monitor a channel for new publications?

Yes. Paste the channel link, turn on **Incremental mode**, and schedule the run. Each run then returns only new and updated records, and unchanged ones are not billed.

#### Why did my run fail instead of returning an empty dataset?

If Dzen refuses every request, the run stops with a clear message so "no results" is never confused with "nothing could be read". Run it again in a few minutes.

#### Can I use it with AI agents or MCP?

Yes. Call it from any Apify integration or MCP client, and use the connector field to push results into Notion, Linear or Airtable.

### Send results into your apps (MCP connectors)

Optionally pipe results into Notion, Linear, Airtable or Apify through Model Context Protocol connectors. Authorize a connector under Apify, Settings, API & Integrations, then select it in `mcpConnectors`. The dataset output is never changed by this.

# Actor input Schema

## `mode` (type: `string`):

Choose 'feed' to walk Dzen's discovery feed page by page, 'url' to scrape pasted Dzen links (article, video, channel), or 'search' to search Dzen by keyword.

## `searchQueries` (type: `array`):

Keywords or phrases to search Dzen for, one per row (Russian works best). Each query's results are returned until Max records or Max pages is reached.

## `searchType` (type: `string`):

Which kind of results to return: everything, only articles, only videos (including shorts), or only channels. Channels found by search carry a rounded subscriber count (for example 15.3K).

## `urls` (type: `array`):

URL mode only. Paste article links (dzen.ru/a/...), video links (dzen.ru/video/watch/...), or channel links (dzen.ru/<channel-name>). Mixed sets are fine.

## `channelSort` (type: `string`):

For channel links: which publications come first. Most popular first is the channel's own top list.

## `channelContent` (type: `string`):

For channel links: return the channel's articles (each read in full), its videos, or its shorts (both returned straight from the channel's listing, which makes them much faster).

## `fetchArticleText` (type: `boolean`):

Off (default): a channel's articles come from the channel's listing: title, lead paragraph, comment count, read time, date and cover image. That is fast and costs only the record. On: each article is read in full (complete text and all its images), charged as Full page read on top of the record. Pasted article and video links are always read in full and charged the same way.

## `includeComments` (type: `boolean`):

Turn on to collect each article's and video's comments (text, author, date, likes, dislikes, reactions, reply count). Channels have no comments. Replies to comments are counted in commentReplies but not returned.

## `maxCommentsPerPost` (type: `integer`):

Upper limit of comments per article or video. Does not count toward Max records. With Most popular first, Dzen returns at most 20 comments per post.

## `commentsSort` (type: `string`):

Which comments come first. Most popular first returns at most 20 per post; Newest first and Oldest first can return all of them.

## `maxItems` (type: `integer`):

Maximum number of records to return across the whole run. This is the run's cap. Use 0 for unlimited.

## `maxPages` (type: `integer`):

Maximum number of result pages read per feed or search query. 0 means no limit: the run then stops only at Max records.

## `resumeFromRunId` (type: `string`):

Paste a previous run ID or dataset ID to continue a large pull without returning records already collected there. For recurring monitoring instead, use Incremental changes below (the two are mutually exclusive once Incremental has saved state).

## `incrementalMode` (type: `boolean`):

Turn this on for daily or recurring monitoring. The first run returns every matching record as NEW. Later runs normally return only NEW, UPDATED and REAPPEARED records, which is how a like-count jump or a new channel publication is spotted. Turn on Emit unchanged or Emit expired only when you also want those rows returned (and billed).

## `stateKey` (type: `string`):

Optional. Name this monitoring campaign to keep its state stable, or to deliberately share state across differently configured runs. Leave empty to let the actor derive a key automatically from the link set.

## `emitUnchanged` (type: `boolean`):

Off by default. Turn on to also return records that have not changed since the last run, marked UNCHANGED. This returns, and bills, extra rows you already have, so leave it off unless you want the full snapshot every run.

## `emitExpired` (type: `boolean`):

Off by default. Turn on to also return records that were present in a previous run but are no longer found, marked EXPIRED. Only produced once a run has fully scanned every link, so never when Max records capped it or when Resume was used.

## `ignoreFieldsForChanges` (type: `array`):

Output field names that should NOT make a record count as UPDATED. views is always ignored, because view counts rise on almost every run. Add likes, commentsCount or subscribers here if you only want to hear about new content, not engagement changes. Ignored fields are still returned in the output.

## `proxy` (type: `object`):

The residential RU connection is the default and recommended.

## `mcpConnectors` (type: `array`):

Optional. Each selected connector receives a condensed summary per record; the full record always stays in the dataset. Leave empty to skip.

## `notionParentPageUrl` (type: `string`):

URL (or id) of the Notion page under which record pages are created. Required to enable the Notion export; ignored by other connectors.

## `maxNotifyListings` (type: `integer`):

Cap on records written to each connector per run. Does not affect the dataset.

## Actor input object example

```json
{
  "mode": "feed",
  "searchQueries": [
    "грибы"
  ],
  "searchType": "all",
  "urls": [
    {
      "url": "https://dzen.ru/tass"
    }
  ],
  "channelSort": "newest",
  "channelContent": "articles",
  "fetchArticleText": false,
  "includeComments": false,
  "maxCommentsPerPost": 20,
  "commentsSort": "top",
  "maxItems": 20,
  "maxPages": 0,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyCountry": "RU"
  },
  "mcpConnectors": [],
  "maxNotifyListings": 50
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

## `comments` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "feed",
    "searchQueries": [
        "грибы"
    ],
    "urls": [
        {
            "url": "https://dzen.ru/tass"
        }
    ],
    "fetchArticleText": false,
    "includeComments": false,
    "incrementalMode": false,
    "emitUnchanged": false,
    "emitExpired": false,
    "proxy": {
        "useApifyProxy": true,
        "apifyProxyCountry": "RU"
    },
    "mcpConnectors": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("abotapi/dzen-ru-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "feed",
    "searchQueries": ["грибы"],
    "urls": [{ "url": "https://dzen.ru/tass" }],
    "fetchArticleText": False,
    "includeComments": False,
    "incrementalMode": False,
    "emitUnchanged": False,
    "emitExpired": False,
    "proxy": {
        "useApifyProxy": True,
        "apifyProxyCountry": "RU",
    },
    "mcpConnectors": [],
}

# Run the Actor and wait for it to finish
run = client.actor("abotapi/dzen-ru-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "feed",
  "searchQueries": [
    "грибы"
  ],
  "urls": [
    {
      "url": "https://dzen.ru/tass"
    }
  ],
  "fetchArticleText": false,
  "includeComments": false,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "proxy": {
    "useApifyProxy": true,
    "apifyProxyCountry": "RU"
  },
  "mcpConnectors": []
}' |
apify call abotapi/dzen-ru-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,abotapi/dzen-ru-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/IBVFySvVFF3gzjdni/builds/l8HVDCjC7u4fCxc46/openapi.json
