# YouTube Live Scraper - Bulk Export to CSV, JSON, API (`reapx/youtube-live-scraper`) Actor

Pull youtube, crawler, video, alternative, limits in bulk. Every row carries quotas, channel, name, likes, number, views, subscribers, public, page. Ready for CSV, Excel, JSON or the API.

- **URL**: https://apify.com/reapx/youtube-live-scraper.md
- **Developed by:** [Tarek Etman](https://apify.com/reapx) (community)
- **Categories:** Social media, Videos, Marketing
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.50 / 1,000 live streams

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Youtube Live Scraper

Youtube Live Scraper extracts structured records in bulk and exports them for analysis, enrichment
and downstream pipelines. It covers youtube, public, page, beyond, performing, video, channel, playlist, stream, shorts, search-result, information, crawls, pages, parses, metadata, on-page, content, produce, containing.

Built for teams that need titles, descriptions, durations, publish, dates, view, counts, like without maintaining scrapers, proxies or browser
infrastructure themselves.

### Quick start (SDK examples)

#### Python

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("youtube-live-scraper").call(run_input={"targets": ["<target>"], "maxResults": 100})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

#### JavaScript

```javascript
import { ApifyClient } from "apify-client";

const client = new ApifyClient({ token: "YOUR_APIFY_TOKEN" });
const run = await client.actor("youtube-live-scraper").call({ targets: ["<target>"], maxResults: 100 });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

#### cURL

```curl
curl -X POST "https://api.apify.com/v2/acts/youtube-live-scraper/runs?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"targets":["<target>"],"maxResults":100}'
```

### Fields returned

| field | description | type |
|---|---|---|
| `title` | title returned for every record | string |
| `url` | url returned for every record | string |
| `author` | author returned for every record | string |
| `viewers` | viewers returned for every record | string |
| `publishedAt` | publishedAt returned for every record | string |
| `image` | image returned for every record | string |
| `mode` | mode returned for every record | string |
| `format` | format returned for every record | string |
| `storage_type` | storage type returned for every record | string |
| `storage_config` | storage config returned for every record | string |
| `scrapedAt` | scrapedAt returned for every record | string |

### What it does

- Extract youtube, public, page, beyond, performing, video into structured rows.
- Enrich each record with channel, playlist, stream, shorts, search-result, information.
- Bulk export covering crawls, pages, parses, metadata, on-page, content.
- Pipeline integration for produce, containing, titles, descriptions, durations, publish.
- Downstream analysis across dates, view, counts, like, comment, thumbnails.
- Recurring monitoring of urls, hashtags, name, location, join, date.
- Deduplicated output keyed on the record identifier.
- Configurable result caps and runtime bounds.

### Use cases

- **Lead generation** — build contactable lists covering youtube, public, page, beyond, performing
- **Data enrichment** — attach video, channel, playlist, stream, shorts to an existing record set
- **Market research** — map search-result, information, crawls, pages, parses across a category or region
- **Competitive monitoring** — track metadata, on-page, content, produce, containing over time on a schedule
- **AI and RAG pipelines** — feed clean structured rows into embeddings and retrieval
- **Warehousing** — land titles, descriptions, durations, publish, dates into BigQuery, Snowflake or Postgres

### Input

Provide `targets` as a list of URLs or identifiers, one per line.

| input | purpose |
|---|---|
| `targets` | URLs or identifiers to process, one per line |
| `maxResults` | hard cap on returned rows |
| `maxSeconds` | runtime bound for the run |
| `includeEmpty` | return rows that resolved to no data, or skip them |

### Output

Every run writes a dataset exportable as CSV, Excel, JSON, or readable directly from the Apify API. Attach a webhook to push results into your own system as soon as a run finishes.

### Integrations

Works with Zapier, Make, n8n, Google Sheets, Slack, and any HTTP endpoint via webhooks. The Apify MCP server exposes this Actor to AI agents directly.

### Performance and limits

Runs are concurrent and bounded by `maxResults` and `maxSeconds`. Proxy rotation and retry handling are managed for you. Failed targets are reported rather than silently dropped.

### Frequently asked questions

##### Do I need an account or cookies?

No. The Actor reads public data only and requires no login, cookies or personal API keys.

##### What formats can I export?

CSV, Excel, JSON, or read the dataset straight from the Apify API.

##### What does a row contain?

Every row carries youtube, public, page, beyond, performing, video, channel, playlist where available.

##### Can I schedule it?

Yes. Attach a schedule or a webhook and the dataset is produced on your cadence.

##### How do I limit cost?

Use `maxResults` to cap returned rows and `maxSeconds` to bound runtime.

##### Is the output stable?

Field names are fixed by the dataset schema, so downstream pipelines do not break between runs.

### Field glossary

**`title`** — the title associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`url`** — the url associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`author`** — the author associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`viewers`** — the viewers associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`publishedAt`** — the publishedAt associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`image`** — the image associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`mode`** — the mode associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`format`** — the format associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`storage_type`** — the storage type associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`storage_config`** — the storage config associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.
**`scrapedAt`** — the scrapedAt associated with the record. Present on every row where the source exposes it; absent values are returned as null rather than omitted, so column order stays stable across runs and downstream schemas do not drift.

### Troubleshooting

- **Empty dataset** — Check that `targets` contains reachable identifiers and that `includeEmpty` is set the way you expect.
- **Run times out** — Lower `maxResults` or raise `maxSeconds`; very large target lists are better split across scheduled runs.
- **Missing fields** — Not every source exposes every field. Absent values are returned as null so the schema stays stable.
- **Rate limiting** — Proxy rotation is automatic. If a source throttles hard, reduce concurrency and retry.
- **Duplicate rows** — Output is deduplicated on the record identifier; duplicates across separate runs are expected by design.

### Data quality notes

Records are parsed from public sources covering youtube, public, page, beyond, performing, video, channel, playlist, stream, shorts. Values are returned exactly as published rather than normalised or inferred, so you can audit any row back to its source URL. Timestamps are ISO-8601 UTC. Numeric counters are integers. No field is synthesised when the source does not publish it.

### Scheduling and automation

Attach a schedule to run this Actor hourly, daily or weekly. Combine it with a webhook to push each finished dataset into your warehouse, CRM or Slack channel automatically. Runs are idempotent with respect to their input, so a repeated schedule produces a comparable dataset rather than a drifting one.

### Support

Open an issue on the Actor's Issues tab. Include the run ID and the input used so it can be reproduced.

# Actor input Schema

## `targets` (type: `array`):

URLs, handles or search terms — any mix. Each one is resolved to the cheapest route that returns data. Provide `targets` as a list, one per line (for example `example-value`). This is the work list the run iterates over.

## `maxResults` (type: `integer`):

Stop after this many rows. You are charged per row returned. Provide `maxResults` as a whole number between 1 and 100000 (for example `100`). Use it to bound both runtime and spend on open-ended inputs.

## `maxSeconds` (type: `integer`):

Stop cleanly after this long and keep the rows already found. Provide `maxSeconds` as a whole number between 30 and 3600 (for example `100`). Use it to stop a run before it exceeds your time budget.

## `includeEmpty` (type: `boolean`):

Adds an unbilled row for each target with no data. Never charged. Toggle `includeEmpty` on or off. Turn it on when you need a row for every input, even empty ones.

## Actor input object example

```json
{
  "targets": [
    "https://www.youtube.com/watch?v=21X5lGlDOfg"
  ],
  "maxResults": 100,
  "maxSeconds": 240,
  "includeEmpty": false
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "targets": [
        "https://www.youtube.com/watch?v=21X5lGlDOfg"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("reapx/youtube-live-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "targets": ["https://www.youtube.com/watch?v=21X5lGlDOfg"] }

# Run the Actor and wait for it to finish
run = client.actor("reapx/youtube-live-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "targets": [
    "https://www.youtube.com/watch?v=21X5lGlDOfg"
  ]
}' |
apify call reapx/youtube-live-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,reapx/youtube-live-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CWs6AE3piAemhce9g/builds/8qZEpSolt86VSYr28/openapi.json
