YouTube Scraper — Search, Channels, Videos & Playlists avatar

YouTube Scraper — Search, Channels, Videos & Playlists

Pricing

$0.40 / 1,000 youtube results

Go to Apify Store
YouTube Scraper — Search, Channels, Videos & Playlists

YouTube Scraper — Search, Channels, Videos & Playlists

All-in-one YouTube data Actor: keyword search, per-video detail (exact player fields when served, else exact-id list metadata), channel videos/shorts/live and whole playlists — one flat schema, no API key, no login, no proxy. Pay per result: 1 unit per row; failed or empty lookups never billed.

Pricing

$0.40 / 1,000 youtube results

Rating

0.0

(0)

Developer

Datamule

Datamule

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

An all-in-one Actor for public YouTube data. No API key, no Google quota, no login, no cookies, no proxy required. Four modes over one flat output schema, so a search row, a channel row and a playlist item all land in the same dataset with a type field telling you what each one is.

Modes

ModeInput fieldWhat you get
searchsearchQueriesKeyword search results, paginated — one row per video
videosvideoUrlsOne row per resolvable exact video ID (an unavailable ID yields no row and is never charged): exact player detail (exact view/like counts, publish date, category, keywords) when YouTube serves the player, otherwise exact-ID list metadata
channelchannelUrlsA channel's Videos, Shorts and Live tabs, plus a channel summary row
playlistplaylistUrlsEvery item of a playlist, in order, with its position

Handles (@name), /channel/UC…, /c/…, /user/…, youtu.be/…, /shorts/…, ?v=… and bare 11-character IDs are all accepted and normalised for you.

Forgot to switch Mode? If the selected mode's field is empty but exactly one of the URL fields (videos, channels, playlists) has entries (e.g. channel URLs pasted with Mode still on search), the run uses that mode and logs a warning instead of failing. The same applies when the search box still holds only the untouched sample query (python tutorial): your pasted URLs win. It never switches to search, so leftover sample queries are never billed. If the sample query is untouched and URLs are in two or more of those fields, the run stops with an input error naming them (nothing fetched or charged): pick the Mode you want or keep URLs in one field.

Pricing — pay per result, nothing else

There is no actor-start fee. You are charged only for rows actually delivered: every video, short, channel or playlistItem row is one result event at $0.0004 ($0.40 per 1,000).

Rows delivered (maxItems)Actor charge (USD)
25$0.01
100 (default)$0.04
1,000$0.40

A video that is unavailable, a channel that will not resolve, or a playlist that returns nothing emits no row and costs you nothing. maxItems is a hard ceiling on rows (and so on cost) for the whole run. If you also set a positive Apify maximum charge per run, the Actor stops when it is reached and keeps the rows already delivered.

Example input

{
"mode": "search",
"searchQueries": ["python tutorial"],
"maxItems": 25,
"language": "en",
"region": "US"
}

Every video of a channel's Shorts and Videos tabs:

{
"mode": "channel",
"channelUrls": ["@YouTube"],
"channelTabs": ["videos", "shorts"],
"maxItems": 100
}

Use it from your own code (Python, standard library only)

Latest 10 videos of a channel into your pipeline: start a run with an explicit Mode, a maxItems ceiling and a per-run cost cap (maxTotalChargeUsd), wait for it to finish, then read the dataset. Set APIFY_TOKEN in your environment. The run stops when the cap is reached, so raise maxTotalChargeUsd if you raise maxItems.

import json, os, time, urllib.request
TOKEN = os.environ["APIFY_TOKEN"]
API = "https://api.apify.com/v2"
def call(method, path, body=None):
req = urllib.request.Request(
API + path, method=method,
data=None if body is None else json.dumps(body).encode(),
headers={"Authorization": f"Bearer {TOKEN}", "Content-Type": "application/json"})
with urllib.request.urlopen(req) as r:
return json.load(r)
run_input = {"mode": "channel", "channelUrls": ["@YouTube"],
"channelTabs": ["videos"], "maxItems": 10, "includeChannelInfo": False}
# Start the run; maxTotalChargeUsd caps what this run can cost you.
run = call("POST", "/acts/datamule~youtube-data-scraper/runs?maxTotalChargeUsd=0.01", run_input)["data"]
# Poll until the run finishes (waitForFinish holds each request up to 60 s).
while run["status"] in ("READY", "RUNNING"):
run = call("GET", f"/actor-runs/{run['id']}?waitForFinish=60")["data"]
time.sleep(1)
if run["status"] != "SUCCEEDED":
raise SystemExit(f"run {run['id']} ended {run['status']}")
items = call("GET", f"/datasets/{run['defaultDatasetId']}/items?clean=true&fields=id,title,url,viewCount,publishedText")
for v in items:
print(v.get("publishedText"), v.get("viewCount"), v["title"], v["url"])

Real output from a test run on 2026-09-27 (first 3 of 10 rows; 10 rows = $0.004 at the base price, less on volume tiers):

1 month ago 2200000 Keira Riff reacts to her Watch History https://www.youtube.com/watch?v=wGA27zJEnaU
1 month ago 4099999 Sam Smith reacts to their Watch History https://www.youtube.com/watch?v=NDHk8bGGq74
1 month ago 7700000 Marina Sena reacts to her Watch History https://www.youtube.com/watch?v=FNoK3gc-Jk0

Optional: on repeat runs, print only newly observed videos

Append this to the recipe above. The first run for an input only saves a baseline. Later runs print IDs not seen before for that exact run_input. The snapshot is a local file you own. It is updated atomically, and only after a successful, non-empty result.

import hashlib
def newly_observed(run_input, items, path="seen_videos.json"):
ids = [v.get("id") for v in items]
if not ids or not all(isinstance(i, str) and i for i in ids):
raise SystemExit("empty or invalid result; snapshot left unchanged")
key = hashlib.sha256(json.dumps(run_input, sort_keys=True).encode()).hexdigest()[:16]
state = {}
if os.path.exists(path):
with open(path) as f:
state = json.load(f) # a corrupt file stops here, unchanged
first_run = key not in state
seen = set(state.get(key, []))
new = [] if first_run else [v for v in items if v["id"] not in seen]
state[key] = sorted(seen | set(ids))
with open(path + ".tmp", "w") as f:
json.dump(state, f)
os.replace(path + ".tmp", path) # atomic: old snapshot survives a crash
return first_run, new
first_run, new = newly_observed(run_input, items)
print("baseline saved, nothing reported" if first_run else f"{len(new)} newly observed")
for v in new:
print(v["id"], v["title"], v["url"])

"Newly observed" means that this ID was not in any earlier result for this input (changing any input field, even maxItems, starts a new baseline). It does not mean the video was newly published. Pinned or re-ranked older videos can show up this way. A video that is missing from a later run has not necessarily been deleted, and a run limited by maxItems or the cost cap only covers the channel's most recent window, not its full history. Run it often enough that the channel posts fewer than maxItems new videos between runs, or you will miss some.

Output

Every row carries type, id, url, title and provenance (_mode, _source, _input, _language, _region). Depending on the type you also get channelId / channelName / channelUrl, durationSeconds / durationText, viewCount / viewCountText, likeCount, publishedAt / publishedText / publishedApprox + publishedApproxPrecision, category, keywords, thumbnailUrl and playlistId / playlistPosition.

Counts that YouTube publishes only as human text ("1.2M views") are parsed to an integer and kept verbatim, so you never have to guess how a number was rounded.

Notes and honest limits

  • Only public data is read. Private, members-only, age-gated-login and region-blocked videos are skipped, not faked.
  • Search and channel-tab list rows carry YouTube's own approximate view counts and relative publish text ("8 months ago").
  • To sort or filter those rows by date in a spreadsheet, use the Published (approx.) column (publishedApprox): the relative text turned into a date on the run day — 2024 for "2 years ago", 2026-07 for "2 months ago", 2026-09-26 for "3 days ago". publishedApproxPrecision says which unit it is (YouTube rounds ages down, so it can be off by one unit), or exact when the exact publishedAt date is known. It is filled for English (language: "en", the default) and costs nothing extra.
  • mode: "videos" returns one row for each resolvable ID you list, and every delivered row belongs to the exact ID you asked for. It returns YouTube's exact player fields (viewCount, likeCount, publishedAt, uploadedAt, category, keywords) when YouTube serves the player payload to the run's IP. When the player is refused — which is common from plain datacenter IPs — the row falls back to the same exact-ID list metadata as search (viewCountText, publishedText) and the exact-only fields are simply omitted rather than guessed. Either way the row is real data for the ID you asked for, never a substitute video.
  • If a run resolves zero rows it exits non-zero instead of reporting an empty success, so a broken input never looks like a clean run.

This is an independent implementation built against public YouTube client behaviour. It is not affiliated with, endorsed by, or derived from YouTube, Google, or any other Apify Actor.