YouTube Scraper — Search, Channels, Videos & Playlists
Pricing
$0.40 / 1,000 youtube results
YouTube Scraper — Search, Channels, Videos & Playlists
All-in-one YouTube data Actor: keyword search, per-video detail (exact player fields when served, else exact-id list metadata), channel videos/shorts/live and whole playlists — one flat schema, no API key, no login, no proxy. Pay per result: 1 unit per row; failed or empty lookups never billed.
Pricing
$0.40 / 1,000 youtube results
Rating
0.0
(0)
Developer
Datamule
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
An all-in-one Actor for public YouTube data. No API key, no Google quota, no
login, no cookies, no proxy required. Four modes over one flat output schema, so
a search row, a channel row and a playlist item all land in the same dataset with
a type field telling you what each one is.
Modes
| Mode | Input field | What you get |
|---|---|---|
search | searchQueries | Keyword search results, paginated — one row per video |
videos | videoUrls | One row per resolvable exact video ID (an unavailable ID yields no row and is never charged): exact player detail (exact view/like counts, publish date, category, keywords) when YouTube serves the player, otherwise exact-ID list metadata |
channel | channelUrls | A channel's Videos, Shorts and Live tabs, plus a channel summary row |
playlist | playlistUrls | Every item of a playlist, in order, with its position |
Handles (@name), /channel/UC…, /c/…, /user/…, youtu.be/…, /shorts/…,
?v=… and bare 11-character IDs are all accepted and normalised for you.
Forgot to switch Mode? If the selected mode's field is empty but exactly one
of the URL fields (videos, channels, playlists) has entries (e.g. channel URLs pasted
with Mode still on search), the run uses that mode and logs a warning instead of
failing. The same applies when the search box still holds only the untouched sample
query (python tutorial): your pasted URLs win. It never switches to search, so
leftover sample queries are never billed.
If the sample query is untouched and URLs are in two or more of those fields, the
run stops with an input error naming them (nothing fetched or charged): pick the
Mode you want or keep URLs in one field.
Pricing — pay per result, nothing else
There is no actor-start fee. You are charged only for rows actually
delivered: every video, short, channel or playlistItem row is one
result event at $0.0004 ($0.40 per 1,000).
Rows delivered (maxItems) | Actor charge (USD) |
|---|---|
| 25 | $0.01 |
| 100 (default) | $0.04 |
| 1,000 | $0.40 |
A video that is unavailable, a channel that will not resolve, or a playlist that
returns nothing emits no row and costs you nothing. maxItems is a hard
ceiling on rows (and so on cost) for the whole run. If you also set a positive
Apify maximum charge per run, the Actor stops when it is reached and keeps the
rows already delivered.
Example input
{"mode": "search","searchQueries": ["python tutorial"],"maxItems": 25,"language": "en","region": "US"}
Every video of a channel's Shorts and Videos tabs:
{"mode": "channel","channelUrls": ["@YouTube"],"channelTabs": ["videos", "shorts"],"maxItems": 100}
Use it from your own code (Python, standard library only)
Latest 10 videos of a channel into your pipeline: start a run with an explicit
Mode, a maxItems ceiling and a per-run cost cap (maxTotalChargeUsd), wait for
it to finish, then read the dataset. Set APIFY_TOKEN in your environment.
The run stops when the cap is reached, so raise maxTotalChargeUsd if you raise
maxItems.
import json, os, time, urllib.requestTOKEN = os.environ["APIFY_TOKEN"]API = "https://api.apify.com/v2"def call(method, path, body=None):req = urllib.request.Request(API + path, method=method,data=None if body is None else json.dumps(body).encode(),headers={"Authorization": f"Bearer {TOKEN}", "Content-Type": "application/json"})with urllib.request.urlopen(req) as r:return json.load(r)run_input = {"mode": "channel", "channelUrls": ["@YouTube"],"channelTabs": ["videos"], "maxItems": 10, "includeChannelInfo": False}# Start the run; maxTotalChargeUsd caps what this run can cost you.run = call("POST", "/acts/datamule~youtube-data-scraper/runs?maxTotalChargeUsd=0.01", run_input)["data"]# Poll until the run finishes (waitForFinish holds each request up to 60 s).while run["status"] in ("READY", "RUNNING"):run = call("GET", f"/actor-runs/{run['id']}?waitForFinish=60")["data"]time.sleep(1)if run["status"] != "SUCCEEDED":raise SystemExit(f"run {run['id']} ended {run['status']}")items = call("GET", f"/datasets/{run['defaultDatasetId']}/items?clean=true&fields=id,title,url,viewCount,publishedText")for v in items:print(v.get("publishedText"), v.get("viewCount"), v["title"], v["url"])
Real output from a test run on 2026-09-27 (first 3 of 10 rows; 10 rows = $0.004 at the base price, less on volume tiers):
1 month ago 2200000 Keira Riff reacts to her Watch History https://www.youtube.com/watch?v=wGA27zJEnaU1 month ago 4099999 Sam Smith reacts to their Watch History https://www.youtube.com/watch?v=NDHk8bGGq741 month ago 7700000 Marina Sena reacts to her Watch History https://www.youtube.com/watch?v=FNoK3gc-Jk0
Optional: on repeat runs, print only newly observed videos
Append this to the recipe above. The first run for an input only saves a
baseline. Later runs print IDs not seen before for that exact run_input. The
snapshot is a local file you own. It is updated atomically, and only after a
successful, non-empty result.
import hashlibdef newly_observed(run_input, items, path="seen_videos.json"):ids = [v.get("id") for v in items]if not ids or not all(isinstance(i, str) and i for i in ids):raise SystemExit("empty or invalid result; snapshot left unchanged")key = hashlib.sha256(json.dumps(run_input, sort_keys=True).encode()).hexdigest()[:16]state = {}if os.path.exists(path):with open(path) as f:state = json.load(f) # a corrupt file stops here, unchangedfirst_run = key not in stateseen = set(state.get(key, []))new = [] if first_run else [v for v in items if v["id"] not in seen]state[key] = sorted(seen | set(ids))with open(path + ".tmp", "w") as f:json.dump(state, f)os.replace(path + ".tmp", path) # atomic: old snapshot survives a crashreturn first_run, newfirst_run, new = newly_observed(run_input, items)print("baseline saved, nothing reported" if first_run else f"{len(new)} newly observed")for v in new:print(v["id"], v["title"], v["url"])
"Newly observed" means that this ID was not in any earlier result for this
input (changing any input field, even maxItems, starts a new baseline). It does not mean the video was newly published. Pinned or re-ranked older
videos can show up this way. A video that is missing from a later run has not
necessarily been deleted, and a run limited by maxItems or the cost cap only
covers the channel's most recent window, not its full history. Run it often
enough that the channel posts fewer than maxItems new videos between runs,
or you will miss some.
Output
Every row carries type, id, url, title and provenance (_mode,
_source, _input, _language, _region). Depending on the type you also get
channelId / channelName / channelUrl, durationSeconds / durationText,
viewCount / viewCountText, likeCount, publishedAt / publishedText /
publishedApprox + publishedApproxPrecision,
category, keywords, thumbnailUrl and playlistId / playlistPosition.
Counts that YouTube publishes only as human text ("1.2M views") are parsed to an integer and kept verbatim, so you never have to guess how a number was rounded.
Notes and honest limits
- Only public data is read. Private, members-only, age-gated-login and region-blocked videos are skipped, not faked.
- Search and channel-tab list rows carry YouTube's own approximate view counts and relative publish text ("8 months ago").
- To sort or filter those rows by date in a spreadsheet, use the Published
(approx.) column (
publishedApprox): the relative text turned into a date on the run day —2024for "2 years ago",2026-07for "2 months ago",2026-09-26for "3 days ago".publishedApproxPrecisionsays which unit it is (YouTube rounds ages down, so it can be off by one unit), orexactwhen the exactpublishedAtdate is known. It is filled for English (language: "en", the default) and costs nothing extra. mode: "videos"returns one row for each resolvable ID you list, and every delivered row belongs to the exact ID you asked for. It returns YouTube's exact player fields (viewCount,likeCount,publishedAt,uploadedAt,category,keywords) when YouTube serves the player payload to the run's IP. When the player is refused — which is common from plain datacenter IPs — the row falls back to the same exact-ID list metadata as search (viewCountText,publishedText) and the exact-only fields are simply omitted rather than guessed. Either way the row is real data for the ID you asked for, never a substitute video.- If a run resolves zero rows it exits non-zero instead of reporting an empty success, so a broken input never looks like a clean run.
This is an independent implementation built against public YouTube client behaviour. It is not affiliated with, endorsed by, or derived from YouTube, Google, or any other Apify Actor.