YouTube Search Results Scraper
Pricing
from $2.99 / 1,000 search results
YouTube Search Results Scraper
Collect current YouTube search results for videos, Shorts, channels, and playlists with paginated, deduplicated metadata and locale-aware source context.
Pricing
from $2.99 / 1,000 search results
Rating
0.0
(0)
Developer
w3crawler
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
YouTube Public Search Scraper
Collect deduplicated public YouTube search results for videos, Shorts, channels, and playlists. The actor requests the ordinary rendered search page and parses its embedded public ytInitialData; it does not call the private YouTube search API or rotate client fingerprints.
Dataset
Every result includes the submitted query, stable rank, page number, actual result type, title, public URL, source location, locale, enrichment provenance, and scrape time. When YouTube exposes the fields, the record also includes video/playlist/channel IDs, channel links, verification badges, displayed counts plus parsed counts, duration, live/upcoming flags, descriptions, accessibility text, thumbnails, playlist previews, search filters, and the exact public renderer source.
Blocked pages, unavailable embedded metadata, and empty matches are represented by exactly four diagnostic fields: url, error, errorCode, and scrapedAt. Per-run counts and proxy state are written to the OUTPUT key-value record.
Input
| Field | Default | Description |
|---|---|---|
query | required | Public YouTube keyword or phrase, up to 500 characters. |
maxItems | 50 | Maximum unique results, 1–500. |
maxPages | 3 | Bounded public page requests, 1–30. The actor uses the normal page query parameter and stops when a page is empty or duplicate. |
resultType | all | all, video, short, channel, or playlist. |
sortBy | relevance | relevance, upload_date, view_count, or rating; type-specific searches use relevance. |
languageCode / countryCode | en / US | Locale context passed to the public search page. |
proxyConfiguration | none | Optional account-authorized Apify Proxy or credential-free HTTP/SOCKS URLs. |
enableProxyFallback | true | Try one ordinary Apify Proxy request after a direct public-page failure. |
includeDiagnostics | true | Keep bounded diagnostics in the dataset. |
requestDelayMs / requestTimeoutSecs | 250 / 60 | Pacing and bounded request timeout. |
Example
{"query": "machine learning","maxItems": 25,"resultType": "video","sortBy": "relevance","maxPages": 2,"countryCode": "US","enableProxyFallback": true,"includeDiagnostics": true,"proxyConfiguration": {"useApifyProxy": false}}
Reliability and limits
- Uses a fixed ordinary browser User-Agent and public HTML only.
- Proxy requests are account-authorized and are summarized as booleans in
OUTPUT; proxy URLs, credentials, cookies, and request identities are never written to result rows. - Does not implement fingerprint spoofing, stealth patches, CAPTCHA/login bypass, private or alternate YouTube clients, or hidden API calls.
- Search results, ranking, locale, badges, counts, and availability are point-in-time public observations.
maxPagesis a safety limit, not a guarantee that YouTube will expose that many public pages.
Local verification
npm testnpx --yes apify validate-schema$env:APIFY_LOCAL_STORAGE_DIR = 'storage'npx --yes apify run --purge --input-file INPUT.jsonnode validate-datasets.js
The checked-in storage/ sample contains the latest local run. Before replacing it, preserve the existing store as a backup.