# Changelog of YouTube Channel Email & Sponsor Lead Scraper API (`trakk/youtube-channel-email-sponsor-leads`) Actor

- **URL**: https://apify.com/trakk/youtube-channel-email-sponsor-leads/changelog.md
- **Full Actor documentation**: https://apify.com/trakk/youtube-channel-email-sponsor-leads.md

## Changelog

### 1.1.118 — A filled order stops apologising

- **A restarted run that delivered everything asked for still reported a shortfall.** It compared the pass against the order: a run that wrote 184 leads, migrated, and closed the remaining 816 announced "Saved 816 of the 1,000 requested" beside a dataset of exactly 1,000. The order is filled when the dataset holds what was asked for.
- **The e-mail count is only stated when it describes the whole dataset**, for the same reason.

### 1.1.116 — Reports that read like sentences

- **Counts are worded, not bracketed.** Every number a run reports used the form-field spelling - "1 lead(s)", "2 search(es)" - including the status line and the shortfall explanation. They now read as English.
- **Queries the run wrote itself are counted apart from the ones you typed.** A 60-query run that wrote 12 of its own reported "58 of 72 searches have no pages left", which reads like a miscount to somebody who entered 60. The sentence now names both numbers.

### 1.1.115 — The status line counts the same thing as the result count

- **A restarted run read as "756 lead(s) saved" next to 970 results.** The summary was already fixed to describe the dataset; the status line beside it in the run list was not, and still counted only the pass after the migration. It now reports the dataset total, and says how the two numbers relate.

### 1.1.114 — A page that repeats itself is not a page

- **Expansion never ran on a wide query set.** New queries were written only once no query had pages left at all — but a query that still lists a next page is not the same as a query that still holds creators. On a 60-query run, two searches kept reporting further pages while every one of those pages repeated channels already examined, so the run read them round after round and wrote no queries of its own: 591 leads of 1,000 requested, `queriesWritten: 0`. A round that returns nobody new now hands straight over to expansion.

### 1.1.113 — A restarted run no longer overfills the order

- **A run whose container migrated could deliver more than was asked for, and charge for it.** Apify moves a run between machines and starts the Actor again with a fresh counter; the dataset keeps every row written before the move. Counting only the new pass, a request for 1,000 leads wrote 1,020 rows and billed all 1,020. Rows already in this run's own dataset now come out of the request, and a run that finds the order already filled ends immediately. Rows carried from a named lead memory are untouched by this: that caller is asking for leads they do not yet have.

### 1.1.112 — Running out of creators is a moment, not a verdict

- **A run that went dry stopped asking, for good.** The first replacement round that returned nobody new ended top-up permanently — but the next queries are written from the creators the run has kept, so every further lead is new material to search with. One measured run went dry at 409 of 1,000 requested and saved 133 more leads afterwards without ever asking again. It now asks again after every 25 further leads.

### 1.1.110 — Totals that describe the dataset, and a search that never gives up

- **A run that was restarted mid-way reported only its second half.** Apify migrates a container between machines and starts the Actor again with fresh counters, while the dataset keeps every row written before the move: one measured run delivered 356 leads and reported 72. Nothing was lost or double-charged — the run reads its own dataset on startup, so not one of those 356 was a duplicate and not one was billed twice — but the summary described the pass rather than the dataset. It now reports `datasetTotal` and `savedBeforeThisPass`, and says so in the message.
- **Expansion no longer gives up when it has little to learn from.** Queries are built from the channel keywords of the creators a run kept, and a run can be short of those exactly when it needs expanding most: a single query that saved one lead produced no usable keywords, and the run finished at 1 of 500 requested. It now falls back to the caller's own wording, paired with the phrases creators title with. The same input now returns 356 leads from that one query.

### 1.1.105 — Twice as fast, and cheaper per lead

- **Runs go about twice as fast.** Measured four ways on the same input in the same minute: the old settings managed 3.55 creators/second, and 7.13 with 2 GB of memory, 24 creators in parallel and a 1,200/minute request ceiling. Memory was the wall, not YouTube — the 1 GB run peaked at exactly its limit, and the source answered 952 requests a minute with the same 0.3% retry rate it gave 545. Going wider still (1,800/min, width 32) added nothing, so the source is not pushed harder than it answers. No request was removed: the same pages are read, just without the Actor holding itself back.
- **A lead costs less.** Compute is billed by time, so halving the time halves what each creator costs to examine: $0.00082 per lead before, $0.00060-0.00069 after, with margin rising from 74% to 79-81%. A run now pays for itself at a keep rate of 1.1%, where it used to need 1.6%.
- **The spending estimate reads the memory the run was given.** It assumed 1 GB, so a 2 GB run would have understated its own compute bill by half — on the one number the cost ceiling is made of.

### 1.1.103 — Asking for a thousand leads no longer means writing a hundred queries

- **A run that exhausts the queries it was given writes its own.** The caller asks for a number of leads, not for a number of queries, but a YouTube search stops at roughly 500 creators per axis, so a selective filter set could spend every page of every query and still finish short — sixty queries yielded 16,890 creators and 786 leads, ninety yielded 25,228 and 711. When all three axes of every query are spent, the run now builds further queries from the channel keywords of the creators it has actually kept — the source's own words for the niche the caller is really after — and those queries then get the same treatment: further pages, the view-count axis, the video axis. The run reports how many it wrote in `candidateSupply.queriesWritten`.

- **Nothing but the cost ceiling limits how large a run may grow.** Two counters did, and neither guarded anything. The replacement-round allowance was half a round per requested lead, so a run for 1,000 leads stopped at 500 of 500 rounds having examined 20,373 creators and saved 548, while 63 of its searches still held unread pages. The candidate allowance was capped at 40 creators examined per requested lead, and a selective search — emailOnly over a subscriber band, measured at 2.7% kept — needs 37, so the cap sat exactly where such a run had to work. Both are gone: a round ends by itself when discovery returns nothing, and the allowance now follows the run's own keep rate with no ceiling.

- **Adding queries no longer kills the run inside its own search.** Searching pays for itself, but the credit only arrived once the searches were over — while they ran, the opening $0.008 had to cover them. Sixty queries put $0.0055 on the wire and fitted; ninety cost $0.0083 and did not, so a 90-query run stopped at `cost_limit` fifty seconds in, having saved nothing, with all 89 searches still holding pages. Adding queries is the one real answer to a shortfall, and it was the one thing guaranteed to produce one. The credit now extends as the searching happens, still bounded, and only what the searches put on the wire is credited — waiting buys nothing.

- **Topping up counts as work, not as idling.** The same guard that killed a run during its opening search also killed it during a replacement round: a run stopped at `no_progress_limit` with 152 leads saved, six freshly written queries in flight and ten searches still holding pages. The idle clock now stops while any search runs and restarts when it returns.

- **A wide search is no longer mistaken for a stalled run.** The no-progress guard counts candidates finished, and nothing reports the creators discovery is finding, so a run whose searches took longer than a minute — sixty queries, or any search slowed by retries — was killed the moment it started working, with zero leads saved. Measured: the same sixty queries took 19 s on a quiet source and 64 s under load. The idle clock now starts at the first candidate; during discovery the searches are bounded by their own cost credit and the run deadline.

- **A spent query set now says so, instead of blaming a limit.** When the searches answer a replacement round with only creators the run has already examined, top-up ends — and it used to end by setting the round counter to its budget, so the run reported having used all 100,000 of its replacement rounds and sent the buyer looking for a limit to raise. It now reports what actually happened: the searches still list further pages, but those pages repeat channels already seen, so the query set is spent and the answer is more queries or looser filters. `sourceExhausted` counts that case as exhausted, with `pagesLeftButRepeating` to tell it apart from a source with no pages left at all.

- **A run may take up to six hours.** The Actor's own time guard already scaled with the candidates a run was allowed to examine, but the platform stopped every run at three hours regardless. Rows are written as they are found, so a run that is stopped keeps everything it had saved.

### 1.1.87 — Runs of up to 5,000 leads, and contacts that belong to the creator

- **`maxResults` now goes to 5,000.** Raising the number alone would have achieved nothing: a run was also bounded by a flat one-hour deadline, thirty thousand requests and one gibibyte of traffic, all sized for a thousand results, and by a second copy of the old ceiling inside the runner that quietly shrank any larger request back to 1,000. All four now scale with what was asked for, and the revenue-linked cost ceiling remains what actually decides whether a run is worth continuing. Rows are written as they are found, so a run that is stopped or times out keeps everything it had saved.
- **A creator's affiliate links are no longer mistaken for their own site.** Creators put referral links to the tools they promote in their channel links, and the crawl was reading the vendor's site as theirs: one tool's support desk arrived as the contact for five unrelated creators in a single run. Links carrying a referral code, and the hosts behind them, are now refused at the point every URL passes through, so following one from a creator's own page cannot reach them either. An address found in a page's markup is also only accepted when it belongs to that page's own domain, which is how a landing page carrying a promoted tool's structured data was handing out that tool's address.

### 1.1.71 — Continuing a search, without paying for the same creator twice

- **Lead memory.** Name a search and reuse that name across runs: creators already delivered under it are skipped before any page is fetched, so extending a search with new keywords costs nothing for leads you already own. Reported by a customer who had been told to resurrect a run instead — advice that was wrong, because a resurrected run repeats its own original input and ignores later edits to the task, so his edited keyword list never reached it.
- **Creators already delivered are now dropped before enrichment, not after.** The duplicate check ran at the save step, which meant a continued run paid in full to enrich every creator it already had and then discarded the result.
- **A continuation starts with an allowance drawn from what it has already delivered.** The spending guard counted only leads saved *in the current run*, so a run carrying 550 existing leads began on the bare opening budget, spent it on creators it already owned or on the thin remainder of its earlier passes, and stopped having saved nothing. The credit is a small share of what the carried leads earned, capped, and the revenue-linked ceiling still governs the rest of the run.

### 1.1.66 — Selective runs deliver what was asked for

- **How many creators a run may examine now follows how selective the filters turn out to be.** It was a flat four per requested lead, which is generous when most creators are kept and nowhere near enough for an emailOnly run over a subscriber band, where one in five to one in ten survives. Such a run stopped at a fraction of the requested count while its searches still held unread pages. Measured on the same input, same moment, 60 leads requested with emailOnly over a 1k-500k band: 37 delivered before, 56 after, at an identical $0.00014 per saved lead. Spending is still bounded by the revenue-linked cost ceiling, which tightens precisely when a run saves nothing.

### 1.1.64 — Honest shortfall reporting, a buyer-supplied blacklist, and no paying twice to continue

- **A run that falls short now names the constraint that actually stopped it.** It used to say the source had run out of channels while every query still reported unread pages, which sent people off widening filters that were never the problem. The summary now carries `queriesWithUnreadPages`, `sourceExhausted` and `topUpRoundsAllowed` alongside `candidatePoolExhausted`, and the warning distinguishes a genuinely exhausted source from the run hitting its own replacement-round allowance or its candidate ceiling.
- **Replacement rounds scale with the request.** A flat six rounds was sized for runs asking for tens of leads; a request for a thousand got one round per 170 leads and stopped far short. The allowance is now a quarter of the requested count, at least six, and the spending guard and candidate ceiling remain the real limits.
- **`emailBlacklist`: addresses or whole domains that must never be saved as a contact.** A bare domain covers its subdomains. Built-in coverage was extended too, after a customer kept receiving `help@skool.com` as the contact for unrelated creators: the crawl follows a creator's community or tooling link and lands on that platform's help desk. Community and course platforms, developer tooling and site builders are now skipped by default, alongside the merch shops and tip jars added previously.
- **A continued run no longer re-saves and re-charges creators already in its dataset.** Resurrecting a run to keep working through a large search used to enrich and bill the same creators again; the run now reads the keys already stored and skips them, reporting the total as `carriedOverLeads`.

### 1.1.61 — Deep is the sponsorship mode, and now says so

- **Deep reads 20 recent videos instead of 12.** Sponsor evidence is published in video descriptions, so that number is what the mode is actually worth. Measured on the same creators: sponsor evidence on 29% of them against 26% at twelve videos and 16% on Balanced, e-mail coverage 49% against 44%, about eleven seconds slower per forty creators — and the cost per creator carrying sponsor evidence went down, because the extra videos ride along with waits the run already pays for.
- **The depth setting now describes what each mode returns rather than how hard it tries.** Deep was presented as a deeper search for e-mail addresses, which it is not: an address nobody published cannot be found by looking harder, and both modes land near 44-49%. It is presented as what it does deliver, sponsorship history, so the choice between Balanced and Deep is now a real one.
- Two passes that did not pay for themselves were measured and dropped: reading the Community tab (creators in this size range post there rarely enough that it found nothing across 80 channels) and running the archive and contact-path passes for every creator without an address, which doubled the cost of a run and moved e-mail coverage by less than the run-to-run spread. Those passes still run where they pay — for creators whose address YouTube keeps behind its protected-email button.

### 1.1.57 — Selective runs stop paying for creators they were always going to reject

- **The subscriber band is now applied before a creator's channel page is bought.** Search already reports each creator's subscriber total, but the band was only enforced after the page had been fetched and enriched, so a run with a narrow band paid for hundreds of creators it discarded. On a 5,000–49,500 band over broad niches, measured across three paired runs per build: requests fell 57% and traffic 69% for the same creators examined. A creator is only rejected early when the search figure is clearly outside the band; borderline figures still go through the full check, and a direct comparison of 44 candidates found no case where the early decision differed from the full one.
- **Keyword runs accept up to 200 niches and lookalike up to 25 example channels**, raised from 20 and 5. Searching is one request per query while the expensive enrichment stays capped by the candidate pool and the cost budget, so the old ceiling protected nothing and forced anyone covering many niches to split the work across runs and pay the fixed start cost each time.

### 1.1.44 — Replacement works on wide searches; every rejection names its filter

- **Fixes the replacement of filtered creators on multi-query runs.** The allowance was compared against the *whole* candidate pool, so a run that searched several niches at once started above it and never ran a single replacement round: seven queries produce about 210 candidates, past the 80 allowed for `maxResults: 20`. A customer asking for 20 leads received 5, twice, and the summary blamed the filters. The allowance is now counted on top of whatever the first discovery pass found, and `RunBudget` remains the spending guard it was always meant to be.
- Every filter now records which one rejected a creator. The run summary gained `filterBreakdown`, and the warnings name the filter responsible for the most rejections along with the setting to change.
- **The no-progress and time ceilings no longer end a selective run.** A run that steadily examines and rejects creators is working, not stuck, yet the 60-second no-progress guard counted only *saved* leads, and the time ceiling was 30 seconds plus 3 per requested lead — 90 seconds for twenty. Both now measure the work the run actually has to do: progress counts every candidate finished, and the time and request ceilings are sized for the candidates the run is allowed to examine. The cost ceiling is unchanged and remains what decides when spending stops being worth it.
- Warns when creators **with an email** are requested while website scanning and social-profile scanning are both switched off, which removes the two places an address is usually published.

### 1.1 — Requested counts, deliverable emails and honest enrichment status

- Replaces creators rejected by the filters with fresh candidates from the searches that still hold unread pages, so a filtered run works at `maxResults` instead of stopping at the first pass (an email-only run delivered 33 of 60 before, 60 of 60 now). Replacement is bounded by six rounds and a pool of four times `maxResults`, and it resumes the existing search cursors rather than refetching page one. `candidateSupply` reports it.
- The email-only filter now requires an address whose domain can receive mail. Domains without mail records and disposable domains no longer satisfy it, and such an address can no longer become `primaryEmail` while a deliverable one exists. `primaryEmailValidationScope` states on every row that validation covers the domain's mail records, not the individual mailbox.
- Drops the video player endpoint for the rest of a run once it stops answering, instead of paying for a refusal plus the fallback on every video. Real player requests are now reported in `playerRequests`/`playerUsableResponses`; the previous status counts included the circuit breaker's own synthetic answers.
- `videoEnrichmentStatus` reports `completed` when every requested video was delivered. Videos recovered through the fallback route are counted in `videoSourceFallbacks` instead of marking the record partial, and `enrichmentQuality` in the run summary names which enrichment fell short.
- Skips YouTube's auto-generated `- Topic` channels during discovery in both search modes, and explains lookalike results in full: shared topics, the seed-derived query, the subscriber comparison, and a per-signal `similarityBreakdown`.

### 1.1 — Deep batch reliability and cost control

- Corrects the traffic allowance for full video pages, funded by saved results, so Deep does not inherit the light-mode one-MiB-per-result cap. Monetary and residential caps are unchanged.
- Sizes the Deep request guard for full-page attempts and fallback without changing lighter modes.
- Retries a page once on a fresh primary route if the automatic residential allowance is too small, without increasing the costly-response limit.
- Calibrates the residential cost reserve against observed platform transfer overhead, without raising event prices or the monetary ceiling.
- Avoids re-downloading confirmed bot-challenge pages and exposes actual traffic/request ceilings in the run summary.

### 1.1 — Deep restored as a heavy mode

- Restores full watch-page metadata and lower default concurrency (9) for Deep.
- Allows 90 seconds for video enrichment and 45 seconds per website phase; Deep has no one-minute throughput target.
- Gives Deep 120 seconds without a saved result while retaining the same monetary, traffic, residential and 1 GB protections.
- Fast, Balanced and pricing stay unchanged.

### 1.1 — resource protection and enrichment throughput

- Checks the remaining paid-result allowance before discovery and limits in-flight enrichment to results the run can afford.
- Bounds traffic, response size, idle work and estimated cost; residential fallback grows with saved leads instead of the requested limit.
- Uses smaller YouTube metadata responses, preserves native language and pinned-comment entry points, and streams completed leads.
- Overlaps full linked-site crawling, latest pinned comments and bounded MX validation with video work; shares destination limits and cancels unused requests.
- Loads linked pages alongside video analysis, avoids repeat requests to inaccessible public profiles, and improves HTML and language processing.
- Preserves available candidates when a later search page fails, handles duplicate channel aliases before enrichment, and reports query and partial-enrichment diagnostics.
- Excludes checkout and newsletter forms from outreach forms. Prices and the Actor-start event are unchanged.

### 1.0.55 — faster batches and results-first output

- Reduced repeated connection setup so large creator batches finish substantially faster.
- Streamed completed creator records immediately and added bounded discovery backfill so failed or filtered candidates do not consume the saved-result limit.
- Added request, discovery, processing and first-result timings to the run summary.
- Tuned Fast mode for 40–50 creator batches at 1 GB while keeping Balanced and Deep enrichment explicit.
- Opens the Leads dataset first after a run; Run summary remains available as the final output view.

### 1.0.50 — protected-email public-source recovery

- Deep mode now checks linked secondary YouTube channels and bounded searches of older contact-oriented uploads when a channel advertises a protected business email but no public address is otherwise found.
- Linked websites now expose emails from JSON-LD/structured state, email-bearing data attributes, common base64/ROT13/RTL obfuscation, Cloudflare protection, and public contact forms.
- Website crawling probes bounded common contact/about/business paths when navigation does not link them.
- Added `protectedEmailStatus`, alternative-source diagnostics, linked-channel and archived-video scan status, and contact-form fields.
- Exact CAPTCHA-protected YouTube values remain untouched; stale third-party lists and generated mailbox guesses are rejected.

### 1.0.46 — Verified regions, content language, and CRM lead tiers

- 🌍 Tested all 30 advertised YouTube markets and clarified that localization is a ranking hint, not proof of creator location.
- 📍 Added optional strict regional matching against the channel's public YouTube country plus explicit match evidence on every row.
- 🗣️ Added content-language fields from recent YouTube video metadata, with a no-request local text fallback plus explicit source, confidence, and coverage.
- 🏆 Added CRM-friendly A–D lead tiers, lead status, recommended contact method, and human-readable qualification reasons without extra requests.
- 🛟 Unknown API region codes now fall back safely to US instead of silently pretending to be supported.

### 1.0.44 — Simpler input, explainable matching, and evidence

- ⚡ Replaced the crowded default Input with Fast, Balanced, and Deep enrichment presets while keeping every advanced API field backward-compatible.
- 🎯 Made contactable creators the safe default so completely contactless rows are not saved unless requested.
- 🧬 Added a dedicated Lookalike creators view with a similarity score, shared topics, audience-size comparison, and human-readable matching reasons.
- 🤝 Added structured sponsor evidence, source video URLs, confidence, evidence snippets, and sponsored-video rate.
- 📘 Rebuilt the README with click-by-click guidance, complete input/output schemas, fictional examples, pricing, and troubleshooting.
- 🖼️ Added a compact original cover in the existing Store visual style.

### 1.0.38 — Reliability, truthfulness, and run experience

- ✅ Added terminal status messages and a structured run summary with contact coverage, warnings, and next steps.
- 🛟 Invalid YouTube sources now finish cleanly with `INPUT_NEEDS_ATTENTION` and an actionable explanation.
- ✔️ Updated verified-channel detection for YouTube's current page-header format.
- 🔎 Made search subscriber parsing compatible with both current and legacy channel renderers.
- 📞 Parse local phone numbers using the channel country or selected region instead of always assuming the US.
- 🎯 Prevented sponsor websites from being mistaken for creator contact sources and improved sponsor-brand deduplication.
- 📧 Prefer contact evidence published directly on the channel when choosing the primary email.
- 💸 Aligned the source manifest with the lightweight 512 MB Cloud default, moved the standard route to a cheaper datacenter proxy, and prevented optional website/social scans from escalating to residential proxy traffic.

### 1.0.20 — Full-mode QA and data quality

- 🧪 Re-tested direct, keyword, and lookalike discovery modes in Apify Cloud.
- ✉️ Added Cloudflare-protected public email decoding for linked contact pages.
- 🎯 Filtered seed impersonators and extremely small false-positive lookalikes.
- 🌐 Prevented sponsor links from being reported as a creator's website.
- 📊 Replaced absent table cells with explicit `null` or empty collections.
- 🔎 Corrected subscriber and video counts parsed from channel search results.

### 1.0.18 — Store experience refresh

- 🎨 Added a polished, guided Input UI with clear sections, examples, and friendly labels.
- 📘 Rebuilt the README with quick-start recipes, output documentation, status explanations, and API examples.
- ✨ Added structured progress logs with readable milestones and a final run summary.
- 🔎 Improved Store title, description, SEO metadata, and discovery categories.
- 🖼️ Added an original Actor icon designed for the Apify Store.

### 1.0.17 — Performance and enrichment

- ⚡ Improved 50-query performance with parallel enrichment and bounded time budgets.
- 🎬 Added richer recent-video analysis, upload frequency, and sponsor signals.
- 📌 Added pinned-comment contact discovery.
- 🌐 Preserved partial website contacts when a page reached its scan deadline.
- 🧹 Hardened URL handling, retry behavior, proxy fallback, and sponsor normalization.

## 1.0.49

- Added channel-link labels, website/social URL lists, contact status, contact counts, and source-scan diagnostics.
- Improved public email coverage by checking creator-owned domains referenced in video descriptions while excluding unrelated sponsor sites.
- Prioritized official and contact-oriented pages during bounded website enrichment.
