TechCrunch Scraper — Startup Funding Leads
Pricing
from $20.00 / 1,000 results
TechCrunch Scraper — Startup Funding Leads
Turn fresh TechCrunch coverage into ready-to-work startup leads. Get company, raise amount, stage, investors plus website, email, phone and socials for every company. Five collection sections, real-time streaming to your dataset, and webhooks for CRM or Slack.
Pricing
from $20.00 / 1,000 results
Rating
0.0
(0)
Developer
Emmanuel
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
TechCrunch Real-Time Data Scraper
Turn TechCrunch funding coverage into ready-to-work startup leads — in real time, straight to your dataset.
Every run watches fresh startup and venture coverage and converts it into structured lead records: company name, raise amount (normalized to USD), funding stage, investor list, the exact funding sentence, story link, publish date — plus the company's website, public email, phone and social profiles when Enable lead details is on.
Results stream to your dataset the moment each company is found, so you can pipe them into a CRM, Google Sheet, or webhook while the run is still going.
⚠️ Paid only: free Apify accounts are limited to a small sample of results per run (2 by default). To unlock full, unlimited exports you need a paid Apify plan. See FAQ.
✨ What it does
| Outcome | How it helps you |
|---|---|
| Fresh funding rounds, structured | Company, amount in USD, stage (Seed → Series G, Growth), and investor names extracted from every funding announcement. |
| Lead details on every company | With Enable lead details on, each record carries the company's website, public email(s), phone number(s) and social profiles when they can be found. |
| Everything is exported | No filters, no silent drops: every company found is saved, even when no contact details exist. Output per run is complete and predictable. |
| Real-time streaming | Records are pushed to the dataset one by one as they are found. Start a 10,000-item run and work your CRM from minute one. |
| Five independent sections | Recently funded, startup news, keyword search, investor-portfolio prospecting, and single-company research — each with its own toggle, keywords and company cap. |
| Respects your budget | The actor honors your run spending limit and stops early instead of overspending. |
| Webhooks | Push every record to Slack, Zapier, Make, n8n or your own endpoint the moment it is saved. |
Example record
{"featureType": "funding","company": "Valar Atomics","amountUsd": 130000000,"amountDisplay": "$130 million","stage": "Series A","investors": ["Founders Fund", "Paradigm"],"summary": "Valar Atomics, a nuclear startup, closed on $130 million in new funding.","title": "Valar Atomics raises $130M Series A to build nuclear reactors","url": "https://techcrunch.com/2026/09/20/valar-atomics-raises-130m-series-a/","articleSummary": "The startup builds small nuclear reactors…","publishedAt": "2026-09-20T12:00:00.000Z","author": "Margaux MacColl","category": "Climate & Energy","source": "TechCrunch","region": "US","website": "https://www.valaratomics.com","websiteDomain": "valaratomics.com","websiteConfidence": "high","email": "hello@valaratomics.com","emails": ["hello@valaratomics.com"],"phone": null,"phones": [],"socials": { "linkedin": "https://www.linkedin.com/company/valar-atomics", "twitter": "https://x.com/valaratomics" },"linkedinUrl": "https://www.linkedin.com/company/valar-atomics","twitterUrl": "https://x.com/valaratomics","enriched": true,"hasLeadDetails": true,"confidence": "high","scrapedAt": "2026-09-25T14:00:00.000Z"}
👥 Who it is for
- VCs & analysts — track every fresh raise in your thesis space the day it happens.
- Sales & SDR teams — companies that just raised money have budget. Get the company, the round, the investors and the contact details in one row.
- Recruiters & executive search — funded companies hire aggressively. Reach founders the week the announcement lands.
- Growth & partnerships teams — discover startups entering your category, with their investors attached.
- Journalists & analysts — a structured, filterable record of who raised what, from whom.
- Lead-gen agencies — resell fresh, contact-complete startup lists with predictable run costs.
- Fundraisers — see which investors are actively writing checks and in which stages.
- Fractional CFOs & advisors — monitor newly funded companies in your niche for outreach.
🧭 Use cases
- Morning funding digest — run daily with Recently funded and pipe results to Slack via webhook.
- "Just raised" outbound — SDR teams filter for
stage = SeedoramountUsd > 1,000,000and contact founders while intent is hot. - Investor portfolio mapping — enable Backed by investor, list your target funds, and get every portfolio company that appears in the news.
- Competitive intelligence — Company research mode follows specific competitors and surfaces their coverage.
- Market maps — Search by keyword ("fintech", "defense tech", "climate") builds category snapshots with funding data attached.
- CRM enrichment — webhook pushes land directly in HubSpot/Pipedrive/Airtable with funding signals on the contact.
- Google Sheets dashboards — point a webhook at Zapier/Make and watch the sheet fill in real time.
- Weekly investor activity report — track which investors are leading rounds across hundreds of stories.
- Recruiting triggers — funded companies = hiring budgets. Sync to your ATS or LinkedIn workflow.
- M&A sourcing — track fast-growing companies and their backers over time.
- Fundraising benchmarks — filter by stage to build a dataset of comparable rounds (amount, investors, timing).
- PR prospecting — see who covers your niche and which companies they write about.
🚀 Quick start
- Add the actor to your Apify account.
- Leave the defaults — Recently funded + Enable lead details are pre-tuned.
- Run. Records stream into your dataset as they are found.
- (Optional) Set a Webhook URL to push each record to your CRM or Slack while the run continues.
Free plan note: free Apify accounts get a small sample per run (2 results by default). Upgrade to a paid Apify plan for full, unlimited output.
📥 Input schema (every field)
What to collect — 5 independent sections
Each section works on its own: toggle it, optionally focus it with keywords, and cap it with its own max-companies value. Enable any combination — every enabled section exports up to its cap.
| # | Section | Toggle | Keywords (optional) | Cap |
|---|---|---|---|---|
| 1 | Recently funded | enableRecentlyFunded | recentlyFundedKeywords | maxRecentlyFunded |
| 2 | Startup news | enableStartupNews | startupNewsKeywords | maxStartupNews |
| 3 | Keyword search | enableSearch | keywords (required here) | maxSearch |
| 4 | Backed by investor | enableBackedBy | backedByKeywords (+ investors, required) | maxBackedBy |
| 5 | Company research | enableCompanyResearch | companyResearchKeywords (+ companies, required) | maxCompanyResearch |
1. Recently funded — on by default
| Field | Type | Default | Description |
|---|---|---|---|
enableRecentlyFunded | boolean | true (prefill) | Collect companies from recent funding-announcement coverage (the funding & exits section). |
recentlyFundedKeywords | string list | [] | Optional. Focus the section on stories matching these keywords (e.g. "Series A", "AI"). Empty = all funding coverage. |
maxRecentlyFunded | integer | 10 | Maximum companies exported from this section. |
2. Startup news
| Field | Type | Default | Description |
|---|---|---|---|
enableStartupNews | boolean | false | Collect companies from general startup coverage (products, launches, milestones). |
startupNewsKeywords | string list | [] | Optional. Focus the section on stories matching these keywords (e.g. "launch", "app"). Empty = all startup coverage. |
maxStartupNews | integer | 10 | Maximum companies exported from this section. |
3. Keyword search
| Field | Type | Default | Description |
|---|---|---|---|
enableSearch | boolean | false | Find stories about specific keywords and extract the companies they cover. Requires keywords. |
keywords | string list | ["artificial intelligence"] (prefill) | Required. Topics to search for (e.g. "fintech", "climate"). Up to 8 keywords per run are used. |
maxSearch | integer | 10 | Maximum companies exported from this section. |
4. Backed by investor
| Field | Type | Default | Description |
|---|---|---|---|
enableBackedBy | boolean | false | Find companies backed by specific investors. Requires investors. |
investors | string list | [] | Required. Investor names to look for (e.g. "Sequoia", "a16z", "Y Combinator"). Up to 8 per run. |
backedByKeywords | string list | [] | Optional. Extra search keywords to widen the investor search (e.g. "raises", "portfolio"). |
maxBackedBy | integer | 10 | Maximum companies exported from this section. |
5. Company research
| Field | Type | Default | Description |
|---|---|---|---|
enableCompanyResearch | boolean | false | Follow coverage of specific companies. Requires companies. |
companies | string list | [] | Required. Company names to research (e.g. "Stripe", "Ramp"). Up to 8 per run. |
companyResearchKeywords | string list | [] | Optional. Extra search keywords to widen the company coverage (e.g. "funding", "partnership"). |
maxCompanyResearch | integer | 10 | Maximum companies exported from this section. |
Lead details
| Field | Type | Default | Description |
|---|---|---|---|
enableLeadDetails | boolean | true | Add each company's public website, email, phone and social profiles to its record. Every company is still exported — companies where details cannot be found are included with empty fields. Adds a little extra time per company. |
Output & limits
Each section above carries its own max-companies cap — there is no global cap. Total output per run = the sum of the enabled sections' caps (lower for free accounts — see FAQ).
| Field | Type | Default | Description |
|---|---|---|---|
maxAgeDays | integer | 30 (prefill) | Skip coverage older than this many days. 0 = no age limit. |
region | string | "US" | Region label stored on each record (e.g. US, EU, GLOBAL). |
Webhook
| Field | Type | Default | Description |
|---|---|---|---|
webhookUrl | string | "" | Optional. Each record is always saved to the dataset; when set, it is also POSTed to this URL as JSON in real time. Works with Slack incoming webhooks, Zapier, Make, n8n, or your own endpoint. |
webhookFormat | select | "json" | json = full record object. slack = Slack-friendly message payload. |
Connection
| Field | Type | Default | Description |
|---|---|---|---|
proxyConfiguration | object | Apify residential, US | Apify residential proxy is enabled by default for reliable collection. Change only if you need a different country or custom proxy URLs. |
Section behavior details
- Recently funded also reads the latest-news feed so roundups that mention multiple companies produce one row per company.
- Startup news (alone) extracts the primary company from non-funding headlines too.
- Backed by investor keeps only companies whose investor list matches your names (substring match, case-insensitive).
- Company research keeps only your listed companies — and still surfaces a company if it appears in the story even without a funding sentence.
- Sections combine: e.g. Recently funded + Backed by investor gives you funded companies filtered to your investor list.
- Keywords within a section focus it: when keyword-filtered stories exist, only those are read; if none match, the section falls back to its unfiltered coverage so it still exports.
- The free-tier cap (2 by default) is scaled across your enabled sections so every enabled section still exports at least one company.
- Every record carries a
sectionfield naming the collection section that produced it — filter the dataset on it to isolate one mode (the By section view groups results this way).
📤 Output schema (field by field)
Each dataset row is one company. Fields marked (lead details) are filled when enableLeadDetails is on and the information is found — they stay null/empty otherwise, and the record is still exported.
| Field | Type | Description |
|---|---|---|
featureType | string | Always "funding". Lets you filter dataset views. |
section | string | Collection section that produced this row: recentlyFunded, startupNews, search, backedBy or companyResearch. Filter on it to isolate one mode (see the By section view). |
company | string | Company name as written in the coverage. |
amountUsd | integer | null | Raise amount normalized to US dollars (130000000). |
amountDisplay | string | null | Amount as written ("$130 million", "$1.3 billion"). |
stage | string | null | Canonical stage: Pre-Seed, Seed, Series A…Series G, Growth, Debt, Grant. |
investors | string[] | Investor names mentioned with the raise (lead investor first when the headline names one). |
summary | string | null | The exact sentence describing this company's raise. |
title | string | Title of the story the company was found in. |
url | string | Link to that story. |
articleSummary | string | null | Short description of the story. |
publishedAt | string | null | ISO timestamp of the story. |
author | string | null | Story author. |
category | string | null | Story category/section. |
source | string | Always "TechCrunch". |
region | string | Your region input value. |
website (lead details) | string | null | The company's own website. |
websiteDomain (lead details) | string | null | Domain of that website. |
websiteConfidence (lead details) | string | null | high / medium / low — how confident the match is. |
email (lead details) | string | null | Best public contact email found. |
emails (lead details) | string[] | All public emails found. |
phone (lead details) | string | null | Best public phone number found. |
phones (lead details) | string[] | All public phone numbers found. |
socials (lead details) | object | null | Public profile URLs keyed by platform (linkedin, twitter, facebook, instagram, youtube, tiktok, yelp). |
linkedinUrl (lead details) | string | null | LinkedIn company page. |
twitterUrl (lead details) | string | null | X/Twitter profile. |
enriched (lead details) | boolean | null | true when a website and/or contact details were found. |
hasLeadDetails (lead details) | boolean | null | Same as enriched — convenient for filtering. |
confidence (lead details) | string | null | Overall confidence in the lead details (high/medium/low). |
scrapedAt | string | ISO timestamp when this record was created. |
Run summary (Key-Value store, key OUTPUT)
| Field | Type | Description |
|---|---|---|
companiesExported | integer | Rows written to the dataset. |
perSection | object | Companies exported per section, keyed by section id — e.g. { "recentlyFunded": 2, "search": 1 }. |
articlesRead | integer | Stories read successfully. |
articlesSkipped | integer | Stories unavailable and skipped. |
leadDetailsEnabled | boolean | Whether lead details were requested. |
companiesWithLeadDetails | integer | Companies that got website/contact details. |
spendingLimitReached | boolean | True when the run stopped because your spending limit was hit. |
errors | string[] | Fixed, human-readable issues (never internal details). |
paywall.detected | boolean | Platform pay-status variables were present. |
paywall.isPaying | boolean | The user is on a paid Apify plan. |
paywall.pricingTier | string | null | FREE/BRONZE/…/DIAMOND when known. |
paywall.blocked | boolean | Run was blocked (free account + block mode). |
paywall.limited | boolean | Free-tier result cap was applied. |
🔔 Webhook setup
- Copy your endpoint URL — Slack incoming webhook, Zapier Catch Hook, Make webhook, n8n webhook, or your own API.
- Paste it into Webhook URL in the actor input.
- Choose the format:
json(full record — best for CRMs and automation tools) orslack(readable message — best for channels). - Run. Every record is POSTed as JSON the moment it is saved, so downstream systems stay in sync during long runs.
JSON payload = the full record object (see Output schema).
Slack payload example:
{ "text": ":rocket: *Valar Atomics*\n*Round:* Series A • *Amount:* $130 million\n*Investors:* Founders Fund, Paradigm\n*Website:* https://www.valaratomics.com\n<https://techcrunch.com/...|Source coverage>" }
Notes:
- Webhook delivery is best-effort: a failed delivery never blocks the dataset write or the run.
- Payloads contain only the record fields documented above — no internal or debug data.
🤖 MCP usage
The actor works with Apify's MCP Server so AI agents can run it and read results.
- In your MCP client (e.g. Claude Desktop), connect to Apify's MCP Server.
- Add this actor to the server's tool list (
actors/toolfolio/techcrunch-real-time-data-scraper). - The agent can then start runs and read the dataset: “Find startups that raised a Series A in the last 14 days and give me the ones with a contact email.”
Tips for agent workflows:
- Keep the per-section caps modest (10–50 companies total) for quick interactive runs.
- Read the run's OUTPUT record for a machine-readable summary (counts + status).
- Use
view=funding/view=leads/view=bySectiondataset links from the output for targeted reads.
⚡ Performance & reliability
- Streaming pipeline — records are pushed one at a time; memory stays flat even on 10,000-item runs.
- Batched lead details — companies are enriched in parallel batches, which keeps the per-company overhead small.
- Proxy-backed lead details — website and contact lookups run through the same residential proxy as collection, with session rotation per attempt.
- Parallel story reading — stories are fetched four at a time through a bounded worker pool, while polite pacing keeps the outbound rate steady (roughly a 4× faster run at equal politeness).
- Timeout — the actor's platform timeout is set to 10,000 seconds; long runs are fine.
- Budget-aware — the actor checks your spending limit after every record and stops gracefully when reached.
- Predictable runtime — because results are never filtered by completeness, runtime scales linearly with the number of stories/companies you request.
❓ FAQ
Is this really paid-only? Free Apify accounts can run it but only get a small sample of results per run (2 by default) — enough to verify data quality. To get full, unlimited data you need a paid Apify plan. This is stated in the input description too.
Why did my free run stop after 2 results? That's the free-tier cap. The log shows a warning and the run ends gracefully with a status message. Upgrade your Apify plan for full output.
How current is the data?
As current as the coverage: each run reads the latest published stories, and maxAgeDays can restrict everything to the last N days. Records include the story's publish date.
Will I get every company mentioned in a roundup? Yes — multi-company stories produce one row per distinct company found.
What if a company has no website or email?
It is still exported — website, email and related fields stay null/empty and hasLeadDetails is false. Lead details are added when found; they never filter results.
How accurate are the lead details?
Each website match carries a websiteConfidence (high/medium/low) and an overall confidence on the record. Email/phone values are only reported when actually published on the company's own web presence — nothing is invented.
Can I track multiple investors at once? Yes — list up to 8 investor names with Backed by investor enabled.
Can I run several modes in one run? Yes, all modes combine. Results are deduplicated per story.
How does pricing work? The actor charges per result (pay-per-event), so you pay for what you get. The actor always respects your run spending limit: when the limit is reached, it stops cleanly instead of charging beyond it.
The run says "spending limit reached" — what now? You set a maximum cost for the run (or your plan did). The actor stopped at that point. Raise the run's max cost or lower the per-section caps to fit.
Something failed mid-run — do I lose everything? No. Records already pushed to the dataset are safe. Individual stories that fail are skipped with a friendly note in the log.
Do you support languages other than English? Coverage is English-language tech news. Company names and contact details are returned as published.
🛠️ Owner configuration (environment variables)
These are for the actor owner, settable in Apify Console → Actor → Environment variables, no code change needed:
| Variable | Default | Description |
|---|---|---|
FREE_TIER_MODE | limit | limit = free accounts get capped output; block = free accounts get no output and an upgrade message. |
FREE_TIER_MAX_ITEMS | 2 | Dataset cap for free (non-paying) accounts in limit mode. |
PPE_EVENT_NAME | result | Pay-per-event event name charged per exported record (must match your monetization settings). |
📚 Resources
- Apify Console — run the actor and read your datasets.
- Apify SDK docs — dataset, webhooks and monetization details.
- Dataset views: Overview, Funding rounds, Leads with details, By section — switch between them on the run's Storage tab.