Crunchbase Scraper — Funded Companies & Leads
Pricing
from $3.00 / 1,000 results
Crunchbase Scraper — Funded Companies & Leads
Collect recently funded companies from Crunchbase news in real time — amount, stage, investors, plus websites, emails and phones. Track keywords, investors, or competitors. Results stream to your dataset as JSON. Free plans are limited; upgrade for full data.
Pricing
from $3.00 / 1,000 results
Rating
0.0
(0)
Developer
Emmanuel
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
Crunchbase Real-Time Data
Real-time Crunchbase funding intelligence, streamed to your dataset as it is collected. Track recently funded companies, monitor startup news, search funding coverage by keyword, investor, or company — and optionally add each company's public website, email, phone, and social profiles. One flat JSON row per company, ready for your CRM, spreadsheet, database, or AI pipeline.
⚡ Streaming output • 🎯 Optional lead details • 🧩 Flat, predictable JSON • 💸 Pay only for what you collect
🔒 Paid-only note: on an Apify free plan this Actor exports a small sample of results (default: 2 per run) and finishes with a clear message. Upgrade to any paid Apify plan for full, unlimited data. See Free tier.
What it does
The Actor reads public Crunchbase funding-news coverage and turns every company mentioned in it into a structured lead: company name, funding amount, stage, investors, the announcement it came from, and (with lead details enabled) the company's own website and contact points.
| Crunchbase Real-Time Data | Typical alternative scrapers | |
|---|---|---|
| Speed | Stories are read in parallel batches; results stream immediately | Serial, slow, one item at a time |
| Memory | 512 MB default | 2–4 GB+ |
| Output | Structured JSON, LLM-ready | Often messy HTML |
| Runtime predictability | Every company is saved — nothing filtered out | Unpredictable when filters drop rows |
Streaming by design: each company is pushed to the dataset the moment it is processed — you watch results arrive in real time, and RAM stays light even on very long runs.
Who it is for
- Sales teams & SDRs — reach founders within hours of their funding announcement, when budgets are fresh and inboxes are open.
- Growth & RevOps teams — build a always-fresh pipeline of funded companies matching your ICP.
- VC & PE analysts — monitor portfolios, track investor activity, and map market segments.
- Recruiters — companies that just raised are hiring; get the list first.
- Agencies & consultants — deliver funding-tracker reports for clients without manual research.
- Data engineers & AI builders — clean, stable JSON schema for pipelines, warehouses, and RAG.
Use cases
- 🎯 Funding-based lead generation — collect companies that raised in the last N days, with contact details attached, and reach out while the raise is news.
- 📈 Investor portfolio discovery — ask "show me everything backed by Y Combinator / a16z / Sequoia this month".
- 🔎 Market & keyword monitoring — follow coverage of "ai infrastructure", "fintech brazil", "climate robotics" and capture every named company.
- 🏢 Company tracking — watch a list of competitors or prospects and get a row for every news mention.
- 📰 Daily funding digest — schedule the actor each morning and pipe results to Slack or Sheets via webhook.
- 🧠 AI agent / RAG feeds — flat
_externalId-keyed rows upsert cleanly into vector stores and CRMs. - 📊 Investor activity analytics — aggregate rounds by stage, size, and lead investor over time.
- 🌍 Regional deal flow — tag every row with your own region label and split datasets by geography.
- 🔁 CRM enrichment — match
companyagainst your existing accounts and backfill funding data (amountUsd,stage,investors). - 🏆 Top-of-funnel for agencies — prospect businesses by funding event instead of static firmographics.
Collection modes — combine any of them
| Mode | Turn on with | You get |
|---|---|---|
| Recently funded (default) | enableRecentlyFunded: true | Companies from the latest funding announcements, with amount, stage, and investors. |
| Startup news | enableStartupNews: true | Companies named in general startup-news coverage (even without a round). |
| Keyword search | enableSearch: true + keywords[] | Companies from coverage matching your keywords. |
| Backed by investor | enableBackedBy: true + investors[] | Companies backed by the investors you list. |
| Company research | enableCompanyResearch: true + companies[] | Every news mention of the companies you track. |
All enabled modes run in one pass — stories are deduplicated across modes and every distinct company is saved once.
Lead details (contact enrichment)
Enable Lead details (enableLeadDetails, on by default) and the Actor adds each company's public website, email, phone, and social profiles (LinkedIn, X, Facebook, Instagram, and more) where findable.
Important: this is an enable option, not a filter. Every collected company goes to the dataset — companies where no details can be found still appear, with empty contact fields. Nothing is dropped, so runtime per 1,000 companies stays predictable. Lead discovery adds a little extra time per company.
Output fields added by lead details: website, email, emails[], phone, phones[], socials{}, linkedinUrl, facebookUrl, instagramUrl, twitterUrl, enriched, enrichmentLevel, confidence-adjacent level via enrichmentLevel.
Quick start
- Open the Actor in Apify Console.
- Leave Recently funded checked (or pick other modes).
- Keep Enable lead details checked to get websites, emails, and phones.
- Set the per-mode max controls — start with
10to test, then raise (1000+) for real runs. Your run's total is the sum of the budgets you enable. - Press Start and watch rows stream into the dataset.
- Export as JSON / CSV / Excel, or pull via the Apify API, integrations, or MCP.
Example — this morning's funding rounds with contacts
{"enableRecentlyFunded": true,"enableLeadDetails": true,"maxRecentlyFunded": 50}
Example — companies backed by specific investors
{"enableBackedBy": true,"investors": ["a16z", "Sequoia", "Y Combinator"],"enableLeadDetails": true,"maxPerInvestor": 100}
Example — keyword monitoring for a niche
{"enableSearch": true,"keywords": ["ai infrastructure", "robotics"],"maxAgeDays": 7,"maxPerKeyword": 50}
Example — track a competitor list
{"enableCompanyResearch": true,"companies": ["Stripe", "Ramp", "Brex"],"enableLeadDetails": false,"maxPerCompany": 200}
Input reference (full schema)
The input form is organized into sections — one per collection mode, then lead details, scope & limits, and connection.
🔎 Recently funded / 🔑 Keyword search / 🏛️ Backed by investor / 🏢 Company research / 📰 Startup news
Each collection mode has its own section with its own "max to collect" control — your run's total is the sum of the budgets you set in the sections you enable. There is no separate global cap.
| Field | Type | Default | Description |
|---|---|---|---|
enableRecentlyFunded | boolean | true | Collect companies from the latest funding announcements — with amount, stage, and investors. |
maxRecentlyFunded — Max companies from latest funding news | integer | 10 | How many companies to collect from the latest announcements. |
enableSearch | boolean | false | Search funding coverage for your keywords. |
keywords | string[] | prefill: ai infrastructure, fintech | One keyword per row (max 8 used per run). |
maxPerKeyword — Max companies per keyword | integer | 10 | Per-keyword budget before moving to the next keyword. 3 keywords × 10 = 30 companies. |
enableBackedBy | boolean | false | Find companies backed by the listed investors (portfolio discovery). |
investors | string[] | prefill: a16z, Y Combinator | One investor per row (max 8 used per run). |
maxPerInvestor — Max companies per investor | integer | 10 | Per-investor budget before moving to the next investor. |
enableCompanyResearch | boolean | false | Track specific companies in the news. |
companies | string[] | prefill: Stripe, Ramp | One company name per row — every news mention becomes a row. |
maxPerCompany — Max mentions per company | integer | 10 | Per-company budget of news mentions. |
enableStartupNews | boolean | false | Also collect companies named in general startup-news coverage, even without a round. Companies found here count toward the recently-funded budget. |
🎯 Lead details
| Field | Type | Default | Description |
|---|---|---|---|
enableLeadDetails | boolean | true | Adds each company's public website, email, phone, and social profiles where available. Additive only — never filters. Adds a little extra time per company. |
⚙️ Scope & limits
| Field | Type | Default | Description |
|---|---|---|---|
maxArticles — Max stories to read | integer | 100 | Maximum news stories to read (1–500). Each story may yield several companies. |
maxAgeDays — Only stories from the last N days | integer | 0 | Only stories published within the last N days. 0 = no age limit. |
region — Region tag | string | "US" | Free-text label saved on every row for splitting datasets later. |
webhookUrl — Webhook URL | string | "" | Optional; each record is POSTed there as it is saved (in addition to the dataset). |
webhookFormat — Webhook format | json | slack | "json" | json = full record; slack = Slack message payload. |
🌐 Connection
| Field | Type | Default | Description |
|---|---|---|---|
proxyConfiguration | object | Apify residential US | Proxy settings; Apify residential (US) is enabled by default. |
At least one collection mode must be enabled, and the lists required by the enabled modes (keywords / investors / companies) must be non-empty.
Output reference (full schema)
Every row is a flat JSON object. One row per company.
| Field | Type | Description |
|---|---|---|
_externalId | string | Stable deduplication key — safe for CRM/warehouse upserts. |
featureType | string | Always "funding". |
mode | string | Collection mode tag ("funding"). |
company | string | Company name. |
title | string | Headline of the story the company came from. |
url | string | null | Link to the announcement. |
source | string | Data source label ("Crunchbase News"). |
publishedAt | string | null | Story publish time (ISO 8601). |
amountUsd | number | null | Round size normalized to USD (e.g. 25000000). |
amountDisplay | string | null | Human-readable amount (e.g. "$25 million"). |
stage | string | null | Normalized stage (Seed, Series A, Series B, Growth, …). |
investors | string[] | Investor names extracted from the announcement. |
summary | string | null | The funding sentence the company was extracted from. |
articleSummary | string | null | Article description/summary. |
category | string | null | Story category. |
author | string | null | Story author. |
region | string | Your region tag from input. |
website | string | null | Company website (lead details). |
email | string | null | Primary contact email (lead details). |
emails | string[] | All found emails (lead details). |
phone | string | null | Primary phone (lead details). |
phones | string[] | All found phones (lead details). |
socials | object | null | Map of social profile links (lead details). |
linkedinUrl / facebookUrl / instagramUrl / twitterUrl | string | null | Direct social links (lead details). |
enriched | boolean | null | true when at least one contact detail was found. |
enrichmentLevel | string | null | "partial" (some details) or "list_only" (none found). |
scrapedAt | string | ISO 8601 UTC time the row was collected. |
Example row:
{"_externalId": "crunchbase-acme-ai","featureType": "funding","mode": "funding","company": "Acme AI","title": "Acme AI raises $25M Series B to scale agents","url": "https://news.crunchbase.com/startups/acme-ai-series-b/","source": "Crunchbase News","publishedAt": "2026-09-24T08:00:00.000Z","amountUsd": 25000000,"amountDisplay": "$25 million","stage": "Series B","investors": ["Felicis", "Founders Fund"],"summary": "Acme AI raised $25 million in a Series B led by Felicis.","articleSummary": "The round brings Acme AI's total funding to $40M.","region": "US","website": "https://acme.ai","email": "hello@acme.ai","emails": ["hello@acme.ai"],"phone": "(415) 555-0134","socials": { "linkedin": "https://www.linkedin.com/company/acme-ai" },"linkedinUrl": "https://www.linkedin.com/company/acme-ai","enriched": true,"enrichmentLevel": "partial","scrapedAt": "2026-09-25T12:00:00.000Z"}
Run summary (OUTPUT tab): after every run the OUTPUT key-value store holds a summary — modes run, stories read, companies collected, the applied cap, webhook setting, any non-fatal errors, and a paywall object (detected, isPaying, pricingTier, limited, blocked, freeTierMaxItems) so the applied restrictions are transparent.
Webhooks
Set Webhook URL in the input and every record is POSTed to your endpoint in real time, in addition to the dataset — perfect for Zapier, Make, n8n, Slack incoming webhooks, or your own API.
- Records are always saved to the dataset first; webhook delivery is best-effort and never blocks or fails the run.
- Failed deliveries are retried once and then skipped with a single log line — your data is still safe in the dataset.
- Format is chosen with Webhook format.
JSON payload — the full record object (same shape as the dataset rows above).
Slack payload:
{"text": ":moneybag: *Acme AI*\n*Round:* Series B • *Amount:* $25 million\n*Investors:* Felicis, Founders Fund\n*Email:* hello@acme.ai\nhttps://acme.ai"}
Setup: paste an HTTPS URL into the Webhook URL field (e.g. a Zapier catch hook, a Slack incoming-webhook URL, or your own endpoint) and run. Test the endpoint with any POST receiver before production runs.
Webhook payloads contain only the documented record fields — nothing else is attached.
Apify API & MCP usage
Fetch results via the Apify API
# Using an Apify API tokencurl -H "Authorization: Bearer $APIFY_TOKEN" \"https://api.apify.com/v2/acts/YOUR_ACTOR_ID/runs/last/dataset/items?clean=true&format=json"
Run from the API
curl -X POST "https://api.apify.com/v2/acts/YOUR_ACTOR_ID/runs" \-H "Authorization: Bearer $APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{ "enableRecentlyFunded": true, "enableLeadDetails": true, "maxRecentlyFunded": 50 }'
MCP (Model Context Protocol)
The actor works with Apify's MCP server so AI agents (Claude, Cursor, custom agents) can call it as a tool:
- Connect to
https://mcp.apify.com(see Apify MCP docs for auth). - Call
act_YOUR_ACTOR_ID(or search the Store) with the input schema above. - Results come back as dataset items — flat JSON your agent can read directly, then fetch the full dataset for the rest.
Example agent call:
{"enableRecentlyFunded": true,"enableLeadDetails": true,"maxAgeDays": 3,"maxRecentlyFunded": 20}
Scheduling & automation
- Schedules: run every morning at 8:00 to get the last 24 hours of funding news — combine with
maxAgeDays: 2and the webhook for a hands-off digest. - Integrations: the default dataset plugs into Google Sheets, Airtable, Excel, and 100+ Apify integrations.
- Deduplication:
_externalIdis stable per company — store rows with upsert semantics to keep a rolling, deduplicated funding feed.
Controlling scope & cost
| Field | Effect |
|---|---|
Per-mode budgets (maxRecentlyFunded, maxPerKeyword, maxPerInvestor, maxPerCompany) | The cost levers. Your run's total is the sum of the budgets in the sections you enable. |
maxArticles | How many stories to read. Each story may yield several companies. |
maxAgeDays | Keep only recent stories — smaller windows finish faster. |
enableLeadDetails | Turn off to skip contact discovery (faster, still full company output). |
Free tier (Apify free plan)
On an Apify free plan, this Actor exports a small sample of results — a maximum of 2 companies per run (owner-configurable) — then finishes gracefully with a clear message telling you to upgrade. The run never errors and your sample stays in the dataset.
| Plan | What you get |
|---|---|
| Apify free plan | ⚠️ Capped sample (default 2 results per run), then a clean stop with an upgrade message. |
| Any paid Apify plan | ✅ Full, uncapped output up to your per-mode budgets. |
Upgrade to a paid Apify plan to get full, unlimited data. The cap is applied automatically from your plan — you don't configure anything, and paying users see one confirmation line in the log ("Paying user — full output.") and no caps.
The owner can switch free handling between a capped sample and a hard block (and tune the cap) via the Actor's environment variables — end users never need to configure anything.
FAQ
Which data source does this use? Public Crunchbase News funding coverage — the same announcements journalists and investors read, turned into structured leads.
Do I need Crunchbase account or API key? No. The actor works from public coverage and needs no credentials, login, or keys.
How fresh is the data?
Stories are read live during your run — with maxAgeDays you can restrict everything to the last 24 hours, week, or any window.
Will one unreadable story break my run? No. Stories that can't be read are skipped with a log line and the run continues.
Why do some rows have no email or phone? With lead details enabled, every company is saved even when its public web presence doesn't expose contacts — nothing is filtered out, which keeps runtime per 1,000 companies predictable. Companies with findable details get them filled in.
Can I get all companies from a day's news?
Yes — leave maxAgeDays at 1–2 and raise the per-mode budgets (e.g. maxRecentlyFunded: 1000). maxArticles bounds how many stories are read.
How fast is it? Stories are read in parallel and results stream as they are processed — a 50-company run with lead details typically finishes in a few minutes. Lead details add a little extra time per company.
What formats can I export? JSON, CSV, Excel, HTML, or RSS — plus the Apify API, webhooks, and integrations.
Does the free plan work? Free-plan runs are capped to a tiny sample (see Free tier). Upgrade to any paid Apify plan for full output.
Can I schedule it? Yes — use Apify Schedules to run on any cron you like, with the webhook for real-time delivery.
Is there an item cap? Each mode's budget can go up to 10,000, and the actor supports the platform's maximum run timeout (10,000 seconds) for very large jobs.
Reliability & error handling
- Every failure is reported as a fixed, human-readable sentence — the actor never dumps technical errors into logs or output, so runs stay clean and supportable.
- Temporary blocks and busy responses are retried automatically with a fresh connection; the run continues from where it stopped.
- A blocked or unavailable listing never stops the whole run — other modes and stories keep going.
- Webhook delivery failures never lose data: the dataset row is always written first.
Support
Open an issue on the Actor page or contact the publisher with your run ID and your (redacted) input JSON.
Contact me
Need something built beyond this Actor? I take on custom projects — from Apify scrapers and data pipelines to full-stack web apps.
| dubem115@gmail.com | |
| GitHub | github.com/DrunkCodes |