Crunchbase Scraper — Funded Companies & Leads avatar

Crunchbase Scraper — Funded Companies & Leads

Pricing

from $3.00 / 1,000 results

Go to Apify Store
Crunchbase Scraper — Funded Companies & Leads

Crunchbase Scraper — Funded Companies & Leads

Collect recently funded companies from Crunchbase news in real time — amount, stage, investors, plus websites, emails and phones. Track keywords, investors, or competitors. Results stream to your dataset as JSON. Free plans are limited; upgrade for full data.

Pricing

from $3.00 / 1,000 results

Rating

0.0

(0)

Developer

Emmanuel

Emmanuel

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Crunchbase Real-Time Data

Real-time Crunchbase funding intelligence, streamed to your dataset as it is collected. Track recently funded companies, monitor startup news, search funding coverage by keyword, investor, or company — and optionally add each company's public website, email, phone, and social profiles. One flat JSON row per company, ready for your CRM, spreadsheet, database, or AI pipeline.

⚡ Streaming output • 🎯 Optional lead details • 🧩 Flat, predictable JSON • 💸 Pay only for what you collect

🔒 Paid-only note: on an Apify free plan this Actor exports a small sample of results (default: 2 per run) and finishes with a clear message. Upgrade to any paid Apify plan for full, unlimited data. See Free tier.


What it does

The Actor reads public Crunchbase funding-news coverage and turns every company mentioned in it into a structured lead: company name, funding amount, stage, investors, the announcement it came from, and (with lead details enabled) the company's own website and contact points.

Crunchbase Real-Time DataTypical alternative scrapers
SpeedStories are read in parallel batches; results stream immediatelySerial, slow, one item at a time
Memory512 MB default2–4 GB+
OutputStructured JSON, LLM-readyOften messy HTML
Runtime predictabilityEvery company is saved — nothing filtered outUnpredictable when filters drop rows

Streaming by design: each company is pushed to the dataset the moment it is processed — you watch results arrive in real time, and RAM stays light even on very long runs.


Who it is for

  • Sales teams & SDRs — reach founders within hours of their funding announcement, when budgets are fresh and inboxes are open.
  • Growth & RevOps teams — build a always-fresh pipeline of funded companies matching your ICP.
  • VC & PE analysts — monitor portfolios, track investor activity, and map market segments.
  • Recruiters — companies that just raised are hiring; get the list first.
  • Agencies & consultants — deliver funding-tracker reports for clients without manual research.
  • Data engineers & AI builders — clean, stable JSON schema for pipelines, warehouses, and RAG.

Use cases

  • 🎯 Funding-based lead generation — collect companies that raised in the last N days, with contact details attached, and reach out while the raise is news.
  • 📈 Investor portfolio discovery — ask "show me everything backed by Y Combinator / a16z / Sequoia this month".
  • 🔎 Market & keyword monitoring — follow coverage of "ai infrastructure", "fintech brazil", "climate robotics" and capture every named company.
  • 🏢 Company tracking — watch a list of competitors or prospects and get a row for every news mention.
  • 📰 Daily funding digest — schedule the actor each morning and pipe results to Slack or Sheets via webhook.
  • 🧠 AI agent / RAG feeds — flat _externalId-keyed rows upsert cleanly into vector stores and CRMs.
  • 📊 Investor activity analytics — aggregate rounds by stage, size, and lead investor over time.
  • 🌍 Regional deal flow — tag every row with your own region label and split datasets by geography.
  • 🔁 CRM enrichment — match company against your existing accounts and backfill funding data (amountUsd, stage, investors).
  • 🏆 Top-of-funnel for agencies — prospect businesses by funding event instead of static firmographics.

Collection modes — combine any of them

ModeTurn on withYou get
Recently funded (default)enableRecentlyFunded: trueCompanies from the latest funding announcements, with amount, stage, and investors.
Startup newsenableStartupNews: trueCompanies named in general startup-news coverage (even without a round).
Keyword searchenableSearch: true + keywords[]Companies from coverage matching your keywords.
Backed by investorenableBackedBy: true + investors[]Companies backed by the investors you list.
Company researchenableCompanyResearch: true + companies[]Every news mention of the companies you track.

All enabled modes run in one pass — stories are deduplicated across modes and every distinct company is saved once.


Lead details (contact enrichment)

Enable Lead details (enableLeadDetails, on by default) and the Actor adds each company's public website, email, phone, and social profiles (LinkedIn, X, Facebook, Instagram, and more) where findable.

Important: this is an enable option, not a filter. Every collected company goes to the dataset — companies where no details can be found still appear, with empty contact fields. Nothing is dropped, so runtime per 1,000 companies stays predictable. Lead discovery adds a little extra time per company.

Output fields added by lead details: website, email, emails[], phone, phones[], socials{}, linkedinUrl, facebookUrl, instagramUrl, twitterUrl, enriched, enrichmentLevel, confidence-adjacent level via enrichmentLevel.


Quick start

  1. Open the Actor in Apify Console.
  2. Leave Recently funded checked (or pick other modes).
  3. Keep Enable lead details checked to get websites, emails, and phones.
  4. Set the per-mode max controls — start with 10 to test, then raise (1000+) for real runs. Your run's total is the sum of the budgets you enable.
  5. Press Start and watch rows stream into the dataset.
  6. Export as JSON / CSV / Excel, or pull via the Apify API, integrations, or MCP.

Example — this morning's funding rounds with contacts

{
"enableRecentlyFunded": true,
"enableLeadDetails": true,
"maxRecentlyFunded": 50
}

Example — companies backed by specific investors

{
"enableBackedBy": true,
"investors": ["a16z", "Sequoia", "Y Combinator"],
"enableLeadDetails": true,
"maxPerInvestor": 100
}

Example — keyword monitoring for a niche

{
"enableSearch": true,
"keywords": ["ai infrastructure", "robotics"],
"maxAgeDays": 7,
"maxPerKeyword": 50
}

Example — track a competitor list

{
"enableCompanyResearch": true,
"companies": ["Stripe", "Ramp", "Brex"],
"enableLeadDetails": false,
"maxPerCompany": 200
}

Input reference (full schema)

The input form is organized into sections — one per collection mode, then lead details, scope & limits, and connection.

🔎 Recently funded / 🔑 Keyword search / 🏛️ Backed by investor / 🏢 Company research / 📰 Startup news

Each collection mode has its own section with its own "max to collect" control — your run's total is the sum of the budgets you set in the sections you enable. There is no separate global cap.

FieldTypeDefaultDescription
enableRecentlyFundedbooleantrueCollect companies from the latest funding announcements — with amount, stage, and investors.
maxRecentlyFunded — Max companies from latest funding newsinteger10How many companies to collect from the latest announcements.
enableSearchbooleanfalseSearch funding coverage for your keywords.
keywordsstring[]prefill: ai infrastructure, fintechOne keyword per row (max 8 used per run).
maxPerKeyword — Max companies per keywordinteger10Per-keyword budget before moving to the next keyword. 3 keywords × 10 = 30 companies.
enableBackedBybooleanfalseFind companies backed by the listed investors (portfolio discovery).
investorsstring[]prefill: a16z, Y CombinatorOne investor per row (max 8 used per run).
maxPerInvestor — Max companies per investorinteger10Per-investor budget before moving to the next investor.
enableCompanyResearchbooleanfalseTrack specific companies in the news.
companiesstring[]prefill: Stripe, RampOne company name per row — every news mention becomes a row.
maxPerCompany — Max mentions per companyinteger10Per-company budget of news mentions.
enableStartupNewsbooleanfalseAlso collect companies named in general startup-news coverage, even without a round. Companies found here count toward the recently-funded budget.

🎯 Lead details

FieldTypeDefaultDescription
enableLeadDetailsbooleantrueAdds each company's public website, email, phone, and social profiles where available. Additive only — never filters. Adds a little extra time per company.

⚙️ Scope & limits

FieldTypeDefaultDescription
maxArticles — Max stories to readinteger100Maximum news stories to read (1–500). Each story may yield several companies.
maxAgeDays — Only stories from the last N daysinteger0Only stories published within the last N days. 0 = no age limit.
region — Region tagstring"US"Free-text label saved on every row for splitting datasets later.
webhookUrl — Webhook URLstring""Optional; each record is POSTed there as it is saved (in addition to the dataset).
webhookFormat — Webhook formatjson | slack"json"json = full record; slack = Slack message payload.

🌐 Connection

FieldTypeDefaultDescription
proxyConfigurationobjectApify residential USProxy settings; Apify residential (US) is enabled by default.

At least one collection mode must be enabled, and the lists required by the enabled modes (keywords / investors / companies) must be non-empty.


Output reference (full schema)

Every row is a flat JSON object. One row per company.

FieldTypeDescription
_externalIdstringStable deduplication key — safe for CRM/warehouse upserts.
featureTypestringAlways "funding".
modestringCollection mode tag ("funding").
companystringCompany name.
titlestringHeadline of the story the company came from.
urlstring | nullLink to the announcement.
sourcestringData source label ("Crunchbase News").
publishedAtstring | nullStory publish time (ISO 8601).
amountUsdnumber | nullRound size normalized to USD (e.g. 25000000).
amountDisplaystring | nullHuman-readable amount (e.g. "$25 million").
stagestring | nullNormalized stage (Seed, Series A, Series B, Growth, …).
investorsstring[]Investor names extracted from the announcement.
summarystring | nullThe funding sentence the company was extracted from.
articleSummarystring | nullArticle description/summary.
categorystring | nullStory category.
authorstring | nullStory author.
regionstringYour region tag from input.
websitestring | nullCompany website (lead details).
emailstring | nullPrimary contact email (lead details).
emailsstring[]All found emails (lead details).
phonestring | nullPrimary phone (lead details).
phonesstring[]All found phones (lead details).
socialsobject | nullMap of social profile links (lead details).
linkedinUrl / facebookUrl / instagramUrl / twitterUrlstring | nullDirect social links (lead details).
enrichedboolean | nulltrue when at least one contact detail was found.
enrichmentLevelstring | null"partial" (some details) or "list_only" (none found).
scrapedAtstringISO 8601 UTC time the row was collected.

Example row:

{
"_externalId": "crunchbase-acme-ai",
"featureType": "funding",
"mode": "funding",
"company": "Acme AI",
"title": "Acme AI raises $25M Series B to scale agents",
"url": "https://news.crunchbase.com/startups/acme-ai-series-b/",
"source": "Crunchbase News",
"publishedAt": "2026-09-24T08:00:00.000Z",
"amountUsd": 25000000,
"amountDisplay": "$25 million",
"stage": "Series B",
"investors": ["Felicis", "Founders Fund"],
"summary": "Acme AI raised $25 million in a Series B led by Felicis.",
"articleSummary": "The round brings Acme AI's total funding to $40M.",
"region": "US",
"website": "https://acme.ai",
"email": "hello@acme.ai",
"emails": ["hello@acme.ai"],
"phone": "(415) 555-0134",
"socials": { "linkedin": "https://www.linkedin.com/company/acme-ai" },
"linkedinUrl": "https://www.linkedin.com/company/acme-ai",
"enriched": true,
"enrichmentLevel": "partial",
"scrapedAt": "2026-09-25T12:00:00.000Z"
}

Run summary (OUTPUT tab): after every run the OUTPUT key-value store holds a summary — modes run, stories read, companies collected, the applied cap, webhook setting, any non-fatal errors, and a paywall object (detected, isPaying, pricingTier, limited, blocked, freeTierMaxItems) so the applied restrictions are transparent.


Webhooks

Set Webhook URL in the input and every record is POSTed to your endpoint in real time, in addition to the dataset — perfect for Zapier, Make, n8n, Slack incoming webhooks, or your own API.

  • Records are always saved to the dataset first; webhook delivery is best-effort and never blocks or fails the run.
  • Failed deliveries are retried once and then skipped with a single log line — your data is still safe in the dataset.
  • Format is chosen with Webhook format.

JSON payload — the full record object (same shape as the dataset rows above).

Slack payload:

{
"text": ":moneybag: *Acme AI*\n*Round:* Series B • *Amount:* $25 million\n*Investors:* Felicis, Founders Fund\n*Email:* hello@acme.ai\nhttps://acme.ai"
}

Setup: paste an HTTPS URL into the Webhook URL field (e.g. a Zapier catch hook, a Slack incoming-webhook URL, or your own endpoint) and run. Test the endpoint with any POST receiver before production runs.

Webhook payloads contain only the documented record fields — nothing else is attached.


Apify API & MCP usage

Fetch results via the Apify API

# Using an Apify API token
curl -H "Authorization: Bearer $APIFY_TOKEN" \
"https://api.apify.com/v2/acts/YOUR_ACTOR_ID/runs/last/dataset/items?clean=true&format=json"

Run from the API

curl -X POST "https://api.apify.com/v2/acts/YOUR_ACTOR_ID/runs" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "enableRecentlyFunded": true, "enableLeadDetails": true, "maxRecentlyFunded": 50 }'

MCP (Model Context Protocol)

The actor works with Apify's MCP server so AI agents (Claude, Cursor, custom agents) can call it as a tool:

  1. Connect to https://mcp.apify.com (see Apify MCP docs for auth).
  2. Call act_YOUR_ACTOR_ID (or search the Store) with the input schema above.
  3. Results come back as dataset items — flat JSON your agent can read directly, then fetch the full dataset for the rest.

Example agent call:

{
"enableRecentlyFunded": true,
"enableLeadDetails": true,
"maxAgeDays": 3,
"maxRecentlyFunded": 20
}

Scheduling & automation

  • Schedules: run every morning at 8:00 to get the last 24 hours of funding news — combine with maxAgeDays: 2 and the webhook for a hands-off digest.
  • Integrations: the default dataset plugs into Google Sheets, Airtable, Excel, and 100+ Apify integrations.
  • Deduplication: _externalId is stable per company — store rows with upsert semantics to keep a rolling, deduplicated funding feed.

Controlling scope & cost

FieldEffect
Per-mode budgets (maxRecentlyFunded, maxPerKeyword, maxPerInvestor, maxPerCompany)The cost levers. Your run's total is the sum of the budgets in the sections you enable.
maxArticlesHow many stories to read. Each story may yield several companies.
maxAgeDaysKeep only recent stories — smaller windows finish faster.
enableLeadDetailsTurn off to skip contact discovery (faster, still full company output).

Free tier (Apify free plan)

On an Apify free plan, this Actor exports a small sample of results — a maximum of 2 companies per run (owner-configurable) — then finishes gracefully with a clear message telling you to upgrade. The run never errors and your sample stays in the dataset.

PlanWhat you get
Apify free plan⚠️ Capped sample (default 2 results per run), then a clean stop with an upgrade message.
Any paid Apify plan✅ Full, uncapped output up to your per-mode budgets.

Upgrade to a paid Apify plan to get full, unlimited data. The cap is applied automatically from your plan — you don't configure anything, and paying users see one confirmation line in the log ("Paying user — full output.") and no caps.

The owner can switch free handling between a capped sample and a hard block (and tune the cap) via the Actor's environment variables — end users never need to configure anything.


FAQ

Which data source does this use? Public Crunchbase News funding coverage — the same announcements journalists and investors read, turned into structured leads.

Do I need Crunchbase account or API key? No. The actor works from public coverage and needs no credentials, login, or keys.

How fresh is the data? Stories are read live during your run — with maxAgeDays you can restrict everything to the last 24 hours, week, or any window.

Will one unreadable story break my run? No. Stories that can't be read are skipped with a log line and the run continues.

Why do some rows have no email or phone? With lead details enabled, every company is saved even when its public web presence doesn't expose contacts — nothing is filtered out, which keeps runtime per 1,000 companies predictable. Companies with findable details get them filled in.

Can I get all companies from a day's news? Yes — leave maxAgeDays at 1–2 and raise the per-mode budgets (e.g. maxRecentlyFunded: 1000). maxArticles bounds how many stories are read.

How fast is it? Stories are read in parallel and results stream as they are processed — a 50-company run with lead details typically finishes in a few minutes. Lead details add a little extra time per company.

What formats can I export? JSON, CSV, Excel, HTML, or RSS — plus the Apify API, webhooks, and integrations.

Does the free plan work? Free-plan runs are capped to a tiny sample (see Free tier). Upgrade to any paid Apify plan for full output.

Can I schedule it? Yes — use Apify Schedules to run on any cron you like, with the webhook for real-time delivery.

Is there an item cap? Each mode's budget can go up to 10,000, and the actor supports the platform's maximum run timeout (10,000 seconds) for very large jobs.


Reliability & error handling

  • Every failure is reported as a fixed, human-readable sentence — the actor never dumps technical errors into logs or output, so runs stay clean and supportable.
  • Temporary blocks and busy responses are retried automatically with a fresh connection; the run continues from where it stopped.
  • A blocked or unavailable listing never stops the whole run — other modes and stories keep going.
  • Webhook delivery failures never lose data: the dataset row is always written first.

Support

Open an issue on the Actor page or contact the publisher with your run ID and your (redacted) input JSON.


Contact me

Need something built beyond this Actor? I take on custom projects — from Apify scrapers and data pipelines to full-stack web apps.

Emaildubem115@gmail.com
GitHubgithub.com/DrunkCodes