Carbon Lead Scanner avatar

Carbon Lead Scanner

Under maintenance

Pricing

Pay per usage

Go to Apify Store
Carbon Lead Scanner

Carbon Lead Scanner

Under maintenance

Carbon Lead Scanner discovers and qualifys local business leads. It searches selected regions and business categories, collects official websites and publicly listed contact details, crawls websites for business information and email addresses, and verifies email domains using DNS/MX checks.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

max

max

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

0

Monthly active users

8 days ago

Last modified

Share

Carbon Lead Scanner turns a business category and a target region into a qualified, evidence-backed lead list. It combines local listings, search results, supplied websites, and directory pages; identifies official business domains; analyzes those websites; extracts publicly visible contact information; checks website and email-domain infrastructure; and ranks each result with transparent scoring.

Use it for local prospecting, market research, sales-list preparation, and website intelligence. Every run is configurable: choose the market, geography, target email-lead volume, discovery sources, crawl depth, verification rules, score threshold, and exports instead of being locked to one niche or location.

What it does

  1. Finds candidate companies through the selected discovery sources.
  2. Normalizes domains and removes duplicates.
  3. Inspects official websites and selected high-value pages.
  4. Extracts public email addresses, phone numbers, addresses, descriptions, and social-profile links.
  5. Checks HTTPS availability, DNS resolution, and optional email-domain MX records.
  6. Detects website technologies, advertising pixels, and booking systems when requested.
  7. Scores leads on a 0-100 scale and records the evidence behind the score.
  8. Writes a dataset and the selected professional export files.

Configure a run

Only four fields are required: industry, region, qualified-lead target, and discovery sources. The scanner provides sensible defaults for crawl limits, verification, filtering, reporting, discovery budgets, and the Google Maps Actor ID; every optional setting can still be changed in the Apify form.

If you choose Provided business websites or Seed or directory pages, provide at least one corresponding URL. Brave Search requires a Brave API key. The Google Maps Actor ID is only used when you select Apify Google Maps Actor; otherwise you can ignore it. A platform APIFY_TOKEN can be used for a nested Maps Actor when running on Apify; local runs can supply apifyToken directly.

Important settings

SettingDefaultWhat it controls
targetCountRequiredTarget number of returned leads. With the default requireEmail: true, this means public business email leads, not phone-only businesses. The Actor checks candidates until it reaches this number or the configured sources/budgets are exhausted.
maxCandidatesPerSource100Discovery buffer per source. It is automatically raised to at least targetCount x 5, because many listings have no usable website or contact data.
maxDiscoveryRounds3Reserved compatibility setting for older custom Maps Actors. The maintained default Maps Actor runs each configured keyword/location job once because it already searches the geographic area comprehensively.
googleMapsRunTimeoutSecs1200Maximum time for one nested Apify Maps scan. If it reaches the limit, already-written map records are kept and the scanner proceeds to website checks and exports; this prevents the parent run from reaching Apify's one-hour timeout.
maxDiscoveryRequests40Hard cap across discovery API, directory, and nested Maps requests. Use it to control third-party usage and cost.
maxTotalCandidates200Maximum unique official websites inspected after discovery. It is always raised to at least targetCount x 5, so the scan has enough candidates to reach the target where sources permit.
additionalRegionsEmptyDistricts, nearby towns, or municipalities to search separately. Useful for a metropolitan area such as Munich plus Schwabing, Garching, and Unterhaching.
queryVariantsEmptyIndustry synonyms, local-language terms, or related categories. Each is combined with the selected region(s).
overpassEndpointsBuilt-in fallbacksOptional preferred Overpass endpoints. Enter your own endpoint here for higher reliability or local, cost-free OSM use.
wikidataRadiusKm25Radius used by the built-in Wikidata nearby-business query.
crawlDepth1How deeply official websites are checked; see the crawl-level table below.
maxPagesPerWebsite4Absolute maximum pages fetched for one business, regardless of crawl depth.
concurrency6Parallel website checks. Lower this for cautious, slower crawling; raise it carefully for faster runs.
requestTimeoutSecs20Per-request timeout. Slow or unavailable pages are skipped after this period.
minFitScore50Minimum 0-100 evidence score a candidate must achieve to be exported.
enableGeminiEnrichmentfalseOpt-in, bounded website intelligence. No public website text is sent to Gemini unless you enable it and provide a key.
geminiModelgemini-2.5-flash-liteStable low-cost default for high-volume structured website extraction. Select gemini-2.5-flash only when stronger reasoning is worth the higher cost.
geminiMaxEnrichments50Hard cap on eligible websites sent to Gemini. Leads beyond the cap remain fully usable with deterministic crawl data.
geminiMaxInputChars12000Maximum cleaned public website text sent for one lead. Lower it to reduce spend.
geminiConcurrency2Independent cap for simultaneous Gemini calls, protecting against rate limits and bursts.
requireVerifiedMxfalseKeep only public email domains with usable MX records and at least one resolvable advertised mail server.
requireEmailtrueRequire a usable public business email in every returned target lead. Disable only if phone-only leads are acceptable.
genericEmailsOnlyfalseKeep only role inboxes such as info@, contact@, or office@.
requireValidPhonefalseReject a candidate only when its discovered phone has an implausible format. This is not a live-call check.
excludeBookingSystemsfalseExclude businesses where a booking tool is detected; useful when selling booking software.

Crawl-depth levels

LevelPages inspectedBest use
0Homepage onlyFast technical triage; may miss emails listed only on Contact or Imprint pages.
1Homepage plus linked high-value pages such as Contact, Imprint, and AboutRecommended default for lead research. It usually finds public contact details without excessive crawling.
2Level 1 plus one additional level of linked high-value internal pagesDeep research for smaller lead targets. Slower and subject to maxPagesPerWebsite.

The Actor stays on the same website origin, honors robots.txt when enabled, and prioritizes contact-oriented pages rather than crawling an entire site.

Discovery sources

SourceCredentialSuitable forConsiderations
OpenStreetMapNoneLocal businesses with mapped websitesCoverage and tags vary by place and industry.
Wikidata knowledge graphNoneSupplementary official websites, phones, addresses, and descriptions for notable/local entitiesCoverage is selective. Use a custom SPARQL query for a precise niche.
Custom open-data or directory JSONOptional source-specific keyMunicipal data portals, trade associations, tourism lists, permitted directory APIsYou configure the endpoint and field mapping. It must contain official business website URLs.
Provided business websitesNoneDirect research listsYou supply the official websites to analyze.
Seed or directory pagesNoneAssociations, directories, and curated listsCrawl only pages you are permitted to use.
Brave Search APIBrave API keyNiche and broader web discoveryAPI quota and cost apply.
Official Google Places APIGoogle Cloud API key with billingHigh-quality local business discovery with official website, address, phone, business status, and Maps evidenceUses Google Maps Platform pay-as-you-go billing; set a strict request budget and API-key restrictions.
Apify Google Maps ActorApify tokenLocation-based business researchNested Actor usage and Store pricing apply.

Source setup guides

OpenStreetMap and your own Overpass endpoint

No key is required. For a normal city search, leave OpenStreetMap fallback searches empty and the scanner uses geocoding plus Overpass. For a large region, enter several precise Nominatim queries or use Additional nearby regions and Industry synonyms. Public Overpass services are shared infrastructure, so keep the request budget conservative. If you operate an Overpass instance yourself, add its full interpreter URL in Custom Overpass endpoints; it is tried before public fallbacks.

Wikidata

No key is required. Select Wikidata knowledge graph, choose a radius, and the scanner performs a nearby-business query around the selected region. It only accepts results that expose an official website, then crawls that website exactly like every other lead.

For a specialist category, use Custom Wikidata SPARQL queries. Each query must return a website variable. It can also return item, itemLabel, phone, address, and description; these become source-provenance fields in the result. Run and validate an advanced query in the Wikidata Query Service before placing it in the input form.

Custom municipal, association, or open-data JSON

Select Custom open-data or directory JSON and add one object per permitted endpoint in Custom public-directory JSON sources. No key is required by the scanner itself. If that endpoint requires an API key, add its documented header in headers; use only the key issued by that data provider, never an Apify token.

[
{
"url": "https://data.example.gov/api/businesses",
"recordsPath": "results.items",
"websiteField": "website",
"nameField": "name",
"phoneField": "telephone",
"emailField": "email",
"addressField": "address",
"descriptionField": "category",
"headers": { "X-API-Key": "YOUR_PROVIDER_KEY" }
}
]

recordsPath is the location of the record array inside the JSON response. Leave it blank if the response itself is an array. The websiteField is mandatory because the scanner intentionally crawls official company websites, not directory-listing pages.

Brave Search API

Create a Brave Search API key, enter it in Brave Search API key, then select Brave Search API. The key can alternatively be provided as BRAVE_API_KEY for local runs. Use maxDiscoveryRequests to limit the number of paid requests. Query variants and additional regions are automatically combined into multiple precise searches.

Google Maps through Apify

Select Apify Google Maps Actor only when you explicitly want this paid optional source. The default Actor ID is microworlds/crawler-google-places; it is not required for any other source. On the Apify platform, the platform token is supplied automatically. For a local run, provide apifyToken or set APIFY_TOKEN.

The default Maps Actor expects keywords and location (or custom_geolocation), not searchStringsArray. The scanner now supplies keywords: [industry] and location: region itself. Add queryVariants and additionalRegions to make separate keyword/location jobs. For advanced use, googleMapsInput may contain keywords, location, or custom_geolocation; it is passed using the current documented contract. The optional Maps Actor contact add-on is deliberately disabled: this scanner crawls and verifies each official website itself, so it does not silently add contact-enrichment fees.

Maps scans can be slow for large cities because the nested Actor divides the area into many map segments. googleMapsRunTimeoutSecs defaults to 20 minutes so the parent Actor always has time left to crawl, validate, score, and export results. A timed-out nested scan is not discarded: the scanner reads its already-written dataset records and continues. Raise this limit only when you also increase the parent run timeout in Apify's Run options.

Select Official Google Places API to use Google’s documented Places API directly inside this scanner. This is the recommended replacement for the nested public Apify Google Maps Actor when you are on Apify’s Creator plan: it does not call another Apify Store Actor.

  1. Create or select a Google Cloud project.
  2. Enable Places API (New) and attach a billing account.
  3. Create an API key, restrict it to Places API (New), and ideally restrict it to the Apify runtime only if your Google Cloud setup supports a safe server-side restriction.
  4. Paste the key into Google Places API key or set GOOGLE_PLACES_API_KEY for a local run.
  5. Select Official Google Places API, set a country code such as DE or US, and keep maxDiscoveryRequests within your intended Google API budget.

The scanner uses the official Text Search (New) endpoint and asks only for the business name, address, official website, public phone, business status, primary type, and Maps evidence URL. It skips permanently closed businesses and records only entries with an official website, which the normal website crawler then verifies and enriches. Google bills Places requests by requested fields and usage, so this source is not free; the response field mask and request budget are intentionally explicit. Follow Google Maps Platform terms, including applicable limits on storing or displaying Google-sourced content.

Gemini website enrichment — optional, bounded, and evidence-grounded

Gemini is never required for discovery, qualification, email extraction, DNS/MX checks, or scoring. Enable it only when you want an additional structured reading of already-crawled public website text. It returns a concise summary, explicit services, explicit customer segments, a cautious outreach angle, confidence, and up to three evidence quotes that must occur verbatim in the supplied website text.

  1. Create a Gemini API key in Google AI Studio.
  2. Enable Gemini website enrichment in the Actor input.
  3. Paste the key into Gemini API key. It is secret in Apify; for local runs, use GEMINI_API_KEY instead.
  4. Keep the default Gemini 2.5 Flash-Lite for the strongest low-cost, high-volume option. Choose Gemini 2.5 Flash only if you accept a higher cost for deeper reasoning.
  5. Start with geminiMaxEnrichments: 10, geminiConcurrency: 1, and geminiMaxInputChars: 8000; inspect the gemini section of SUMMARY.json before increasing limits.

The integration uses Google’s Generate Content API with JSON-only output, low temperature, bounded input/output sizes, independent concurrency, timeout handling, and retries only for temporary rate-limit or server errors. Every result is application-validated after JSON parsing. If Gemini fails, times out, is rate-limited, produces invalid output, or reaches its per-run cap, the lead remains available and deterministic crawl results are retained. AI output never contributes points to fitScore, never changes the acceptance decision, and never creates or verifies contact details.

Gemini receives only the bounded public website text already fetched by the crawler, plus the company name, URL, industry, and region. Do not enable it for pages or data you are not permitted to send to Google. The AI Intelligence Excel sheet, raw JSON, and dataset ai view preserve the status, model, structured output, grounded evidence quotes, and a non-sensitive failure note.

Source choices that are intentionally not built in

Yelp’s current free access is an evaluation trial and its listings do not consistently provide an official website suitable for compliant website-first qualification. OpenCorporates is useful for legal-entity enrichment, but its standard limits and reuse licence make it unsuitable as a broad free website-discovery source. Common Crawl and Internet Archive data can be stale and therefore are not used to claim a current contact method. These can be connected through a permitted custom JSON source only when you have a compliant, current data feed.

Outputs

The Actor always writes qualified leads to the Apify dataset. Select any combination of the following standalone exports for the run:

FilePurpose
RESULTS.jsonClean lead-record array for integrations.
REPORT.jsonDetailed structured report containing run metadata, summary, leads, and optional rejected candidates.
RESULTS.csvUniversal table export for simple imports.
RESULTS.xlsxFormatted workbook with an executive summary, qualified leads, technical intelligence, raw data, and rejected-candidate audit.
REPORT.htmlResponsive, searchable report for browser review.
REPORT.pdfPresentation-ready PDF with executive metrics, score distribution, qualification funnel, lead directory, and website-intelligence table.
SUMMARY.jsonCompact metrics and limitation notes for the run.

The Excel, HTML, and PDF files are designed for review as well as export: they use clear hierarchy, score-focused visual treatment, structured tables, and explanatory notes. The PDF is generated only when Executive PDF intelligence report is selected.

Run locally

Requirements: Node.js 20 or newer.

npm install
cp examples/local-input.json examples/my-input.json
# Review and set every field in examples/my-input.json for your own run.
npm run build
node dist/src/main.js --input examples/my-input.json --output-dir output

The local input file is an explicit example. Generated files are written to the directory passed with --output-dir.

For a no-paid-source Munich configuration with OpenStreetMap, Wikidata, regional expansion, MX checks, and all report formats, start from examples/munich-hair-salons-local.json. Review its industry and region settings before use; it deliberately has no Brave, Maps, or custom-directory key.

Run locally with Docker

docker build -t carbon-lead-scanner .
docker run --rm \
-v "$PWD/examples:/usr/src/app/examples:ro" \
-v "$PWD/output:/usr/src/app/output" \
carbon-lead-scanner \
node dist/src/main.js --input examples/local-input.json --output-dir output

Deploy to Apify

  1. Create a new TypeScript Actor in Apify Console.
  2. Replace the generated source with this project, keeping .actor/, src/, examples/, Dockerfile, package.json, and tsconfig.json.
  3. Build the Actor.
  4. Open Input, set the four required fields, adjust optional settings when needed, and start it.
  5. Review the dataset and selected file links under Output.

You can also deploy from a local checkout:

npm install -g apify-cli
apify login
apify push

Understanding verification and scoring

Email-domain verification

emailReachability = verified-domain means all of the following checks passed:

  1. The public email address has a syntactically valid domain.
  2. The domain is not on the disposable-email list.
  3. DNS returns one or more MX records for the domain.
  4. The MX records do not contain a null-MX entry, which explicitly states that the domain does not receive email.
  5. The MX hostname is not recognized as a parked or suspicious mail host.
  6. At least one advertised MX hostname resolves to an IPv4 or IPv6 address.

This verifies that the domain publishes working mail-routing infrastructure. It does not prove that the exact mailbox exists, is monitored, accepts unsolicited messages, or will deliver a message. The Actor deliberately does not issue SMTP RCPT TO probes because they are unreliable, frequently blocked, and can create legal or operational problems.

Fit score: exact values

The Actor starts with discovery confidence worth 40% of the score, then adds evidence from the official website. The final score is capped at 100. Every accepted lead includes a scoreSignals array that lists the points actually awarded.

EvidencePoints
Discovery confidenceUp to 40 (sourceConfidence x 0.40)
Public usable business email found+20
Email domain matches the official website domain+5
Verified mail-domain infrastructure+10
Public phone number found+10
Public address found+5
Description with at least 40 characters+5
Public social profile found+3
Website technology identified+5
Advertising or tracking signal identified+3
Working HTTPS+2
DNS resolution confirmed+2

Typical discovery-confidence contributions are approximately 33 for Apify Google Maps, 29 for OpenStreetMap, 28 for provided URLs, and 27 for Brave Search or Nominatim fallback results.

Example: a Google Maps business with a company-domain email, verified mail infrastructure, phone, address, HTTPS, and DNS earns roughly 33 + 20 + 5 + 10 + 10 + 5 + 2 + 2 = 87 before any extra technology, description, social, or advertising signals.

Quality tiers

TierMeaning
STRICTA usable public email was found.
MODERATENo email, but a phone was found and the score is at least 75.
RELAXEDA phone was found, but the score is below 75.
ULTRA-RELAXEDNo usable public email or phone was found; these normally disappear when requireContactMethod is enabled.

targetCount is a completion target, not a promise that the real web contains that many businesses matching the selected filters. By default every returned lead has a public email. The scanner continues through its discovery and website-inspection buffer until it obtains the requested count, then stops. If it still cannot reach the count, SUMMARY.json and the run log report target-not-reached with the genuine shortfall and source health details rather than presenting a partial set as successful completion. Increase source coverage or the explicitly shown discovery/candidate budgets only when you want the Actor to spend more time or paid-source usage.

Built-in reliability and intelligence improvements

  • True target-oriented processing: discovery gathers a larger website pool, then keeps qualifying leads until the target is reached or the configured candidate budget is exhausted.
  • Cost guardrails: maxDiscoveryRequests caps external discovery calls and maxTotalCandidates caps website inspections. Both values appear in the summary limitations.
  • Source health report: SUMMARY.json includes discoverySourceStats with requests, candidates, status, and a failure or budget note for every selected source.
  • Field provenance: each lead includes fieldSources, showing whether the company name, email, phone, address, description, or website originated from a selected source or the official website crawl.
  • Phone plausibility: results include phoneValid; when enabled, the scanner rejects obvious invalid formats. It does not call or otherwise contact the number.
  • Multi-source deduplication: official domains are normalized and deduplicated across all sources, with the strongest discovery confidence retained.
  • Evidence-first exports: JSON, CSV, Excel, HTML, and PDF keep source, evidence URLs, score signals, infrastructure checks, and rejected-candidate reasons.

Responsible use

Only crawl pages and collect data you are permitted to access. Respect robots.txt, site terms, rate limits, database rights, privacy law, and outreach rules such as GDPR and ePrivacy requirements. Publicly visible contact data is not automatic permission for unsolicited marketing. Retain evidence URLs and establish an appropriate lawful basis before outreach.

Development checks

npm test
npm run build

The test suite verifies input strictness, source credential rules, URL normalization, and scoring behavior.