# Company Finder – B2B Companies by Industry, Size & Location (`inovaflow/company-finder`) Actor

Find companies by industry, headcount and location, or enrich a company list: every row has the website domain, industry, size band, HQ, founded year, type, description, specialties and socials, deduplicated by domain. B2B company database and list builder for outbound. Dataset-only, MCP-ready.

- **URL**: https://apify.com/inovaflow/company-finder.md
- **Developed by:** [inovaflow](https://apify.com/inovaflow) (community)
- **Categories:** Lead generation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 company records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

**Find the B2B companies you should be selling to — by industry, headcount and location — and get back a clean, deduplicated company list with a website domain and real firmographics on every row.** Describe the niche (`software companies`, `marketing agencies`, `logistics`, `fintech`) and where (`Austin, TX`, `Berlin`, `Netherlands`), optionally a headcount band, and every company comes back with its **domain**, **industry**, **size band and headcount range**, **headquarters**, **founded year**, **company type**, **description**, **specialties**, **social profiles**, phone and address. Or paste a list of company names, domains or profile URLs and get the same firmographics for each.

Company data on the market is fragmented: one tool searches a professional network, another scrapes a map, a third resells a stale database dump. None of them gives you a single row per company that your next steps can chain on. This Actor assembles the list the way a good SDR would — business listings, public company profiles and the company's own website, merged and deduplicated by domain — and returns only what passes your filters.

### Company finder: what you get

One row per company:

| Field | What it tells you |
| --- | --- |
| `name`, `domain`, `website`, `linkedinUrl` | The company and the two keys everything else chains on |
| `industry`, `industries[]`, `specialties[]`, `listingCategory` | What the company does — its declared industry, listing categories and specialties |
| `companySize`, `employeesMin`, `employeesMax`, `employeesOnProfile` | Headcount band (`1-10` … `10000+`), the range behind it, and how many people list the company as their employer |
| `headquarters`, `city`, `region`, `country`, `hqMatchesLocation` | Where it is headquartered (ISO country code when known) and whether that is inside the location you searched — listings also return companies with an office or landing page in a city |
| `foundedYear`, `companyType` | How old it is and how it is owned (`Privately Held`, `Public Company`, `Nonprofit`, …) |
| `description`, `tagline`, `followers`, `logoUrl` | For the first line of your e-mail and for qualification |
| `phone`, `address`, `googleMapsUrl`, `rating`, `reviewsCount` | Listing details when the company has a physical presence |
| `socials{}` | Profile, X/Twitter, Facebook, Instagram, YouTube |
| `sources[]`, `matchedQueries[]`, `status` | Where each row came from (`listing`, `profile-search`, `profile`, `website`, `input`), which of your searches found it, and whether the profile could be read (`enriched`) or only the website (`website_only`) / the listing (`listing_only`) |

Three dataset views: **Companies** (one line per company), **Firmographics** (size, founded, type, HQ, industry), **Contact channels** (website, phone, address, socials).

### Two ways in

- **Search** — `keywords` × `locations`: `["software companies", "SaaS"]` in `["Austin, TX", "Denver, CO"]`. Optional filters: `companySizes` (headcount bands), `foundedAfter` / `foundedBefore`, `companyTypes`, `requireIndustryMatch`, `requireHqInLocation`. Companies from every source are merged and deduplicated by domain; `maxCompanies` caps the list (and the price).
- **Enrich** — `companies`: names (`Gong`), domains (`gong.io`) or profile URLs, mixed. Every entry comes back with the same firmographics — the enrichment step for a list you already have (for example the domains another Actor produced).

### How the list is built

1. **Discovery** — two built-in sources, both on by default: business listings for the niche in each location (strong for anything with an address) and public company profiles found through web search (strong for remote-first, B2B and tech companies with no storefront). Turn either off under *Sources & enrichment*.
2. **Website** — each company's home page is read for the description, social profiles (including the company's profile link), phone and the structured organization data sites publish.
3. **Public profile** — the company's public profile is read for industry, headcount band, headquarters, type, founded year, specialties, description and follower count. A profile found by name is kept only when it is verifiably the same company (same website or matching name), never a namesake.
4. **Filters and dedup** — size, founded, type and industry filters are applied to the enriched row; rows are deduplicated by domain (then profile); only delivered rows are charged. When a filter is set and the fact is unknown, the company is dropped unless **Keep unknown** is on — and every drop is counted by reason in the run summary, so a thin result is explainable.

No login, no cookies, no third-party data vendors, no nested scrapers — every source is built in.

### Chain it into a prospecting pipeline

Company Finder is the top of an outbound funnel. Its `domain` column is exactly what the next steps take:

- **People** — pass the domains to a decision-maker finder to get the people to contact at each company by title and seniority.
- **E-mails** — pass name + domain to an e-mail finder & verifier.
- **Tech stack** — pass the domains to a technology lookup to filter by CRM, marketing automation, ecommerce platform.
- **Hiring signals** — pass the companies to a hiring-intent scraper as a watchlist to see who is scaling which team.

Every row is flat JSON with stable field names, so an agent can map columns without a transform step.

### Who uses it

- **Outbound / GTM agents** — a keyword-discoverable, MCP-callable tool that builds a target-account list unattended and returns a typed dataset.
- **SDR and RevOps teams** — ICP lists by niche and city with headcount and founded-year filters already applied, ready for a sequencer.
- **Agencies and consultants** — local and regional company lists for a vertical, with the website, phone and socials on the same row.
- **Data teams** — a firmographics enrichment step for any list of domains.

### Set it up in a minute

1. Enter one or more **Industry or niche keywords** and **Locations** (or paste **Companies to enrich**).
2. Optionally pick **Company size** bands and open **More filters** for founded year and company type.
3. Start. Rows arrive as each company is finished; the run summary is in the `OUTPUT` record.

Advanced settings (collapsed) control the sources, enrichment depth, per-query limits, country/language, concurrency and the proxy used for company profiles (residential by default — profiles rate-limit datacenter traffic).

### Use it from an agent or the API

```json
{ "keywords": ["software companies"], "locations": ["Austin, TX"], "companySizes": ["11-50", "51-200"], "maxCompanies": 100 }
```

```json
{ "companies": ["gong.io", "lemlist", "https://www.linkedin.com/company/stripe"] }
```

Agents may also pass `industries`, `location` (single), `domains`, `websites` or `urls`, or objects `{ "name": "...", "domain": "...", "location": "..." }`. Results are in the default dataset (`?view=companies`, `?view=firmographics`, `?view=contacts`); the run summary (candidates per source, drops by reason, charges) is in the `OUTPUT` record of the run's key-value store. Through the Apify MCP server, call `inovaflow/company-finder` with the same input.

### Output example

```json
{
  "name": "Gong",
  "domain": "gong.io",
  "website": "https://www.gong.io/",
  "linkedinUrl": "https://www.linkedin.com/company/gong-io",
  "industry": "Software Development",
  "industries": ["Software Development"],
  "specialties": ["Revenue Intelligence", "Sales Enablement", "Conversation Analytics"],
  "companySize": "1001-5000",
  "employeesMin": 1001,
  "employeesMax": 5000,
  "employeesOnProfile": 1834,
  "headquarters": "San Francisco, California",
  "city": "San Francisco",
  "region": "California",
  "country": "us",
  "hqMatchesLocation": true,
  "description": "Transform how your revenue team wins with AI. The Gong Revenue AI Operating System brings customer data, insights, and workflows together in one trusted place.",
  "tagline": "AI OS for Revenue Teams",
  "foundedYear": 2015,
  "companyType": "Privately Held",
  "followers": 345251,
  "socials": { "linkedin": "https://www.linkedin.com/company/gong-io", "twitter": "https://twitter.com/gong_io", "facebook": null, "instagram": null, "youtube": "https://www.youtube.com/c/GongIo" },
  "sources": ["profile-search", "website", "profile"],
  "matchedQueries": ["sales software · San Francisco, CA"],
  "status": "enriched"
}
```

### Pricing

Pay per event: **$0.01 per company record** delivered with a domain, plus a small per-run start fee. Companies without a website, duplicates and companies your filters drop are never charged — so a 500-company sweep filtered down to 80 matches costs $0.80.

### Good to know

- **Size, founded and type come from the public company profile.** Companies without a profile (or with an unreadable one) keep their listing and website facts and are reported with `status: website_only` / `listing_only`; with a size filter on they are dropped unless you turn on *Keep unknown*.
- **Listings need a location.** A keyword without any location is searched through company profiles only.
- **Web search coverage** for profile discovery is a few dozen companies per keyword × location; use several keywords (synonyms, sub-niches) and several locations to build big lists.
- Results reflect what companies publish about themselves; `sources[]` and `status` tell you what was read for each row.

# Actor input Schema

## `keywords` (type: `array`):

What kind of companies, one per line — an industry, niche or category as you would describe it: `software companies`, `marketing agencies`, `logistics`, `dental clinics`, `fintech`, `cybersecurity consulting`. Each keyword is searched in each location.

## `locations` (type: `array`):

Cities, regions or countries, one per line (`Austin, TX`, `Berlin, Germany`, `London`, `Netherlands`). Leave empty for a location-free search (works for well-defined niches, e.g. `"sales engagement platform"`).

## `companySizes` (type: `array`):

Keep only companies in these headcount bands. Empty = any size. Companies whose size is unknown are dropped when a band is set (see "Keep unknown" under More filters).

## `maxCompanies` (type: `integer`):

Stop after this many companies have been delivered (also caps what you pay).

## `companies` (type: `array`):

One per line: a company name (`Gong`), a domain (`gong.io`) or a company profile URL. Also accepted as `domains`, `websites` or `urls`, or objects {name, domain, website}.

## `foundedAfter` (type: `integer`):

Keep only companies founded in or after this year (e.g. 2015 for young companies).

## `foundedBefore` (type: `integer`):

Keep only companies founded in or before this year.

## `companyTypes` (type: `array`):

Keep only these ownership types.

## `requireIndustryMatch` (type: `boolean`):

Keep only companies whose reported industry, category, specialties or description mentions one of your keywords. Off by default: discovery is already keyword-scoped, and this drops companies that describe themselves with other words.

## `requireWebsite` (type: `boolean`):

Drop companies with no website (a domain is the key the other steps of a prospecting pipeline chain on).

## `requireHqInLocation` (type: `boolean`):

Keep only companies headquartered in the searched location (city or country). Off by default: listings also return companies with an office or a landing page in the location — every row carries `hqMatchesLocation` so you can filter afterwards.

## `keepUnknown` (type: `boolean`):

When a size / founded / type filter is set, also keep companies for which that fact could not be determined.

## `useBusinessListings` (type: `boolean`):

Discover companies from map/business listings (name, category, website, phone, address, rating). Strong for any niche with a physical presence.

## `useCompanyProfiles` (type: `boolean`):

Discover companies through their public professional-network profiles (industry, headcount, headquarters, founded year). Strong for B2B / tech / remote-first companies with no storefront.

## `enrichWithProfile` (type: `boolean`):

Read every company's public profile for industry, size band, headquarters, type, founded year, specialties, description and follower count. Needed for the size / founded / type filters.

## `enrichWithWebsite` (type: `boolean`):

Read the company home page for the description, social profiles, phone and the profile link.

## `maxPerQuery` (type: `integer`):

How many candidates each keyword × location may add from each source before filtering.

## `countryCode` (type: `string`):

Two-letter country code for the listing searches (us, gb, de, …). Derived from the location when it names a country.

## `language` (type: `string`):

Two-letter language code for the listing searches.

## `maxConcurrency` (type: `integer`):

How many companies are enriched at once.

## `proxyConfiguration` (type: `object`):

Proxy used to read public company profiles (they rate-limit datacenter IPs — Apify residential proxy is recommended and is the default). Listing and web searches use Apify's datacenter proxy; company websites are fetched directly with an automatic proxy fallback.

## Actor input object example

```json
{
  "keywords": [
    "SaaS",
    "marketing agency",
    "commercial real estate"
  ],
  "locations": [
    "New York",
    "Toronto"
  ],
  "companySizes": [],
  "maxCompanies": 100,
  "companies": [
    "gong.io",
    "lemlist",
    "https://www.linkedin.com/company/stripe"
  ],
  "companyTypes": [],
  "requireIndustryMatch": false,
  "requireWebsite": true,
  "requireHqInLocation": false,
  "keepUnknown": false,
  "useBusinessListings": true,
  "useCompanyProfiles": true,
  "enrichWithProfile": true,
  "enrichWithWebsite": true,
  "maxPerQuery": 100,
  "countryCode": "us",
  "language": "en",
  "maxConcurrency": 8,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `companies` (type: `string`):

One row per company: name, domain, industry, size band, headquarters, founded, type, description.

## `firmographics` (type: `string`):

Size, headcount range, founded year, type, headquarters, industry and specialties per company.

## `contacts` (type: `string`):

Website, phone, address, social profiles and listing link per company.

## `summary` (type: `string`):

Counts per source, filters, charges.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": [
        "software companies"
    ],
    "locations": [
        "Austin, TX"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("inovaflow/company-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": ["software companies"],
    "locations": ["Austin, TX"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("inovaflow/company-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": [
    "software companies"
  ],
  "locations": [
    "Austin, TX"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call inovaflow/company-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,inovaflow/company-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/oCec5BLC3g5ut6thQ/builds/DJhAR0LmgLZs7WoNl/openapi.json
