# Lookalike Company Finder — Similar Companies to Your Customers (`inovaflow/lookalike-company-finder`) Actor

Paste 3–20 customer domains and get similar companies back, scored 0–100 with the reasons: industry, keywords, size band, tech overlap, geography. Domain, description, HQ, LinkedIn per row. No login, dataset-only, MCP-ready.

- **URL**: https://apify.com/inovaflow/lookalike-company-finder.md
- **Developed by:** [inovaflow](https://apify.com/inovaflow) (community)
- **Categories:** Lead generation, Business, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $30.00 / 1,000 lookalikes

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Your best next customers look like your best current customers. **Lookalike Company Finder** takes a few example domains — three customers, a segment, a competitor's logos page — builds a profile of what they have in common, and returns **similar companies scored 0–100 with the reasons**: shared keywords, industry, size band, tech overlap, geography, and how many competitor lists name them together. Every row carries the domain, description, HQ, employee band and LinkedIn page, read from the company's own website and public LinkedIn page. No login, no API key, no data vendor.

- **Sales ops & ABM** — turn "accounts like Acme" into a scored target list for the territory.
- **SDRs & founders** — expand a hand-picked list of 5 ideal customers to 50 more of the same shape.
- **Agencies** — find every local business that looks like the three you already serve in a city.
- **Investors & researchers** — map the competitive set around any company from its domain.

### What does Lookalike Company Finder do?

It is a **similar-companies finder driven by example domains**, not by filters you have to guess:

1. **Profiles the seeds.** Reads each seed's website (title, description, keywords, headline text, structured data, tech stack), its public LinkedIn company page (industry, company size, headquarters, specialties) and its Product Hunt product (topics, team size).
2. **Finds candidates.** Web searches for `X competitors`, `X alternatives`, `companies like X` (Google, with Bing / DuckDuckGo fallbacks), the "10 best X alternatives" articles those searches surface (the companies they link and name), the Product Hunt topics the seeds sit in, and keyword searches built from the profile (`sales crm companies`, `drain cleaning in Denver`).
3. **Reads every candidate the same way** and scores it against the profile. The top `maxResults` are delivered, best first, each with `matchedOn` — the evidence in plain words — and a `scoreBreakdown`.

### Why use this lookalike company finder?

- **Explained scores.** `matchedOn: ["keywords: sales crm, pipeline, deals", "industry: Software Development", "tech: Segment, Intercom", "named in 3 competitor lists"]` — you can see why a company is on the list before you spend a credit on it.
- **Real fields or null.** Industry, size and HQ come from the company's public LinkedIn page or its own structured data; nothing is inferred from the name. Signals without evidence are left out of the score, not guessed.
- **Works beyond SaaS.** Local services, agencies, manufacturers: the profile uses the seeds' own words and locations, so "three Denver plumbers" finds Denver plumbers.
- **Pay per delivered row.** Seeds, excluded domains, duplicates, unreachable sites and everything under your minimum score cost nothing.
- **CRM-ready.** Dataset + `LOOKALIKES.csv`; run it from the API, a schedule, Zapier / Make / n8n, or an AI agent through MCP.

### What data does it extract?

| Field | Description |
| --- | --- |
| `company`, `domain`, `website` | Company name (LinkedIn → site name → domain), registrable domain (the row key), home page URL |
| `similarityScore` | 0–100 against the seed profile |
| `matchedOn` | Plain-words evidence: keywords, industry, tech, size, geography, competitor lists / topics naming it |
| `scoreBreakdown` | Per-signal scores 0–1 (`keywords`, `industry`, `tech`, `size`, `geography`, `cooccurrence`), `null` where there was no evidence, plus `coverage` |
| `industry`, `industrySource` | LinkedIn industry label (`linkedin`) or keyword-classified label (`keywords`) |
| `employeeBand`, `employeesOnLinkedIn` | `1-10` … `10001+` from LinkedIn / structured data; the LinkedIn head-count when shown |
| `hq`, `country` | Headquarters (LinkedIn or JSON-LD address) and ISO country |
| `techOverlap`, `techStack` | Business technologies shared with the seeds; every technology fingerprinted on the site |
| `description`, `tagline`, `keywords`, `specialties` | What the company says it does; the terms it shares with the profile; LinkedIn specialties |
| `linkedin`, `twitter`, `productHuntUrl`, `foundedYear` | Public profiles and founding year when published |
| `seedsMatched`, `foundVia`, `listicleMentions`, `sources` | Which seeds it was linked to, how it was found (`serp`, `listicle`, `producthunt`, `keyword`), how many competitor lists name it, every source URL / query |
| `scrapedAt` | Provenance |

### How to find lookalike companies

1. Open the Actor and click **Try for free**.
2. Under **Example customer domains** paste 3–20 domains of companies you want more of.
3. Optionally set **Locations** (`United States`, `Berlin, Germany`), **Company sizes** (`51-200`, `mid-market`), **Only companies in the seeds' industry**, and **Exclude domains** (existing customers).
4. Click **Start**. A 20-row run takes 3–6 minutes (it reads ~80 pages); results arrive sorted by score.
5. Export CSV / Excel / JSON or open `LOOKALIKES.csv`. The `OUTPUT` record holds the **seed profile** (industries, key phrases, shared tech, sizes, countries) so you can check what the run understood your seeds to be.

#### Tips for better lookalikes

- **Homogeneous seeds.** Five sales-CRM vendors give a sharp profile; a CRM, a bakery and a law firm give mush. Run one segment at a time.
- **Add the odd one out to `excludeDomains`** and re-run — the profile is rebuilt every run.
- **Raise `minSimilarityScore` to 50** when you only want the obvious peers; lower it to 15 for a wide net you will curate by hand.
- **Local segments:** give the seeds' cities in **Locations** (or let the run pick them up from the seeds' HQ) so keyword searches stay local.

### How much does it cost?

Pay-per-event, no subscription: **$0.03 per lookalike company delivered** plus a small run-start fee. 50 lookalikes ≈ $1.50. Rows filtered out by your settings, duplicates, seeds and low scores are free. Apify's free plan covers a few test runs a month.

### Input

Only **Example customer domains** is required. Example:

```json
{
    "seedDomains": ["pipedrive.com", "close.com", "folk.app"],
    "maxResults": 20,
    "locations": ["United States"],
    "companySizes": ["11-50", "51-200", "201-500"],
    "excludeDomains": ["salesforce.com"]
}
```

### Output

One dataset item per company:

```json
{
    "company": "Freshsales",
    "domain": "freshworks.com",
    "website": "https://www.freshworks.com/",
    "similarityScore": 74,
    "matchedOn": ["keywords: sales crm, pipeline, deals", "industry: Software Development", "tech: Segment", "size: 5001-10000 employees", "geography: US", "named in 4 competitor lists; search result for a seed; linked to 2 seeds"],
    "scoreBreakdown": { "keywords": 0.81, "industry": 1, "tech": 0.4, "size": 0.25, "geography": 1, "cooccurrence": 0.9, "coverage": 1 },
    "industry": "Software Development",
    "industrySource": "linkedin",
    "employeeBand": "5001-10000",
    "hq": "San Mateo, California",
    "country": "US",
    "techOverlap": ["Segment"],
    "description": "Freshsales is an AI-powered sales CRM …",
    "linkedin": "https://www.linkedin.com/company/freshworks-inc",
    "seedsMatched": ["close.com", "pipedrive.com"],
    "foundVia": ["listicle", "serp"],
    "listicleMentions": 4,
    "sources": ["bing:Pipedrive alternatives", "listicle:zapier.com/blog/pipedrive-alternatives", "https://www.freshworks.com/", "https://www.linkedin.com/company/freshworks-inc"],
    "scrapedAt": "2026-09-26T14:02:11.000Z"
}
```

The key-value store also holds **`LOOKALIKES.csv`** and **`OUTPUT`** (run summary: seed profile, candidates found / read / dropped per filter, search and transport statistics, notices).

### FAQ

#### Where do the similar companies come from?

From the public web, found the way an analyst would: competitor / alternative searches, the articles that list alternatives, Product Hunt topics and keyword searches. There is no vendor database behind it — which is why every row also carries the source it was found through.

#### Why is `employeeBand` or `hq` sometimes null?

They are read from the company's public LinkedIn page (linked from its website) or its structured data. A company without a linked LinkedIn page or with a LinkedIn auth wall keeps those fields null; the score is then computed on the remaining signals and `scoreBreakdown.coverage` shows it.

#### Does it need my LinkedIn or Google account?

No. LinkedIn company pages are read as a logged-out visitor sees them; web search uses Apify's Google SERP proxy with Bing / DuckDuckGo fallbacks.

#### Can I feed it company names instead of domains?

Not yet — domains are unambiguous. Use the company's main website domain.

#### Is this legal?

The Actor reads what companies publish about themselves on their own websites, their public LinkedIn company page and Product Hunt. No personal data is collected. Check your local rules before contacting the companies.

### Support

Open an issue in the **Issues** tab with your run ID and the seeds that misbehaved. The **API** tab shows how to call this Actor from code or from any MCP-capable AI agent.

# Actor input Schema

## `seedDomains` (type: `array`):

3–20 domains of companies you want more of — your best customers, or a segment (`hubspot.com`, `pipedrive.com`, `close.com`). One domain works but gives a thin profile; three or more is where it gets sharp. Full URLs are fine.

## `maxResults` (type: `integer`):

How many companies to deliver, best score first. About three times as many candidates are read to pick them.

## `locations` (type: `array`):

Countries or cities the lookalikes should be in — `United States`, `Germany`, `Denver, CO`. Companies whose known HQ is elsewhere are dropped; companies with no known location are kept and scored without geography. Cities also steer the keyword searches (useful for local-service segments).

## `companySizes` (type: `array`):

Employee bands to keep: `1-10`, `11-50`, `51-200`, `201-500`, `501-1000`, `1001-5000`, `5001-10000`, `10001+` (or `small`, `mid-market`, `enterprise`). Companies whose known band is outside the list are dropped; unknown sizes are kept. Leave empty to match the seeds' own sizes.

## `mustMatchIndustry` (type: `boolean`):

Drop candidates whose industry (public LinkedIn label, or keyword-classified) is neither a seed's industry nor its family.

## `excludeDomains` (type: `array`):

Existing customers, competitors or anything else you never want back. The seeds themselves are always excluded.

## `minSimilarityScore` (type: `integer`):

Companies scoring below this (0–100) are not delivered or charged. 25 keeps clear lookalikes; 50 keeps only the close ones.

## `enrichResults` (type: `boolean`):

On: every candidate's public LinkedIn page (industry, size, HQ, specialties) and about page are read too — fuller rows, a few more requests. Off: only the home page is read; industry and size then come from the page text and structured data where present.

## `maxSearchQueriesPerSeed` (type: `integer`):

How many competitor / alternative searches to run for each seed (`X competitors`, `X alternatives`, `companies like X`, `X vs`).

## `maxListicles` (type: `integer`):

Search results that are articles listing alternatives ("10 best X alternatives") are read for the companies they name. This caps how many per run.

## `language` (type: `string`):

Language for web search results, e.g. `en`, `de`, `fr`.

## `countryCode` (type: `string`):

Two-letter country for web search defaults (`us`, `gb`, `de`). Does not filter the companies — use Locations for that.

## `maxConcurrency` (type: `integer`):

How many candidate websites are read at once.

## `proxyConfiguration` (type: `object`):

Fallback transport for company websites, LinkedIn and Product Hunt (every host is read directly first; the proxy is used only when a host blocks that). Residential is the default. Web search uses the Google SERP proxy group automatically, with Bing / DuckDuckGo as fallbacks.

## Actor input object example

```json
{
  "seedDomains": [
    "pipedrive.com",
    "close.com",
    "folk.app"
  ],
  "maxResults": 8,
  "mustMatchIndustry": false,
  "minSimilarityScore": 25,
  "enrichResults": true,
  "maxSearchQueriesPerSeed": 4,
  "maxListicles": 30,
  "language": "en",
  "countryCode": "us",
  "maxConcurrency": 8,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `lookalikes` (type: `string`):

One row per similar company: score, reasons, industry, size, HQ, tech overlap, description, LinkedIn.

## `csv` (type: `string`):

Spreadsheet / CRM-ready CSV of the delivered rows.

## `summary` (type: `string`):

Seed profile (industries, key phrases, shared tech, sizes, countries), candidates found / read / dropped, search engine and transport stats, notices.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "seedDomains": [
        "pipedrive.com",
        "close.com",
        "folk.app"
    ],
    "maxResults": 8,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("inovaflow/lookalike-company-finder").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "seedDomains": [
        "pipedrive.com",
        "close.com",
        "folk.app",
    ],
    "maxResults": 8,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("inovaflow/lookalike-company-finder").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "seedDomains": [
    "pipedrive.com",
    "close.com",
    "folk.app"
  ],
  "maxResults": 8,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call inovaflow/lookalike-company-finder --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,inovaflow/lookalike-company-finder"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wObwyvk3hbxxrx1zu/builds/EHi6LbhfajXwi48HH/openapi.json
