# Company Enrichment by Domain: Logo, Social Profiles, Contacts (`datagleaner/company-enrichment-by-domain`) Actor

Company enrichment by domain for agents and CRMs, a Clearbit alternative: name, description, logo, official social profiles (LinkedIn company page, X, Facebook, Instagram, YouTube, TikTok, GitHub), email, phone, address, founding date. $0.80 per 1,000 enriched companies; misses free.

- **URL**: https://apify.com/datagleaner/company-enrichment-by-domain.md
- **Developed by:** [Data Gleaner](https://apify.com/datagleaner) (community)
- **Categories:** Lead generation, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.80 / 1,000 enriched companies

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Company Enrichment by Domain: logo, social profiles and contacts from any website

**Company enrichment by domain, built for AI agents and CRM pipelines.** Give it a list of domains and get back one company profile per domain: **company name, description, logo, the company's own official social profiles (LinkedIn company page, X, Facebook, Instagram, YouTube, TikTok, GitHub, Pinterest), general email and phone, address, country, founding date and light tech hints**. It answers the lookups agents make before outreach: *company enrichment by domain*, *find company socials*, *domain to social profiles*, *company social media links from website*, *company info from domain*, and *a Clearbit alternative* that needs no API key.

**$0.80 per 1,000 enriched companies ($0.0008 each).** You pay only for a domain that returns a company name plus at least one social profile or contact. Unreachable, parked and empty domains are free.

### What it does

For each domain it reads the home page, plus the contact and about pages the home page links to (at most 3 pages by default), over plain HTTP with no browser. From those pages it builds one profile:

- **Name, legal name, description and logo** from the site's schema.org Organization data (JSON-LD), Open Graph tags, the meta description and the page title. The logo is the JSON-LD logo, else the apple-touch-icon or the largest declared icon.
- **Official social profiles**, one per platform. A profile counts only when the site itself claims it: in its JSON-LD `sameAs`, in its header, footer or navigation, or under a handle that names the brand or domain. Profiles linked from the page body (customer testimonials, blog authors, quoted tweets) are left out, and LinkedIn returns company pages only, never a person's `/in/` profile.
- **General email and phone**: the site's own addresses (role addresses such as info@, hello@ and sales@ first) and phone numbers validated and normalised to E.164, home-country numbers first. Addresses on other domains (partners, demo data) are dropped.
- **Address, country and founding date** from JSON-LD. Country comes from the address, else the country TLD, else the main phone's country code.
- **`sameAs`**: every link the company lists about itself, which often includes Wikipedia, Wikidata and Crunchbase.
- **Tech hints** read from the HTML and response headers at no extra cost: Shopify, WordPress, WooCommerce, Webflow, Wix, Squarespace, Framer, HubSpot, Intercom, Zendesk, Google Analytics, Google Tag Manager, Meta Pixel, Stripe, Klaviyo, Marketo, Next.js, Cloudflare, Vercel and more.
- **Parked domains** (for-sale and parking pages) are flagged as `parked` and are free.

### Use cases

- **AI agents** that need "who is this company, and where are its official accounts" before writing an outreach email or filling a CRM record.
- **CRM and lead enrichment**: turn a column of domains into names, logos, LinkedIn company pages and contact routes.
- **Social profile discovery**: find a company's official X, Instagram, TikTok or YouTube from its website.
- **List cleaning**: spot dead and parked domains in a lead list at no cost.
- **Light technographics**: which leads run Shopify, WordPress or HubSpot.

### Use it from n8n, Make, Zapier or an AI agent

Actor ID: `datagleaner/company-enrichment-by-domain`

Minimal input:

```
{"domains": ["stripe.com", "hubspot.com"]}
```

Each tool below runs this Actor with your own Apify API token.

- **n8n:** add the **Apify** node (`@apify/n8n-nodes-apify`). On n8n Cloud you install it from the community node registry. Choose **Run an Actor and get dataset**, set Actor to `datagleaner/company-enrichment-by-domain` and paste the input above.

- **Make:** use the Apify app's **Run an Actor** module, then **Get Dataset Items** to read the results. **Watch Actor Runs** can trigger a scenario when a run finishes.

- **Zapier:** use the Apify action **Run Actor**, then the search **Fetch dataset items**. The trigger **Finished Actor run** starts a Zap when a run ends.

- **AI agents (MCP):** connect to `https://mcp.apify.com/?tools=datagleaner/company-enrichment-by-domain`. In Claude Code:

  ```
  claude mcp add --transport http apify "https://mcp.apify.com/?tools=datagleaner/company-enrichment-by-domain"
  ```

  Then run `/mcp` to sign in to Apify in your browser. Other clients can sign in with OAuth or send the header `Authorization: Bearer YOUR_APIFY_TOKEN`. Clients that run local MCP servers can use Apify's package (`@apify/actors-mcp-server`, run with `npx -y` and `APIFY_TOKEN` set) instead. Then ask the agent in plain words, for example:

  > Enrich stripe.com and hubspot.com: company name, LinkedIn page and contact email. Put the results in a table.

- **LangChain (Python):**

```python
## pip install langchain-apify, then set APIFY_TOKEN in your environment
import json
from langchain_apify import ApifyActorsTool
tool = ApifyActorsTool("datagleaner/company-enrichment-by-domain")
result = tool.invoke({"run_input": json.loads('{"domains": ["stripe.com", "hubspot.com"]}')})
```

Through the API, one call returns the rows:

```bash
curl -X POST "https://api.apify.com/v2/acts/datagleaner~company-enrichment-by-domain/run-sync-get-dataset-items?token=<APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"domains": ["stripe.com", "linear.app"]}'
```

Each domain is one row with the same fields every time, so an agent can read `socials.linkedin` or `emails[0]` without parsing. A typical domain takes 1 to 3 seconds, and a failure costs nothing.

### Input

```json
{
  "domains": ["stripe.com", "https://www.hubspot.com", "plausible.io"],
  "maxPagesPerDomain": 3,
  "includeTechHints": true,
  "maxConcurrency": 10,
  "delaySecs": 0.3
}
```

| Field | Meaning |
|---|---|
| `domains` | Domains or URLs, one per line. Any path is ignored: the lookup starts at the home page. |
| `maxPagesPerDomain` | Home page plus linked contact and about pages (default 3, max 10). `1` reads the home page only. |
| `includeTechHints` | Return `techHints` (default true). Costs no extra requests. |
| `maxConcurrency` | Domains looked up in parallel (default 10). |
| `delaySecs` | Pause between two requests to the same site (default 0.3). |
| `requestTimeoutSecs`, `respectRobotsTxt`, `proxyConfiguration` | Network options. Leave the proxy empty: sites are read directly, and a site that blocks with HTTP 403 or 429 is retried through Apify Proxy automatically. Set a proxy only to send every request through it. |

Leave `domains` empty to run a 3-domain example.

### Output

One row per domain. Fields that were not found are empty or `null`, never guessed.

```json
{
  "input": "stripe.com",
  "domain": "stripe.com",
  "finalUrl": "https://stripe.com/",
  "status": "ok",
  "enriched": true,
  "companyName": "Stripe",
  "legalName": "Stripe, LLC",
  "description": "Stripe powers online and in-person payment processing and financial solutions for businesses of all sizes.",
  "logoUrl": "https://images.stripeassets.com/.../favicon.svg",
  "foundingDate": null,
  "address": "920 5th Avenue, Suite 1900, Seattle, WA, 98104, US",
  "country": "US",
  "phones": ["+18889262289"],
  "emails": [],
  "socials": {
    "linkedin": "https://www.linkedin.com/company/stripe",
    "x": "https://x.com/stripe",
    "facebook": "https://www.facebook.com/StripeHQ",
    "instagram": "https://www.instagram.com/stripehq",
    "youtube": "https://www.youtube.com/@stripe",
    "tiktok": null,
    "github": "https://github.com/stripe",
    "pinterest": null
  },
  "sameAs": ["https://twitter.com/stripe", "https://en.wikipedia.org/wiki/Stripe,_Inc.", "https://www.crunchbase.com/organization/stripe", "https://www.wikidata.org/wiki/Q7624104"],
  "techHints": ["Next.js"],
  "language": "en-US",
  "contactPageUrl": "https://stripe.com/contact/sales",
  "pagesCrawled": 3,
  "pageUrls": ["https://stripe.com/", "https://stripe.com/contact/sales", "https://stripe.com/about"],
  "error": null,
  "scrapedAt": "2026-10-09T14:37:48+00:00"
}
```

`status` is `ok` (the site was read), `unreachable` (network or HTTP error), `parked` (parking or for-sale page), `invalidDomain`, `blockedByRobots` or `error`. `enriched` tells you whether the row was billed. The **Company profiles** view in the Output tab shows one column per social platform, ready for CSV or Google Sheets.

### Pricing

Pay per event, one event:

- **`company-enriched`: $0.0008** for one domain that returned a company name plus at least one social profile or contact (email, phone or address).
- Unreachable, parked and empty domains are **free**, and still returned so you can see why.

**Worked example.** 5,000 CRM domains, of which about 80% return a usable profile: 4,000 x $0.0008 = **$3.20**. Platform usage is small: a domain takes about five plain HTTP requests (robots.txt, the home page and up to two linked pages).

**Compared with others.** Company enrichment Actors on the Store charge $0.001 to $0.02 per company, and a LinkedIn company finder charges $0.008 for the LinkedIn URL alone; here $0.0008 returns the LinkedIn page plus the other socials, the logo, contacts and description, and a miss costs nothing.

Set a maximum total charge on the run and the Actor stops cleanly when it is reached.

### What to expect

In our 20-domain test (SaaS, e-commerce and Japanese companies), 16 domains were enriched. Of the 19 sites read, 18 returned a name and description, 17 a logo, 13 an X profile, 12 Facebook, 11 YouTube, 9 Instagram and 8 a LinkedIn company page. Emails (7), phones (6) and founding dates (5) appear only where the site publishes them.

### Limits

- **No JavaScript rendering.** A site that draws its footer or contact details with scripts only (Notion, Canva) returns its name and logo but no socials, and is then free.
- **Published data only.** Founding date, address and phone come from what the site publishes; nothing is looked up in a third-party database, and employee counts or revenue are not estimated.
- **Sites that block bots.** A site that answers HTTP 403 or 429 (many Shopify stores do) is retried once through Apify Proxy, which usually gets through. Sites behind a hard bot wall (Canva, for example) still come back `unreachable`, and are free.
- **Geo-redirects** (a store that sends visitors to a regional checkout) may return a thin page.
- At most `maxPagesPerDomain` pages per domain, and only pages the home page links to.

### FAQ

**Is this a Clearbit alternative?** For the public part of a company profile, yes: name, logo, description, socials and contacts from the company's own website, with no API key, no subscription and no charge for misses. It does not estimate headcount or revenue.

**How do I find a company's LinkedIn page from its domain?** Put the domain in `domains`; `socials.linkedin` holds the company page the site links to or lists in its structured data. Person profiles (`/in/`) are never returned.

**Why is a social profile missing when I can see it on the site?** Either the site renders it with JavaScript, or it appears only in the page body under a handle that does not name the brand, which this Actor treats as someone else's profile. That rule is what keeps testimonial and author accounts out.

**Do I need an API key or login?** No. Only public pages are read.

**Can I get every email and contact form on the site instead?** Use [Website Contact Details Scraper](https://apify.com/datagleaner/website-contact-details-scraper): it reads up to 30 pages per site and returns every email, phone, contact form and social profile with the page it was found on. This Actor is the fast one-row-per-company lookup.

**Can I export to CSV or a CRM?** Yes: download CSV, Excel or JSON, or connect Google Sheets, Zapier, Make, Clay or your own code through Apify integrations and the API.

### Related Actors

- [Website Contact Details Scraper](https://apify.com/datagleaner/website-contact-details-scraper): every email, phone, social profile and contact form on a website, page by page.

### Responsible use

This Actor reads publicly available web pages only. Company emails and phone numbers can still be personal data. You are responsible for complying with the law that applies to you, including data protection rules (GDPR, CCPA and similar) and anti-spam law, when you store or contact the companies in the results.

# Actor input Schema

## `domains` (type: `array`):

Company domains or website URLs, one per line. `stripe.com`, `www.stripe.com` and `https://stripe.com/pricing` all work; the lookup always starts at the home page. Leave empty to run a 3-domain example.

## `maxPagesPerDomain` (type: `integer`):

The home page, plus the contact and about pages it links to. 3 covers almost every site; 1 reads the home page only (fastest).

## `includeTechHints` (type: `boolean`):

Name the platforms the site's HTML reveals (Shopify, WordPress, Webflow, Wix, HubSpot, Intercom, Google Analytics, Stripe and more). Costs no extra requests.

## `maxConcurrency` (type: `integer`):

How many domains are looked up at the same time.

## `delaySecs` (type: `number`):

Polite pause between two requests to the same website. Different domains run in parallel regardless.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for each page before retrying.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that the site's robots.txt disallows.

## `proxyConfiguration` (type: `object`):

Optional. Leave empty: sites are read directly, and a site that blocks with HTTP 403 or 429 is retried through Apify Proxy automatically. Set a proxy here only to send every request through it.

## Actor input object example

```json
{
  "domains": [
    "stripe.com",
    "hubspot.com",
    "plausible.io"
  ],
  "maxPagesPerDomain": 3,
  "includeTechHints": true,
  "maxConcurrency": 10,
  "delaySecs": 0.3,
  "requestTimeoutSecs": 15,
  "respectRobotsTxt": true
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "stripe.com",
        "hubspot.com",
        "plausible.io"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("datagleaner/company-enrichment-by-domain").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "stripe.com",
        "hubspot.com",
        "plausible.io",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("datagleaner/company-enrichment-by-domain").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "stripe.com",
    "hubspot.com",
    "plausible.io"
  ]
}' |
apify call datagleaner/company-enrichment-by-domain --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datagleaner/company-enrichment-by-domain"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4YZPuXN2ejg4bpTba/builds/8e6GVplUhRiR3sqxz/openapi.json
