# Company Enrichment from Domain: Logo, Socials, Hiring (`moha-tah/website-company-profile`) Actor

Turn a list of company domains into clean company profiles: name, description, logo, social profiles, key pages (pricing, careers, docs), schema.org company facts, hiring signal with ATS and open-job count, email provider and DMARC. No personal data. Pay per domain.

- **URL**: https://apify.com/moha-tah/website-company-profile.md
- **Developed by:** [Mohamed T.](https://apify.com/moha-tah) (community)
- **Categories:** Lead generation, Business, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**Turn a list of company domains into clean, CRM-ready company profiles.** For each website you get the company name, description, **logo**, **social profiles**, **key pages** (pricing, careers, docs, contact, blog…), **schema.org company facts** (legal name, founding year, HQ address), a **hiring signal** (careers page, applicant tracking system and **number of open jobs**) and the **email setup** (Google Workspace vs Microsoft 365, SPF, DMARC policy).

A pay-per-domain Clearbit-style enrichment built on public website data and public DNS. **No personal data**: no employee names, no personal emails, no phone numbers.

### What does it do?

For every domain the Actor loads the homepage (plain HTTP, robots.txt respected) and extracts:

- **Identity**: name, page title, meta description, language, hreflang locales, country hint.
- **Brand assets**: logo URL (from schema.org, header logo or touch icon), favicon, Open Graph image.
- **Social profiles**: LinkedIn company page, X/Twitter, Facebook, Instagram, YouTube, GitHub, TikTok, Discord, Crunchbase, G2, Trustpilot… (company pages only).
- **Key pages**: about, careers, contact, pricing, blog, docs, login, privacy, terms.
- **Company facts** from schema.org Organization markup: legal name, founding date, HQ address, `sameAs` links, employee count when published.
- **Hiring**: careers page, detected ATS (Greenhouse, Lever, Ashby, Workable, SmartRecruiters, Recruitee, Personio, Teamtailor, BambooHR, Workday, Welcome to the Jungle), and **open-job count** from the ATS's public job-board API (Greenhouse, Lever, Ashby, Workable, Recruitee, Personio).
- **Email setup** from DNS: email provider (MX), SPF record, DMARC policy (`none` / `quarantine` / `reject`).

### Why use it?

- **Enrich leads and CRM accounts** in bulk before outreach (HubSpot, Salesforce, Clay, Sheets).
- **Hiring = buying signal**: prioritize companies with many open roles, or a specific ATS.
- **Qualify by stack**: Google Workspace vs Microsoft 365 is a strong ICP filter for B2B SaaS.
- **Clean data for directories and marketplaces**: logos, descriptions and socials for thousands of companies.
- **Email deliverability and security audits**: spot domains without DMARC or with `p=none`.

### How to use it

1. Paste company websites into **Domains or URLs** (one per line).
2. Keep the defaults, or switch off *Count open jobs* / *Email provider & security* for faster runs.
3. Click **Start** (about 200 domains per minute), then export to CSV/Excel/JSON or pull via API.

### Input example

```json
{
  "domains": ["stripe.com", "apify.com", "qonto.com"],
  "detectHiring": true,
  "countJobs": true,
  "emailSetup": true
}
```

### Output example

```json
{
  "domain": "apify.com",
  "url": "https://apify.com/",
  "ok": true,
  "name": "Apify",
  "description": "Thousands of tools to automate your business. Get real-time web data, track competitors, generate leads, and integrate your apps and AI agents.",
  "logoUrl": "https://apify.com/img/apify-logo/apify-symbol-200x200.svg",
  "language": "en",
  "countryHint": "CZ",
  "organization": {
    "name": "Apify", "legalName": "Apify Technologies s.r.o.", "foundingDate": "2015",
    "address": {"streetAddress": "Na Příkopě 959/27", "locality": "Prague", "postalCode": "11000", "country": "CZ"}
  },
  "socialProfiles": {"linkedin": "http://linkedin.com/company/apify/", "x": "https://x.com/apify", "github": "https://github.com/apify/crawlee", "youtube": "https://www.youtube.com/apify", "discord": "https://discord.com/invite/jyEM2PRvMU"},
  "keyPages": {"about": "https://apify.com/about", "careers": "https://apify.com/jobs", "pricing": "https://apify.com/pricing", "contact": "https://apify.com/contact", "login": "https://console.apify.com/sign-in"},
  "hiring": {"careersPage": "https://apify.com/jobs", "atsProvider": "ashby", "atsBoard": "apify", "openJobs": 10, "isHiring": true},
  "email": {"hasMx": true, "providers": ["Google Workspace"], "spf": "v=spf1 a mx include:_spf.google.com include:mailgun.org include:amazonses.com ... -all", "dmarcPolicy": "reject"}
}
```

You can download the dataset in various formats such as JSON, CSV, Excel, XML or HTML. The **Company profiles** view shows logos inline.

### How much does it cost?

You pay **per successfully enriched domain**. Sites that are unreachable or block automated visitors are **not charged**. See the **Pricing** tab.

### Known limits

- Some sites block automated visitors (HTTP 403). They're returned with `ok: false` and are free.
- Sites built entirely with client-side JavaScript may expose fewer links and no logo in their HTML.
- Open-job counts come only from ATSs with a public job-board API. For other ATSs (Workday, SmartRecruiters, Teamtailor…) the provider is detected but the count is `null`.
- `countryHint` comes from schema.org address or the country-code domain. `.com` sites without markup get `null`.

### FAQ

**Is this legal / GDPR-friendly?**
It reads public company websites and public DNS, respects robots.txt, and deliberately skips personal data: person profiles (e.g. `linkedin.com/in/…`), personal names in schema.org (founders, authors), emails and phone numbers are never extracted.

**Can I call it from Clay, Make, Zapier or n8n?**
Yes. Use the Apify API "run Actor synchronously and get dataset items" with one domain per call, or batch many domains in one run.

**A field is wrong for my domain?**
Open an issue on the **Issues** tab with the domain. Heuristics are improved every week.

# Actor input Schema

## `domains` (type: `array`):

Company websites, one per line (stripe.com or https://stripe.com/ both work).

## `detectHiring` (type: `boolean`):

Finds the careers page and the applicant tracking system (Greenhouse, Lever, Ashby, Workable, SmartRecruiters, Recruitee, Personio, Teamtailor, BambooHR, Workday…).

## `countJobs` (type: `boolean`):

Counts open positions through the ATS's public job-board API when available (Greenhouse, Lever, Ashby, Workable, Recruitee, Personio).

## `emailSetup` (type: `boolean`):

Reads public DNS: email provider from MX records, SPF record and DMARC policy.

## `respectRobotsTxt` (type: `boolean`):

Skip pages disallowed for crawlers.

## `maxConcurrency` (type: `integer`):

Domains processed in parallel.

## `requestTimeoutSecs` (type: `integer`):

Per request; retries with backoff are automatic.

## Actor input object example

```json
{
  "domains": [
    "stripe.com",
    "apify.com",
    "qonto.com"
  ],
  "detectHiring": true,
  "countJobs": true,
  "emailSetup": true,
  "respectRobotsTxt": true,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "stripe.com",
        "apify.com",
        "qonto.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("moha-tah/website-company-profile").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "domains": [
        "stripe.com",
        "apify.com",
        "qonto.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("moha-tah/website-company-profile").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "stripe.com",
    "apify.com",
    "qonto.com"
  ]
}' |
apify call moha-tah/website-company-profile --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,moha-tah/website-company-profile"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QjDeW81qsO99rYPho/builds/XVwfHxmVzLQLbSrgP/openapi.json
