# Company Enrichment Scraper — Address, Emails & Tech Stack (`tactful_anvil/company-website-enrichment`) Actor

Turn a list of domains into company profiles: legal name, office address, emails, phones, LinkedIn & socials, tech stack, email provider, SaaS tools from DNS, domain age, VAT/registration numbers, hiring. HTTP-only. $3.20/1,000 companies.

- **URL**: https://apify.com/tactful_anvil/company-website-enrichment.md
- **Developed by:** [Mr Zack](https://apify.com/tactful_anvil) (community)
- **Categories:** Lead generation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.20 / 1,000 company enricheds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Company Enrichment Scraper — Address, Emails & Tech Stack

Give it a list of **company websites or domains** and get back one clean **company profile per domain**: company & legal name, **office address** (street, city, postal code, country), emails, phones, LinkedIn and other socials, **VAT & registration numbers**, **tech stack** (CMS, e-commerce, analytics, ad pixels, chat, payments), **email provider** and **SaaS tools verified in DNS**, **domain age**, careers page & ATS.

**$3.20 per 1,000 companies.** Websites that can't be reached, invalid inputs and duplicates are **never charged**. HTTP-only — no browser, no login, no API keys.

### Who this is for

- **Sales & SDR teams** — turn a raw domain list (CRM export, Google Maps results, event attendee list) into sales-ready rows: who they are, where they sit, how to reach them, what they already use.
- **Agencies** — qualify prospects in one pass: `runsPaidAds` (Meta/Google/LinkedIn/TikTok pixels present) + `cms` + `ecommercePlatform` tell you who needs what.
- **SaaS go-to-market** — find companies on a competitor's stack (`techStack`, `verifiedTools`, `emailSendingServices`) or on Google Workspace vs Microsoft 365 (`emailProvider`).
- **Data & RevOps** — fill missing firmographics (legal entity, country, founded year, domain age) and de-duplicate accounts by registration number.
- **AI agents (MCP)** — *"who is behind acme.de, where are they based, and how do I contact them?"* is one call with a predictable price.

### What you get per company

| Field | Example | How it's found |
|---|---|---|
| `companyName` / `legalName` | `Octopus Energy` / `Octopus Energy Limited` | JSON-LD, site name, imprint, © line |
| `address`, `street`, `city`, `region`, `postalCode`, `country` | `182 Oxford Street, London W1D 1NN` | JSON-LD / microdata / imprint & contact pages (US, UK, EU, CA, AU, ID formats) |
| `primaryEmail`, `emails`, `emailDetails` | `hello@octopus.energy` | mailto, text, Cloudflare-protected, JSON-LD — ranked: company-domain first, real inboxes (`info@`, `sales@`, people) before `privacy@`/`legal@` |
| `primaryPhone`, `phones` | `+448081966842` | tel: links, JSON-LD, phone lines — normalised to E.164 when the country is known |
| `linkedinUrl`, `facebookUrl`, `instagramUrl`, `xUrl`, `youtubeUrl`, `tiktokUrl`, `githubUrl`, `otherSocials` | | links + JSON-LD `sameAs` |
| `vatId`, `registrationNumber`, `registrationCourt` | `DE176055816`, `HRB 263370`, `Stuttgart` | imprint/legal pages (EU VAT, German HR, UK Companies House, KvK, SIREN, ABN, CNPJ, NIB…) |
| `description`, `businessType`, `foundedYear`, `founders`, `logoUrl`, `faviconUrl`, `language` | | meta tags + JSON-LD |
| `techStack`, `cms`, `ecommercePlatform`, `analytics`, `advertisingPixels`, `runsPaidAds`, `marketingTools`, `supportTools`, `paymentProviders`, `hosting` | `Shopify`, `["Meta Pixel","Google Ads"]` | 135+ fingerprints on scripts, headers and cookies, plus the site's public Google Tag Manager container for ad pixels fired through GTM — **never** on the visible text (a site that *mentions* Salesforce is not reported as *using* it) |
| `emailProvider`, `mxHosts` | `Google Workspace` | DNS MX |
| `emailSendingServices` | `["HubSpot","SendGrid"]` | DNS SPF |
| `verifiedTools` | `["Atlassian","Stripe","OpenAI"]` | DNS TXT verification records — SaaS accounts the company has set up |
| `dmarcPolicy` | `reject` | DNS DMARC |
| `domainCreatedAt`, `domainAgeYears`, `domainExpiresAt`, `registrar` | `1995-09-12`, `31` | RDAP (not every country TLD publishes it, e.g. `.de`) |
| `contactPageUrl`, `aboutPageUrl`, `legalPageUrl`, `teamPageUrl`, `careersPageUrl`, `atsProvider`, `isHiring` | | page discovery + ATS fingerprints (Greenhouse, Lever, Ashby…) |
| `dataCompleteness` | `88` | 0–100: how much of the profile was filled |
| `status`, `error`, `pagesCrawled`, `crawlErrors`, `checkedAt` | | transparency |

Every field is always present — empty when the website doesn't publish it, never guessed. Unreachable websites produce a free row with `status: "failed"` and the reason (`domain does not exist`, `HTTP 403`…), so you can see exactly what happened to every input.

### Input

| Field | Default | What it does |
|---|---|---|
| `websites` | 3 examples | Domains or URLs, one per line. Emails (`jane@acme.com`) become their domain |
| `datasetId` + `websiteField` | — / `website` | Enrich another Actor's results directly (Google Maps scrapers, lead lists…). Nested fields as dot paths: `company.website` |
| `maxPagesPerSite` | `6` | Homepage + contact, imprint/legal, about, team, privacy, careers. More pages never cost more |
| `includeDnsIntel` | `true` | Email provider, SPF tools, TXT-verified SaaS, DMARC |
| `includeDomainAge` | `true` | RDAP registration date & registrar |
| `maxConcurrency` | `10` | Websites in parallel |
| `proxyConfiguration` | Apify Proxy | Only used when a website blocks the direct request |

```json
{ "websites": ["octopus.energy", "personio.de", "https://www.allbirds.com"] }
```

### Schedule it / chain it

- **Chain after a lead source:** run a Google Maps or directory scraper, then this Actor with `datasetId` = that run's dataset and `websiteField` = the field holding the website. Every business becomes a full company profile.
- **Keep your CRM fresh:** schedule a monthly run over your account list (Console → Schedules) and watch `techStack`, `emailProvider`, `isHiring` and `runsPaidAds` change over time.
- **Verify the emails:** pipe `primaryEmail` into our **Bulk Email Verifier** before sending.

### For AI agents & MCP

- **Minimal input:** `{"websites": ["<domain>"]}`.
- **Output:** one flat JSON object per domain, stable field names (table above), arrays for multi-value fields.
- **Price is predictable:** $0.0032 per company profile; failed websites cost $0. Crawl depth never changes the price.
- **Honest failures:** a website that can't be reached is a free `status: "failed"` row with a reason — never a fake profile.

### Honest limits

Websites that render everything with JavaScript (some single-page apps) or that block automated requests expose less data — you still get DNS intel, domain age and a `failed`/sparse row, not invented values. Addresses are read in US, UK, Canadian, Australian, European and Indonesian formats; other formats may be missed. Pages in legacy encodings (Shift_JIS, windows-1251, Latin-1…) are decoded correctly. A subdomain (`blog.acme.com`) gets the DNS intel of the company domain (`acme.com`). Links to files (PDF, zip…), parked or for-sale domains, and addresses that resolve to private/internal networks are skipped for free. `founders` exists only when the site publishes them in structured data.

### Related Actors

- **Contact Details Scraper** (`tactful_anvil/contact-details-scraper`) — deeper email/phone/social extraction with MX-verified emails.
- **Bulk Email Verifier** (`tactful_anvil/bulk-email-verifier`) — verify the emails you found.
- **Company Hiring Signals** (`tactful_anvil/company-hiring-signals-scraper`) — every open role from the company's ATS.
- **Ad Budget Signals** (`tactful_anvil/ad-budget-signals`) — does a domain run Google/LinkedIn ads?

### Changelog

- **0.1.1** (2026-10-01) — legacy page encodings decoded (Japanese Shift_JIS sites no longer show garbled names); subdomains get the company domain's email provider & SaaS tools; short/redirecting domains named after their real brand (`fb.com` → Facebook); far-away servers with IPv6 records no longer time out; huge pages/file links read only as far as needed (steady memory); max-cost limit no longer stops early when sites in progress fail; **parked / for-sale domains are free failed rows** (were billed as sparse profiles); `paymentProviders` also from DNS proof (Stripe verification) + Adyen, Braintree, Affirm, Square, Mollie, Razorpay, Midtrans; alias domains without mail use the real site's DNS; long lists on small memory finish cleanly before the timeout.
- **0.1.0** (2026-09-25) — first release.

# Actor input Schema

## `websites` (type: `array`):

Company domains or URLs, one per line — e.g. stripe.com, https://www.acme.de. Emails (jane@acme.com) are turned into their domain. Duplicates are skipped for free.

## `datasetId` (type: `string`):

Optional. Pick a dataset from another Actor's run (Google Maps, lead lists, our Contact Details Scraper): the website is read from the field below in every item. Use the picker — this Actor runs with limited permissions and can only read datasets you select here.

## `websiteField` (type: `string`):

Field that holds the website/domain in the dataset items (e.g. website, url, domain). Nested fields as dot paths, e.g. company.website.

## `maxPagesPerSite` (type: `integer`):

Homepage + up to this many company-fact pages (contact, imprint/legal, about, team, privacy, careers). More pages never cost more.

## `includeDnsIntel` (type: `boolean`):

Reads public DNS: MX → email provider (Google Workspace, Microsoft 365…), SPF → sending tools (HubSpot, SendGrid…), TXT verifications → SaaS accounts (Atlassian, Stripe, OpenAI…), DMARC policy.

## `includeDomainAge` (type: `boolean`):

Domain registration date, age in years, expiry and registrar via RDAP (not every ccTLD publishes it).

## `maxConcurrency` (type: `integer`):

Websites processed at the same time.

## `requestTimeoutSecs` (type: `integer`):

Per-page timeout.

## `proxyConfiguration` (type: `object`):

Used only as a fallback when a website blocks the direct request (403/429/challenge).

## Actor input object example

```json
{
  "websites": [
    "stripe.com",
    "allbirds.com",
    "basecamp.com"
  ],
  "websiteField": "website",
  "maxPagesPerSite": 6,
  "includeDnsIntel": true,
  "includeDomainAge": true,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 15,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `profiles` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "stripe.com",
        "allbirds.com",
        "basecamp.com"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("tactful_anvil/company-website-enrichment").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "websites": [
        "stripe.com",
        "allbirds.com",
        "basecamp.com",
    ],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("tactful_anvil/company-website-enrichment").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "stripe.com",
    "allbirds.com",
    "basecamp.com"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call tactful_anvil/company-website-enrichment --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tactful_anvil/company-website-enrichment"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3AJ8C5j1bzcjw6P0y/builds/6MR7btj5Wq9ByrQSO/openapi.json
