# Company Data Enrichment API — Domain to Firmographics (`jungle_synthesizer/company-domain-firmographic-enrichment-scraper`) Actor

Turn company domains into firmographic data — legal name, HQ, employee bands, live hiring board, self-serve vs sales-led pricing, and compliance badges. Company data enrichment for any domain, with the ATS hiring signal and pricing model most tools skip, each field traced to its source page.

- **URL**: https://apify.com/jungle\_synthesizer/company-domain-firmographic-enrichment-scraper.md
- **Developed by:** [BowTiedRaccoon](https://apify.com/jungle_synthesizer) (community)
- **Categories:** Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.80 / 1,000 record scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Company Data Enrichment API — Domain to Firmographics Scraper

Turn a list of company domains into firmographic rows read straight from each company's own site. Give it `twilio.com`, `zendesk.com`, or any domain with a corporate web presence, and get back legal name, headquarters, headcount signals, a live hiring board, pricing model, and compliance badges — every field traced to the page it came from.

***

### Company Domain Firmographic Enrichment Scraper Features

- Extracts legal name and entity suffix (Inc, LLC, Ltd, GmbH, SAS...) from JSON-LD or the site's own footer
- Resolves the company's live careers board (Greenhouse, Lever, Ashby, SmartRecruiters) for an open-role count and hiring departments
- Flags self-serve vs sales-led pricing — the single most-used B2B ICP filter, and most enrichment tools don't bother
- Reads the trust/security page for SOC 2, ISO 27001, HIPAA, PCI DSS, GDPR, and FedRAMP badges
- Collects HQ address, every office location on the contact page, phone numbers, and role email addresses
- Pulls social profiles (LinkedIn, Twitter/X, Facebook, Instagram, YouTube, GitHub, Crunchbase) straight off the homepage
- Every populated field ships with an `evidence_urls` entry naming the exact page it came from — nothing unsourced

***

### Who Uses Company Firmographic Data?

- **RevOps and growth teams** — attach firmographic columns to a domain list before it's segmented or routed to reps
- **Sales development** — filter prospects by hiring velocity and department, or by self-serve vs sales-led pricing before outreach
- **GRC and vendor-risk teams** — screen vendors by published compliance badges before a security review even starts
- **Market researchers** — track headcount bands, funding mentions, and office footprint across a competitor set
- **Data teams** — join the output on `linkedin_slug` to enrich an existing company table without re-deriving the join key

***

### How Company Data Enrichment Works

1. Submit a list of company domains — bare hosts or full URLs both work.
2. The actor fetches each company's homepage, follows its own nav and footer links to find the about, contact, careers, pricing, and trust pages, then reads them.
3. Fields are pulled from JSON-LD structured data first, with page text and markup as fallback — nothing is guessed when neither signal is present.
4. You get one row per domain, always, even when a page couldn't be found or a field isn't published. The output tells you what it knows and where it read it.

***

### Input

```json
{
  "domains": ["twilio.com", "zendesk.com", "docusign.com"],
  "resolveHiringSignal": true,
  "maxItems": 10
}
```

| Field                 | Type    | Default | Description |
|-----------------------|---------|---------|---------------------------------------------------------------------------------------------------------------------|
| `domains`             | array   | —       | Required. Company domains to enrich — bare domain or full URL, `https://` and `www.` are normalized away. |
| `resolveHiringSignal` | boolean | `true`  | Resolve the company's own careers board for a live open-role count and hiring departments. Adds a premium charge only when a board is actually found and read. |
| `maxItems`            | integer | `10`    | Maximum number of domains to process in this run. |

***

### Company Domain Firmographic Enrichment Scraper Output Fields

```json
{
  "domain": "figma.com",
  "final_url": "https://www.figma.com/",
  "company_name": "Figma",
  "legal_name": "Figma, Inc.",
  "legal_entity_suffix": "Inc.",
  "industry": "Software / SaaS",
  "founded_year": 2012,
  "hq_address": { "street": "760 Market St", "locality": "San Francisco", "region": "CA", "postcode": "94102", "country": "US" },
  "employee_count_band": null,
  "open_role_count": 163,
  "hiring_departments": ["engineering", "sales", "product", "design"],
  "ats_vendor": "greenhouse",
  "careers_url": "https://www.figma.com/careers/",
  "linkedin_company_url": "https://www.linkedin.com/company/figma",
  "linkedin_slug": "figma",
  "has_pricing_page": true,
  "has_trust_security_page": true,
  "compliance_badges": ["soc2", "gdpr"],
  "evidence_urls": { "company_name": "https://www.figma.com/", "open_role_count": "https://boards-api.greenhouse.io/v1/boards/figma/jobs" },
  "scraped_at": "2026-09-26T14:20:55.296Z"
}
```

| Field                     | Type    | Description                                                                                      |
|---------------------------|---------|--------------------------------------------------------------------------------------------------|
| `domain`                  | string  | The normalized input domain (bare host, no scheme/www).                                          |
| `final_url`               | string  | The URL the homepage actually resolved to, after any redirect.                                   |
| `company_name`            | string  | Brand/company name.                                                                              |
| `legal_name`              | string  | Registered legal name, from JSON-LD or the footer copyright line.                                |
| `legal_entity_suffix`     | string  | The jurisdiction tell parsed off the legal name (Inc, LLC, Ltd, GmbH, SAS, Pty Ltd, KK, ...).    |
| `tagline`                 | string  | Short marketing line, when one is published.                                                     |
| `description`             | string  | Meta description / social preview description.                                                   |
| `long_description`        | string  | About-page lede paragraph text.                                                                  |
| `industry`                | string  | Coarse industry label, classified from the site's own description text.                          |
| `industry_taxonomy`       | object  | Sector-level classification alongside the coarse label.                                          |
| `founded_year`            | integer | Founding year.                                                                                   |
| `hq_address`              | object  | Headquarters address, when published.                                                            |
| `all_office_locations`    | array   | Every address block found on the contact/locations page.                                         |
| `country`                 | string  | HQ country, mirrored from `hq_address.country`.                                                  |
| `phone_numbers`           | array   | Phone numbers published on the contact page.                                                     |
| `contact_emails`          | array   | Role email addresses published on the site.                                                      |
| `employee_count_stated`   | integer | Employee count as directly stated on the site.                                                   |
| `employee_count_band`     | string  | `1-10` through `5000+`, derived only when the site states a number.                              |
| `employee_count_evidence` | string  | Which signal produced the band — `stated` or null. Never a guess without evidence.               |
| `open_role_count`         | integer | Live posting count read from the company's own ATS board.                                        |
| `hiring_departments`      | array   | Distinct hiring functions with at least one open posting.                                        |
| `ats_vendor`              | string  | The careers-board vendor, detected from the careers link.                                        |
| `careers_url`             | string  | The company's own careers page URL.                                                              |
| `linkedin_company_url`    | string  | LinkedIn company page URL.                                                                       |
| `linkedin_slug`           | string  | The LinkedIn company slug — a join key for LinkedIn-sourced datasets.                            |
| `twitter_url`             | string  | X/Twitter profile URL.                                                                           |
| `facebook_url`            | string  | Facebook page URL.                                                                               |
| `instagram_url`           | string  | Instagram profile URL.                                                                           |
| `youtube_url`             | string  | YouTube channel URL.                                                                             |
| `github_org`              | string  | GitHub organization slug.                                                                        |
| `crunchbase_url`          | string  | Crunchbase organization URL.                                                                     |
| `funding_mentions`        | array   | Funding-round phrases found on the about/company page.                                           |
| `is_public_company`       | boolean | True when a stock ticker pattern is found on the site.                                           |
| `ticker`                  | string  | Stock ticker symbol, when `is_public_company` is true.                                           |
| `leadership`              | array   | Name + title pairs read off the leadership/team page.                                            |
| `has_pricing_page`        | boolean | True when a public pricing page was found — self-serve vs sales-led.                             |
| `pricing_tiers`           | array   | Tier name + headline price, where published.                                                     |
| `currencies_accepted`     | array   | Currency codes/symbols found on the pricing page.                                                |
| `languages_offered`       | array   | Locale codes from the site's own locale switcher — the international-footprint signal.           |
| `has_status_page`         | boolean | True when a status page is published.                                                            |
| `has_trust_security_page` | boolean | True when a dedicated trust/security page was found.                                             |
| `compliance_badges`       | array   | `soc2`, `iso27001`, `hipaa`, `pci_dss`, `gdpr`, `fedramp` — found on the homepage or trust page. |
| `privacy_policy_url`      | string  | Privacy policy URL.                                                                              |
| `terms_url`               | string  | Terms of service URL.                                                                            |
| `jsonld_organization`     | object  | The raw schema.org Organization data, passed through when present.                               |
| `evidence_urls`           | object  | Per-field map of which page produced the value.                                                  |
| `scraped_at`              | string  | Emission timestamp.                                                                              |

***

### FAQ

#### How do I enrich a company by domain?

Company Data Enrichment API needs nothing but a domain list. Submit `domains` and a row comes back for each one — legal name, HQ, hiring signal, pricing model, and compliance badges, all read from the company's own site.

#### What data can I get from a company domain?

Company Data Enrichment API returns firmographics most tools skip: a live count of open roles by department from the company's own careers board, whether the site runs self-serve or sales-led pricing, and which compliance badges (SOC 2, ISO 27001, HIPAA, PCI DSS, GDPR, FedRAMP) it publishes — on top of the standard name, industry, HQ, and social profiles.

#### Do I need an account or API key for the target company's site?

No. Company Data Enrichment API reads only what a company already publishes on its own public pages — no login, no API key, no special access.

#### Can I turn off the hiring-signal lookup?

Yes. Set `resolveHiringSignal` to `false` and the actor skips the careers-board lookup entirely, along with the premium charge that comes with a successful one.

#### How much does Company Data Enrichment API cost to run?

Every domain gets a base charge for its saved row. A second, premium charge applies only on domains where `resolveHiringSignal` actually resolved a supported careers board and read its live job list — a domain with no reachable board, or the flag turned off, bills at the base rate only.

***

### Need More Features?

Need additional fields, a different ATS vendor supported, or bulk-scale tuning? [File an issue](https://console.apify.com/actors/issues) or get in touch.

### Why Use Company Data Enrichment API?

- **Reads past the homepage** — legal name, HQ, pricing model, and compliance badges come from the pages a company actually maintains, not a single scraped landing page.
- **The hiring signal most enrichment tools don't have.** A live open-role count and department breakdown from the company's own careers board doubles as a free stand-in for headcount and a timed buying signal.
- **Every field is traceable.** `evidence_urls` names the exact page each value came from, so a value with no source page comes back null instead of guessed.

# Actor input Schema

## `sp_intended_usage` (type: `string`):

What will this data feed? E.g. lead lists, KYB checks, price tracking.

## `sp_improvement_suggestions` (type: `string`):

Provide any feedback or suggestions for improvements.

## `sp_contact` (type: `string`):

We'll personally help with your use case. No spam.

## `domains` (type: `array`):

Company domains to enrich (bare domain or full URL — https:// and www. are normalized away). One row is returned per domain, even when the site could not be reached or a field could not be found.

## `resolveHiringSignal` (type: `boolean`):

Resolve the company's own careers board (Greenhouse, Lever, Ashby or SmartRecruiters) for a live open-role count and hiring departments. Adds 1-3 requests per domain and a premium charge only when a board is actually found and read.

## `maxItems` (type: `integer`):

Maximum number of domains to process in this run.

## Actor input object example

```json
{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "domains": [
    "twilio.com",
    "zendesk.com",
    "docusign.com"
  ],
  "resolveHiringSignal": true,
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "domains": [
        "twilio.com",
        "zendesk.com",
        "docusign.com"
    ],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("jungle_synthesizer/company-domain-firmographic-enrichment-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "domains": [
        "twilio.com",
        "zendesk.com",
        "docusign.com",
    ],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("jungle_synthesizer/company-domain-firmographic-enrichment-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "domains": [
    "twilio.com",
    "zendesk.com",
    "docusign.com"
  ],
  "maxItems": 10
}' |
apify call jungle_synthesizer/company-domain-firmographic-enrichment-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jungle_synthesizer/company-domain-firmographic-enrichment-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aG2wQ9TQx9IEae2uS/builds/N9mFnXJAJASoG3Ivl/openapi.json
