# Domain Tech Stack, Contacts & SEO Profiler (`tinlark/domain-intelligence-profiler`) Actor

One profile per company domain: technology stack, published business contacts, social links, SEO meta, sitemap size, security headers, DNS, email authentication and domain age. Reads public pages only where robots.txt allows.

- **URL**: https://apify.com/tinlark/domain-intelligence-profiler.md
- **Developed by:** [Tinlark](https://apify.com/tinlark) (community)
- **Categories:** Lead generation, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Domain Tech Stack, Contacts & SEO Profiler

Paste a list of company domains and get one structured profile per domain: the **technology stack**, **published business contact details**, **social profile links**, **SEO meta data**, sitemap size, HTTP **security header grade**, and optionally **DNS, email-authentication (SPF, DMARC) and domain registration (RDAP)** data.

The Actor reads each site's home page and a few standard pages (contact, imprint, about), checks the site's robots.txt first, identifies itself with a `TinlarkBot` User-Agent, and never logs in, solves captchas or tries to get around blocks.

### Use cases

- **Bulk tech stack lookup.** Which CMS, shop platform, framework, analytics, tag manager, CRM, chat widget, payment provider, CDN or server does each company use? About 160 technologies are recognised, each with a category, a version when the site shows one, a confidence level and the evidence found.
- **Company contact info from the company's own pages.** Role addresses such as `info@`, `sales@` or `support@` on the company's own domain, and phone numbers the site publishes. The default policy is business-only.
- **Domain social links finder.** LinkedIn company page, X, Facebook, Instagram, YouTube, GitHub and TikTok links that the site itself links to.
- **Website SEO meta checker.** Title, description, canonical URL, Open Graph tags, H1 count, language, schema.org types, hreflang count, `noindex` flag, generator.
- **Sitemap size.** Number of URLs in the XML sitemap (index files are followed for up to five child sitemaps).
- **Security headers.** HSTS, CSP, framing, content-type, referrer and permissions policy, with a simple A to F grade.
- **DNS and domain age (optional).** Nameservers, mail provider, SPF and DMARC policy, registrar, registration and expiry dates, domain age in days.

### Who it is for

Sales and RevOps teams enriching account lists, agencies and SEO freelancers auditing prospects, and analysts who want a technology footprint for many companies at once. One run replaces several single-purpose tools and gives one schema.

### Input

| Field | What it does |
|---|---|
| `domains` (required) | One domain per entry, such as `notion.so`. URLs are accepted; only the host name is used. |
| `modules` | `tech`, `contacts`, `socials`, `meta`, `sitemap`, `securityHeaders`, `dnsRdap`. Default: `tech`, `contacts`, `socials`, `meta`. |
| `maxPagesPerDomain` | Home page plus contact, imprint and about pages linked from it, up to this total. Default 4. Used for contacts and socials. |
| `contactPolicy` | `business-only` (default) keeps role addresses on the company's own domain. `all-published` returns every address the pages publish. |
| `concurrency` | Domains processed in parallel (default 15). Requests to one site are always at least one second apart. |
| `timeoutSecs` | Per-request timeout (default 20). |

#### Example input

```json
{
  "domains": ["notion.so", "linear.app", "figma.com", "vercel.com", "cohere.com"],
  "modules": ["tech", "contacts", "socials", "meta"]
}
```

### Output

One row per domain in the default dataset (JSON, CSV, Excel, XML or API). The *Overview*, *SEO meta*, *DNS and registration* and *Technologies* views show the main fields. Fields of modules you did not switch on are left out.

#### Example row (real output, trimmed)

```json
{
  "recordType": "profile",
  "domain": "hubspot.com",
  "finalUrl": "https://www.hubspot.com/",
  "status": "ok",
  "httpStatus": 200,
  "robotsAllowed": true,
  "pagesFetched": 4,
  "title": "HubSpot | Software & Tools for your Business - Homepage",
  "metaDescription": "HubSpot's customer platform includes all the marketing, sales, customer service, and CRM software you need to grow your business.",
  "h1Count": 1,
  "lang": "en",
  "schemaTypes": ["AggregateRating", "Brand", "Organization", "Product", "WebSite"],
  "technologies": [
    { "name": "Cloudflare", "category": "CDN", "version": null, "confidence": "high", "evidence": ["header:server", "header:cf-ray", "cookie"] },
    { "name": "HubSpot CMS", "category": "CMS", "version": null, "confidence": "high", "evidence": ["html:/hubfs/", "meta:generator"] }
  ],
  "contacts": { "emails": [], "phones": ["18884827768"], "contactPageUrl": "https://offers.hubspot.com/contact-sales", "policy": "business-only" },
  "socials": { "linkedinCompany": "https://www.linkedin.com/company/hubspot", "x": "https://x.com/HubSpot", "github": null },
  "sitemap": { "url": "https://www.hubspot.com/sitemap.xml", "urlCount": 3066, "lastmod": "2026-10-02", "isIndex": false },
  "securityHeaders": { "hsts": true, "csp": true, "xFrameOptions": "DENY", "score": 7, "grade": "A" },
  "dns": { "nameservers": ["jerry.ns.cloudflare.com", "yolanda.ns.cloudflare.com"], "mxProvider": "Google Workspace", "dmarcPolicy": "reject" },
  "rdap": { "registrar": "MarkMonitor Inc.", "createdAt": "2005-02-06T20:02:28Z", "expiresAt": "2027-02-06T20:02:28Z", "domainAgeDays": 7907 }
}
```

#### Field notes

- `status` is `ok`, `robots-disallowed` (the site's robots.txt does not allow TinlarkBot), `http-error` (the home page answered 4xx or 5xx, for example 403 to automated clients), `rate-limited` (HTTP 429), `unreachable`, `not-html`, or `invalid-domain`. The `error` field says why.
- `technologies[].confidence`: `high` when a response header, cookie or meta generator tag shows it or at least two signals agree, `medium` for a single page-source signal, `low` for technologies implied by another one (for example React when Next.js is found).
- Detection reads the HTML the server sends. It does not run JavaScript, so tools that only appear after scripts execute, or are loaded by a tag manager, can be missed.
- `contacts` emails come from `mailto:` links, visible text and Cloudflare-obfuscated addresses on the pages read. Phone numbers come from `tel:` links on any page and from numbers labelled phone, tel or call on contact and imprint pages, written as the site writes them. Fax numbers are skipped.
- `socials` shows link presence only. The Actor never opens the social networks. Personal LinkedIn profiles (`/in/`) and share links are ignored.
- `dns.spfPolicy` is the `all` qualifier of the SPF record: `fail`, `softfail`, `neutral` or `pass`. `securityHeaders.grade` is a simple Tinlark scale (HSTS 2 points, CSP 2, framing protection 1, nosniff 1, referrer policy 1, permissions policy 1; A is 7 or more).
- The `SUMMARY` record of the run's key-value store has counts per status, requests and run time.

### Pricing

**Free during launch (until 31 October 2026).** You pay only Apify's own platform usage for your runs.

Measured platform usage: about $0.13 per 1,000 domains with the default modules and $0.14 with all modules on our measurements (a run of 50 domains cost $0.006 in platform usage and took about 3.5 minutes). On the Apify free plan that usage is covered by your monthly credit.

**From 1 November 2026: pay per event.** Planned prices, shown here so you can budget:

| Event | Price | When it is charged |
|---|---|---|
| Domain profile | $8.00 per 1,000 ($0.008 each) | Each domain whose home page was read. Includes tech, contacts, socials, meta, sitemap and securityHeaders. |
| DNS and registration lookup | $2.00 per 1,000 ($0.002 each) | Each domain with the `dnsRdap` module on, when at least one of the two lookups returned data. |

Not charged: domains that are unreachable, blocked by robots.txt, answer an HTTP error, serve no HTML, or are invalid input. Volume discounts apply on higher Apify plans (see the event prices on this page once they are active).

Examples at the planned prices: 100 domains with the default modules cost $0.80. The same 100 domains with `dnsRdap` added cost $0.20 more. Set *Maximum cost per run* in the run options to cap spending once pricing is active; the Actor stops cleanly at the cap.

### Limits and honest notes

- **Some sites will not answer.** In a test of 98 well-known company domains, 82 returned a profile. Seven answered 401 or 403 to automated clients, six disallow crawlers in robots.txt, one answered HTTP 429, one did not respond and one served markdown instead of HTML. The Actor reports these with a status and does not charge for them. It does not use proxies or other means to get past blocks.
- **Contacts are a best effort.** In a test of 50 well-known software company domains, a business email was found for 18, a phone number for 11 and a LinkedIn company link for 33. Many sites publish contacts only in forms, scripts or images.
- **Only the pages named above are read.** The Actor reads the home page, up to `maxPagesPerDomain - 1` contact, imprint or about pages linked from it, `robots.txt`, and the sitemap. It does not crawl a site.
- **The technology list is Tinlark's own** (about 160 technologies). It is wide, not exhaustive, and sees only what the HTML, headers and cookies reveal.
- **RDAP coverage depends on the registry.** Some country domains (for example `.de`) publish no RDAP service, so `rdap` is empty and `rdapNote` says why. Registrant details are never read. Registry servers are queried at most once per second, so a list of `.com` domains with `dnsRdap` runs at about one domain per second.
- **Use the registrable domain** (`example.com`) for DNS and RDAP. For sub-domains, the Actor looks up the parent domain for RDAP.
- **Pace.** Requests to one site are at least one second apart, and many sites answer slowly to automated clients. In cloud runs on 512 MB, 5 domains took about 11 seconds and 50 varied domains took 3.5 to 4 minutes, with the default modules or with all of them. Plan about 70 minutes for 1,000 domains and raise the run *Timeout* above the one-hour default for long lists.

### Data sources, terms and your responsibility

- **The listed sites themselves**: their public pages, headers, robots.txt and XML sitemaps, read only where robots.txt allows.
- **Cloudflare public DNS over HTTPS** (`cloudflare-dns.com`, JSON API) for DNS, MX, SPF and DMARC records when `dnsRdap` is on. The domain names you list are therefore sent to that resolver.
- **IANA RDAP bootstrap and the registries' own RDAP servers** for registration dates, registrar and status.

The Actor looks for contact details a company publishes about itself. The default policy keeps role addresses on the company's own domain and drops named personal addresses. Business contact details can still be personal data in some jurisdictions (a sole trader's imprint, for example). You are responsible for having a lawful basis for any outreach, for honouring opt-outs, and for following the terms of the sites you list. Do not use the output for spam. This Actor is not affiliated with any of the technologies or companies it detects.

### FAQ

**Does it use a proxy or a browser?** No. It makes plain HTTP requests, which keeps runs cheap. Sites that need JavaScript to show their content are read as the server sends them.

**Why is my domain marked `http-error` with HTTP 403?** The site refuses automated clients. The Actor does not try to evade that.

**Why were no emails found?** Many companies publish only a contact form. With `contactPolicy: business-only`, personal-looking addresses and addresses on other domains are also dropped. Use `all-published` to see everything the pages show.

**Can I run it on a schedule or from code?** Yes: Apify Schedules, the API, client libraries or the Apify MCP server. Input and output schemas are defined, so the fields are self-describing.

**Can I choose only DNS and domain age?** Yes: set `modules` to `["dnsRdap"]`. No site is visited in that case, and only the DNS and registration event is charged.

**Disclaimers and legality: is this allowed?** The Actor reads public pages that the site owner serves to any visitor, only where robots.txt allows it, at a low request rate and under a declared `TinlarkBot` User-Agent. It does not log in, solve captchas or get around blocks, and it keeps to business contact details by default. Whether you may use a given output, especially contact details, depends on your jurisdiction and on the terms of the sites you list; you are responsible for a lawful basis, for honouring opt-outs and for not using the data for spam. Tinlark is not affiliated with the sites or technologies it detects, and this is not legal advice.

### Related Tinlark Actors

- [ATS Jobs Scraper](https://apify.com/tinlark/ats-jobs-hiring-signals): give it the same company domains and it finds each company's job board and returns the open jobs (Greenhouse, Lever, Ashby, Workable, Recruitee, Personio, Workday). Use it to add hiring data to a profiled domain list.
- [App Store Ratings & Reviews Monitor](https://apify.com/tinlark/app-store-ratings-reviews-monitor): ratings, versions and reviews of a company's iOS and Mac apps, by country.

### Changelog

- 2026-10-03: added the Related Tinlark Actors section.

### Support

Something wrong or missing? Open an issue on this Actor's Issues tab with your input (the domain) and what you expected. Missing technologies can be added to the fingerprint list.

# Actor input Schema

## `domains` (type: `array`):

One domain per entry, such as notion.so. URLs are accepted; the host name is used. Use the registrable domain (example.com) rather than a sub-domain for DNS and RDAP results.

## `modules` (type: `array`):

tech: technology stack. contacts: published business emails and phone numbers. socials: social profile links. meta: title, description, canonical, Open Graph, headings, schema types. sitemap: URL count. securityHeaders: HTTP security header grade. dnsRdap: DNS, mail authentication and domain registration data (charged separately). Empty means tech, contacts, socials and meta.

## `maxPagesPerDomain` (type: `integer`):

Home page plus up to this many pages in total, chosen from the contact, imprint and about pages the home page links to. Only used for contacts and socials.

## `contactPolicy` (type: `string`):

business-only keeps addresses such as info@, sales@ or support@ on the company's own domain and drops named personal addresses. all-published returns every address the pages publish; you are responsible for how you use them.

## `concurrency` (type: `integer`):

Domains processed at the same time. Requests to one site are always at least one second apart.

## `timeoutSecs` (type: `integer`):

Give up on a single request after this many seconds.

## Actor input object example

```json
{
  "domains": [
    "notion.so",
    "https://www.figma.com/about"
  ],
  "modules": [
    "tech",
    "contacts",
    "socials",
    "meta"
  ],
  "maxPagesPerDomain": 4,
  "contactPolicy": "business-only",
  "concurrency": 15,
  "timeoutSecs": 20
}
```

# Actor output Schema

## `profiles` (type: `string`):

One profile per domain.

## `technologies` (type: `string`):

Detected technologies, one row per technology.

## `summary` (type: `string`):

Counts, statuses and request totals.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "notion.so",
        "linear.app",
        "figma.com",
        "vercel.com",
        "cohere.com"
    ],
    "modules": [
        "tech",
        "contacts",
        "socials",
        "meta"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tinlark/domain-intelligence-profiler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": [
        "notion.so",
        "linear.app",
        "figma.com",
        "vercel.com",
        "cohere.com",
    ],
    "modules": [
        "tech",
        "contacts",
        "socials",
        "meta",
    ],
}

# Run the Actor and wait for it to finish
run = client.actor("tinlark/domain-intelligence-profiler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "notion.so",
    "linear.app",
    "figma.com",
    "vercel.com",
    "cohere.com"
  ],
  "modules": [
    "tech",
    "contacts",
    "socials",
    "meta"
  ]
}' |
apify call tinlark/domain-intelligence-profiler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tinlark/domain-intelligence-profiler"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/jy1OSzyDsP8iDgW2j/builds/SLl4ZDMRO9KDjei5Z/openapi.json
