# Site Health Audit: Website Security, SEO & GDPR (`fewbricks/site-health-audit`) Actor

Audit any website from its URL: speed, security headers, SSL, email deliverability (SPF, DKIM, DMARC), WordPress exposure, cookies and legal pages, on-page SEO and AI visibility. Get scored findings with a plain-language fix, plus a client-ready PDF report in 5 languages.

- **URL**: https://apify.com/fewbricks/site-health-audit.md
- **Developed by:** [Frédéric Oudry](https://apify.com/fewbricks) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $30.00 / 1,000 website auditeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Site Health Audit do?

Site Health Audit checks any website **from its URL alone** and returns a scored diagnosis: **security, email deliverability, GDPR and legal pages, SEO, speed and AI visibility**. No access to the site, no plugin, nothing to install on the owner's side.

It is built for **web agencies, freelancers and AI agents** that need a reliable website audit before quoting or starting work, and a report they can hand to the site owner:

- one **score out of 100 and a grade from A to E** per website, with a score per area;
- every finding with its **impact in plain language** and the **fix**;
- a **client-ready PDF and HTML report** under your own name;
- findings and report in **English, French, German, Spanish or Italian**;
- legal checks that follow the **country of the website** (European Economic Area, United Kingdom, Switzerland).

Because it runs on Apify, you can also call it through the **API**, **schedule** it to re-audit a list of sites every week, send results to other tools through **integrations**, and use it from an AI assistant through **MCP**.

### What does the website audit check?

| Area | Checks |
|---|---|
| **Speed** | response time, HTML weight, compression, HTTP/2, redirect chains, http → https |
| **Security** | SSL certificate (validity, expiry, protocol), HSTS, CSP, clickjacking protection, X-Content-Type-Options, Referrer-Policy, Permissions-Policy, exposed server versions, CAA record |
| **Email** | MX, **SPF**, **DKIM** (25 common selectors, with a guard against DNS wildcards), **DMARC** and its policy |
| **CMS** | WordPress, PrestaShop, Shopify, Wix, Webflow, Squarespace, Joomla, Drupal, HubSpot. For WordPress: version, visible plugins and theme, `xmlrpc.php`, `readme.html`, user enumeration through the REST API (checked on the response content) |
| **Privacy and legal** | cookies set on first load, third-party trackers (Google Analytics / Tag Manager, Meta Pixel, Hotjar, LinkedIn, TikTok, Clarity), consent manager detection (about 30 vendors), links to the legal notice, privacy policy and terms, in the languages of the covered countries |
| **SEO** | title, meta description, `<h1>`, canonical, `lang`, viewport, noindex, alt attributes, Open Graph, hreflang, `robots.txt`, sitemap |
| **AI visibility** | access of AI crawlers in `robots.txt` (OpenAI, Anthropic, Perplexity, Google, Mistral AI, Apple, Meta, Amazon, Common Crawl), answer bots told apart from training bots; `/llms.txt`; schema.org structured data (valid JSON-LD, Organization or LocalBusiness entity, `sameAs`); `nosnippet` / `noai` directives; text readable **without JavaScript** |

### What do you get?

One dataset item per address:

- `reachable`: whether the website could be read. Only reached websites are graded, reported and charged;
- `score`: the **score out of 100**, the **grade from A to E** and the score per area;
- `all_findings`, sorted by severity. Each finding has a stable `code`, a `severity` (`critical`, `major`, `minor`), the technical `message`, the `impact` in plain language for the site owner, and the `fix`;
- a one-sentence `summary`;
- the `jurisdiction` applied (country, rules, and how the country was determined);
- `sections`: the raw data of every area;
- `report_pdf_url` and `report_html_url`: signed links to the **client-ready report**.

In the run's **Output** tab, the *Overview* view shows one row per website and the *Findings* view one row per finding. The reports are listed under *Client-ready reports*.

#### The client-ready report

The report is written for a business owner who is not technical. Page one gives the grade and the priorities. Each area follows, with what needs fixing and what is already in order. A technical appendix closes the report for whoever does the work. Your name, contact details and accent color appear on it when you provide them. The HTML version is a single self-contained file that reads well on a phone; the PDF is A4.

### How to audit a website

1. Open the **Input** tab and enter one or more addresses in **Websites to audit** (full URL or bare domain).
2. Choose the **language** of the findings and report, and set the **country** of the website if you know it.
3. Optional: add your name, contact details and accent color to brand the report, or set **Client-ready report** to *None* to get the JSON data only.
4. Click **Start**. Websites are audited in parallel; one website never takes more than two minutes.
5. Read the results in the **Output** tab, download the reports, or export the dataset as JSON, CSV or Excel.

### Input

```json
{
  "urls": ["https://example.de", "another-site.com"],
  "language": "en",
  "country": "auto",
  "report": "pdf+html",
  "brandName": "Example Studio",
  "brandContact": "hello@studio.example",
  "brandAccent": "#1f3a8a",
  "concurrency": 3
}
```

Only `urls` is required. Bare domains are accepted. Duplicates are audited once. Up to 500 addresses per run.

- `language`: `en` (default), `fr`, `de`, `es`, `it`.
- `country`: `auto` (default) or an ISO code such as `FR`, `DE`, `GB`.
- `report`: `pdf+html` (default), `html`, `pdf` or `none`.
- `concurrency`: websites audited at the same time, 1 to 10 (default 3).

Only public websites on the standard ports (80 and 443) are accepted. An address that is not a public website (`localhost`, a private IP address, another scheme than http or https) gives an item with `"error": "invalid_url"` and is not charged.

### Output

```json
{
  "url": "https://example.de",
  "host": "example.de",
  "audited_at": "2026-10-03T06:59:05+00:00",
  "language": "en",
  "reachable": true,
  "score": {
    "total": 72,
    "grade": "C",
    "by_section": { "performance": 95, "security": 60, "email": 70, "privacy": 70, "seo": 80, "llm": 55 },
    "cap": null
  },
  "jurisdiction": { "country": "DE", "gdpr": true, "consent_required": true, "legal_notice_required": true, "legal_notice_name": "Impressum", "law_ref": "§ 5 DDG", "authority": null, "basis": "tld" },
  "all_findings": [
    {
      "section": "email",
      "code": "dmarc_missing",
      "severity": "major",
      "message": "No DMARC record on _dmarc.example.de",
      "impact": "Without DMARC, Gmail, Yahoo and Outlook treat your email with suspicion, and you cannot tell who is spoofing your address.",
      "fix": "Publish a DMARC record, starting in monitoring mode (p=none) with a reporting address."
    }
  ],
  "summary": "example.de: grade C (72/100). 12 issues found: 4 important, 8 to improve. Start with: …",
  "report_html_url": "https://api.apify.com/v2/key-value-stores/<store-id>/records/report-example.de.html?signature=<signature>",
  "report_pdf_url": "https://api.apify.com/v2/key-value-stores/<store-id>/records/report-example.de.pdf?signature=<signature>"
}
```

The excerpt leaves out `cms` and `sections`. `score.cap` is `null`, or `{ "limit": 59, "code": "<finding code>" }` when a ceiling lowered the score (see the grade rules below).

The report links are signed: anyone who has the link can open the report, without an Apify account. They work as long as the run's storage is kept.

A website that could not be reached still gives an item, with `"reachable": false`, the reason in `summary`, and the DNS, email and certificate findings that could be established. It has no report links. An address that could not be audited at all gives `"reachable": false` with an `error` code (`invalid_url`, `audit_timeout` or `audit_failed`) and an `error_message` in English.

### How much does a website audit cost?

Site Health Audit uses pay-per-event pricing. Platform usage is included: you pay only for these events.

| Event | Price | Charged |
|---|---|---|
| Website audited (`site_audited`) | $0.03 | once per website reached, when its result is in the dataset |
| Report generated (`report_generated`) | $0.15 | once per website reached, when its report is stored. Not charged when the report is set to *None* |
| Actor start (`apify-actor-start`) | $0.00005 | once per run, by the platform (once per GB of memory above 1 GB) |

One website costs **$0.03 for the data alone and $0.18 with the client-ready report**. Auditing 100 websites with reports costs $18; the data alone costs $3.

You are only charged for websites the Actor could reach: an unreachable site, a site that blocks the check, or an invalid address costs nothing and produces no report. If a report cannot be generated, the audit is delivered and only the $0.03 is charged.

#### Can I cap the cost of a run?

Yes. Set a maximum cost for the run in Apify Console or through the API. Before starting a website, the Actor checks that the remaining budget covers its full price ($0.03, or $0.18 with a report). Websites that do not fit are skipped without any request, are not charged, and are listed in the run's final status message and in the log. A limit of $0.20 covers one website with its report.

### Countries and legal checks

Three findings depend on the country: a missing legal notice, trackers loaded without a recognized consent manager, and a missing privacy policy.

- **Country detection**: your `country` input first, then a national domain (`.fr`, `.de`, `.co.uk`), then the region of the page language (`de-AT`). A language alone (`lang="fr"` on a `.com`) never triggers a national obligation, because French is also spoken in Belgium, Switzerland and Canada. **Set `country` when you know it.**
- **Legal notice**: reported as mandatory only where a national text and an established named page were sourced: France (mentions légales), Germany, Austria and Liechtenstein (Impressum), Spain (aviso legal). Elsewhere the audit stays silent.
- **Tracker consent**: reported across the EEA and the UK, except where the national rule could not be confirmed (Bulgaria, Estonia, Iceland) and not in Switzerland. In the UK, analytics-only sites are not flagged.
- **Outside Europe**: no legal claim is made; a missing privacy policy is reported as a good-practice point.

#### Where do the legal rules come from?

A legal obligation is stated for a country only when it was traced to a source: in most cases the national law on its official publication site or the country's data protection authority, alongside the European texts (GDPR, ePrivacy Directive, e-Commerce Directive); a few rules rest on secondary sources. When a rule could not be confirmed, the audit says nothing rather than guess. This is not legal advice, and the rules have not been reviewed by a lawyer.

### How is the grade computed?

- Each area starts at 100 and loses **40 points per critical finding, 15 per major one, 5 per minor one**.
- Missing security headers cost **at most 15 points together**, however many are missing.
- The overall score weights the six areas: security 25%, speed 20%, email 15%, privacy 15%, AI visibility 13%, SEO 12%.
- Grades: **A** from 90, **B** from 75, **C** from 60, **D** from 40, **E** below.
- A critical finding caps the overall score at 59 (grade D at best). An invalid SSL certificate caps it at 39 (grade E).
- A website that could not be reached gets **no grade**: the summary and the item say why (`reachable` is `false`).

### FAQ

#### Is it legal to audit a website I do not own?

The audit only reads what a website publishes to every visitor and to public DNS: it makes a small number of ordinary, unauthenticated requests, identifies itself in its User-Agent, never tries to log in and never exploits anything. Even so, **audit only websites you own or are authorised to assess**, and follow the laws that apply to you. You are responsible for how you use the results.

#### Why does a website show as unreachable or blocked?

Some websites refuse automated visits (HTTP 401, 403 or 429), sit behind a bot-protection service, or answer too slowly. The item then has `"reachable": false` and its `summary` gives the reason. You are not charged. If the site is yours, allow the visit in your firewall or bot protection, or try again later; a temporary refusal often clears on a second run.

#### Why is a finding marked as a point to confirm?

Trackers and consent banners are read from the initial HTML, without running scripts. A banner loaded by another script can be missed, so the finding asks for a check in a browser before you act on it.

#### Can I audit a specific page instead of the home page?

Yes. The audit reads the address you give. Files such as `robots.txt` are always read at the root of the site.

#### What personal data is stored?

The results describe the website, not people. WordPress usernames are never stored, only their number. Cookies are stored by name, without their value. No email address is extracted from pages, and the report addresses of DMARC records are masked. From structured data, only the public entity the site declares about itself is kept (name, address of its page, public profiles), without its phone number. Page titles and descriptions are stored as published.

#### Where do I get help or report a problem?

Open an issue in the **Issues** tab of this Actor. Give the run ID and the address concerned.

### Limits

- **Read-only**: unauthenticated GET requests only, no login attempt, no form submission, no write. Per website, the audit reads the address you give, checks whether `http://` redirects to `https://`, and reads `/robots.txt`, `/sitemap.xml`, `/llms.txt` and `/llms-full.txt`, the public DNS records of the domain and the SSL certificate. On WordPress sites it also requests `/wp-json/`, `/wp-json/wp/v2/users`, `/xmlrpc.php`, `/readme.html` and `/wp-login.php`, once each, to check whether they are exposed. Nothing else is requested.
- One page per website: this is not a crawler, and internal pages are not visited.
- Trackers are read from the initial HTML, without running scripts: a consent banner loaded by another script can be missed.
- WordPress detection relies on public traces; a hardened site can hide its plugins and version.
- DKIM is looked up on 25 common selectors; a custom selector goes unnoticed, and the finding says so.
- AI visibility measures what can be verified from the site (access, structure, readability). It does not tell whether an assistant actually cites the site. `llms.txt` is a proposed standard with uneven adoption; its absence is reported as a minor point.
- Speed checks measure the HTML response, not a full browser rendering (no Core Web Vitals).
- An audit that has not finished after 120 seconds is stopped and not charged.
- Asian markets and their platforms are out of scope.

### Use it from the API, schedules, integrations and AI agents

The Actor ID is `fewbricks/site-health-audit`.

#### API

Run the audit and get the results in one call:

```bash
curl -X POST "https://api.apify.com/v2/acts/fewbricks~site-health-audit/run-sync-get-dataset-items?token=<YOUR_API_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://example.com"], "language": "en", "report": "pdf"}'
```

#### Schedules and integrations

Save your list of websites as a task and add a schedule to audit it every week or every month. Use Apify integrations and webhooks to send each result to your own tools when a run finishes.

#### AI agents (MCP)

```json
{ "mcpServers": { "apify": { "url": "https://mcp.apify.com?tools=fewbricks/site-health-audit" } } }
```

The dataset fields are described in the Actor's schemas, so an agent knows what each field means before it runs the audit.

# Actor input Schema

## `urls` (type: `array`):

One or more addresses (full URL or bare domain, for example https://example.com or example.com). One result per address; duplicates are audited once. Up to 500 per run.

## `language` (type: `string`):

Language used for every text in the output and in the client-ready report.

## `country` (type: `string`):

Decides which legal requirements are checked (legal notice, cookie consent, privacy policy). Leave on automatic to infer it from the domain and the page language.

## `report` (type: `string`):

A report written for the site owner, in the chosen language. Saved to the run's key-value store; its signed link is added to each result. Charged per website reached; choose None for the JSON data only.

## `brandName` (type: `string`):

Shown in the report header. Leave empty for an unbranded report.

## `brandContact` (type: `string`):

Email or phone number shown at the end of the report.

## `brandAccent` (type: `string`):

Hex color such as #1f3a8a. A color too light to read is darkened automatically.

## `concurrency` (type: `integer`):

Number of sites audited at the same time (1 to 10).

## Actor input object example

```json
{
  "urls": [
    "https://example.com"
  ],
  "language": "en",
  "country": "auto",
  "report": "pdf+html",
  "concurrency": 3
}
```

# Actor output Schema

## `audits` (type: `string`):

One item per website: score, grade, findings with impact and fix, and signed links to the reports.

## `reports` (type: `string`):

PDF and HTML reports, one per website, stored under keys starting with report-.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("fewbricks/site-health-audit").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://example.com"] }

# Run the Actor and wait for it to finish
run = client.actor("fewbricks/site-health-audit").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com"
  ]
}' |
apify call fewbricks/site-health-audit --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,fewbricks/site-health-audit"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/kh7soHPNSD1Hm6zyp/builds/TN2C5rDTrH4MLBhOj/openapi.json
