# Content-Type & X-Content-Type-Options Auditor (`phoenix2810/content-type-auditor`) Actor

Audit one public URL for Content-Type accuracy, X-Content-Type-Options nosniff, and MIME confusion risk via magic-byte sniffing.

- **URL**: https://apify.com/phoenix2810/content-type-auditor.md
- **Developed by:** [Sanskar Jaiswal](https://apify.com/phoenix2810) (community)
- **Categories:** Developer tools, SEO tools, Open source
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Content-Type & X-Content-Type-Options Auditor

Fetches one public URL and inspects its Content-Type header, X-Content-Type-Options header, and magic-byte sniff to detect MIME confusion risk. Returns a readiness score and recommendations. Built for security teams, devops engineers, site migration QA, and frontend platform teams.

### Use cases

- Verify X-Content-Type-Options: nosniff on static assets after a CDN or origin cutover.
- Detect script or stylesheet responses with a Content-Type that disagrees with the actual bytes.
- Catch missing or generic Content-Type headers (e.g. application/octet-stream on a JS file) before production deploy.
- Run scheduled checks on critical resources to catch configuration drift on edge servers.
- Feed structured results into security QA dashboards or CI pipelines.

### Input

| Field | Type | Description |
| --- | --- | --- |
| `startUrl` | string | Public HTTP or HTTPS URL to audit. URLs with credentials and private network targets are rejected. |
| `timeoutSeconds` | integer | Request timeout from 3 to 30 seconds. Defaults to 10. |
| `maxBodyBytes` | integer | Maximum response body bytes to read for magic-byte sniffing. Defaults to 64 KB; capped at 512 KB. |

### Output

The actor pushes one dataset item per run.

| Field | Type | Description |
| --- | --- | --- |
| `inputUrl` | string | Original URL from input. |
| `normalizedInputUrl` | string | Normalized input URL after defaulting the scheme. |
| `finalUrl` | string | Final page URL after redirects. |
| `ok` | boolean | True when the fetch succeeded. |
| `checkedAt` | string | ISO timestamp for the audit. |
| `httpStatus` | integer or null | HTTP status code from the response. |
| `declaredContentType` | string or null | Raw Content-Type header value. |
| `declaredType` | string or null | Parsed media type (lowercased, without parameters). |
| `declaredCharset` | string or null | Charset from the Content-Type header, if present. |
| `xContentTypeOptions` | string or null | Raw X-Content-Type-Options header value. |
| `hasNosniff` | boolean | True when X-Content-Type-Options contains nosniff. |
| `sniffedType` | string or null | Magic-byte sniffed media type, or null if unrecognized. |
| `typeMismatch` | boolean | True when the sniffed type conflicts with the declared Content-Type. |
| `confusionRisk` | string | MIME confusion risk level: low, medium, or critical. |
| `score` | integer | Content-Type readiness score from 0 to 100. |
| `grade` | string | Letter grade from A to F. |
| `issues` | array | Human-readable issues. |
| `recommendations` | array | Suggested fixes. |
| `error` | string or null | Fetch-level error, if the request failed. |

### Example input

```json
{
  "startUrl": "https://example.com/app.js",
  "timeoutSeconds": 10
}
```

### Example output

```json
{
  "inputUrl": "https://example.com/app.js",
  "normalizedInputUrl": "https://example.com/app.js",
  "finalUrl": "https://example.com/app.js",
  "ok": true,
  "checkedAt": "2025-01-01T00:00:00.000Z",
  "httpStatus": 200,
  "declaredContentType": "application/javascript; charset=utf-8",
  "declaredType": "application/javascript",
  "declaredCharset": "utf-8",
  "xContentTypeOptions": "nosniff",
  "hasNosniff": true,
  "sniffedType": null,
  "typeMismatch": false,
  "confusionRisk": "low",
  "score": 100,
  "grade": "A",
  "issues": [],
  "recommendations": [
    "Content-Type declaration, charset, and X-Content-Type-Options look correct for this resource."
  ],
  "error": null
}
```

### Security

- Only public HTTP and HTTPS URLs are fetched.
- URLs with usernames or passwords are rejected.
- Private IPv4, private IPv6, localhost, link-local, and private DNS resolutions are blocked before fetching.
- Redirect destinations are revalidated before they are followed.
- Body reads are capped to limit memory use during magic-byte sniffing.
- The actor does not require logins, cookies, browser sessions, or credentials.

### Pricing

| Event | Suggested price |
| --- | ---: |
| Actor start | `$0.005` |
| URL audited | `$0.01` |

Suggested launch price: about `$0.015` per audited URL. Teams can schedule the actor for recurring checks on important resources after deploys and CDN cutovers.

### FAQ

#### Does this actor crawl multiple URLs or a whole site?

No. It fetches one URL per run. This keeps runs cheap and predictable for CI and scheduled monitoring.

#### How does the magic-byte sniff work?

The actor reads the first bytes of the response body and compares them against known file signatures (image, PDF, gzip, WOFF fonts, etc.). It also checks text prefixes for HTML, JSON, XML, and SVG. If the sniffed type disagrees with the declared Content-Type, the actor flags a mismatch.

#### What makes a risk critical?

A type mismatch on a script or stylesheet resource without X-Content-Type-Options: nosniff is critical, because browsers may execute content with a different type than intended, creating a MIME-sniffing attack vector.

#### Why is X-Content-Type-Options important?

Without nosniff, browsers may sniff the response body and interpret it as a different type than declared in the Content-Type header. For scripts and stylesheets, this can allow cross-site script execution. nosniff instructs the browser to respect the declared type.

#### How does the score work?

The score starts at 100 and is reduced for: missing nosniff (-10, -30 for scripts/stylesheets), missing Content-Type (-25), generic application/octet-stream with detectable bytes (-10), type mismatch (-15, -35 for scripts/stylesheets), and missing charset on HTML (-5). The resulting letter grade reflects overall Content-Type posture.

# Actor input Schema

## `startUrl` (type: `string`):

Public HTTP or HTTPS URL to audit.

## `timeoutSeconds` (type: `integer`):

Request timeout from 3 to 30 seconds.

## `maxBodyBytes` (type: `integer`):

Maximum response body bytes to read for magic-byte sniffing. Defaults to 65 KB and is capped at 512 KB.

## Actor input object example

```json
{
  "startUrl": "https://example.com/",
  "timeoutSeconds": 10,
  "maxBodyBytes": 65536
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrl": "https://example.com/"
};

// Run the Actor and wait for it to finish
const run = await client.actor("phoenix2810/content-type-auditor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrl": "https://example.com/" }

# Run the Actor and wait for it to finish
run = client.actor("phoenix2810/content-type-auditor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrl": "https://example.com/"
}' |
apify call phoenix2810/content-type-auditor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,phoenix2810/content-type-auditor"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/cSqnvIlZq1UMnjDl4/builds/ySE8RbWD7RNMSEUO5/openapi.json
