# Universal Screenshot API (`elab/universal-screenshot-api`) Actor

Batch website screenshot and PDF capture with full-page-by-default rendering, dark mode, ad/tracker/cookie-banner blocking, SSRF-safe URL validation, and an explicit error taxonomy.

- **URL**: https://apify.com/elab/universal-screenshot-api.md
- **Developed by:** [Barak Eliov](https://apify.com/elab) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $6.00 / 1,000 screenshots

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Universal Screenshot API

Batch website screenshot and PDF capture. Full-page-by-default rendering, dark mode, ad/tracker/cookie-banner
blocking, SSRF-safe URL validation, and an explicit error taxonomy with no charge on failure.

This is v1 (P0 feature set only, per the approved spec in `SPEC.md`). See "What's not in v1 yet" below.

### What this Actor does

Give it a list of URLs and it captures a PNG/JPEG/WebP screenshot and/or a PDF for each one, using a real
headless Chromium browser (Crawlee's `PlaywrightCrawler`). Every input URL always produces at least one
dataset row per requested output type - successes and failures alike - so you never get a silently missing
result.

### Input example

```json
{
  "urls": ["https://apify.com", "https://example.com"],
  "outputs": ["screenshot"],
  "fullPage": true,
  "format": "png",
  "colorScheme": "no-preference"
}
```

Full P0 input fields (all have defaults; see `.actor/input_schema.json` for the authoritative list):
`urls`, `outputs`, `format`, `quality`, `fullPage`, `device`, `viewportWidth`, `viewportHeight`,
`deviceScaleFactor`, `waitUntil`, `waitForSelector`, `delayMs`, `timeoutSecs`, `scrollToBottom`,
`colorScheme`, `selectorsToHide`, `blockAds`, `blockTrackers`, `hideCookieBanners`, `stealthMode`,
`proxyConfiguration`. Plus `cookies` (P1, pre-approved - see below).

### Output example

One row per URL per requested output type (real output from a local run against `https://example.com`):

```json
{
  "url": "https://example.com",
  "finalUrl": "https://example.com/",
  "status": "success",
  "outputType": "screenshot",
  "device": "desktop",
  "format": "png",
  "width": 1280,
  "height": 800,
  "pdfPageCount": null,
  "fileSizeBytes": 11423,
  "assetUrl": "https://api.apify.com/v2/key-value-stores/.../records/screenshot-...png",
  "mimeType": "image/png",
  "captureTimeMs": 1083,
  "cacheHit": false,
  "waitForSelectorTimedOut": false,
  "fullPageHeightCapped": false,
  "consoleErrorCount": null,
  "extractedText": null,
  "extractedLinks": null,
  "extractedMetadata": null,
  "errorCategory": null,
  "errorMessage": null,
  "retryCount": 0,
  "scrapedAt": "2026-09-30T17:40:12.858Z"
}
```

A failed URL gets `"status": "error"`, `errorCategory` set to one of `blocked` / `not_found` / `timeout` /
`render_error` / `invalid_url` / `upstream_changed`, and every capture-specific field left `null`. Failed
rows are **never charged**.

### Errors (never silent, never charged)

| Category | When |
|---|---|
| `invalid_url` | Malformed URL, non-http(s) URL, or the URL resolves to a private/link-local/loopback/cloud-metadata IP address (SSRF protection, always on) |
| `not_found` | DNS failure, connection refused, 404/410 |
| `blocked` | 401/403/429 or a bot-challenge page |
| `timeout` | Navigation/render exceeded `timeoutSecs` |
| `render_error` | Browser crash or an unclassified render failure |
| `upstream_changed` | Reserved for graceful-degradation cases (page structure changed) |

A `waitForSelector` that never appears is **not** an error: the page is still captured and
`waitForSelectorTimedOut: true` is reported (and the successful output is charged normally).

### SSRF protection (always on)

Before any browser context is opened for a URL, its hostname is DNS-resolved and every resolved IP is
checked against RFC1918/link-local/loopback/CGNAT/cloud-metadata ranges (IPv4 and IPv6). A hit is rejected
as `invalid_url` with zero navigation attempted. This is mandatory and not a togglable input.

**Known limitation:** this is a pre-navigation check, not a pinned connection. A target that changes DNS
or redirects to a private address *after* the check passed would not be caught by this v1 (a real gap
shared with most screenshot tools; closing it fully needs IP-pinned egress, out of scope for v1). An
earlier design also re-validated the connected IP after navigation as extra defense-in-depth, but this was
removed after local testing showed it produces false positives on any network path that goes through a
local/corporate HTTP proxy or VPN (common setups) - it broke normal captures on non-proxied networks too
aggressively to justify keeping.

### Ad/tracker blocking and cookie banners

`blockAds`/`blockTrackers` intercept requests to a small, curated list of well-known ad and tracker domains
(see `src/blocklist.ts`). `hideCookieBanners` hides a small list of well-known cookie-consent banner
selectors via CSS. Both lists are intentionally short for v1 and easy to extend - they are not a full
filter-list replacement (e.g. EasyList).

### Cookies (capture your own authenticated pages)

The `cookies` input (P1, pre-approved) lets you supply session cookies to capture pages behind your own
login, e.g.:

```json
"cookies": [{ "name": "session", "value": "abc123", "domain": "example.com", "path": "/" }]
```

Supply only credentials you own or are authorized to use. This does not scrape third-party logged-in
areas; it lets you authenticate your own session before capture.

### Devices

`device` selects a realistic viewport + pixel-density + user-agent preset: `desktop`, `desktop_hd`,
`laptop`, `tablet`, `mobile`, `iphone_15`, `pixel_8`, `ipad`, or `custom` (use with `viewportWidth`/
`viewportHeight`). `deviceScaleFactor` left at its default (`1`) uses the preset's own pixel density;
any other value overrides it.

### Limits

- Max 100 URLs per run (internal safety cap; there is no `maxItems` input field in v1).
- Max 120s `timeoutSecs`, max 60s `delayMs` per URL. `timeoutSecs` bounds the *entire* per-URL pipeline
  (navigation, `waitForSelector`, auto-scroll, and the final screenshot/PDF capture) - exceeding it at
  any stage reports `errorCategory: "timeout"` for that URL and is not charged.
- Full-page height is internally capped at 20,000px, a cost/reliability safety limit (not a Chromium
  limit - Chromium itself can render far taller pages). Pages taller than the cap are captured up to
  20,000px and `fullPageHeightCapped: true` is set in the output row; shorter pages are unaffected and
  report `fullPageHeightCapped: false`.
- WebP output is produced by capturing PNG and converting with `sharp` (Chromium's native screenshot API
  only supports png/jpeg).
- `stealthMode` masks common automation fingerprints; it does not defeat CAPTCHAs or paywalls.
- No login/CAPTCHA solving beyond the `cookies` feature above (your own sessions only).

### Use with AI agents / MCP

Call the Actor with `urls` as an array (never a comma-separated string). Check `status` per row; only
`status: "success"` rows have a populated `assetUrl`. `outputType` tells you whether a row is a
`screenshot` or a `pdf`. Batch multiple URLs in one call instead of one run per URL.

### FAQ

**Am I charged for failed URLs or a `waitForSelector` timeout?** No charge for failed rows. A
`waitForSelectorTimedOut: true` row that still successfully produced an image/PDF **is** charged - a
screenshot was in fact delivered.

**Does this bypass CAPTCHAs or paywalls?** No. `stealthMode` only reduces the chance of being blocked by
basic bot detection.

**Can I capture a page that needs a login?** Only your own pages, via the `cookies` input. This Actor does
not scrape third-party logged-in areas.

### What's not in v1 yet (P1/P2, explicitly deferred)

`devices` multi-device fan-out, `selectorsToBlur`, `elementSelector` (single-element capture), `customCss`,
`omitBackground`, `acceptCookieConsent`, `blockImages`, `userAgent`/`locale`/`timezone`/`geolocation`
overrides, `extractData` (text/links/metadata/console extraction - the output fields exist and are always
`null` for forward compatibility), `pdfFormat`/`pdfLandscape`/`pdfPrintBackground` (PDF is currently always
A4/portrait/backgrounds-on), and a user-configurable `maxConcurrency` (internally fixed at 5). `cookies` is
the one P1 field implemented in this build (explicitly pre-approved - see `SPEC.md`).

### Local development

```bash
npm install
npx playwright install chromium   # one-time, downloads a browser to the shared Playwright cache
npm run typecheck
npm test
npm run build
apify run --input-file smoke-input.json --purge
```

# Actor input Schema

## `urls` (type: `array`):

Page URLs to capture. Each URL produces one dataset row per requested output type (screenshot and/or pdf). Max 100 URLs per run.

## `outputs` (type: `array`):

What to produce for each URL. You are charged only for outputs that succeed.

## `format` (type: `string`):

Image format when 'screenshot' is requested. PNG and JPEG are captured natively; WebP is produced by converting the PNG capture. Ignored for PDF output.

## `quality` (type: `integer`):

Compression quality 1-100 for jpeg/webp. Ignored for png.

## `fullPage` (type: `boolean`):

Capture the full scrollable page, not just the viewport. On by default (most competing tools default this off). Capture height is internally capped at 20,000px for cost/reliability safety (not a Chromium limit) - pages taller than that are captured up to the cap and the dataset row's 'fullPageHeightCapped' field is set to true.

## `device` (type: `string`):

Realistic viewport + pixel-density + user-agent preset. Use 'custom' with Viewport width/height below to override.

## `viewportWidth` (type: `integer`):

Used only when Device preset is 'custom'.

## `viewportHeight` (type: `integer`):

Used only when Device preset is 'custom' and Full page is false.

## `deviceScaleFactor` (type: `integer`):

Pixel density override: 2 or 3 for retina/high-DPI captures. Leave at the default (1) to use the selected device preset's own value.

## `waitUntil` (type: `string`):

Navigation readiness before capture. 'networkidle' is slower but far more reliable for JS-heavy pages than 'load'.

## `waitForSelector` (type: `string`):

Optional CSS selector to wait for before capturing. If it never appears within Timeout, the page is still captured and 'waitForSelectorTimedOut' is set true in the output rather than failing the run.

## `delayMs` (type: `integer`):

Extra wait (ms) after the page is ready, for animations or late content. Capped at 60000 ms.

## `timeoutSecs` (type: `integer`):

Max seconds to wait for navigation and rendering before this URL is reported as a timeout error. Capped at 120s.

## `scrollToBottom` (type: `boolean`):

Auto-scroll through the page before capturing to trigger lazy-loaded images/sections, then scroll back to top.

## `colorScheme` (type: `string`):

Emulate prefers-color-scheme. Set to 'dark' to force dark-mode rendering on sites that support it.

## `selectorsToHide` (type: `array`):

CSS selectors to hide (visibility:hidden) before capturing, e.g. for chat widgets or popups.

## `blockAds` (type: `boolean`):

Block requests to a small, curated list of known ad-serving domains before capture.

## `blockTrackers` (type: `boolean`):

Block requests to a small, curated list of known analytics/tracker domains before capture.

## `hideCookieBanners` (type: `boolean`):

Hide known cookie-consent banner overlays via CSS before capturing (visual cleanup only; does not grant consent).

## `stealthMode` (type: `boolean`):

Mask common automation fingerprints so target sites are less likely to block the request. Does not bypass CAPTCHAs or paywalls.

## `proxyConfiguration` (type: `object`):

Apify Proxy configuration for the capture request.

## `cookies` (type: `array`):

Session cookies for capturing your OWN authenticated/paywalled pages, e.g. \[{ "name": "session", "value": "...", "domain": "example.com", "path": "/" }]. Supply only credentials you own or are authorized to use.

## Actor input object example

```json
{
  "urls": [
    "https://apify.com"
  ],
  "outputs": [
    "screenshot"
  ],
  "format": "png",
  "quality": 80,
  "fullPage": true,
  "device": "desktop",
  "viewportWidth": 1280,
  "viewportHeight": 800,
  "deviceScaleFactor": 1,
  "waitUntil": "networkidle",
  "waitForSelector": "",
  "delayMs": 0,
  "timeoutSecs": 30,
  "scrollToBottom": true,
  "colorScheme": "no-preference",
  "selectorsToHide": [],
  "blockAds": true,
  "blockTrackers": true,
  "hideCookieBanners": true,
  "stealthMode": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "cookies": []
}
```

# Actor output Schema

## `results` (type: `string`):

One row per URL per requested output type: status (success/error), outputType (screenshot/pdf), dimensions, errorCategory, and assetUrl pointing to the captured file.

## `files` (type: `string`):

The actual screenshot (PNG/JPEG/WebP) and PDF files captured for each URL, referenced by assetUrl in each dataset row.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("elab/universal-screenshot-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://apify.com"],
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("elab/universal-screenshot-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call elab/universal-screenshot-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,elab/universal-screenshot-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pVHVrq1TFo0bA9Mfo/builds/12TTELoTz2Evzzfev/openapi.json
