# Full Page Screenshot & URL to PDF · No Cookie Banners (`thequietstack/website-screenshot-pdf`) Actor

Full-page website screenshots and PDFs that really capture the whole page: lazy-loaded images scrolled in, cookie banners hidden (never accepted), desktop/laptop/tablet/mobile. Optional visual diff vs. the last run. Pay per delivered file - blocked, 404 and dead URLs are free.

- **URL**: https://apify.com/thequietstack/website-screenshot-pdf.md
- **Developed by:** [TheQuietStack](https://apify.com/thequietstack) (community)
- **Categories:** Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $20.00 / 1,000 screenshots

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Full Page Screenshot & URL to PDF · No Cookie Banners

**Full page screenshot and URL to PDF:** website capture that scrolls lazy-loaded images in and hides cookie banners, desktop to mobile.

**Website Screenshot & PDF** turns a list of URLs into **full-page screenshots (PNG/JPEG) and PDFs that show the
whole page** - including the images that only load while you scroll, and without the cookie banner covering the
content. Every URL gets a row: a link to the file, or the exact reason it could not be captured.
**You pay only for files that were delivered.** Blocked pages, 404s, dead domains and timeouts are free.

### What you get

- A **full-page screenshot** (PNG or JPEG) and/or a **PDF** per URL and device (desktop, laptop, tablet, mobile).
- One **dataset row per capture**: status, HTTP status, page title, file links, file sizes, real page height,
  whether the image was cut, how many consent elements were hidden, load time.
- Optional **visual diff** against the last run: changed-pixel percentage plus a diff image with the changes in red.
- A `SUMMARY` record with counts per status and what was charged.

### Use cases

- **Archiving and evidence**: keep a dated full-page copy of a landing page, a price page or terms of service.
- **Website change monitoring**: schedule the Actor with `compareWithPrevious` and get `changed: true` when a page moves.
- **QA and visual regression**: capture the same URLs on desktop and mobile after every deploy.
- **SEO and competitor reports**: full-page screenshots of SERP landing pages or competitor homepages for a client deck.
- **URL to PDF**: print web pages to A4/Letter PDFs with backgrounds, in the screen or print stylesheet.

### Why this Actor

| Common problem | What this Actor does |
|---|---|
| "Only screenshots above the fold" | Captures the real document height; scrolls the page step by step first so lazy-loaded images and sections render; removes the scroll lock that consent banners put on the page. The row reports the page height and whether it was cut at your `maxHeightPx`. |
| Cookie / consent banner over the content | Hides the banners of OneTrust, Cookiebot, Usercentrics, Didomi, Quantcast, TrustArc, Sourcepoint, Borlabs, Complianz, CookieYes, Iubenda, Osano, Klaro, consentmanager and common generic patterns. **Hidden, never clicked** - no consent is given on your behalf. Add your own selectors for chat widgets or promo bars. |
| "It didn't work and still used my credits" | Pay per delivered file. Failed, blocked, invalid and duplicate URLs are never charged and are listed with the reason. |
| Unpredictable compute cost | A fixed price per screenshot and per PDF, plus an optional hard cap `maxCostUsd`. |
| Did the page change? | Optional pixel diff against the previous run of the same URL and device: `diffPercent`, `changed`, and a highlighted diff image. Included in the screenshot price. |

### Input

```json
{
  "urls": ["https://example.com", "https://en.wikipedia.org/wiki/Screenshot"],
  "outputs": ["screenshot", "pdf"],
  "devices": ["desktop", "mobile"],
  "fullPage": true,
  "maxHeightPx": 15000,
  "screenshotFormat": "png",
  "hideCookieBanners": true,
  "hideSelectors": ["#intercom-container"],
  "compareWithPrevious": false,
  "maxCostUsd": "1.00"
}
```

| Field | Default | Notes |
|---|---|---|
| `urls` | - | Required. `example.com` is read as `https://example.com`. Duplicates are removed before charging. |
| `outputs` | `["screenshot"]` | `screenshot`, `pdf` or both. |
| `devices` | `["desktop"]` | `desktop` 1920x1080, `laptop` 1366x768, `tablet` iPad (gen 7), `mobile` iPhone 13. One capture per URL and device. |
| `fullPage` / `maxHeightPx` | `true` / 15000 | Longer pages are cut at `maxHeightPx` and marked `truncated: true`. |
| `screenshotFormat` / `jpegQuality` | `png` / 80 | PNG is lossless, JPEG is smaller. |
| `hideCookieBanners` / `hideSelectors` | `true` / `[]` | See above. |
| `scrollForLazyLoad` | `true` | Scroll through the page before capturing. |
| `waitUntil`, `waitForSelector`, `delayMs`, `navigationTimeoutSecs` | `load`, -, 0 (max 10,000 ms), 30 (max 60 s) | Timing controls. A `waitForSelector` that never appears gives `selector-not-found` (free). |
| `pdfFormat`, `pdfLandscape`, `pdfMedia`, `pdfPrintBackground` | A4, false, `screen`, true | `screen` makes the PDF look like the website; `print` uses the site's print stylesheet. |
| `compareWithPrevious`, `minChangePercent`, `baselineStoreName` | false, 0.5, `screenshot-baselines` | Visual diff (PNG only). The last screenshot per URL and device is kept in a named key-value store in your account. The baseline only moves forward for captures you were charged for. |
| `maxCostUsd` | empty | Hard cap. The run stops before starting a capture it could not pay for. |
| `maxConcurrency` | 3 | Parallel pages. Raise the run memory before raising this. |

### Output

Files go to the run's key-value store; one dataset row per URL and device. A real row from a run on the Apify
platform (01.10.2026, heise.de, screenshot + PDF; store ID and URL signature shortened). The page was
33,182 px long, so the screenshot was cut at the default `maxHeightPx` of 15,000 and marked `truncated`; the PDF
is never cut (54 pages):

```json
{
  "url": "https://www.heise.de/",
  "device": "desktop",
  "status": "ok",
  "httpStatus": 200,
  "finalUrl": "https://www.heise.de/",
  "title": "heise online - IT-News, Nachrichten und Hintergründe | heise online",
  "screenshotUrl": "https://api.apify.com/v2/key-value-stores/<storeId>/records/screenshot-www.heise.de-b8529010-desktop.png?signature=...",
  "screenshotKey": "screenshot-www.heise.de-b8529010-desktop.png",
  "screenshotBytes": 6112966,
  "pdfUrl": "https://api.apify.com/v2/key-value-stores/<storeId>/records/pdf-www.heise.de-b8529010-desktop.pdf?signature=...",
  "pdfKey": "pdf-www.heise.de-b8529010-desktop.pdf",
  "pdfBytes": 14578578,
  "pdfPages": 54,
  "pageHeightPx": 33182,
  "capturedHeightPx": 15000,
  "truncated": true,
  "viewportWidthPx": 1920,
  "consentElementsHidden": 1,
  "lazyLoadScrollSteps": 18,
  "diffStatus": null,
  "blockedBy": null,
  "error": null,
  "loadMs": 28591,
  "charged": true,
  "chargedEvents": ["screenshot", "pdf"],
  "capturedAt": "2026-10-01T11:00:56.001Z"
}
```

A URL that could not be captured (same day, control probe; not charged):

```json
{ "url": "https://qs-control-probe-does-not-exist-20261001.com/", "status": "dns-error", "error": "Domain does not resolve (page.goto: net::ERR_NAME_NOT_RESOLVED at https://qs-control-probe-does-not-exist-20261001.com/)", "loadMs": 69, "charged": false, "chargedEvents": [] }
```

Statuses: `ok`, `blocked` (bot-protection challenge, with `blockedBy`), `http-error` (4xx/5xx, with
`httpStatus`), `dns-error`, `network-error`, `tls-error`, `timeout` (retried once), `selector-not-found`,
`invalid-url`, `capture-error`, `storage-error`. Only `ok` is charged.

With `compareWithPrevious`: `diffStatus` (`baseline-created` on the first run, then `compared`),
`diffPercent` (share of pixels that differ, 4 decimals), `changedPixels` (exact count), `changed`,
`heightChangePx`, `baselineCapturedAt`, `diffImageUrl` (changed pixels in red).

The `SUMMARY` record has counts per status, charged events, the stop reason and the URLs not reached
because of a cost limit.

### How much does it cost?

Pay per event. Platform usage (compute, storage) is included in the price; you pay only for delivered files.

| Event | Charged when | Price |
|---|---|---|
| `screenshot` | a screenshot file was stored for one URL and device (visual diff included) | $0.02 |
| `pdf` | a PDF file was stored for one URL and device | $0.02 |
| Actor start | once per run | $0.00005 |

Examples:

- **1,000 full-page screenshots** = 1,000 x $0.02 = **$20**.
- **100 URLs as screenshot + PDF** = 200 files = **$4**.
- **Weekly change check of 50 URLs** (screenshot + diff) = **$1 per week**.
- 100 URLs of which 7 are blocked, 404 or dead = 93 screenshots = **$1.86**; the 7 failed rows are free.

No charge for failed, blocked, invalid, duplicate or not-reached URLs. No per-row dataset charge.
`maxCostUsd` and the platform's "Max cost per run" are both respected: a capture only starts if it can
be paid within the limit.

### Limits

- **Time budget per page:** loading, scrolling and waiting end at most 60 s after the start, the capture itself gets 20 s. Slower pages end as `timeout` (free) instead of running for minutes.
- Pages behind a bot challenge (Cloudflare "Just a moment...", DataDome, PerimeterX, captcha walls) are
  reported as `blocked`. This Actor does not try to get past them and uses no proxies.
- Pages that need a login are captured as the logged-out visitor sees them.
- Elements with `position: fixed` (sticky headers, promo bars) appear once, where they sit on the first
  screen; hide them with `hideSelectors` if needed.
- Chromium only. Very long pages are cut at `maxHeightPx` (max 30,000 px); PDFs are not cut.
- Cookie banner hiding is pattern based; an unknown banner can stay visible - add its selector to
  `hideSelectors`.

### Use it via API, integrations and AI agents

Run it and get the rows back in one call (waits up to 300 s; use the normal run endpoint for long lists):

```bash
curl -X POST "https://api.apify.com/v2/acts/thequietstack~website-screenshot-pdf/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://example.com"],"outputs":["screenshot","pdf"]}'
```

Python (`pip install apify-client`):

```python
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("thequietstack/website-screenshot-pdf").call(
    run_input={"urls": ["https://example.com"], "devices": ["desktop", "mobile"]}
)
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(row["url"], row["device"], row["status"], row["screenshotUrl"])
```

- **Schedules**: run it daily or weekly from Apify Schedules with `compareWithPrevious: true` for change monitoring.
- **Integrations**: Make, Zapier, n8n, Google Drive, Slack and webhooks via the Apify integrations tab.
- **AI agents (MCP)**: add `thequietstack/website-screenshot-pdf` as a tool in the Apify MCP server
  (`https://mcp.apify.com?tools=thequietstack/website-screenshot-pdf`).

### FAQ

**Does it capture the whole page, not just the first screen?** Yes. It scrolls through the page first so lazy-loaded
content renders, then captures the full document height up to `maxHeightPx` (default 15,000 px, max 30,000).
The row says `truncated: true` when a page was longer. PDFs are not cut.

**Does it accept cookie banners for me?** No. Known consent banners are hidden with CSS, never clicked, so no
consent is given on your behalf. Unknown banners can be added with `hideSelectors`.

**What happens with Cloudflare or captcha pages?** They are reported as `blocked` (with `blockedBy`) and are not
charged. The Actor does not try to get past bot protection and uses no proxies.

**Can I take mobile screenshots?** Yes, `devices: ["mobile"]` emulates an iPhone 13; `tablet` an iPad (gen 7).

**How do I detect that a page changed?** Turn on `compareWithPrevious`. The first run stores a baseline, the next runs
return `diffPercent`, `changed` and a red-marked diff image.

**Where are the files and how long are they kept?** In the run's key-value store, linked from every row. Default run
storage follows your Apify plan's data retention; copy the files elsewhere (e.g. Google Drive integration) to keep them.

**Can I take screenshots of pages behind a login?** No. Pages are captured as a logged-out visitor sees them.

### Data license, third-party content and legal

- **You are responsible for having the rights to capture, store and use the pages you submit.** The
  screenshots and PDFs reproduce third-party websites, which may be protected by copyright, trademark
  and database rights and by the sites' terms of use. Keep attribution (the source URL is in every row)
  and use the files only where the law or the site owner allows it (e.g. your own sites, QA,
  archiving for evidence, internal monitoring).
- Do not use this Actor to capture personal data without a legal basis (GDPR/CCPA), or to circumvent
  access controls. It never clicks consent buttons and never bypasses bot protection.
- The Actor code is licensed under the Apache License 2.0. The output files are not covered by that
  license - their rights stay with the owners of the captured websites.

# Actor input Schema

## `urls` (type: `array`):

Pages to capture. "example.com" is read as https://example.com. Duplicates are removed before anything is charged.

## `outputs` (type: `array`):

Screenshot, PDF or both. Each delivered file is one charged event.

## `devices` (type: `array`):

One capture per URL and device. desktop 1920x1080, laptop 1366x768, tablet iPad (gen 7), mobile iPhone 13 (390 px wide, 3x pixel density).

## `fullPage` (type: `boolean`):

Capture the whole page height, not just the first screen.

## `maxHeightPx` (type: `integer`):

Very long pages are cut at this height; the row says so (truncated: true) and reports the real page height.

## `screenshotFormat` (type: `string`):

PNG is lossless; JPEG files are much smaller for long pages.

## `jpegQuality` (type: `integer`):

1-100, only used for JPEG.

## `hideCookieBanners` (type: `boolean`):

Hides consent banners of OneTrust, Cookiebot, Usercentrics, Didomi, Quantcast, TrustArc, Sourcepoint, Borlabs, Complianz, CookieYes, Iubenda and generic patterns, and removes their scroll lock. Banners are hidden, never clicked - no consent is given on your behalf.

## `hideSelectors` (type: `array`):

Extra elements to hide, e.g. a chat widget or a sticky promo bar: "#intercom-container", ".newsletter-popup".

## `scrollForLazyLoad` (type: `boolean`):

Scrolls down step by step before capturing so lazy-loaded images and sections render, then scrolls back to the top.

## `waitUntil` (type: `string`):

When the page counts as loaded before scrolling and capturing. 'networkidle' is safest for heavy pages but slower.

## `waitForSelector` (type: `string`):

Optional. Capture only after this element is visible. If it never appears the URL is reported as selector-not-found (not charged).

## `delayMs` (type: `integer`):

Extra wait before capturing, e.g. for animations.

## `navigationTimeoutSecs` (type: `integer`):

Maximum time to load one page (max 60 s). Loading, scrolling and waiting for one page are capped at 60 s in total. URLs that time out are reported and never charged.

## `pdfFormat` (type: `string`):

Paper size of the PDF output.

## `pdfLandscape` (type: `boolean`):

Print the PDF in landscape orientation.

## `pdfMedia` (type: `string`):

screen = looks like the website in the browser. print = uses the site's print stylesheet.

## `pdfPrintBackground` (type: `boolean`):

Include background colors and images in the PDF.

## `compareWithPrevious` (type: `boolean`):

Compares each PNG screenshot pixel by pixel with the last capture of the same URL and device (stored in a named key-value store in your account). Rows get diffPercent, changed and a highlighted diff image. Included in the screenshot price.

## `minChangePercent` (type: `number`):

A row is marked changed: true when this share of pixels differs.

## `baselineStoreName` (type: `string`):

Named key-value store that keeps the last screenshot per URL and device. Use different names for separate projects.

## `maxCostUsd` (type: `string`):

The run stops before starting a capture it could not pay for within this amount. Uses the live event prices. Leave empty for no cap (the platform's Max cost per run still applies).

## `maxConcurrency` (type: `integer`):

Pages captured at the same time. 3 fits 2 GB memory; raise memory before raising this.

## Actor input object example

```json
{
  "urls": [
    "https://example.com",
    "https://en.wikipedia.org/wiki/Screenshot"
  ],
  "outputs": [
    "screenshot"
  ],
  "devices": [
    "desktop"
  ],
  "fullPage": true,
  "maxHeightPx": 15000,
  "screenshotFormat": "png",
  "jpegQuality": 80,
  "hideCookieBanners": true,
  "scrollForLazyLoad": true,
  "waitUntil": "load",
  "delayMs": 0,
  "navigationTimeoutSecs": 30,
  "pdfFormat": "A4",
  "pdfLandscape": false,
  "pdfMedia": "screen",
  "pdfPrintBackground": true,
  "compareWithPrevious": false,
  "minChangePercent": 0.5,
  "baselineStoreName": "screenshot-baselines",
  "maxConcurrency": 3
}
```

# Actor output Schema

## `captures` (type: `string`):

One row per URL and device: status, screenshot/PDF link, size, page height, diff result, and the reason for every URL that was not captured.

## `files` (type: `string`):

All screenshots, PDFs and diff images of this run (key-value store).

## `summary` (type: `string`):

Counts per status, charged events, cost-cap stop and URLs not reached.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com",
        "https://en.wikipedia.org/wiki/Screenshot"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("thequietstack/website-screenshot-pdf").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://example.com",
        "https://en.wikipedia.org/wiki/Screenshot",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("thequietstack/website-screenshot-pdf").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com",
    "https://en.wikipedia.org/wiki/Screenshot"
  ]
}' |
apify call thequietstack/website-screenshot-pdf --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,thequietstack/website-screenshot-pdf"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/MBp3vEr58NgjjNbQZ/builds/TotuGT4zQPX8ipbCf/openapi.json
