# Website Screenshot - Full Page and Lazy Content (`jake_codes/website-screenshot`) Actor

Screenshot any public web page as PNG, JPEG or WebP. Loads lazy content before capturing, hides cookie banners and floating bars, lets you choose full page, first screen or an exact height, and never charges for a failed capture.

- **URL**: https://apify.com/jake\_codes/website-screenshot.md
- **Developed by:** [Jacob Knowles](https://apify.com/jake_codes) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Screenshot — Full Page and Lazy Content

> **Free beta.** Tested on Apify against a fixed set of 20 public pages chosen to be difficult (lazy loading,
> sticky headers, cookie banners, emoji, very long pages) and on an automated test suite. That is a small,
> hand-picked set: some sites will still render wrongly. Known problems are listed below, and every image
> comes with warnings when something looked off.

Screenshot public web pages as PNG, JPEG or WebP: the whole page, the first screen, or an exact height. One URL
or up to 1,000 per run.

### Faithful by default; cleanup only if you ask

**Faithful capture (default).** The image shows the page as a visitor sees it after it loads: nothing is hidden,
clicked or accepted. Cookie banners stay visible. Floating bars and widgets appear once, at the top of the image,
where a visitor first sees them. They are not repeated down the page.

**Optional cleanup (off by default).** You can hide cookie banners (common consent tools, plus banners that name
themselves "cookie-banner", "consent" or "notice"), hide floating bottom bars and chat widgets, or hide elements by
CSS selector. Cleanup **changes the page and can remove real content** that uses the same names. Every hidden
element is counted in the result's `warnings`.

### What it does

- **Lazy content:** scrolls through the page once so lazy-loaded images and scroll-triggered sections load, waits
  for images, then captures. Images still loading at capture time are reported in `warnings`.
- **Long pages:** captured screen by screen and joined. If the page changes height while it is being captured, it
  is captured again after it settles. If it never settles, the result says `layoutChanged: true`.
- **Sticky headers:** converted to normal elements, so they appear once rather than over the content lower down.
- **Emoji and CJK text:** color emoji and Chinese/Japanese/Korean fonts are installed.
- **Hard limits per page:** a page that never finishes loading is captured as it is, with a warning. A page that
  hangs is stopped and reported. One slow page does not hold up the rest of the run.

### What you pay for

This Actor is free during the beta: you pay only Apify's normal platform usage for the run (compute and storage),
like any Actor without its own price. Failed pages (unreachable site, HTTP error, timeout, refused address) appear
in the results with an `errorCode` and are cheap, because they stop early.

**Planned pricing.** Pay-per-event pricing may be added later: $0.002 per stored screenshot up to 5,000 px tall,
plus $0.0015 for each further 5,000 px. Failed pages would never be charged. Apify notifies users before any pricing
change takes effect. Each row's `extraHeightBlocks` already shows how many extra-height blocks an image would count
as.

### Output

One dataset row per URL:

- **Every row:** `url`, `status` (`ok` / `failed`), `screenshotUrl`, `screenshotKey`, `width`, `height`, `format`,
  `bytes`, `httpStatus`, `finalUrl`, `title`, `pageHeight`, `truncated`, `layoutChanged`, `extraHeightBlocks`, `warnings`, `loadMs` and
  `captureMs`.
- **Failed rows also have:** `errorCode` and `error`.

Images are stored in the run's key-value store. The `SUMMARY` record gives the run outcome (`complete`, `partial`,
`stopped_by_limit` or `failed`) and lists the failures.

**Check `warnings`, `truncated` and `layoutChanged` before using an image.** They are how the Actor tells you
an image may not match the page.

### Known limitations

- **Bot-protected sites** (Cloudflare challenges, login walls) are not bypassed. There is no proxy. When the
  page looks like a challenge or access-denied page, the image comes with a warning.
- **Pages taller than 30,000 px** are cut at 30,000 px and flagged `truncated`.
- **Pages that keep growing** (infinite feeds, some news front pages) stop after the scroll time limit and are
  flagged. Such an image is not a complete page.
- **Elements that change position on timers or scroll handlers** can occasionally appear twice or cover a heading.
  Seen on apple.com, where a dark bar covers a section heading. This is unresolved.
- **Pages that add or remove content after loading** (promotions, personalised boxes) are captured as they
  settled, which can differ between runs. The warnings say when this happened.
- **Local and private network addresses** are refused, including redirects to them.

### Develop

```
npm install          # on a machine where Chromium can run
npm test             # 35 tests: pixel-checked captures + end-to-end Actor runs with simulated billing
```

# Actor input Schema

## `urls` (type: `array`):

One or more public web pages (up to 1,000 per run). 'example.com' works too.

## `mode` (type: `string`):

Full page scrolls the page first so lazy-loaded images appear.

## `height` (type: `integer`):

Used only with 'Exact height from the top'.

## `width` (type: `integer`):

Use 390 for a phone-width layout.

## `viewportHeight` (type: `integer`):

Height of the first screen.

## `format` (type: `string`):

JPEG and WebP are much smaller. Very tall pages fall back to PNG when the format cannot hold them.

## `quality` (type: `integer`):

Ignored for PNG.

## `scale` (type: `integer`):

2 gives a sharper, 'retina' image at twice the pixel size.

## `maxHeight` (type: `integer`):

Taller pages are cut here and flagged as truncated.

## `loadLazyContent` (type: `boolean`):

Scroll through the page once, wait for images, then return to the top before capturing.

## `delayMs` (type: `integer`):

For pages that animate in after loading.

## `darkMode` (type: `boolean`):

Ask the page for its dark colour scheme.

## `waitUntil` (type: `string`):

If the page never gets there, it is captured anyway with a warning.

## `navigationTimeoutSecs` (type: `integer`):

A page that has started rendering is still captured when this runs out.

## `scrollBudgetSecs` (type: `integer`):

Upper bound for the lazy-content scroll, so endless pages cannot stall a run.

## `concurrency` (type: `integer`):

Pages captured at the same time (1-3). More than 1 can cause timeouts on long pages at the default 2 GB memory.

## `captureErrorPages` (type: `boolean`):

Off: an error page counts as a failed capture (and is never charged under pay-per-event pricing).

## `ignoreHttpsErrors` (type: `boolean`):

For sites with expired or self-signed certificates.

## `hideCookieBanners` (type: `boolean`):

Hides banners from common consent tools and banners that name themselves 'cookie-banner/consent/notice'. Nothing is clicked or accepted. Can hide real content that uses the same names.

## `fixedElements` (type: `string`):

Floating elements are always shown in the first screen only, where a visitor sees them. The hiding options remove them entirely.

## `hideSelectors` (type: `array`):

Optional, e.g. #newsletter-popup

## Actor input object example

```json
{
  "urls": [
    "https://apify.com"
  ],
  "mode": "fullPage",
  "height": 2000,
  "width": 1280,
  "viewportHeight": 800,
  "format": "png",
  "quality": 80,
  "scale": 1,
  "maxHeight": 30000,
  "loadLazyContent": true,
  "delayMs": 0,
  "darkMode": false,
  "waitUntil": "load",
  "navigationTimeoutSecs": 30,
  "scrollBudgetSecs": 12,
  "concurrency": 1,
  "captureErrorPages": false,
  "ignoreHttpsErrors": false,
  "hideCookieBanners": false,
  "fixedElements": "keep"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `images` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("jake_codes/website-screenshot").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://apify.com"] }

# Run the Actor and wait for it to finish
run = client.actor("jake_codes/website-screenshot").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com"
  ]
}' |
apify call jake_codes/website-screenshot --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jake_codes/website-screenshot"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/SA3RXPxLd3lQYMADz/builds/nY895ke54ngDIUn0t/openapi.json
