# Website Screenshot & PDF - Fixed Price per File (`tidyfetch/screenshot-fixed-price`) Actor

Bulk website screenshots (PNG/JPEG) and PDFs at a fixed price per file. Full-page or element capture, devices, best-effort cookie handling.

- **URL**: https://apify.com/tidyfetch/screenshot-fixed-price.md
- **Developed by:** [Tidyfetch](https://apify.com/tidyfetch) (community)
- **Categories:** Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 screenshot storeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Screenshot & PDF, fixed price per file

Give it a list of URLs and get back a PNG, JPEG or PDF for each one at a flat price per stored file:
**$3 per 1,000 PNG/JPEG files or $5 per 1,000 PDF files, plus $0.005 per run.** A multi-page PDF counts as one file. Platform usage is included, so a screenshot costs the same whether the page loads in 1 second or 40.

It was built around problems that the most-used screenshot Actor on the Store has left open since 2024: sticky headers repeated across a full-page capture, cookie banners covering the content, emoji rendered as boxes, no way to limit the height, no way to capture only part of a page, and usage-based charges for URLs that produced no file. This Actor addresses each of them; cookie-banner handling is best-effort (see Limits). Login, CAPTCHA solving and proxies are not supported.

### What you get

For every URL, one file in the run's key-value store and one row in the dataset:

```json
{
  "url": "https://example.com",
  "key": "0001-example-com.png",
  "downloadUrl": "https://api.apify.com/v2/key-value-stores/<storeId>/records/0001-example-com.png",
  "format": "png",
  "width": 1280,
  "height": 4321,
  "bytes": 183920,
  "status": "ok",
  "error": null,
  "durationMs": 2310,
  "notes": ["cookie-banner:selector:#onetrust-accept-btn-handler", "frozen-elements:2"]
}
```

`status` is `ok`, `error` (that URL failed, including HTTP 4xx/5xx answers; the others continue), `skipped-robots` (robots.txt disallows the path, or the site's robots.txt could not be fetched) or `skipped-budget` (your run's maximum charge was reached). Only `ok` rows are charged, and an `ok` row always carries a stored file. If the platform's charging call itself fails after the file was stored, the row stays `ok` with `charged: 0` and a `chargeError` field, so you never lose an output you did not pay for.

### How it differs from `apify/screenshot-url`

- Fixed price per file (see below) instead of paying for platform usage; a slow page costs the same as a fast one.
- Cookie banners are accepted or hidden automatically: OneTrust, Cookiebot, Quantcast, TrustArc, Didomi, Usercentrics, Sourcepoint, Osano, CookieYes, Klaro, iubenda, HubSpot, Complianz, Borlabs, Termly, Axeptio, tarteaucitron, Shopify, Squarespace, Wix, Webflow and generic `#cookie-banner`-style markup, plus "Accept all"-style buttons in 15 languages. What cannot be clicked is hidden with CSS.
- Sticky and fixed headers appear once, at the top, in full-page captures. Bottom-anchored floating widgets (chat bubbles, cookie bars) are hidden.
- Emoji and CJK text render correctly: the image ships Noto Color Emoji and Noto CJK fonts.
- `maxHeightPx` caps the document height of full-page PNG/JPEG captures (default 20,000 CSS px) so infinite-scroll pages do not produce a giant image. It does not apply to PDFs or `clipSelector` captures, and device scaling can make the image taller in real pixels.
- `clipSelector` captures a single element (`#pricing`, `main article`).
- Device emulation (`iPhone 13`, `Pixel 7`, `iPad Pro 11`, ... any Playwright device name) and dark mode (`prefers-color-scheme: dark`).
- Lazy-loaded images are triggered by scrolling through the page before capture.
- PDF output with paper size, orientation, background printing and screen/print media choice.
- `respectRobotsTxt` (on by default) skips URLs that the site's `robots.txt` disallows for `User-agent: *`, without charging. Only the wildcard group is evaluated (the page is fetched with a normal browser user agent). A `robots.txt` that answers 5xx or times out is treated as "disallow everything for now" as RFC 9309 requires; a 404 means "no rules". If the page redirects, every redirect target is checked against its own site's `robots.txt` before anything is stored; a disallowed target gives a `skipped-robots` row and is neither stored nor charged. The browser follows redirects itself, so that target may still receive the one request that revealed it.
- HTTP error answers (404, 500, ...) are reported as `error`, not stored and not charged. Turn on `captureErrorPages` if you want the error page captured; those captures are charged like any other output.
- One failed URL never fails the run; you get an `error` row and the rest continue.
- Failed or skipped URLs get no per-file charge; the $0.005 run fee applies to every run. Storing and charging are serialised, so the run's maximum charge is a hard cap even with several pages rendering in parallel.
- "Successful" means a file was rendered and stored. The content is not validated: a login screen, a CAPTCHA page or an access notice that answers HTTP 200 is captured and charged like any other page.

### Pricing

| Event | Price | When |
|---|---|---|
| `run-started` | $0.005 | once per run |
| `screenshot` | $0.003 | per PNG or JPEG successfully stored |
| `pdf` | $0.005 | per PDF successfully stored |

Examples: 1 PNG = $0.008; 100 full-page PNGs = $0.005 + 100 x $0.003 = $0.305; 1 PDF = $0.010. A run where every URL fails or is skipped costs only the $0.005 run fee. Set a maximum total charge on the run if you want a hard cap; URLs beyond it are reported as `skipped-budget`.

Prices may be lower on higher Apify plan tiers; the Actor page shows the exact figures for your account.

### Input

| Field | Default | Meaning |
|---|---|---|
| `urls` | required | Array of URLs (strings or `{ "url": ... }`). Up to 1,000 per run. |
| `format` | `png` | `png`, `jpeg` or `pdf`. |
| `fullPage` | `true` | Whole scrollable page instead of the first viewport. |
| `maxHeightPx` | `20000` | Document height cap in CSS px for full-page PNG/JPEG (200-30,000). Ignored for PDF and `clipSelector`. |
| `viewportWidth` / `viewportHeight` | `1280` / `800` | Browser size in CSS px when no device is set. |
| `device` | none | Playwright device name (sets viewport, scale factor, touch, user agent). |
| `darkMode` | `false` | Emulate `prefers-color-scheme: dark`. |
| `clipSelector` | none | CSS selector; only the first match is captured. Not for PDF. |
| `hideCookieBanners` | `true` | Accept or hide consent banners. |
| `freezeStickyElements` | `true` | Fixed/sticky elements become static; bottom-anchored ones are hidden. |
| `delayMs` | `500` | Extra wait after load (0-30,000). |
| `waitUntil` | `load` | `load`, `domcontentloaded` or `networkidle`. |
| `timeoutSecs` | `60` | Navigation and screenshot timeout (5-300); the URL is reported as `error` when exceeded. Rendering may take up to 15 s more; the robots.txt check and storing are outside this limit. |
| `jpegQuality` | `80` | 1-100, JPEG only. |
| `blockResources` | `false` | Drop video/audio, websockets and analytics/ad/chat hosts; images, CSS, fonts and scripts are kept. |
| `respectRobotsTxt` | `true` | Skip URLs (and redirect targets) disallowed for `User-agent: *`, and all URLs of a site whose robots.txt answers 5xx or is unreachable. |
| `captureErrorPages` | `false` | Capture (and charge) pages that answer HTTP 4xx/5xx instead of reporting them as `error`. |
| `scrollForLazyLoad` | `true` | Scroll through the page before a full-page capture. |
| `maxConcurrency` | `2` | Pages rendered in parallel (1-10). |
| `ignoreHttpsErrors` | `true` | Ignore certificate errors while rendering. The robots.txt check does not ignore them, so with `respectRobotsTxt` on, a site with an invalid certificate may come back as `skipped-robots`. |
| `pdfFormat` / `pdfLandscape` / `pdfPrintBackground` / `pdfEmulateScreenMedia` | `A4` / `false` / `true` / `false` | PDF only. |

### Limits and known behaviour

- Pages that need a login, a CAPTCHA solve or geo-specific proxies are out of scope for this version. There is no proxy option yet.
- `maxHeightPx` is limited to 30,000 px; very tall captures use a lot of memory, so raise the run's memory for many parallel tall pages.
- Cookie-banner handling is best-effort. The Actor only presses accept buttons inside a consent banner (a container named after cookies or consent, a dialog whose text is about cookies, or a consent-manager frame), never an "OK" or "I agree" elsewhere on the page. A banner built with an unknown library may remain visible or only be hidden with CSS; send the URL as an issue and it will be added.
- Page content is not checked: login walls, CAPTCHA pages and access notices returned with HTTP 200 are captured and charged.
- Device emulation runs in Chromium even for iPhone/iPad descriptors (Safari-only rendering differences are not reproduced).
- PDF output uses the page's print stylesheet by default; turn on `pdfEmulateScreenMedia` for a what-you-see rendering.
- The download URL points to the run's key-value store, which follows the platform's data retention for your plan.

### Using the result

The `downloadUrl` in each dataset row returns the file directly (`Content-Type` is `image/png`, `image/jpeg` or `application/pdf`). In integrations, read the dataset, filter `status == "ok"` and fetch `downloadUrl`.

### Legal

You are responsible for the URLs you submit. The Actor fetches only the page you give it, does not log in, does not fill forms, and respects `robots.txt` by default.

### About

This Actor was written by an AI coding agent (Claude Code) under human direction and review; a human maintainer reviews issues and the replies to them.

### Changelog

See the Changelog tab (CHANGELOG.md).

# Changelog

This Actor's version history is a separate document: https://apify.com/tidyfetch/screenshot-fixed-price/changelog.md

# Actor input Schema

## `urls` (type: `array`):

Pages to capture. One output per URL. http(s) only.

## `format` (type: `string`):

PNG (lossless), JPEG (smaller, see quality) or PDF (paged, uses print CSS unless 'PDF: emulate screen media' is on).

## `fullPage` (type: `boolean`):

Capture the whole scrollable page instead of only the first viewport. Ignored for PDF and when a clip selector is set.

## `maxHeightPx` (type: `integer`):

Maximum document height in CSS pixels for full-page PNG/JPEG captures, so a 100,000 px infinite-scroll page does not produce a giant image. Ignored for PDF and clipSelector. Device scaling can make the image taller in real pixels; this is not a file-size limit.

## `viewportWidth` (type: `integer`):

Browser width in CSS pixels. Ignored when a device is selected.

## `viewportHeight` (type: `integer`):

Browser height in CSS pixels. This is the image height when 'Full page' is off. Ignored when a device is selected.

## `device` (type: `string`):

Optional Playwright device name, e.g. 'iPhone 13', 'Pixel 7', 'iPad Pro 11', 'Desktop Chrome HiDPI'. Sets viewport, scale factor, touch and user agent. Full list: https://github.com/microsoft/playwright/blob/main/packages/playwright-core/src/server/deviceDescriptorsSource.json

## `darkMode` (type: `boolean`):

Emulate prefers-color-scheme: dark.

## `clipSelector` (type: `string`):

Capture only the first element matching this CSS selector, e.g. '#pricing' or 'main article'. Not available for PDF.

## `hideCookieBanners` (type: `boolean`):

Clicks the 'Accept' button of common consent libraries (OneTrust, Cookiebot, Quantcast, Didomi, Usercentrics, ...) or, failing that, hides the banner with CSS.

## `freezeStickyElements` (type: `boolean`):

Turns position: fixed/sticky headers into normal elements so they appear once at the top of a full-page capture instead of repeating or floating mid-page. Bottom-anchored fixed elements (chat bubbles, cookie bars) are hidden.

## `delayMs` (type: `integer`):

Milliseconds to wait after the page loads (animations, web fonts, late JavaScript).

## `waitUntil` (type: `string`):

Navigation is considered finished at this event. 'networkidle' is slowest but best for JavaScript-heavy pages.

## `timeoutSecs` (type: `integer`):

Navigation and screenshot timeout; when exceeded the URL is reported as an error and the other URLs continue. Rendering may take up to 15 s more, and the robots.txt check and storing are outside this limit.

## `jpegQuality` (type: `integer`):

1-100, only for JPEG.

## `blockResources` (type: `boolean`):

Drops video/audio, websockets and common analytics/ad/chat-widget hosts to load faster. Images, CSS, fonts and scripts are kept.

## `respectRobotsTxt` (type: `boolean`):

If the site's robots.txt disallows the URL path for 'User-agent: \*', the URL is skipped with status 'skipped-robots' and is not charged. Redirect targets are checked against their own site's robots.txt before anything is stored. A robots.txt that answers 5xx or cannot be reached also skips the URL (RFC 9309); a 404 means no rules.

## `captureErrorPages` (type: `boolean`):

Off (default): a page that answers 404, 500 etc. is reported as 'error' and is not stored or charged. On: the error page is captured, stored and charged like any other output (the HTTP status is recorded in 'notes').

## `scrollForLazyLoad` (type: `boolean`):

Before a full-page capture, scroll through the page so lazy-loaded images render.

## `maxConcurrency` (type: `integer`):

How many URLs are rendered at the same time in the shared browser. Raise memory if you raise this.

## `ignoreHttpsErrors` (type: `boolean`):

Ignore certificate errors while rendering (self-signed or expired certificates). The robots.txt check does not ignore them, so with respectRobotsTxt on, such a site may come back as skipped-robots.

## `pdfFormat` (type: `string`):

Only for PDF.

## `pdfLandscape` (type: `boolean`):

Only for PDF.

## `pdfPrintBackground` (type: `boolean`):

Include background colours and images in the PDF.

## `pdfEmulateScreenMedia` (type: `boolean`):

Render the PDF with screen CSS instead of print CSS (closer to what you see in the browser; pages with print stylesheets may look better with this off).

## Actor input object example

```json
{
  "urls": [
    "https://example.com"
  ],
  "format": "png",
  "fullPage": true,
  "maxHeightPx": 20000,
  "viewportWidth": 1280,
  "viewportHeight": 800,
  "darkMode": false,
  "hideCookieBanners": true,
  "freezeStickyElements": true,
  "delayMs": 500,
  "waitUntil": "load",
  "timeoutSecs": 60,
  "jpegQuality": 80,
  "blockResources": false,
  "respectRobotsTxt": true,
  "captureErrorPages": false,
  "scrollForLazyLoad": true,
  "maxConcurrency": 2,
  "ignoreHttpsErrors": true,
  "pdfFormat": "A4",
  "pdfLandscape": false,
  "pdfPrintBackground": true,
  "pdfEmulateScreenMedia": false
}
```

# Actor output Schema

## `results` (type: `string`):

One row per input URL with status, size and a download link.

## `files` (type: `string`):

Screenshots and PDFs in the run's key-value store.

## `summary` (type: `string`):

Counts per status for the run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://example.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tidyfetch/screenshot-fixed-price").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://example.com"] }

# Run the Actor and wait for it to finish
run = client.actor("tidyfetch/screenshot-fixed-price").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://example.com"
  ]
}' |
apify call tidyfetch/screenshot-fixed-price --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tidyfetch/screenshot-fixed-price"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/V67GQ5u5d4eNwL723/builds/pF8CxhnIfHoEBhk15/openapi.json
