# Bulk Website Screenshot & URL to PDF Converter (`oldjard/screenshot-pdf`) Actor

Screenshot a list of websites and save them as PDFs in one run. Full page or first screen, desktop or iPhone/Android, PNG, JPEG or WebP, A4 or Letter PDF. Cookie banners and pop-ups hidden. Failed pages are free. $5 per 1,000 screenshots, $6 per 1,000 PDFs.

- **URL**: https://apify.com/oldjard/screenshot-pdf.md
- **Developed by:** [Joshua White](https://apify.com/oldjard) (community)
- **Categories:** Developer tools, Automation, SEO tools
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk Website Screenshot & URL to PDF Converter

**Take screenshots of a whole list of websites and convert them to PDF**, in one run. Paste URLs, pick desktop or
mobile, full page or first screen, and get a PNG, JPEG or WebP image and/or an A4 or Letter PDF of each page, with a
download link per row. **Cookie banners, sign-up pop-ups and chat bubbles are hidden** (never accepted), lazy images
are loaded, and **pages that fail are never charged**.

**Try it in one click:** the input is prefilled with two pages, as screenshot + PDF. Apify's free plan covers about
1,000 screenshots a month.

### How to screenshot websites or convert a URL to PDF in 3 steps

1. Paste web addresses into **URLs**, one per line (`example.com` works; typos like `htps://` are fixed).
2. Choose **screenshot**, **PDF** or both, a **device** (desktop, laptop, tablet, iPhone, Android) and **full page**
   or first screen.
3. Click **Start**. Each result row links its image and PDF; download the table as CSV, Excel or JSON, or call it
   from the API, Make, Zapier or n8n.

### How much does a website screenshot cost?

**$5 per 1,000 screenshots** and **$6 per 1,000 PDFs**, so Apify's $5 monthly free credit covers about 1,000 screenshots. Failed pages (DNS errors, timeouts, bot checks, HTTP errors,
robots.txt blocks, file URLs, or a "wait for element" that never appeared) are free. 200 full-page screenshots plus
PDFs cost $2.20. If you set a maximum cost for a run, the actor stops cleanly when it reaches it and tells you how
many pages it skipped.

### What you can use bulk screenshots and PDFs for

- **Archiving and evidence.** Keep a dated image or PDF of a page: ads, prices, terms, compliance pages, press
  coverage.
- **Visual monitoring.** Screenshot competitors' homepages, pricing pages or landing pages on a schedule.
- **Reports and decks.** Mobile and desktop screenshots of client or prospect sites for audits, pitches and QA.
- **URL to PDF.** Turn articles, docs, invoices or any web page into a printable PDF, in bulk.
- **Thumbnails and previews.** Small JPEG or WebP images of many sites for a directory, a dashboard or an AI pipeline.

### Features

| | |
|---|---|
| Image formats | PNG (lossless), JPEG and WebP (much smaller files), quality setting |
| Page size | Full scrolling page (up to 30,000 px tall) or just the first screen |
| Devices | Desktop 1920×1080, laptop 1366×768, tablet, iPhone and Android with retina pixel density, or any custom size |
| PDF | A4, Letter, Legal, A3, A5, Tabloid; portrait or landscape; margins; backgrounds; screen or print layout |
| Clean-up | Hides cookie / consent banners (OneTrust, Cookiebot, Didomi, Usercentrics, Quantcast, TrustArc and home-made ones), sign-up modals, sticky bottom bars and chat bubbles; your own CSS selectors to hide; custom CSS |
| Timing | Wait for page load or network quiet, wait for a CSS selector, extra delay |
| Extras | Dark mode, lazy-image scrolling, blank-page detection, honest error codes |

Banners are **hidden, never accepted**: the tool doesn't give consent on anyone's behalf.

### Input example

```json
{
  "urls": ["https://apify.com", "en.wikipedia.org/wiki/Web_scraping"],
  "captureScreenshot": true,
  "capturePdf": true,
  "device": "iphone",
  "fullPage": true,
  "imageFormat": "jpeg",
  "pdfFormat": "Letter"
}
```

### Output

One row per URL. Files are stored in the run's key-value store, and the row links to them.

Pages finish in parallel, so rows don't arrive in input order. **`index`** is each URL's position in your list,
starting at 1 (blank lines and `#` comments aren't counted; a removed duplicate keeps its number, so there's a gap).
Sort by `index` to line the rows up with your list.

```json
{
  "index": 1,
  "url": "https://en.wikipedia.org/wiki/Web_scraping",
  "finalUrl": "https://en.wikipedia.org/wiki/Web_scraping",
  "title": "Web scraping - Wikipedia",
  "statusCode": 200,
  "screenshotUrl": "https://api.apify.com/v2/key-value-stores/<store>/records/en-wikipedia-org-wiki-Web-scraping-51e366a8.png",
  "screenshotKey": "en-wikipedia-org-wiki-Web-scraping-51e366a8.png",
  "screenshotFormat": "png",
  "screenshotWidth": 1920,
  "screenshotHeight": 8745,
  "screenshotBytes": 2398899,
  "screenshotTruncated": false,
  "looksBlank": false,
  "pdfUrl": "https://api.apify.com/v2/key-value-stores/<store>/records/en-wikipedia-org-wiki-Web-scraping-51e366a8.pdf",
  "pdfKey": "en-wikipedia-org-wiki-Web-scraping-51e366a8.pdf",
  "pdfPages": 11,
  "pdfBytes": 1830112,
  "pageHeight": 8745,
  "device": "desktop",
  "cookieBannersHidden": 0,
  "error": null,
  "errorCode": null,
  "capturedAt": "2026-10-05T19:12:03.511Z"
}
```

Files stay in the run's storage for as long as your Apify plan keeps run data. Download the ones you want to keep.

### Errors (not charged)

| `errorCode` | Meaning |
|---|---|
| `DNS` | The domain does not exist |
| `TIMEOUT` | The page didn't start loading within the timeout |
| `CONNECTION`, `TLS`, `TOO_MANY_REDIRECTS` | Network, certificate or redirect-loop problems |
| `HTTP_BLOCKED`, `HTTP_ERROR` | The site answered 401/403/429, or another 4xx/5xx |
| `BOT_CHALLENGE` | The site showed a bot check or waiting room (Cloudflare, Akamai, DataDome, ...) instead of the page |
| `ROBOTS_DISALLOWED` | The site's robots.txt asks automated visitors not to open this page |
| `NOT_A_WEBPAGE` | The URL is a file (PDF, image, download), not a page |
| `WAIT_FOR_SELECTOR` | The element you asked to wait for never appeared |
| `INVALID_INPUT` | The line isn't a web address (for example `not a url`); `url` shows what you typed |

Every line you paste gets a row, so the status counts them all: "Captured 2 of 4 entries; 2 weren't web addresses".

### Tips

- **Memory:** the default 2 GB runs 2 pages at a time (one per ~700 MB). For big lists, raise memory and **Pages at
  once** together (about 700 MB per page).
- **Very long pages** are cut at **Max page height** (15,000 px by default) and the row says
  `screenshotTruncated: true`. WebP images can't be taller than 16,383 pixels, so use PNG or JPEG for very long
  retina captures.
- **Smaller files:** JPEG or WebP at quality 70–80 is usually indistinguishable from PNG for web pages.
- **Charts and animations:** add a few seconds of **Extra delay**, or **Wait for element** with the chart's selector.
- **Cleaner PDFs:** try **PDF layout → Print stylesheet**; many sites drop menus and ads when printing.

### Limits and good behaviour

The tool opens public pages only: no logins, no paywalls. It respects robots.txt and identifies itself in its user
agent (`ScreenshotPdfBot`). Sites protected by aggressive bot checks may show a challenge page instead; those rows
come back as `BOT_CHALLENGE` and are free.

### Ready-made examples

Each one opens this actor with the input already filled in. Click **Try** to run it, or change the input to fit your own list.

- [Full-page screenshot of a website](https://apify.com/oldjard/screenshot-pdf/examples/full-page-website-screenshot)
- [Save a web page as PDF for archiving](https://apify.com/oldjard/screenshot-pdf/examples/save-webpage-as-pdf)
- [Mobile screenshot as an iPhone sees it](https://apify.com/oldjard/screenshot-pdf/examples/mobile-screenshot-iphone)
- [Screenshot several competitor homepages](https://apify.com/oldjard/screenshot-pdf/examples/screenshot-competitor-homepages)
- [Dark mode screenshot of a website](https://apify.com/oldjard/screenshot-pdf/examples/dark-mode-screenshot)
- [Print a web page to PDF (US Letter, print styles)](https://apify.com/oldjard/screenshot-pdf/examples/webpage-to-pdf-print-letter)

### More tools from oldjard

- [Tech Stack Detector](https://apify.com/oldjard/tech-stack-detector): what any list of websites is built with.
- [Sitemap URL Extractor](https://apify.com/oldjard/sitemap-url-extractor): every URL on a website, for RAG and SEO.
- [Shopify Products Scraper & Price Monitor](https://apify.com/oldjard/shopify-products-price-monitor): catalogs and price changes from any Shopify store.
- [Workday, Greenhouse, Lever & Ashby Jobs Scraper](https://apify.com/oldjard/ats-career-site-jobs): every open job from company career sites.
- [AI Web Scraper (your own key)](https://apify.com/oldjard/ai-web-scraper): describe fields in English, get JSON.
- [Website Change Monitor](https://apify.com/oldjard/website-change-monitor): a before/after diff by webhook, Slack or Discord when a page changes.
- [Company Registry Lookup](https://apify.com/oldjard/company-registry-lookup): UK Companies House, Spain, France, Finland and Norway in one schema.
- [UK & EU Public Tenders](https://apify.com/oldjard/uk-eu-public-tenders): Find a Tender and TED notices in one table, with daily only-new alerts.

### Use it from an AI agent or the API

- **Minimal input:** `{"urls": ["https://example.com"]}`. Everything else has a sensible default.
- **Cost:** $0.005 per screenshot and $0.006 per PDF. Pages that fail to load are free. One page as both ≈ $0.011.
- **Run time (our runs):** 20–35 s for 2–6 pages at the default 2 GB, so an agent will usually poll once.
- **Results:** the default dataset, one row per URL with `screenshotUrl` and `pdfUrl` download links; the files are in
  the key-value store, and the run summary is its `OUTPUT` record.
- Works over the Apify MCP server (`search-actors`, then `call-actor`) and is eligible for agentic payments (x402).

### FAQ

**Can it capture pages behind a login?** No, public pages only.

**HTML to PDF?** It converts live URLs to PDF; raw HTML input isn't supported.

**Is there a website screenshot API?** Yes. Call the actor from the Apify API (or Make, Zapier, n8n) with a list of
URLs; each result row has `screenshotUrl` and `pdfUrl` download links.

**Does it click "Accept" on cookie banners?** No. Banners are hidden on the page, never accepted, so the tool gives no
consent on anyone's behalf. The row's `cookieBannersHidden` counts the banner elements it hid.

**How long can a full-page screenshot be?** Up to 30,000 px tall (set **Max page height**; the default is 15,000 px so
endless-scroll pages finish). A page that was cut says `screenshotTruncated: true`.

**Can I take screenshots on a schedule?** Yes. Save your input as a task and schedule it in Apify, daily or weekly.
Each run stores new files, and each row has a `capturedAt` time, so you build a dated archive.

**Where are the image and PDF files?** In the run's key-value store (Storage tab). Each row links its own files.

### Feedback

Found a page that doesn't render right? Open an issue on the Issues tab with the URL and your settings, and it will
be looked at.

# Actor input Schema

## `urls` (type: `array`):

Web pages to capture, one per line. Paste a spreadsheet column straight in: bare domains like "example.com" work, duplicates are removed.

## `captureScreenshot` (type: `boolean`):

Save an image of each page. Charged per screenshot.

## `capturePdf` (type: `boolean`):

Also save each page as a PDF (URL to PDF). Charged per PDF.

## `device` (type: `string`):

Screen size and type to emulate. Phone and tablet presets load the mobile version of the site.

## `fullPage` (type: `boolean`):

Capture the whole scrolling page, not just the first screen.

## `imageFormat` (type: `string`):

PNG is lossless. JPEG and WebP files are much smaller.

## `imageQuality` (type: `integer`):

1–100. Ignored for PNG.

## `viewportWidth` (type: `integer`):

Override the device's screen width in CSS pixels. Leave empty to use the device preset.

## `viewportHeight` (type: `integer`):

Override the device's screen height. For full-page captures this is only the first screen.

## `deviceScaleFactor` (type: `number`):

1 = normal, 2 = retina. Leave empty to use the device preset (phones are 2.6–3×).

## `maxPageHeight` (type: `integer`):

Full-page captures stop here (in CSS pixels), so endless-scroll pages finish. The row says when a page was cut.

## `pdfFormat` (type: `string`):

Paper size for PDFs.

## `pdfLandscape` (type: `boolean`):

Landscape instead of portrait pages.

## `pdfMargin` (type: `string`):

Page margin on all sides, e.g. "1cm", "0.5in", "10mm" or "0".

## `pdfPrintBackground` (type: `boolean`):

Keep background colours and images, as on screen.

## `pdfMediaType` (type: `string`):

'As on screen' looks like the website. 'Print' uses the site's own print styles, if it has them.

## `hideCookieBanners` (type: `boolean`):

Hide cookie and consent pop-ups (OneTrust, Cookiebot, Didomi, Usercentrics and home-made ones). They are hidden, never accepted.

## `closePopups` (type: `boolean`):

Hide sign-up, newsletter and app-install modals that cover the page.

## `hideFixedBottomBars` (type: `boolean`):

Bars and chat bubbles pinned to the bottom of the screen would otherwise float over the middle of a full-page image.

## `hideSelectors` (type: `array`):

CSS selectors of anything else to hide, e.g. ".ad-banner" or "#chat-widget".

## `customCss` (type: `string`):

CSS added to every page before the capture.

## `darkMode` (type: `boolean`):

Ask the site for its dark theme (prefers-color-scheme: dark).

## `blockAds` (type: `boolean`):

Skip the big ad networks and trackers. Pages load faster and ad slots come out empty. Turn off to capture pages exactly as visitors see them, ads included.

## `waitUntil` (type: `string`):

When the page counts as ready. Slow trackers are capped at 15 seconds either way.

## `delaySecs` (type: `number`):

Wait this long before the capture, e.g. for animations or charts.

## `waitForSelector` (type: `string`):

CSS selector to wait for before the capture, e.g. "#chart canvas". If it never shows, the page is reported as failed (and not charged).

## `scrollToLoadLazyContent` (type: `boolean`):

Scroll down the page first so lazy-loaded images and sections appear in full-page captures and PDFs.

## `navigationTimeoutSecs` (type: `integer`):

Give up on a page that does not start loading in this time (not charged).

## `maxConcurrency` (type: `integer`):

Pages captured in parallel. Leave empty for automatic (one per ~700 MB of memory: 2 at the default 2 GB).

## `canaryMinSuccessRate` (type: `number`):

Internal monitoring only; leave empty. 0.8 fails the run if fewer than 80% of pages are captured (blank images count as failures).

## Actor input object example

```json
{
  "urls": [
    "https://apify.com",
    "https://en.wikipedia.org/wiki/Web_scraping"
  ],
  "captureScreenshot": true,
  "capturePdf": true,
  "device": "desktop",
  "fullPage": true,
  "imageFormat": "png",
  "imageQuality": 80,
  "maxPageHeight": 15000,
  "pdfFormat": "A4",
  "pdfLandscape": false,
  "pdfMargin": "1cm",
  "pdfPrintBackground": true,
  "pdfMediaType": "screen",
  "hideCookieBanners": true,
  "closePopups": true,
  "hideFixedBottomBars": true,
  "darkMode": false,
  "blockAds": true,
  "waitUntil": "load",
  "delaySecs": 0,
  "scrollToLoadLazyContent": true,
  "navigationTimeoutSecs": 45
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset, one row per URL: index, url, finalUrl, title, statusCode, screenshotUrl, screenshotFormat, screenshotWidth, screenshotHeight, screenshotBytes, looksBlank, pdfUrl, pdfPages, pdfBytes, device, error, errorCode, capturedAt. Open screenshotUrl or pdfUrl directly to download the file.

## `files` (type: `string`):

Key-value store holding the PNG, JPEG or WebP screenshots and the PDF files.

## `summary` (type: `string`):

Run summary (JSON): pages captured and failed, screenshots and PDFs made, blank pages, invalid inputs and duplicates removed.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://apify.com",
        "https://en.wikipedia.org/wiki/Web_scraping"
    ],
    "capturePdf": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("oldjard/screenshot-pdf").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://apify.com",
        "https://en.wikipedia.org/wiki/Web_scraping",
    ],
    "capturePdf": True,
}

# Run the Actor and wait for it to finish
run = client.actor("oldjard/screenshot-pdf").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://apify.com",
    "https://en.wikipedia.org/wiki/Web_scraping"
  ],
  "capturePdf": true
}' |
apify call oldjard/screenshot-pdf --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,oldjard/screenshot-pdf"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/fALjoK1A85kPiGxxa/builds/zGdJSwaUBLLXB9omC/openapi.json
