# Website Screenshot & PDF: Mobile, Dark Mode, Markdown (`grit-77/website-screenshot-pdf-markdown`) Actor

Website screenshot and PDF capture for supplied public URLs and viewports. Optionally extract Markdown; save files with links and capture metadata in JSON dataset rows.

- **URL**: https://apify.com/grit-77/website-screenshot-pdf-markdown.md
- **Developed by:** [Grit](https://apify.com/grit-77) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Screenshot & PDF: Mobile, Dark Mode, Markdown

Website screenshot scraper returns image or PDF captures of supplied public URLs with capture metadata and stored file links. Web snapshot API output can also include extracted Markdown when `returnMarkdown` is enabled.

### What it does

Supply URLs, output formats, and viewports. The actor stores delivered files in the run key-value store and emits dataset metadata with `public_url`. Markdown extraction is optional and is stored separately. A `truncated` image was cut at the configured height.

### Input example

```json
{
  "urls": [
    "https://example.com"
  ],
  "formats": [
    "png"
  ],
  "viewports": [
    "desktop"
  ],
  "returnMarkdown": true,
  "maxItems": 1
}
```

The example uses fields defined in the actor input schema. Other filters can be set in the Apify input form.

### Output fields

Examples below come from the first record in [sample_output.json](sample_output.json); long strings and arrays are shown as excerpts. Other modes and error rows can have different fields.

| Field | Type in sample | Example value |
|---|---|---|
| `url` | string | `"https://example.com"` |
| `final_url` | string | `"https://example.com/"` |
| `status` | number | `200` |
| `title` | string | `"Example Domain"` |
| `format` | string | `"png"` |
| `viewport` | string | `"desktop"` |
| `color_scheme` | string | `"light"` |
| `width` | number | `1440` |
| `height` | number | `900` |
| `pages` | null | `null` |
| `truncated` | boolean | `false` |
| `bytes` | number | `38162` |
| `key` | string | `"example-com-327c3fda-desktop-light.png"` |
| `public_url` | string | `"https://api.apify.com/v2/key-value-stores/EXAMPLEstoreID/records/example-com-327c…` |
| `duration_ms` | number | `1421` |
| `error` | null | `null` |
| `captured_at` | string | `"2026-10-05T10:52:23+00:00"` |
| `markdown` | string | `"# Example Domain\n\nThis domain is for use in documentation examples without need…` |
| `markdown_key` | string | `"example-com-327c3fda-desktop-light.md"` |
| `markdown_url` | string | `"https://api.apify.com/v2/key-value-stores/EXAMPLEstoreID/records/example-com-327c…` |

### Output example

A trimmed record copied from [sample_output.json](sample_output.json):

```json
{
  "url": "https://example.com",
  "final_url": "https://example.com/",
  "status": 200,
  "title": "Example Domain",
  "format": "png",
  "viewport": "desktop",
  "color_scheme": "light",
  "width": 1440,
  "height": 900,
  "pages": null,
  "truncated": false,
  "bytes": 38162,
  "key": "example-com-327c3fda-desktop-light.png",
  "public_url": "https://api.apify.com/v2/key-value-stores/EXAMPLEstoreID/records/example-com-327c3fda-desktop-light.png",
  "duration_ms": 1421,
  "error": null,
  "captured_at": "2026-10-05T10:52:23+00:00",
  "markdown": "# Example Domain\n\nThis domain is for use in documentation examples without needing permission. This is not a service; avoid relying on it for testing and monitoring purposes.\n\nهذا النطاق مُخصص للاستخدام في أمثلة التوثيق دون الحاجة إلى إذن. هذه ليست خدمة، يُرجى تجنب الاعتماد عليها لأغراض الاختبار والمراقبة.\n\n该域名仅用于文档示例，无需获得许可。这并非一项服务，请勿将其用于测试和监控目的。\n\nL’usage de ce domaine est réservé à des exemples  ...[truncated in this sample file]",
  "markdown_key": "example-com-327c3fda-desktop-light.md",
  "markdown_url": "https://api.apify.com/v2/key-value-stores/EXAMPLEstoreID/records/example-com-327c3fda-desktop-light.md"
}
```

### FAQ

**Can it capture login-protected or CAPTCHA pages?** No. HTTP errors and unreachable pages produce error rows. A site may block datacenter IPs; `proxyConfiguration` is available for public-page access.

**What limits a capture?** `maxItems` limits input URLs. Image height is capped by `maxHeight`; PDFs paginate. The actor reports truncation in the output row.

**What is the pricing model?** The proposed pay-per-event model charges $0.003 per delivered `snapshot` file and $0.001 per delivered `markdown` extraction. Failed or blocked pages are not charged; the owner sets the proposed values in Apify Console.

**How are rate limits and transient failures handled?** Timeouts, network errors, HTTP 429, and server errors are retried up to `maxRetries` with backoff. The default `maxRetries` is 2.

**How fresh is a snapshot?** A run captures the page as it appears during navigation; `captured_at` records that time. Dynamic content, cookie banners, and site responses can change between runs.

### Use with AI agents / MCP

After this actor is available to your account, configure it as an actor tool in the [Apify MCP server](https://docs.apify.com/platform/integrations/mcp), or start a run through the [Apify API](https://docs.apify.com/api/v2) with the same input JSON. Give the agent the actor identifier and access through your Apify credentials. Read the completed run’s default dataset through the returned run or dataset reference, check error fields before using records, and keep source URLs with quoted data. Limit the input to the task the agent is answering.

### Source and reuse notes

Cookie banners are hidden with CSS where recognized; the actor does not click accept. `returnMarkdown` extracts readable content and may omit navigation, forms, and hidden text. `colorScheme` changes appearance only if the site supports the requested scheme.

# Actor input Schema

## `urls` (type: `array`):

Pages to capture. Plain URLs (https:// optional) or {"url": ...} objects.

## `formats` (type: `array`):

One load of the page can produce several files. PDF uses print layout (see pdfMedia).

## `viewports` (type: `array`):

desktop 1440 px, laptop 1280 px, tablet 768 px (iPad emulation), mobile 390 px (iPhone emulation: touch, mobile user agent). Each preset is a separate capture.

## `colorScheme` (type: `string`):

Sets prefers-color-scheme. Dark only changes sites that ship a dark stylesheet. Choose several via colorSchemes in JSON input.

## `fullPage` (type: `boolean`):

Capture the whole scrollable page instead of just the first screen. Ignored for clipSelector.

## `deviceScaleFactor` (type: `number`):

Pixel density: 1 = normal, 2 = retina (4x the pixels and a bigger file).

## `maxHeight` (type: `integer`):

Taller pages are cut at this height and flagged truncated=true. Chromium cannot render beyond ~16000 device pixels, so the cap is also limited by deviceScaleFactor.

## `imageQuality` (type: `integer`):

1-100.

## `returnMarkdown` (type: `boolean`):

Adds the main text of the page (headings, paragraphs, lists, tables, links) as Markdown next to the screenshot - ready for LLM context. Billed as the 'markdown' event.

## `waitUntil` (type: `string`):

Navigation event to wait for. 'networkidle' waits until the network is quiet for 500 ms and can be slow on chatty sites.

## `delayMs` (type: `integer`):

Wait this long after load (and after scrolling) before capturing, for animations and late widgets.

## `waitForSelector` (type: `string`):

CSS selector that must appear before capturing.

## `scrollToBottom` (type: `boolean`):

Scrolls the page in steps so lazy-loaded images and infinite sections render before a full-page capture.

## `clipSelector` (type: `string`):

Capture only this element (first match) instead of the page. PNG / JPEG / WebP only.

## `hideSelectors` (type: `array`):

Elements matching these selectors get display:none before capture, e.g. .chat-widget, #promo-bar.

## `blockCookieBanners` (type: `boolean`):

Hides common consent banners and overlays (OneTrust, Cookiebot, Didomi, Usercentrics, Quantcast, generic patterns). They are only hidden, never clicked, so no consent is given.

## `blockAds` (type: `boolean`):

Aborts requests to common ad and tracker domains. Faster loads and no ad slots; some layouts leave blank space where ads were.

## `pdfFormat` (type: `string`):

Paper size for PDF output.

## `pdfMedia` (type: `string`):

'print' uses the site's print CSS (usual for articles). 'screen' keeps the on-screen look.

## `timeoutSecs` (type: `integer`):

Navigation and selector timeout.

## `maxItems` (type: `integer`):

Only the first N URLs are processed. Each URL x viewport x color scheme is one page load; each requested format is one billed file.

## `concurrency` (type: `integer`):

Pages captured in parallel (each uses roughly 200-400 MB). 3 is safe on 1 GB, use 6-8 with 4 GB.

## `maxRetries` (type: `integer`):

Extra attempts with exponential backoff for timeouts, network errors, 429 and 5xx. 404 and other client errors are not retried.

## `locale` (type: `string`):

e.g. en-US, de-DE. Sets Accept-Language and navigator.language.

## `proxyConfiguration` (type: `object`):

Optional. Leave off unless a site blocks datacenter IPs.

## Actor input object example

```json
{
  "urls": [
    {
      "url": "https://en.wikipedia.org/wiki/Web_scraping"
    }
  ],
  "formats": [
    "png"
  ],
  "viewports": [
    "desktop"
  ],
  "colorScheme": "light",
  "fullPage": true,
  "deviceScaleFactor": 1,
  "maxHeight": 16000,
  "imageQuality": 80,
  "returnMarkdown": false,
  "waitUntil": "load",
  "delayMs": 0,
  "scrollToBottom": false,
  "blockCookieBanners": true,
  "blockAds": false,
  "pdfFormat": "A4",
  "pdfMedia": "print",
  "timeoutSecs": 60,
  "maxItems": 100,
  "concurrency": 3,
  "maxRetries": 2
}
```

# Actor output Schema

## `dataset` (type: `string`):

Dataset with all result rows

## `files` (type: `string`):

Screenshots, PDFs and Markdown files in the key-value store

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        {
            "url": "https://en.wikipedia.org/wiki/Web_scraping"
        }
    ],
    "formats": [
        "png"
    ],
    "viewports": [
        "desktop"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("grit-77/website-screenshot-pdf-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [{ "url": "https://en.wikipedia.org/wiki/Web_scraping" }],
    "formats": ["png"],
    "viewports": ["desktop"],
}

# Run the Actor and wait for it to finish
run = client.actor("grit-77/website-screenshot-pdf-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    {
      "url": "https://en.wikipedia.org/wiki/Web_scraping"
    }
  ],
  "formats": [
    "png"
  ],
  "viewports": [
    "desktop"
  ]
}' |
apify call grit-77/website-screenshot-pdf-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,grit-77/website-screenshot-pdf-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/rtr6QQZHkObhekutQ/builds/q1XVSLGPe95FhHRDX/openapi.json
