# PDF to Images – also PowerPoint & Word (PNG, JPG, WebP) (`mauberme/document-to-images`) Actor

Turn every page of a PDF, PowerPoint or Word file into an image. Use it for website previews, thumbnails, social posts or to feed pages to an AI tool. Paste file links, get one image per page (PNG, JPG or WebP) plus a ZIP. Free during launch.

- **URL**: https://apify.com/mauberme/document-to-images.md
- **Developed by:** [Mauberme](https://apify.com/mauberme) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF to Images – also PowerPoint & Word (PNG, JPG, WebP)

**What is it for?** Need the pages of a document as pictures? Paste links to PDF, PowerPoint or Word files and get **one image per page**, ready for a website preview, a thumbnail, a social media post or an AI tool that reads images. No software to install, no manual screenshots.

Convert **PDF, PowerPoint (PPTX, PPT), Word (DOCX, DOC), Excel (XLSX, XLS) and OpenDocument (ODP, ODT, ODS)** files into **one image per page** in **WebP, PNG or JPG**. Give it document URLs (or upload a file) and get back a dataset with a public link to every page image, its size in pixels and bytes, and optionally a **ZIP with all pages** of each document.

Typical uses: PDF to PNG for previews and thumbnails, slide decks (PPTX to images) for social posts or web galleries, Word to JPG for CMS uploads, page images for OCR or AI vision pipelines, catalogue and brochure pages for e-commerce.

### Features

- **PDF to PNG, PDF to WebP, PDF to JPG** at 72–300 DPI.
- **PowerPoint to images, Word to images, Excel to images**: Office files are converted with LibreOffice (headless) and then rendered page by page.
- **Page ranges**: `"1"`, `"1-3,5"`, `"4-"` (from page 4 to the end).
- **Resize** with `maxWidth` (never upscales) and **trim white margins** around the content.
- **Quality control** for WebP/JPG, lossless PNG.
- **ZIP per document** with all page images.
- One dataset row per page, so a broken page never ruins the whole document.

### Input example

```json
{
  "documents": [
    "https://example.com/brochure.pdf",
    "https://example.com/pitch-deck.pptx"
  ],
  "format": "webp",
  "dpi": 150,
  "pages": "1-3",
  "maxWidth": 1600,
  "trimWhitespace": false,
  "quality": 85,
  "zip": true
}
```

| Field | Type | Default | Description |
|---|---|---|---|
| `documents` | array of URLs | – | Public http(s) URLs of PDF/Office files. |
| `uploadedDocument` | file upload | – | A single file uploaded from the Apify Console. |
| `format` | `webp` | `png` | `jpeg` | `webp` | Output image format. |
| `dpi` | 72–300 | 150 | Rendering resolution. |
| `pages` | string | all | Page selection, e.g. `1-3,5`. |
| `maxWidth` | integer | – | Downscale pages wider than this. |
| `trimWhitespace` | boolean | `false` | Crop white margins. |
| `quality` | 1–100 | 85 | WebP/JPG quality. |
| `zip` | boolean | `false` | Also store a ZIP per document. |
| `maxPagesPerDocument` | 1–500 | 200 | Safety cap per document. |
| `maxFileSizeMb` | 1–100 | 50 | Skip larger files. |

### Output example

One row per page:

```json
{
  "documentUrl": "https://example.com/brochure.pdf",
  "page": 1,
  "totalPages": 12,
  "format": "webp",
  "width": 1240,
  "height": 1754,
  "bytes": 184233,
  "imageUrl": "https://api.apify.com/v2/key-value-stores/<storeId>/records/doc-1-page-001.webp",
  "zipUrl": "https://api.apify.com/v2/key-value-stores/<storeId>/records/doc-1-images.zip",
  "error": null
}
```

If a document or a page fails (corrupt file, password-protected PDF, unsupported type, download error), you get a row with `error` filled in and the run continues with the rest.

Images are stored in the run's default key-value store; `imageUrl` and `zipUrl` are direct download links.

### Pricing

**Free during launch.** This Actor has no usage fee of its own: you only pay Apify's standard platform usage for your runs (on the free Apify plan that is covered by your monthly credit). A pay-per-result price may be introduced later; Apify notifies users in advance of any price change.

### Limits

- Up to 100 documents per run.
- Max file size: 50 MB by default, 100 MB hard cap.
- Max pages per document: 200 by default, 500 hard cap.
- DPI 72–300; very large page formats are automatically rendered at a lower DPI so no side exceeds 10,000 px.
- Office → PDF conversion times out after 3 minutes per file; each page render times out after 2 minutes.
- **Office files are slower**: LibreOffice needs a few seconds to start for every Office document. PDFs start rendering immediately.
- Only public `http`/`https` URLs are accepted. Private networks, localhost and cloud metadata addresses are blocked, including through redirects.
- Password-protected documents are not supported.
- Rendering fidelity of Office files depends on LibreOffice and the fonts available (DejaVu and Noto are installed). Documents using proprietary fonts will fall back to similar fonts.

### FAQ

**Can I convert only the first page (thumbnail)?** Yes: `"pages": "1"`, `"dpi": 72` and, for example, `"maxWidth": 400`.

**Does it keep transparency?** Pages are rendered on a white background, as in a PDF viewer.

**Does it run macros in Office files?** No. Documents are only converted for display.

# Actor input Schema

## `documents` (type: `array`):

Public http(s) URLs of PDF, PPTX, PPT, DOCX, DOC, XLSX, XLS, ODP, ODT or ODS files. Each page becomes one image.

## `uploadedDocument` (type: `string`):

Upload a single PDF or Office document instead of (or in addition to) the URLs above.

## `format` (type: `string`):

Image format of every page.

## `dpi` (type: `integer`):

Rendering resolution. 72 = screen thumbnail, 150 = sharp on screens, 300 = print quality.

## `pages` (type: `string`):

Pages to convert, e.g. "1", "1-3,5" or "2-" (from page 2 to the end). Leave empty for all pages.

## `maxWidth` (type: `integer`):

Optional. Downscale pages wider than this, keeping the aspect ratio. Never upscales.

## `trimWhitespace` (type: `boolean`):

Crop the white border around the page content.

## `quality` (type: `integer`):

Compression quality for WebP and JPG (1-100). Ignored for PNG.

## `zip` (type: `boolean`):

Besides the individual images, store one ZIP with all pages of each document.

## `maxPagesPerDocument` (type: `integer`):

Safety limit: only the first N selected pages of each document are converted. Hard cap 500.

## `maxFileSizeMb` (type: `integer`):

Files larger than this are skipped with an error. Hard cap 100 MB.

## Actor input object example

```json
{
  "documents": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ],
  "format": "webp",
  "dpi": 150,
  "trimWhitespace": false,
  "quality": 85,
  "zip": false,
  "maxPagesPerDocument": 200,
  "maxFileSizeMb": 50
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `files` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "documents": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mauberme/document-to-images").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "documents": ["https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("mauberme/document-to-images").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "documents": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
  ]
}' |
apify call mauberme/document-to-images --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mauberme/document-to-images"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aBs2kaEocThegnyx5/builds/9CPoTDEqyeawOmdqQ/openapi.json
