# Dataset to XLSX with Images (`bgmowl/dataset-to-excel-with-images-beta`) Actor

Compatible with Microsoft Excel. Turn inline JSON or CSV into an XLSX with embedded thumbnails and an image report. Up to 200 rows; public HTTPS PNG, JPEG and WEBP images. Free beta: $0 developer fee; Apify usage applies. Independent BGMOWL tool, not affiliated with Microsoft.

- **URL**: https://apify.com/bgmowl/dataset-to-excel-with-images-beta.md
- **Developed by:** [bgm owl](https://apify.com/bgmowl) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Dataset to XLSX with Images

Turn a small product catalog, listing export, or research shortlist into an XLSX workbook you can actually browse visually. Paste JSON records or CSV, choose the image URL field, and download an `.xlsx` with real embedded thumbnails beside your data.

**Compatible with Microsoft Excel** through the standard XLSX file format. This is an independent tool and is not affiliated with or endorsed by Microsoft. Native Microsoft Excel rendering, sorting, and filtering have not yet been verified for this beta; see the compatibility notes below.

**Free beta:** there is no Actor fee. Apify platform usage, including compute, storage, and data transfer, may still consume credits or incur charges under your plan. This project does not make a run cost-free.

[![Dataset to XLSX with Images demo: turn inline records into a visual workbook](https://api.apify.com/v2/key-value-stores/FGL15twewurY34tAd/records/bgmowl-dataset-to-xlsx.gif?signature=1bgwqOfZ3EZhLzc4eECv6)](https://api.apify.com/v2/key-value-stores/FGL15twewurY34tAd/records/bgmowl-dataset-to-xlsx.gif?signature=1bgwqOfZ3EZhLzc4eECv6)

[Open animated demo](https://api.apify.com/v2/key-value-stores/FGL15twewurY34tAd/records/bgmowl-dataset-to-xlsx.gif?signature=1bgwqOfZ3EZhLzc4eECv6) · The preview on this page is static; the link opens the looping GIF.

*Illustrative layout using real workbook values and images. Up to 200 rows; public HTTPS image URLs.*

### What you get

- One Excel workbook containing every accepted input row, even when an image is missing or fails.
- PNG thumbnails embedded in the workbook. Images do not depend on a live `IMAGE()` formula or an external image fetch when the workbook opens.
- All source data cell values written as literal text, including values beginning with `=`, `+`, `-`, or `@`.
- A detailed image report and simple run totals for automation.

For your own data, this beta accepts **inline data only**. It does not read Apify dataset IDs, download CSV files, connect to databases, or accept API tokens. If your data came from another Actor, pass the selected records into this Actor's `records` input. A separate demo mode uses fictional bundled data.

### Quick start

#### Try the fictional demo

The Console input form initially checks **Try the fictional product demo**. Leave the source fields absent and run it to create a clearly labeled fictional 12-product catalog with 10 bundled product images, one missing-image notice, and one rejected-image notice. It makes no requests to image hosts. Apify platform usage can still apply.

For API or CLI use, explicitly provide:

```json
{ "demoMode": true }
```

Only `demoMode` and an optional `title` are accepted in demo mode. Do not include `records`, `csvText`, `imageColumn`, or `columns`, even as empty fields. An empty input object does not select the demo: `demoMode` defaults to false outside the prefilled Console form.

#### Export your own data

1. Turn demo mode off, or omit `demoMode` in an API input.
2. Provide either `records` or `csvText`. Omit the other field entirely.
3. Set `imageColumn` to the exact source field containing image URLs.
4. Optionally choose the order of data columns and a workbook title.
5. Run the Actor. In the run's **Output** selector, choose **Download files (Excel and report)**, then download **OUTPUT.xlsx**. If that view is unavailable, open **Storage → Key-value store** and download **OUTPUT.xlsx** there.
6. Check the image report before sharing the workbook. A successful run can still have rejected, missing, failed, or skipped images.

#### JSON example

The image URL below is a placeholder. Replace it with a direct, public HTTPS PNG, JPEG, or WEBP URL you are permitted to download.

```json
{
  "records": [
    {
      "sku": "00123",
      "name": "Red mug",
      "price": "12.50",
      "imageUrl": "https://example.com/images/red-mug.png"
    },
    {
      "sku": "00124",
      "name": "Item with no image",
      "price": "9.00",
      "imageUrl": ""
    }
  ],
  "imageColumn": "imageUrl",
  "columns": ["sku", "name", "price", "imageUrl"],
  "title": "Catalog preview"
}
```

#### CSV example

CSV is passed as a JSON string, with newline characters escaped. Use comma-separated CSV with a header row. Keep the `records` field absent.

```json
{
  "csvText": "sku,name,imageUrl\n00123,Red mug,https://example.com/images/red-mug.png\n00124,Item with no image,\n",
  "imageColumn": "imageUrl"
}
```

### Inputs

| Field | Required | Meaning |
| --- | --- | --- |
| `demoMode` | No | Explicit `true` runs the fictional bundled demo. Defaults to `false`; the Console form prefills `true`. |
| `records` | One source in real-data mode | List of 1–200 JSON objects. Omit in demo mode. |
| `csvText` | One source in real-data mode | Inline comma-separated CSV with headers and 1–200 data rows. Omit in demo mode. |
| `imageColumn` | In real-data mode | Exact, case-sensitive existing field name used to find each image URL. Omit in demo mode. |
| `columns` | No | Ordered list of 1–30 distinct source field names to export. If omitted, columns are derived in first-seen order. Omit in demo mode. |
| `title` | No | Non-blank workbook title, 1–100 characters. Defaults to `Dataset to Excel with Images`. |

The Actor enforces conditional demo/real-data requirements, source exclusivity, and the total input and column limits at runtime. Unknown top-level input fields are rejected. Field names must be non-blank and no longer than 128 characters. CSV headers must be unique, and each record must have the same field count as the header; empty physical lines are ignored. Selecting a subset of columns does not raise the 30-column source limit.

#### Data fidelity

Data values are deliberately text, so Excel will not automatically turn identifiers into numbers, dates, formulas, or active URL cells. CSV leading zeros are retained. JSON integers are preserved as received by the Python runtime, including integers larger than Excel's numeric precision.

Missing fields and JSON `null` become empty cells. Booleans become `true` or `false`; nested arrays and objects become compact JSON text. Cells longer than 32,767 characters, non-finite numbers, and unsupported XML control or Unicode characters are rejected rather than silently truncated or altered.

Use JSON strings for identifiers and exact decimal formatting. A browser, upstream JavaScript integration, or other producer may round a large JSON number before it reaches this Actor; the Actor cannot restore digits already lost. Numbers do not retain their original JSON spelling, so use `"12.50"` rather than `12.50` when the trailing zero matters. Numeric calculations in Excel require an explicit conversion after export.

### Files and results

The run's default key-value store contains:

| Key | Content |
| --- | --- |
| `OUTPUT.xlsx` | Binary Excel workbook, MIME type `application/vnd.openxmlformats-officedocument.spreadsheetml.sheet`. |
| `SUMMARY.json` | JSON report with totals and per-row image outcomes. |
| `OUTPUT` | JSON output metadata. |

The default dataset contains **one summary record**, not the original rows. Exporting that dataset to Excel will produce only the summary; download `OUTPUT.xlsx` for the embedded-image workbook.

```json
{
  "version": "0.1.0",
  "demo": false,
  "rowCount": 5,
  "columnCount": 3,
  "embeddedCount": 1,
  "missingCount": 1,
  "rejectedCount": 1,
  "failedCount": 1,
  "skippedCount": 1,
  "workbookKey": "OUTPUT.xlsx",
  "summaryKey": "SUMMARY.json"
}
```

Each row has one image outcome:

- `embedded`: a thumbnail was added.
- `missing`: the source image field is missing or empty.
- `rejected`: the source or image violates a validation or safety limit.
- `failed`: the image could not be downloaded or processed.
- `skipped`: the shared image budget or fetch deadline prevented processing.

Image errors do not remove the row. Invalid input, such as conflicting sources or too many rows, fails the run instead of silently truncating the data.

The workbook has a `Dataset` sheet with a thumbnail, your chosen data columns, `Image status`, and `Image detail`. The detailed JSON report's `rows` array uses a one-based `sourceRow` index into the data records, excluding the CSV header. Report rows contain safe status/detail codes rather than source URLs. The `demo` flag identifies fictional demo output; demo counts are 12 rows, 10 embedded, one missing, one rejected, and no failed or skipped images.

### Beta limits

| Limit | Value |
| --- | --- |
| Data rows | 1–200 |
| Source data columns | 30 maximum |
| Total JSON input | 2 MiB maximum |
| Downloaded image | 2 MiB maximum per image |
| Total image download budget | 20 MiB per run |
| Decoded image | At most 8,000,000 pixels and 4,096 pixels on either side |
| Concurrent image workers | 4 |
| Image-fetch phase | 35 seconds |
| Image formats | Static PNG, JPEG, WEBP; animated images are rejected |
| Image transport | Direct public HTTPS URLs on port 443, at most 4,096 characters, without fragments |

One MiB means 1,048,576 bytes. The image limits are safeguards, not a guarantee that every permitted image will load. Slow, protected, expired, malformed, or unavailable sources may fail. Image servers must return a matching `image/png`, `image/jpeg`, or `image/webp` content type. Remaining images can be skipped when the shared budget or time cap is reached. The 35-second cap applies to fetching, not the complete Actor run, which also needs startup, workbook creation, and storage time.

### Image safety and privacy

- Use only data and image URLs you are allowed to share with Apify and download from their hosts.
- Download requests reveal the requester's IP address and the complete requested URL to the image host. On Apify, this is the Actor's outbound address.
- Do not submit passwords, access tokens, private files, sensitive personal data, or secret image URLs. Query strings, including signed URL tokens, are still disclosed to the host and can remain in the input or exported source data.
- URL usernames/passwords, redirects, private or other non-public IP destinations are forbidden. Authentication, custom request headers, and cookies are not supported.
- This is an export tool, not a privacy scrubber. The workbook can contain your original fields and image URLs. Review it before sharing.

### Spreadsheet compatibility

Excel desktop is recommended. Thumbnails are embedded drawing objects anchored to rows, not native in-cell images. Sorting or rearranging rows is not guaranteed to keep image anchors aligned; verify the result after editing. Google Sheets imports, Excel web, and other spreadsheet applications may display or move images differently.

The beta is best for small, reviewable exports. It is not a full-fidelity backup of the source JSON, a streaming export, or a promise that protected image hosts can be accessed.

### Development

Runtime: Python 3.12. The Docker entry point is `python -m src`. Direct runtime dependencies are pinned in `requirements.txt`: Apify SDK 4.0.2, XlsxWriter 3.2.9, and Pillow 12.3.0.

Install dependencies in a virtual environment with `python -m pip install -r requirements.txt`. To run directly without an Apify account, save a UTF-8 JSON input file and run:

```sh
python -m src --local INPUT.json local-output
```

Or run the bundled fictional demo without image-host requests:

```sh
python -m src --local demo/try-demo.json demo-output
```

The local CLI writes `OUTPUT.xlsx` and `SUMMARY.json` to the chosen directory and prints the run summary. Real-data mode makes outbound requests for valid image URLs. The Apify runtime additionally writes `OUTPUT` metadata and a summary dataset record.

For the Apify local workflow, save your input in `storage/key_value_stores/default/INPUT.json` and use `apify run` from this project directory. Install `requirements-qa.txt` for development, then run the offline tests with `python -m unittest discover -s tests -v`. `python -m tools.make_demo` regenerates the original fictional illustrations and sample workbook; `python -m tools.render_preview` renders file-derived previews. The previews are not screenshots from Microsoft Excel. Keep test inputs free of secrets. Do not publish or run a cloud build without first reviewing platform usage and settings.

Recommended beta run settings: 512 MiB memory and a 60-second timeout. Image processing stops at its own 35-second deadline; run startup and output storage need additional time. See `QA_REVIEW.md` for verification and the remaining platform checks.

Metadata follows Apify's current Actor schema conventions. The input schema deliberately has no top-level `oneOf`, which the Apify input meta-schema does not allow; runtime validation enforces demo exclusivity or exactly one real-data source plus `imageColumn`. `demoMode` combines `prefill: true` with `default: false`, so an API call that omits it never silently switches to fictional data.

#### Reference sources

- [Actor definition](https://docs.apify.com/actors/development/actor-definition/actor-json)
- [Input schema specification](https://docs.apify.com/actors/development/actor-definition/input-schema/specification/v1)
- [Output schema specification](https://docs.apify.com/actors/development/actor-definition/output-schema)
- [Dataset schema specification](https://docs.apify.com/storage/dataset-schema)
- [Official schema source](https://github.com/apify/apify-shared-js/tree/master/packages/json_schemas/schemas)
- [Python Docker images](https://docs.apify.com/actors/development/actor-definition/dockerfile)
- [Apify SDK 4.0.2](https://pypi.org/project/apify/4.0.2/) and [Actor API](https://docs.apify.com/sdk/python/reference/class/Actor)
- [XlsxWriter 3.2.9](https://pypi.org/project/xlsxwriter/3.2.9/) and [Pillow 12.3.0](https://pypi.org/project/pillow/12.3.0/)

Schema and dependency references were checked on 2026-10-08. The beta's contract and tests, rather than later upstream releases, govern this version.

# Actor input Schema

## `demoMode` (type: `boolean`):

When true, generate a clearly labeled fictional 12-product catalog using bundled assets: 10 embedded images, 1 missing-image notice, and 1 rejected-image notice. No image-host requests are made. Only demoMode and optional title may be supplied. To use your own data, turn this off, provide records OR csvText, and set imageColumn. The Console initially prefills true; API and CLI callers must explicitly send true to run the demo.

## `records` (type: `array`):

Your data: turn demoMode off, then supply this OR csvText, never both. Omit this field in demo mode. Accepts 1–200 JSON objects; field names are column names. Use strings for long identifiers to prevent upstream rounding. Do not include secrets or data you are not allowed to process.

## `csvText` (type: `string`):

Your data: turn demoMode off, then supply this OR records, never both. Omit this field in demo mode. Inline comma-separated CSV with headers and 1–200 data rows; leading zeros stay text. Uploads, file URLs, dataset IDs, and API tokens are not supported in this beta.

## `imageColumn` (type: `string`):

Required for your own data; omit in demo mode. Exact, case-sensitive existing source field containing image URLs, up to 128 characters. Missing images retain the row and are reported. URLs must use public HTTPS; redirects, URL credentials, and non-public IP destinations are forbidden. Avoid secret URLs: requests disclose the full URL and requester IP to the image host.

## `columns` (type: `array`):

Optional for your own data; omit in demo mode. Ordered list of 1–30 distinct source field names to include. When omitted, columns follow first-seen input order. The imageColumn still determines the thumbnail source.

## `title` (type: `string`):

Non-blank human-readable title containing 1–100 characters. The default is Dataset to Excel with Images.

## Actor input object example

```json
{
  "demoMode": true,
  "title": "Dataset to Excel with Images"
}
```

# Actor output Schema

## `results` (type: `string`):

A single dataset record with row and image counts plus the workbook and summary storage keys.

## `files` (type: `string`):

Native file listing for the run. Download OUTPUT.xlsx for embedded images and SUMMARY.json for the per-row image report.

## `workbook` (type: `string`):

Direct XLSX download with embedded thumbnails. Use the Download files output for Console download controls; the workbook is a binary file, not an inline browser preview.

## `summary` (type: `string`):

Run totals and per-row image outcomes, including missing, rejected, failed, and skipped images.

## `output` (type: `string`):

Machine-readable output metadata stored under the OUTPUT key.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "demoMode": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("bgmowl/dataset-to-excel-with-images-beta").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "demoMode": True }

# Run the Actor and wait for it to finish
run = client.actor("bgmowl/dataset-to-excel-with-images-beta").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "demoMode": true
}' |
apify call bgmowl/dataset-to-excel-with-images-beta --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bgmowl/dataset-to-excel-with-images-beta"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/0LC4AAdN2Ybv8frU9/builds/p2vESg5gW0mbcNmwc/openapi.json
