# Video Forensic Viewer — Frames & Contact Sheets (`andrew_babo/video-forensic-viewer`) Actor

Turn any video into timestamped frames, contact sheets and a JSON manifest in one fast ffmpeg pass. Built for AI agents that can see images but not video.

- **URL**: https://apify.com/andrew\_babo/video-forensic-viewer.md
- **Developed by:** [Andrew Babo](https://apify.com/andrew_babo) (community)
- **Categories:** AI, Videos, Developer tools
- **Stats:** 7 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Video Forensic Viewer — Timestamped Frames & Contact Sheets

Turn any video into **timestamped frames**, **contact sheets** and a **JSON manifest** — in one ffmpeg pass, in seconds.

Point it at a video URL (or a file you uploaded to a key-value store), choose how often to sample (1 s, 0.5 s, 0.1 s or anything custom) and how many frames belong on one sheet (10, 20, 30 ...). You get back:

- **Contact sheets** — grids of frames, every tile stamped with its exact time (`00:12.4`).
- **Full-resolution frames** — one timestamped JPEG per sample.
- **`manifest.json`** — every frame with its time, which sheet and which row/column it sits in, a download URL, and how much it changed from the previous frame.
- **A dataset row per video** — dimensions, codecs, frame and sheet counts, processing time, and links to everything.

### Why AI agents like it

Models can look at images, not at video. This Actor is the missing step in between.

- One contact sheet shows 20–30 moments in a **single** image, so an agent can scan a whole minute of video for the price of one picture.
- When something interesting shows up in a tile, the manifest maps that tile straight back to its timestamp and to the sharp, full-size frame — the agent can zoom in on exactly the right second.
- Each frame carries a change score, so an agent can skip near-identical frames and spend its budget only where the video actually changes.

### Speed

Everything happens in **one decode pass**: sampling, timestamping, tiling and change detection all run inside a single ffmpeg filter graph, so no frame is written to disk and read back.

Rough guide for the frame pipeline:

| Video | Interval | Typical time |
| --- | --- | --- |
| 5 min | 1 s (300 frames) | ~5 s |
| 5 min | 0.1 s (3 000 frames) | ~7 s |
| 10 min | 1 s (600 frames) | ~10 s |

### Transcript (optional, runs on the machine)

Set `transcribe: true` and the same ffmpeg pass also writes a 16 kHz mono WAV, which is transcribed locally with whisper.cpp — no API key, nothing leaves the run. Audio is split on silence and transcribed in parallel processes, then stitched back onto the original timeline.

- 99 languages, auto-detected (or pin one with `transcriptionLanguage`).
- Word-level timestamps via DTW alignment (`wordTimestamps`, on by default).
- Outputs: `transcript.json` (segments + words + confidence), `.srt`, `.vtt`, `.txt`, plus a `speech` block inside the manifest so text and frames share one timeline.
- Silence is skipped by Silero VAD, so quiet footage costs almost nothing.

Frames and transcript run at the same time, and the audio is split into one chunk per available core (Apify gives 1 core per 4 GB). Measured end to end on a real 3:16 recorded speech, `speedMode: fast`:

| Memory | Whole run (frames + sheets + transcript) | Cost |
| --- | --- | --- |
| 8 GB | 51 s | $0.029 |
| **16 GB (recommended)** | **42 s** (~4.7x faster than real time) | **$0.059** |
| 32 GB | 50 s | $0.076 |

`speedMode` picks the model: `fast` = `base`, `balanced` = `small`, `accurate` = `large-v3-turbo`. Use `balanced` or `accurate` for accented or non-English audio where wording must be exact. Every result row reports `costUsd` and `costPerVideoMinuteUsd` so you can see the price per video.

### Contact sheets are for looking, not for reading text

Tiles are downscaled, so small on-screen text breaks up. Use a sheet to find the moment, then open the full-resolution frame listed in the manifest for that timestamp. This Actor does not run OCR or summarisation — it gives you clean, addressable frames, an aligned transcript, and lets you decide what to run on them.

### Safety caps

`maxFrames` and `maxDurationSecs` keep 0.1 s runs from exploding. When a cap is hit, the analysis window is trimmed, `truncated` is set to `true` and `notes` explains what was cut — the run still succeeds.

### Output example (manifest excerpt)

```json
{
  "settings": { "intervalSecs": 1, "framesPerSheet": 20, "grid": { "cols": 5, "rows": 4 } },
  "sheets": [
    { "sheetIndex": 0, "url": "https://api.apify.com/v2/key-value-stores/.../01-clip-sheet-001.jpg", "startTimeSecs": 0, "endTimeSecs": 19 }
  ],
  "frames": [
    {
      "index": 12,
      "timeSecs": 12.0,
      "timecode": "00:12.0",
      "sheetIndex": 0,
      "row": 2,
      "col": 2,
      "frameUrl": "https://api.apify.com/v2/key-value-stores/.../01-clip-frame-000013.jpg",
      "changeScore": 7.41,
      "significantChange": true
    }
  ]
}
```

### For AI agents (MCP-ready)

This actor is built to be called by AI agents. It works out of the box with the
[Apify MCP Server](https://mcp.apify.com) — add it to Claude Desktop, Cursor or any MCP
client: the agent can inspect video frames on its own.

```json
{
  "mcpServers": {
    "video-forensic-viewer": {
      "url": "https://mcp.apify.com/?actors=andrew_babo/video-forensic-viewer",
      "headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }
    }
  }
}
```

#### Agent skill (paste into your agent's instructions)

```text
Use the "video-forensic-viewer" tool to look INSIDE a video: extract key
frames, build contact sheets and inspect scene changes and technical
metadata. Use it when the user asks what happens in a video, wants
thumbnails/storyboards or needs visual evidence. It does not transcribe
speech — pair it with a speech-to-text actor for that.

HOW TO CALL
- { "videoUrls": ["https://.../clip.mp4"] }
- Limit the analysed duration / frame count when the video is long — that
  is the cost driver.
- Contact sheets are the fastest way to give a human or a vision model an
  overview of the whole video.

OUTPUT CONTRACT
- One row per video: frame image URLs, contact sheet URL, scene timestamps,
  duration, resolution, codec.
- Describe only what the returned frames actually show; never infer events
  between frames as fact.
- noResults: true with errorCode NO_INPUT means no video was supplied;
  a per-row errorCode means that file could not be downloaded or decoded.
```

# Actor input Schema

## `videoUrls` (type: `array`):

Direct links to video files (MP4, MOV, MKV, WebM, AVI ...). Several videos can be processed in one run.

## `keyValueStoreRecords` (type: `array`):

Videos you uploaded to an Apify key-value store, e.g. \[{"storeId":"<store id>","key":"clip.mp4"}]. Omit storeId to read from this run's own store.

## `intervalSecs` (type: `string`):

1 = one frame per second, 0.5 = two per second, 0.1 = ten per second. Any custom value between 0.02 and 60 works.

## `framesPerSheet` (type: `integer`):

How many frames are tiled into one contact sheet, e.g. 10, 20 or 30. The grid shape is chosen automatically from the video aspect ratio.

## `startTimeSecs` (type: `integer`):

Skip the beginning of the video.

## `endTimeSecs` (type: `integer`):

Leave empty to run to the end of the video.

## `maxFrames` (type: `integer`):

Safety cap, mainly for 0.1 s intervals. The window is trimmed and the reason is reported instead of failing.

## `maxDurationSecs` (type: `integer`):

Upper bound on how much of the video is analysed. 0 means no limit.

## `storeFrames` (type: `boolean`):

Save every extracted frame as its own timestamped JPEG. Turn off to get contact sheets and the manifest only.

## `frameWidth` (type: `integer`):

Width of the stored frames. Height follows the original aspect ratio.

## `cellWidth` (type: `integer`):

Width of one tile inside a contact sheet.

## `jpegQuality` (type: `integer`):

ffmpeg quality scale: 2 is best quality and largest files, 15 is smallest.

## `dedupeThreshold` (type: `integer`):

Frames whose visual difference from the previous frame is below this score are flagged as near-duplicates in the manifest. Nothing is deleted; downstream analysis can skip them.

## `concurrency` (type: `integer`):

How many videos are processed at the same time in one run.

## `maxFileSizeMb` (type: `integer`):

Videos larger than this are skipped with an error entry.

## `requestTimeoutSecs` (type: `integer`):

Timeout for downloading a single video URL.

## `speedMode` (type: `string`):

fast = base model (cheapest, quickest). balanced = small model (better for Vietnamese, accents, names). accurate = large-v3-turbo. Only used when transcription is on; an explicit transcriptionModel overrides it.

## `transcribe` (type: `boolean`):

Run on-machine speech recognition (CTranslate2 large-v3-turbo, 99 languages, no external API). Audio is decoded in CPU-optimized batches.

## `transcriptionModel` (type: `string`):

All models are bundled in the image, so nothing is downloaded at run time. Version 0.4 always uses large-v3-turbo in its optimized CTranslate2 path; older choices remain compatible with the fallback engine.

## `transcriptionLanguage` (type: `string`):

'auto' detects the language per chunk. Or give an ISO code such as en, vi, es, ja.

## `wordTimestamps` (type: `boolean`):

Give every word its own start and end time (DTW alignment). Turn off for slightly faster runs.

## `skipSilence` (type: `boolean`):

Use the built-in Silero voice activity detector so silent stretches are never sent through recognition.

## Actor input object example

```json
{
  "videoUrls": [
    "https://download.samplelib.com/mp4/sample-30s.mp4"
  ],
  "intervalSecs": "1",
  "framesPerSheet": 20,
  "startTimeSecs": 0,
  "maxFrames": 1500,
  "maxDurationSecs": 3600,
  "storeFrames": true,
  "frameWidth": 1280,
  "cellWidth": 320,
  "jpegQuality": 4,
  "dedupeThreshold": 3,
  "concurrency": 2,
  "maxFileSizeMb": 2048,
  "requestTimeoutSecs": 300,
  "speedMode": "fast",
  "transcribe": false,
  "transcriptionModel": "turbo",
  "transcriptionLanguage": "auto",
  "wordTimestamps": true,
  "skipSilence": true
}
```

# Actor output Schema

## `dataset` (type: `string`):

One row per video: dimensions, frame and sheet counts, links to every contact sheet, the manifest and the individual frames.

## `keyValueStore` (type: `string`):

Timestamped JPEG frames, contact sheet images and the JSON manifest for each video.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrls": [
        "https://download.samplelib.com/mp4/sample-30s.mp4"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("andrew_babo/video-forensic-viewer").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videoUrls": ["https://download.samplelib.com/mp4/sample-30s.mp4"] }

# Run the Actor and wait for it to finish
run = client.actor("andrew_babo/video-forensic-viewer").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrls": [
    "https://download.samplelib.com/mp4/sample-30s.mp4"
  ]
}' |
apify call andrew_babo/video-forensic-viewer --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,andrew_babo/video-forensic-viewer"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/xPBZHksVg55PM2dyc/builds/pbLIa2Sul0q1vsC3e/openapi.json
