Video Forensic Viewer — Frames & Contact Sheets avatar

Video Forensic Viewer — Frames & Contact Sheets

Pricing

Pay per usage

Go to Apify Store
Video Forensic Viewer — Frames & Contact Sheets

Video Forensic Viewer — Frames & Contact Sheets

Turn any video into timestamped frames, contact sheets and a JSON manifest in one fast ffmpeg pass. Built for AI agents that can see images but not video.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Andrew Babo

Andrew Babo

Maintained by Community

Actor stats

0

Bookmarked

7

Total users

1

Monthly active users

10 days ago

Last modified

Share

Video Forensic Viewer — Timestamped Frames & Contact Sheets

Turn any video into timestamped frames, contact sheets and a JSON manifest — in one ffmpeg pass, in seconds.

Point it at a video URL (or a file you uploaded to a key-value store), choose how often to sample (1 s, 0.5 s, 0.1 s or anything custom) and how many frames belong on one sheet (10, 20, 30 ...). You get back:

  • Contact sheets — grids of frames, every tile stamped with its exact time (00:12.4).
  • Full-resolution frames — one timestamped JPEG per sample.
  • manifest.json — every frame with its time, which sheet and which row/column it sits in, a download URL, and how much it changed from the previous frame.
  • A dataset row per video — dimensions, codecs, frame and sheet counts, processing time, and links to everything.

Why AI agents like it

Models can look at images, not at video. This Actor is the missing step in between.

  • One contact sheet shows 20–30 moments in a single image, so an agent can scan a whole minute of video for the price of one picture.
  • When something interesting shows up in a tile, the manifest maps that tile straight back to its timestamp and to the sharp, full-size frame — the agent can zoom in on exactly the right second.
  • Each frame carries a change score, so an agent can skip near-identical frames and spend its budget only where the video actually changes.

Speed

Everything happens in one decode pass: sampling, timestamping, tiling and change detection all run inside a single ffmpeg filter graph, so no frame is written to disk and read back.

Rough guide for the frame pipeline:

VideoIntervalTypical time
5 min1 s (300 frames)~5 s
5 min0.1 s (3 000 frames)~7 s
10 min1 s (600 frames)~10 s

Transcript (optional, runs on the machine)

Set transcribe: true and the same ffmpeg pass also writes a 16 kHz mono WAV, which is transcribed locally with whisper.cpp — no API key, nothing leaves the run. Audio is split on silence and transcribed in parallel processes, then stitched back onto the original timeline.

  • 99 languages, auto-detected (or pin one with transcriptionLanguage).
  • Word-level timestamps via DTW alignment (wordTimestamps, on by default).
  • Outputs: transcript.json (segments + words + confidence), .srt, .vtt, .txt, plus a speech block inside the manifest so text and frames share one timeline.
  • Silence is skipped by Silero VAD, so quiet footage costs almost nothing.

Frames and transcript run at the same time, and the audio is split into one chunk per available core (Apify gives 1 core per 4 GB). Measured end to end on a real 3:16 recorded speech, speedMode: fast:

MemoryWhole run (frames + sheets + transcript)Cost
8 GB51 s$0.029
16 GB (recommended)42 s (~4.7x faster than real time)$0.059
32 GB50 s$0.076

speedMode picks the model: fast = base, balanced = small, accurate = large-v3-turbo. Use balanced or accurate for accented or non-English audio where wording must be exact. Every result row reports costUsd and costPerVideoMinuteUsd so you can see the price per video.

Contact sheets are for looking, not for reading text

Tiles are downscaled, so small on-screen text breaks up. Use a sheet to find the moment, then open the full-resolution frame listed in the manifest for that timestamp. This Actor does not run OCR or summarisation — it gives you clean, addressable frames, an aligned transcript, and lets you decide what to run on them.

Safety caps

maxFrames and maxDurationSecs keep 0.1 s runs from exploding. When a cap is hit, the analysis window is trimmed, truncated is set to true and notes explains what was cut — the run still succeeds.

Output example (manifest excerpt)

{
"settings": { "intervalSecs": 1, "framesPerSheet": 20, "grid": { "cols": 5, "rows": 4 } },
"sheets": [
{ "sheetIndex": 0, "url": "https://api.apify.com/v2/key-value-stores/.../01-clip-sheet-001.jpg", "startTimeSecs": 0, "endTimeSecs": 19 }
],
"frames": [
{
"index": 12,
"timeSecs": 12.0,
"timecode": "00:12.0",
"sheetIndex": 0,
"row": 2,
"col": 2,
"frameUrl": "https://api.apify.com/v2/key-value-stores/.../01-clip-frame-000013.jpg",
"changeScore": 7.41,
"significantChange": true
}
]
}

For AI agents (MCP-ready)

This actor is built to be called by AI agents. It works out of the box with the Apify MCP Server — add it to Claude Desktop, Cursor or any MCP client: the agent can inspect video frames on its own.

{
"mcpServers": {
"video-forensic-viewer": {
"url": "https://mcp.apify.com/?actors=andrew_babo/video-forensic-viewer",
"headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }
}
}
}

Agent skill (paste into your agent's instructions)

Use the "video-forensic-viewer" tool to look INSIDE a video: extract key
frames, build contact sheets and inspect scene changes and technical
metadata. Use it when the user asks what happens in a video, wants
thumbnails/storyboards or needs visual evidence. It does not transcribe
speech — pair it with a speech-to-text actor for that.
HOW TO CALL
- { "videoUrls": ["https://.../clip.mp4"] }
- Limit the analysed duration / frame count when the video is long — that
is the cost driver.
- Contact sheets are the fastest way to give a human or a vision model an
overview of the whole video.
OUTPUT CONTRACT
- One row per video: frame image URLs, contact sheet URL, scene timestamps,
duration, resolution, codec.
- Describe only what the returned frames actually show; never infer events
between frames as fact.
- noResults: true with errorCode NO_INPUT means no video was supplied;
a per-row errorCode means that file could not be downloaded or decoded.