Video Forensic Viewer — Frames & Contact Sheets
Pricing
Pay per usage
Video Forensic Viewer — Frames & Contact Sheets
Turn any video into timestamped frames, contact sheets and a JSON manifest in one fast ffmpeg pass. Built for AI agents that can see images but not video.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Andrew Babo
Maintained by CommunityActor stats
0
Bookmarked
7
Total users
1
Monthly active users
10 days ago
Last modified
Categories
Share
Video Forensic Viewer — Timestamped Frames & Contact Sheets
Turn any video into timestamped frames, contact sheets and a JSON manifest — in one ffmpeg pass, in seconds.
Point it at a video URL (or a file you uploaded to a key-value store), choose how often to sample (1 s, 0.5 s, 0.1 s or anything custom) and how many frames belong on one sheet (10, 20, 30 ...). You get back:
- Contact sheets — grids of frames, every tile stamped with its exact time (
00:12.4). - Full-resolution frames — one timestamped JPEG per sample.
manifest.json— every frame with its time, which sheet and which row/column it sits in, a download URL, and how much it changed from the previous frame.- A dataset row per video — dimensions, codecs, frame and sheet counts, processing time, and links to everything.
Why AI agents like it
Models can look at images, not at video. This Actor is the missing step in between.
- One contact sheet shows 20–30 moments in a single image, so an agent can scan a whole minute of video for the price of one picture.
- When something interesting shows up in a tile, the manifest maps that tile straight back to its timestamp and to the sharp, full-size frame — the agent can zoom in on exactly the right second.
- Each frame carries a change score, so an agent can skip near-identical frames and spend its budget only where the video actually changes.
Speed
Everything happens in one decode pass: sampling, timestamping, tiling and change detection all run inside a single ffmpeg filter graph, so no frame is written to disk and read back.
Rough guide for the frame pipeline:
| Video | Interval | Typical time |
|---|---|---|
| 5 min | 1 s (300 frames) | ~5 s |
| 5 min | 0.1 s (3 000 frames) | ~7 s |
| 10 min | 1 s (600 frames) | ~10 s |
Transcript (optional, runs on the machine)
Set transcribe: true and the same ffmpeg pass also writes a 16 kHz mono WAV, which is transcribed locally with whisper.cpp — no API key, nothing leaves the run. Audio is split on silence and transcribed in parallel processes, then stitched back onto the original timeline.
- 99 languages, auto-detected (or pin one with
transcriptionLanguage). - Word-level timestamps via DTW alignment (
wordTimestamps, on by default). - Outputs:
transcript.json(segments + words + confidence),.srt,.vtt,.txt, plus aspeechblock inside the manifest so text and frames share one timeline. - Silence is skipped by Silero VAD, so quiet footage costs almost nothing.
Frames and transcript run at the same time, and the audio is split into one chunk per available core (Apify gives 1 core per 4 GB). Measured end to end on a real 3:16 recorded speech, speedMode: fast:
| Memory | Whole run (frames + sheets + transcript) | Cost |
|---|---|---|
| 8 GB | 51 s | $0.029 |
| 16 GB (recommended) | 42 s (~4.7x faster than real time) | $0.059 |
| 32 GB | 50 s | $0.076 |
speedMode picks the model: fast = base, balanced = small, accurate = large-v3-turbo. Use balanced or accurate for accented or non-English audio where wording must be exact. Every result row reports costUsd and costPerVideoMinuteUsd so you can see the price per video.
Contact sheets are for looking, not for reading text
Tiles are downscaled, so small on-screen text breaks up. Use a sheet to find the moment, then open the full-resolution frame listed in the manifest for that timestamp. This Actor does not run OCR or summarisation — it gives you clean, addressable frames, an aligned transcript, and lets you decide what to run on them.
Safety caps
maxFrames and maxDurationSecs keep 0.1 s runs from exploding. When a cap is hit, the analysis window is trimmed, truncated is set to true and notes explains what was cut — the run still succeeds.
Output example (manifest excerpt)
{"settings": { "intervalSecs": 1, "framesPerSheet": 20, "grid": { "cols": 5, "rows": 4 } },"sheets": [{ "sheetIndex": 0, "url": "https://api.apify.com/v2/key-value-stores/.../01-clip-sheet-001.jpg", "startTimeSecs": 0, "endTimeSecs": 19 }],"frames": [{"index": 12,"timeSecs": 12.0,"timecode": "00:12.0","sheetIndex": 0,"row": 2,"col": 2,"frameUrl": "https://api.apify.com/v2/key-value-stores/.../01-clip-frame-000013.jpg","changeScore": 7.41,"significantChange": true}]}
For AI agents (MCP-ready)
This actor is built to be called by AI agents. It works out of the box with the Apify MCP Server — add it to Claude Desktop, Cursor or any MCP client: the agent can inspect video frames on its own.
{"mcpServers": {"video-forensic-viewer": {"url": "https://mcp.apify.com/?actors=andrew_babo/video-forensic-viewer","headers": { "Authorization": "Bearer <YOUR_APIFY_TOKEN>" }}}}
Agent skill (paste into your agent's instructions)
Use the "video-forensic-viewer" tool to look INSIDE a video: extract keyframes, build contact sheets and inspect scene changes and technicalmetadata. Use it when the user asks what happens in a video, wantsthumbnails/storyboards or needs visual evidence. It does not transcribespeech — pair it with a speech-to-text actor for that.HOW TO CALL- { "videoUrls": ["https://.../clip.mp4"] }- Limit the analysed duration / frame count when the video is long — thatis the cost driver.- Contact sheets are the fastest way to give a human or a vision model anoverview of the whole video.OUTPUT CONTRACT- One row per video: frame image URLs, contact sheet URL, scene timestamps,duration, resolution, codec.- Describe only what the returned frames actually show; never infer eventsbetween frames as fact.- noResults: true with errorCode NO_INPUT means no video was supplied;a per-row errorCode means that file could not be downloaded or decoded.