Video Scene Detection, Keyframe Extractor: extract scene frames avatar

Video Scene Detection, Keyframe Extractor: extract scene frames

Pricing

from $20.00 / 1,000 minute of video processeds

Go to Apify Store
Video Scene Detection, Keyframe Extractor: extract scene frames

Video Scene Detection, Keyframe Extractor: extract scene frames

Extract frames from a video file, one per scene, with scene change detection: scene timestamps, JPEG keyframes (40 by default, 100 max) and an ffprobe summary as JSON. No ffmpeg to install, no extra API key. $0.01 a run plus $0.02 a video minute; failed videos are not charged.

Pricing

from $20.00 / 1,000 minute of video processeds

Rating

0.0

(0)

Developer

FrameProbe

FrameProbe

Maintained by Community

Actor stats

0

Bookmarked

40

Total users

25

Monthly active users

8 days ago

Last modified

Share

Give it a video URL. Get back every scene change in that video, a JPEG still for each of up to 40 of those scenes, and a full ffprobe report, as JSON.

Actor idframeprobe/video-scene-splitter
Minimal input{"videoUrls": ["https://api.apify.com/v2/key-value-stores/Bfr82R8zp35dJRkcL/records/autotest-clip.mp4"]}
Cost$0.01 per run, plus $0.02 per minute of video, rounded up, minimum one minute per video. 20 short clips cost $0.41 (raise maxVideos past its default of 10 first). One 10-minute talk costs $0.21. A video that fails to download or decode is not charged.
OutputOne row per video: scenes[] with startSeconds, endSeconds, durationSeconds and keyframeKey, plus sceneCount, cutCadenceSeconds and a ten-field ffprobe summary (durationSeconds, width, height, fps, videoCodec, audioCodec, hasAudio, bitrateKbps, sizeBytes, containerFormat). Every boundary is returned; the stills are capped at maxKeyframes, 40 by default and 100 at most, and a video with more scenes than that gets stills sampled evenly across it. The JPEGs are saved in the run's key-value store.

Click Try for free. The input form arrives with a short demo clip already filled in, so the first run needs no typing and comes back with real scene timestamps and real stills.

  • Split a video into scenes, and get the cut timestamps.
  • Extract keyframes from a video for a contact sheet, a thumbnail grid, or a vision model.
  • Read a file's duration, resolution, frame rate and codecs without installing ffmpeg anywhere.

Extract frames from a video file, one per scene: cut timestamps, JPEG stills (40 by default, 100 at most) and an ffprobe summary, all as structured output your model or pipeline can read.

No model and no extra API key of your own. Deterministic ffmpeg under the hood: the same video gives the same result every time.

Pass direct video URLs, or chain it after TikTok Scraper, Instagram Scraper, or YouTube Scraper with a dataset ID. Pairs with FrameProbe's Reel Teardown when you want an AI read on top.

What you get, per video

  • Scene boundaries. Every cut, with start, end and duration. Contiguous, covering the whole video, no gaps.
  • One keyframe per scene, up to maxKeyframes (40 by default, 100 at most), saved to the key-value store. A video with more scenes than that does not get a still for every scene: the stills are sampled at evenly spaced scenes so coverage stays spread across the whole video, and every boundary is still returned in scenes[] either way. Each still is taken from the middle of its scene, not the first frame, so it is not a dissolve or a transition. If the middle happens to be a black frame, it samples elsewhere in the scene instead of handing your model an empty picture. The key is stored inside each scene object, so you never join two lists by position.
  • An ffprobe summary. Duration, resolution, frame rate, video and audio codec, audio presence, bitrate, file size, container. Ten fields, not ffprobe's full output.

Quickstart

The recommended path is chaining. Social platforms block video fetching from cloud servers, so run a scraper first and hand this Actor its dataset.

  1. Run any TikTok, Instagram or YouTube scraper.
  2. Open its run, copy the dataset ID from the Storage tab.
  3. Start this Actor with that ID in Dataset from a scraper run.
  4. Read scenes[] on each output record; fetch the stills by keyframeKey.

Already have direct media URLs? Put them in Direct video URLs instead and skip steps 1 and 2.

{
"datasetId": "aBcD1234efGh5678",
"maxVideos": 10,
"sceneThreshold": 0.35,
"minSceneSeconds": 0.6,
"maxKeyframes": 40
}

Example output

{
"url": "https://cdn.example.com/video/7234567890.mp4",
"platform": "tiktok",
"videoId": "9f2a1c7d4e6b8a03",
"status": "ok",
"durationSeconds": 52.209,
"width": 1080,
"height": 1920,
"fps": 29.97,
"videoCodec": "h264",
"audioCodec": "aac",
"hasAudio": true,
"bitrateKbps": 2412.6,
"sizeBytes": 15728640,
"containerFormat": "mov,mp4,m4a,3gp,3g2,mj2",
"sceneCount": 14,
"cutCadenceSeconds": 2.4,
"keyframeCount": 14,
"keyframesTruncated": false,
"scenes": [
{"index": 0, "startSeconds": 0.0, "endSeconds": 3.44, "durationSeconds": 3.44,
"keyframeKey": "scene-9f2a1c7d4e6b8a03-0000.jpg"},
{"index": 1, "startSeconds": 3.44, "endSeconds": 5.02, "durationSeconds": 1.58,
"keyframeKey": "scene-9f2a1c7d4e6b8a03-0001.jpg"}
],
"processingSeconds": 8.31,
"processedAt": "2026-08-12T22:49:05+00:00"
}

Input

FieldTypeDefaultWhat it does
videoUrlsarraydemo clipDirect media URLs (.mp4 and similar)
datasetIdstringnoneRead URLs from a scraper's dataset. The recommended input
maxVideosinteger10 (max 100)Ceiling on videos attempted this run
sceneThresholdnumber0.35 (0.10-0.90)How different two frames must be to count as a cut. Lower finds more scenes
minSceneSecondsnumber0.6 (0-10)Cuts closer than this merge into the scene before them
maxKeyframesinteger40 (max 100)Cap on saved stills. Every boundary is still returned
keyframeLongEdgeinteger1024 (240-1920)Longest edge of each still, aspect preserved
includeKeyframesbooleantrueOff gives boundaries and the probe report only. Faster and cheaper

Cost

EventPrice
Actor start$0.01 per run
Minute of video processed$0.02 per minute, rounded up, minimum one per video

Price scales with the work because processing cost does. A 40-second clip is one minute-unit; a 3 minute 10 second video is four.

What that means in practice: 20 short clips cost $0.41. A single 10-minute talk costs $0.21.

Not charged: videos that fail to download, resolve or decode. They still appear in the dataset with a specific reason. A video whose stills were cut short by the time budget is charged; its scene list and probe report are complete and the record flags keyframesTruncated.

Limits, stated plainly

  • Social page links do not work from here. TikTok, Instagram and YouTube block video fetching from datacenter IPs. Paste one and you get a message naming the scraper to chain instead. This is a platform restriction, not a bug in this Actor.

  • Instagram and TikTok media URLs expire. They are short-lived signed links, usually good for hours. Fetch them in the same session you scraped them; a stored URL will 403.

  • Maximum 600 seconds, 200 MB, and 2160 px on either side, per video. Longer, larger or higher resolution fails with a specific message. The pixel cap means 4K (3840 px wide) is refused; up to 2160 px on the long edge is accepted.

  • Links to private or internal addresses are refused, with a gap. A link that is not https, or whose host is or looks up to localhost, a private network, a cloud metadata address or 100.64.x.x, is refused before anything connects to it. Two cases get past that check. A host can give a public address when we check it and a private one when we connect (DNS rebinding). And for page links, the page-to-video lookup can follow a redirect, or fetch a link found in the page, without the check seeing it. Downloads re-check every redirect. The exposure is highest when you feed in a dataset scraped from pages you do not control, because whoever wrote the page wrote those links.

  • Scene detection compares brightness, not colour. Two shots of different colours but similar brightness read as one shot: a cut from navy to dark red is invisible at the default sceneThreshold of 0.35, however obvious it looks to you. sceneThreshold is the lever, and it works: measured on a navy/dark-red clip, 0.35 finds neither cut and 0.2 finds both. Lower it and the detector also fires more readily on overlays and kinetic text in real footage, so expect extra boundaries alongside the missing one. This is a property of the metric, not a defect, and it is pinned by a permanent test so that if it is ever fixed you hear it from the changelog.

  • Cut cadence is a measured proxy. The detector fires on overlays and kinetic text as well as hard cuts, so read it as editing rhythm rather than an exact shot count.

  • A resumed run does not retry failures. If the platform migrates your run between servers, it picks up where it left off and skips everything already in the dataset, including the videos that failed. That is deliberate: it is what stops a migration from charging you twice. To retry a failed video, start a new run with it.

  • Each video gets a processing budget set from its own length and resolution, and a very long, very high-resolution video can exceed it. The budget pays for scene detection and the stills; the download is not counted against it. Scene detection is the expensive part and its cost tracks pixels, so it is roughly ten seconds of processing per minute of 1080p video when yours is the only video in the run, and about twenty when two are being processed at once, which is the normal case for a run with several videos. A square 2160 px frame is a little over twice 1080p, not four times. A video that runs out fails with a message saying so and is not charged.

  • A video can come back with scene timestamps but no stills. If the budget runs out before any still is extracted, the video fails and is not charged, rather than returning a record with an empty image list. If you only want the timestamps, set includeKeyframes to false.

    Before 2026-09-19 this was a flat 25 seconds for every video, which was not enough to decode anything past about three minutes at 1080p, and those videos failed. If you sent a long video and got a timeout, try it again.

Changelog

  • 0.1 First build. Scene boundaries, one keyframe per scene, full ffprobe report. Per-minute pricing. Chains from any scraper's dataset.

Built by FrameProbe. Also see Reel Teardown, which adds a vision read on top of the same pipeline: hooks, structure, on-screen text and why a video worked.