# Video Frame Extractor for AI — No Blur, No Duplicates (`adorable_partial/video-frame-extractor-ai`) Actor

Turn videos into clean image datasets for AI: frames every N seconds, keyframes or one per scene. Skips blurry and near-duplicate frames automatically. Timestamps + sharpness manifest, JPEG/PNG/WebP, ZIP. Pay per frame kept.

- **URL**: https://apify.com/adorable\_partial/video-frame-extractor-ai.md
- **Developed by:** [Leandro Zanatta](https://apify.com/adorable_partial) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 frame extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Video Frame Extractor for AI — Clean Image Datasets (No Blur, No Duplicates)

Turn videos into **clean image datasets** for annotation, model training and computer vision pipelines. Extract frames **every N seconds**, **only keyframes** or **one per scene**, and the Actor automatically **skips blurry frames** and **near-duplicates** (moments where nothing moved), so you don't pay to store, label or train on the same picture hundreds of times.

Every frame comes with its **timestamp**, size and a **sharpness score**, ready for Label Studio, CVAT, Roboflow, a vector database or your own training code.

![25 sampled frames: 15 kept, 3 skipped as blurry, 7 skipped as duplicates; scene mode finds 3 shots](https://api.apify.com/v2/key-value-stores/yW0gCVGwwieGd7SSa/records/demo.png?signature=VDusFjmBuChVmhoG6E4O)

*Real output of this Actor on a 12-second test clip with a motion-blurred stretch and a frozen ending. Footage: Qviri and Minh Nguyen, CC BY-SA 4.0, Wikimedia Commons.*

### Why use it

- 🧹 **Cleaner datasets, less labeling**: plain "one frame per second" exports are full of blurry and identical images. Annotators waste time on them and models overfit to them. This Actor filters both out before you ever see them.
- 💸 **You only pay for useful frames**: skipped frames are free, so cleaning the dataset also lowers the bill.
- 🎯 **Keeps the motion that matters**: the duplicate filter measures how much of the picture changed, so a car crossing a static street scene is kept while the empty street is not.
- 🎞️ **Three sampling strategies**: fixed interval, keyframes (fastest) or scene change (one image per shot).
- 🧾 **Training-ready manifest**: one row per frame with image link, timestamp, width, height and sharpness. Export as JSON, CSV or Excel.
- 🧩 **No FFmpeg or OpenCV setup**: send video URLs, get image links and an optional ZIP.

### Who uses it

| Sector | Typical use |
|---|---|
| **AI / ML teams** | Build object-detection, segmentation and classification datasets from raw video |
| **Data labeling companies** | Pre-filter frames before sending them to annotators (Label Studio, CVAT, Roboflow) |
| **Security & smart cities** | Sample CCTV and traffic footage for detector training and evaluation |
| **Autonomous driving, drones, robotics** | Dashcam and drone flights into image sets for perception models |
| **Retail & manufacturing** | Shelf, conveyor and inspection footage into images for quality-control models |
| **Media & research** | Thumbnails, shot lists and visual indexes of long videos; frames for multimodal LLMs and vector search |
| **Sports & fitness** | Key moments from training videos for pose estimation and analysis |

### How it works

1. Each video is downloaded and decoded inside the run.
2. Candidate frames are picked by the **sampling mode**.
3. Each candidate gets a **sharpness score** (variance of the Laplacian). Frames below the **blur threshold** are skipped.
4. Each remaining frame is compared with the last kept one; if less than **Minimum change (%)** of the picture changed, it is skipped as a near-duplicate.
5. Kept frames are optionally resized, encoded as JPEG / PNG / WebP and stored; one dataset row is written per frame.

### Sampling modes

| Mode | What you get | Best for |
|---|---|---|
| **Every N seconds** (`interval`) | One candidate every *N* seconds (0.04 to any value; e.g. 0.5, 1, 10) | Datasets with even time coverage |
| **Keyframes only** (`keyframes`) | Only the encoder's keyframes (I-frames), usually every 2–10 s | Fast previews of long videos |
| **Scene change** (`scene`) | One frame each time the shot changes | Edited videos, ads, movies, shot lists |

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| **Video URLs** (`videoUrls`) | array | — | Direct links to videos. Public Google Drive and Dropbox share links are converted automatically. Up to 4 GB per file. |
| **Sampling mode** (`mode`) | string | `interval` | `interval`, `keyframes` or `scene` |
| **Every N seconds** (`everySeconds`) | string | `1` | Time between candidates in interval mode |
| **Scene change sensitivity** (`sceneThreshold`) | string | `0.35` | 0.2 = more frames, 0.5 = only big cuts |
| **Skip blurry frames** (`skipBlurry`) | boolean | `true` | Drop motion-blurred or out-of-focus frames |
| **Blur threshold** (`blurThreshold`) | integer | `40` | Minimum sharpness to keep a frame. Every frame's score is in the output so you can tune it |
| **Skip near-duplicate frames** (`skipDuplicates`) | boolean | `true` | Drop frames where almost nothing changed |
| **Minimum change (%)** (`minChangePercent`) | string | `0.5` | Share of the picture that must change since the last kept frame |
| **Max side (px)** (`maxSide`) | integer | original | Downscale so the longest side is at most this (e.g. 640 for YOLO-style models) |
| **Image format** (`outputFormat`) | string | `jpeg` | `jpeg`, `png` or `webp` |
| **Quality** (`quality`) | integer | `90` | JPEG / WebP quality (50–100) |
| **Create ZIP** (`createZip`) | boolean | `false` | Also store one ZIP with all frames of each video |
| **Max frames per video** (`maxFramesPerVideo`) | integer | `200` | Hard limit for cost control |
| **Max minutes per video** (`maxDurationMinutes`) | integer | `60` | Only the first N minutes are scanned and billed |

Example:

```json
{
  "videoUrls": ["https://example.com/camera-01.mp4", "https://example.com/drone-flight.mov"],
  "mode": "interval",
  "everySeconds": "2",
  "maxSide": 640,
  "createZip": true,
  "maxFramesPerVideo": 500
}
```

**Supported inputs:** MP4, MOV, WebM, MKV, AVI and other common containers; H.264, H.265/HEVC, VP8, VP9, AV1 and more.

### Output

One dataset row per kept frame:

```json
{
  "sourceUrl": "https://example.com/camera-01.mp4",
  "videoIndex": 1,
  "frameIndex": 4,
  "timestampSeconds": 3.017,
  "imageUrl": "https://api.apify.com/v2/key-value-stores/.../records/video001-frame-00004.jpg",
  "width": 1920,
  "height": 1080,
  "sharpness": 1234.7
}
```

| Field | Meaning |
|---|---|
| `imageUrl` | Download link of the frame image |
| `timestampSeconds` | Position of the frame in the video |
| `width`, `height` | Size of the stored image (after *Max side*) |
| `sharpness` | Higher = sharper. Clearly blurred frames score below ~40, typical sharp frames several hundred |
| `videoIndex`, `frameIndex` | Which video and which kept frame |
| `error` | Present only when a video failed (failed videos are free) |

A per-video summary is saved in the run's key-value store as `SUMMARY-video001`:

```json
{ "framesExtracted": 15, "candidates": 25, "skippedBlurry": 3, "skippedDuplicates": 7, "durationSeconds": 12.03, "zipUrl": "https://api.apify.com/v2/key-value-stores/.../records/video001-frames.zip" }
```

### Pricing

Pay per **frame kept**, plus a small fee per **started minute of video scanned**, plus a tiny per-run start fee. No subscription. Skipped frames and failed videos are free. Apify plan discounts apply automatically (see the **Pricing** tab).

| Example | Frames kept | Minutes scanned |
|---|---|---|
| 12-second clip, every 0.5 s (image above) | 15 | 1 |
| 10-minute video, every 2 s, mostly static camera | ~100–300 | 10 |
| 1-hour video, scene mode | one per shot | 60 |

Use **Max frames per video**, **Max minutes per video** or the run's *maximum cost* setting to cap spending; the Actor stops cleanly when the limit is reached.

### Use it from code or an AI agent

**Python: download a clean dataset**

```python
import pathlib, urllib.request
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("adorable_partial/video-frame-extractor-ai").call(run_input={
    "videoUrls": ["https://example.com/camera-01.mp4"],
    "mode": "interval", "everySeconds": "1", "maxSide": 640,
})
out = pathlib.Path("dataset/images"); out.mkdir(parents=True, exist_ok=True)
for frame in client.dataset(run["defaultDatasetId"]).iterate_items():
    if "imageUrl" in frame:
        urllib.request.urlretrieve(frame["imageUrl"], out / f"v{frame['videoIndex']}_{frame['timestampSeconds']:.2f}.jpg")
```

**JavaScript**

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('adorable_partial/video-frame-extractor-ai').call({
    videoUrls: ['https://example.com/ad.mp4'],
    mode: 'scene',
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((f) => console.log(f.timestampSeconds, f.imageUrl));
```

**HTTP**

```bash
curl -X POST "https://api.apify.com/v2/acts/adorable_partial~video-frame-extractor-ai/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"videoUrls": ["https://example.com/camera-01.mp4"], "everySeconds": "2"}'
```

**No-code and agents:** use the Apify modules in **Make**, **Zapier** or **n8n**, trigger runs from **webhooks** or **schedules**, or let an AI agent call it through the **Apify MCP server** (for example to give a multimodal LLM a few representative frames of a video).

### FAQ

**How do I get more or fewer frames?** Lower *Every N seconds* for more frames. Raise *Minimum change (%)* (e.g. 2–5) to keep only clearly different images, or set it to 0 to disable the duplicate check while keeping the blur filter.

**Too many frames marked blurry (or not enough)?** Check the `sharpness` values in the output and set *Blur threshold* just below the sharp frames you want to keep. Dark, low-detail scenes naturally score lower.

**Does it upscale?** No. *Max side* only downsizes.

**Can I get every single frame?** Set *Every N seconds* to the frame interval (e.g. 0.04 for 25 fps) and turn off both filters. Mind *Max frames per video*.

**Are faces or license plates blurred?** No. If you need anonymized frames, run the related anonymizer Actor on the output images.

**Is my data kept?** Videos are processed inside your run; frames are stored only in your own Apify storage, under your account's data-retention settings.

### Related Actors

- [Video Optimizer](https://apify.com/adorable_partial/video-optimizer) — reduce FPS, resize and compress videos
- [Image & Video Anonymizer](https://apify.com/adorable_partial/image-anonymizer) — blur faces and license plates (GDPR / LGPD)
- [Audio Extractor](https://apify.com/adorable_partial/audio-extractor) — video to MP3, WAV or FLAC
- [Audio & Video to Text (Whisper)](https://apify.com/adorable_partial/whisper-audio-video-transcriber) — transcripts and subtitles

### Support

Need another sampling strategy or export format (COCO, YOLO folders)? Open an issue in the **Issues** tab.

# Actor input Schema

## `videoUrls` (type: `array`):

Direct links to videos (MP4, MOV, WebM, MKV, AVI...). Public Google Drive and Dropbox share links work too.

## `mode` (type: `string`):

How frames are picked before filtering.

## `everySeconds` (type: `string`):

For 'Every N seconds': time between frames, e.g. 1, 0.5 or 10.

## `sceneThreshold` (type: `string`):

For 'On scene change': 0.2 = sensitive (more frames), 0.5 = only big cuts.

## `skipBlurry` (type: `boolean`):

Drop motion-blurred or out-of-focus frames.

## `blurThreshold` (type: `integer`):

Minimum sharpness score to keep a frame (the score is returned for each frame). Raise it to be stricter.

## `skipDuplicates` (type: `boolean`):

Drop frames that look almost identical to the previous kept frame (static scenes).

## `minChangePercent` (type: `string`):

A frame counts as a near-duplicate when less than this share of the picture changed since the last kept frame. 0.5 = default (keeps a moving car), 0 = keep everything, 5 = only big changes.

## `maxSide` (type: `integer`):

Downscale frames so the longest side is at most this. Leave empty to keep the original size.

## `outputFormat` (type: `string`):

File format of the extracted frames.

## `quality` (type: `integer`):

Compression quality for JPEG and WebP.

## `createZip` (type: `boolean`):

Bundle each video's frames into one ZIP file.

## `maxFramesPerVideo` (type: `integer`):

Stop after this many kept frames per video.

## `maxDurationMinutes` (type: `integer`):

Only the first N minutes of each video are scanned and billed.

## Actor input object example

```json
{
  "videoUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/0/06/Roncesvalles_Ave_crossing_at_Galley_Ave%2C_PXO_Type_A_lights_flashing%2C_July_2026.webm"
  ],
  "mode": "interval",
  "everySeconds": "1",
  "sceneThreshold": "0.35",
  "skipBlurry": true,
  "blurThreshold": 40,
  "skipDuplicates": true,
  "minChangePercent": "0.5",
  "outputFormat": "jpeg",
  "quality": 90,
  "createZip": false,
  "maxFramesPerVideo": 200,
  "maxDurationMinutes": 60
}
```

# Actor output Schema

## `frames` (type: `string`):

One row per extracted frame: image link, timestamp, size and sharpness score.

## `files` (type: `string`):

The frame images and optional ZIP files.

## `summaries` (type: `string`):

Frames kept and how many were skipped as blurry or duplicate, per video.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "videoUrls": [
        "https://upload.wikimedia.org/wikipedia/commons/0/06/Roncesvalles_Ave_crossing_at_Galley_Ave%2C_PXO_Type_A_lights_flashing%2C_July_2026.webm"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("adorable_partial/video-frame-extractor-ai").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "videoUrls": ["https://upload.wikimedia.org/wikipedia/commons/0/06/Roncesvalles_Ave_crossing_at_Galley_Ave%2C_PXO_Type_A_lights_flashing%2C_July_2026.webm"] }

# Run the Actor and wait for it to finish
run = client.actor("adorable_partial/video-frame-extractor-ai").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "videoUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/0/06/Roncesvalles_Ave_crossing_at_Galley_Ave%2C_PXO_Type_A_lights_flashing%2C_July_2026.webm"
  ]
}' |
apify call adorable_partial/video-frame-extractor-ai --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,adorable_partial/video-frame-extractor-ai"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Gw3aoBs3sl7uyvFO5/builds/MJleILyzXpsUF1hiv/openapi.json
