# Face Detection, Tracking & Auto Reframe for Video (`andrew_babo/face-detection-video`) Actor

Detect and track faces in video (SCRFD ONNX), find the active speaker, and get auto-reframe crop keyframes to turn landscape video into 9:16 vertical without cutting faces off.

- **URL**: https://apify.com/andrew\_babo/face-detection-video.md
- **Developed by:** [Andrew Babo](https://apify.com/andrew_babo) (community)
- **Categories:** AI, Videos
- **Stats:** 1,471 total users, 748 monthly users, 97.2% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Face Detection, Tracking & Auto Reframe for Video

Detect and track faces in a video, work out **who is speaking**, and get
**auto-reframe crop keyframes** that turn a landscape recording into a 9:16
vertical clip without cutting anyone's head off.

Runs SCRFD (ONNX) on CPU — no GPU, no torch, no local install. The face model is
baked into the image, so nothing is downloaded at run time.

**Use it for:** vertical clips for short-form video, talking-head reframing,
podcast multi-speaker editing, thumbnail selection, face presence analytics,
smart cropping in automated video pipelines.

### Quick start

```json
{
  "op": "reframe",
  "source": "https://example.com/podcast.mp4",
  "options": { "target_aspect": "9:16" }
}
```

Feature-detect the build (free, a few seconds, needs no source):

```json
{ "op": "capabilities" }
```

> Analyse a **480p proxy**, not the master. The **Video & Audio Toolkit** actor
> (`op: "proxy"`) makes one in seconds, and every result is in normalised
> 0–1 coordinates, so it applies to the full-resolution master unchanged.

### Operations

| `op` | What it returns |
| --- | --- |
| `faces` | face detections per sampled frame, grouped into tracks |
| `reframe` | `faces` + crop keyframes for a target aspect ratio |
| `asd` | active-speaker detection: which tracked face is talking, when |
| `diarize` | visual speaker segmentation helper |
| `capabilities` | ops, features and vCPUs of this build; needs no `source` |

### Options

| Option | Default | Notes |
| --- | --- | --- |
| `sample_fps` | `2` (`8` for `asd`) | analysis frames per second, max 10 |
| `detect_fps` | `2` for `asd` | how often the detector actually runs; boxes are interpolated between |
| `det_size` | `640` (`480` for `asd`) | detector input size |
| `batch_size` | `8` | frames inferred per batch |
| `score_threshold` | `0.5` | minimum detection confidence |
| `start_sec` / `duration_sec` | — | analyse one window; timestamps are offset back to the original timeline |
| `track_iou` | `0.3` | overlap needed to continue a track |
| `track_max_gap_ms` | `700` | longer disappearance splits the track |
| `min_track_ms` | `400` | shorter tracks are discarded |
| `include_frames` | `false` | include per-frame boxes in the artifact |
| `target_aspect` | `9:16` | reframe only |
| `track_id` | most prominent track | reframe only |
| `smooth_window` | `5` | median filter; larger = calmer camera |
| `dead_zone` | `0.02` | ignore movement below 2% of the frame |
| `headroom` | `0.12` | keep the face slightly above centre |
| `interp` | `linear` | `linear` or `hold` between keyframes |
| `silence_ratio` | `0.25` | `asd`: RMS below this share of peak counts as silence |
| `min_segment_ms` | `600` | `asd`: shorter speaking turns are merged away |
| `min_correlation` | `0.05` | `asd`: below this a track is treated as never speaking |
| `faces_url` | — | reuse a previous `faces` result and skip detection entirely |

### Output

```json
{
  "status": "success",
  "op": "reframe",
  "artifacts": [{ "name": "reframe", "kv_key": "reframe.json" }],
  "meta": {
    "engine": "scrfd-onnx",
    "width": 1280, "height": 720, "duration_sec": 61.4,
    "sample_fps": 2, "frame_count": 123, "face_frame_ratio": 0.951,
    "track_count": 2,
    "tracks": [
      { "id": "face_01", "start_ms": 0, "end_ms": 58000, "duration_ms": 58000,
        "sample_count": 112, "avg_area_px": 20480, "avg_score": 0.912 }
    ],
    "reframe": {
      "target_aspect": 0.5625,
      "crop_px": { "w": 405, "h": 720 },
      "track_id": "face_01",
      "face_coverage": 0.951,
      "keyframes": [
        { "t_ms": 0, "rect": { "x": 0.31, "y": 0, "w": 0.316, "h": 1 }, "interp": "linear" }
      ]
    }
  }
}
```

- `rect` is in **normalised 0–1 coordinates** relative to the source frame, so it
  is resolution independent.
- Per-frame boxes live in the artifact JSON, not in the dataset row, to keep
  rows small.
- The keyframes match the crop keyframe format of the **Video Render Engine**
  actor, so they can be pasted straight into an edit timeline.

#### Active speaker detection (`op: "asd"`)

For every face track the actor measures mouth movement over time and correlates
it with the audio envelope. A mouth moving *while there is sound* means that
person is talking — no torch, no heavyweight ASD model.

```json
{
  "segments": [
    { "track_id": "face_01", "start_ms": 0,    "end_ms": 4000 },
    { "track_id": "face_02", "start_ms": 4000, "end_ms": 8000 }
  ],
  "track_scores": [
    { "track_id": "face_01", "speech_score": 0.41, "speaking_ms": 4000 }
  ],
  "speech_ratio": 1.0
}
```

Add `target_aspect` to an `asd` run and the returned crop keyframes **follow
whoever is speaking** instead of one fixed track (`track_id` becomes
`"active_speaker"`).

### Error handling

```json
{ "ok": false, "reason": "BAD_INPUT", "message": "..." }
```

| `reason` | Meaning | What to do |
| --- | --- | --- |
| `BAD_INPUT` | missing/unreadable source | check the URL is publicly reachable |
| `UPSTREAM_BLOCKED` | the host refused the download | host the proxy yourself first |
| `TIMEOUT` | analysis exceeded the run timeout | lower `sample_fps`, analyse a proxy, or shard |
| `OOM_LIMIT` | not enough memory | run with 16 GB |
| `INTERNAL` | unexpected failure | retry; report the run ID |

If no face is visible at all, the run still succeeds and `meta.note` explains
why no reframe could be produced.

### Performance

16 GB run (≈4 vCPU) on a 480p proxy:

| Job | Typical time (10 min source) |
| --- | --- |
| `faces` at `sample_fps: 2` | ~1–2 min |
| `reframe` at `sample_fps: 2` | ~1–2 min |
| `asd` at `sample_fps: 8`, `detect_fps: 2` | ~3–5 min |

Cost scales linearly with `sample_fps`. Fastest recipe: make a 480p proxy, run
`faces` once, then reuse it with `faces_url` for `asd` or a second reframe.

### FAQ

**Do I need a GPU?** No. SCRFD runs on CPU through onnxruntime.

**Will the crop jitter?** Movement smaller than `dead_zone` is ignored and the
path is median-filtered (`smooth_window`), so the virtual camera glides instead
of twitching. Raise `smooth_window` for an even calmer result.

**What if someone turns away?** Occluded or back-turned faces cannot be
detected; those moments keep the nearest keyframe, so the crop simply holds.

**Can I reframe to 1:1 or 4:5?** Yes — set `target_aspect` to any ratio.

**Audio-only speaker detection?** Use the **Speaker Diarization** actor; it
needs no visible face and pairs well with `asd`.

# Actor input Schema

## `op` (type: `string`):

faces = detect + track faces. reframe = the same, plus normalized crop keyframes for a target aspect ratio (editplan-ready). asd = active-speaker detection (mouth motion x audio envelope, no torch needed): per-track speech scores + speaking segments; pass target\_aspect to also get crop keyframes that follow whoever is speaking. diarize = visual diarization: the same ASD pass, returned as speaker turns (spk\_<track>) with speaker\_count/coverage/degraded — use it when audio diarization returns 0 speakers. capabilities = return supported ops/features without running any work (no source needed).

## `source` (type: `string`):

https:// URL, or kv:<storeId>/<key> from an earlier run (e.g. the 480p proxy from video-audio-toolkit — analysing the proxy is much cheaper than the master).

## `options` (type: `object`):

Vision options include sampling/detection controls plus embeddings (default true) and embedding\_samples (default 5, max 8). Face embeddings are 512-d normalized ArcFace vectors computed only from representative aligned samples, without a second video decode.

## `output` (type: `object`):

{ signed\_upload\_url } to PUT the result JSON straight into your own storage.

## `cleanup` (type: `string`):

on\_success = keep only result artifacts in the run's key-value store. always = also drop them after the signed upload. off = keep everything for debugging.

## `callback` (type: `object`):

{ url, secret\_header: { name, value } } — POSTed with the result JSON when the run finishes.

## Actor input object example

```json
{
  "op": "capabilities",
  "cleanup": "on_success"
}
```

# Actor output Schema

## `results` (type: `string`):

Full run result JSON: status, op, artifacts \[{name, kv\_key, url, bytes}], meta (tracks, ASD segments, crop keyframes), timings, errors.

## `resultRecord` (type: `string`):

The same result JSON stored as the RESULT record of the default key-value store.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "op": "capabilities"
};

// Run the Actor and wait for it to finish
const run = await client.actor("andrew_babo/face-detection-video").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "op": "capabilities" }

# Run the Actor and wait for it to finish
run = client.actor("andrew_babo/face-detection-video").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "op": "capabilities"
}' |
apify call andrew_babo/face-detection-video --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,andrew_babo/face-detection-video"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/dUNxVyhAnukQZICjx/builds/BLfXNousNB8jQf60k/openapi.json
