Face Detection, Tracking & Auto Reframe for Video
Pricing
Pay per usage
Face Detection, Tracking & Auto Reframe for Video
Detect and track faces in video (SCRFD ONNX), find the active speaker, and get auto-reframe crop keyframes to turn landscape video into 9:16 vertical without cutting faces off.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Andrew Babo
Maintained by CommunityActor stats
0
Bookmarked
1.5K
Total users
748
Monthly active users
12 days ago
Last modified
Share
Detect and track faces in a video, work out who is speaking, and get auto-reframe crop keyframes that turn a landscape recording into a 9:16 vertical clip without cutting anyone's head off.
Runs SCRFD (ONNX) on CPU — no GPU, no torch, no local install. The face model is baked into the image, so nothing is downloaded at run time.
Use it for: vertical clips for short-form video, talking-head reframing, podcast multi-speaker editing, thumbnail selection, face presence analytics, smart cropping in automated video pipelines.
Quick start
{"op": "reframe","source": "https://example.com/podcast.mp4","options": { "target_aspect": "9:16" }}
Feature-detect the build (free, a few seconds, needs no source):
{ "op": "capabilities" }
Analyse a 480p proxy, not the master. The Video & Audio Toolkit actor (
op: "proxy") makes one in seconds, and every result is in normalised 0–1 coordinates, so it applies to the full-resolution master unchanged.
Operations
op | What it returns |
|---|---|
faces | face detections per sampled frame, grouped into tracks |
reframe | faces + crop keyframes for a target aspect ratio |
asd | active-speaker detection: which tracked face is talking, when |
diarize | visual speaker segmentation helper |
capabilities | ops, features and vCPUs of this build; needs no source |
Options
| Option | Default | Notes |
|---|---|---|
sample_fps | 2 (8 for asd) | analysis frames per second, max 10 |
detect_fps | 2 for asd | how often the detector actually runs; boxes are interpolated between |
det_size | 640 (480 for asd) | detector input size |
batch_size | 8 | frames inferred per batch |
score_threshold | 0.5 | minimum detection confidence |
start_sec / duration_sec | — | analyse one window; timestamps are offset back to the original timeline |
track_iou | 0.3 | overlap needed to continue a track |
track_max_gap_ms | 700 | longer disappearance splits the track |
min_track_ms | 400 | shorter tracks are discarded |
include_frames | false | include per-frame boxes in the artifact |
target_aspect | 9:16 | reframe only |
track_id | most prominent track | reframe only |
smooth_window | 5 | median filter; larger = calmer camera |
dead_zone | 0.02 | ignore movement below 2% of the frame |
headroom | 0.12 | keep the face slightly above centre |
interp | linear | linear or hold between keyframes |
silence_ratio | 0.25 | asd: RMS below this share of peak counts as silence |
min_segment_ms | 600 | asd: shorter speaking turns are merged away |
min_correlation | 0.05 | asd: below this a track is treated as never speaking |
faces_url | — | reuse a previous faces result and skip detection entirely |
Output
{"status": "success","op": "reframe","artifacts": [{ "name": "reframe", "kv_key": "reframe.json" }],"meta": {"engine": "scrfd-onnx","width": 1280, "height": 720, "duration_sec": 61.4,"sample_fps": 2, "frame_count": 123, "face_frame_ratio": 0.951,"track_count": 2,"tracks": [{ "id": "face_01", "start_ms": 0, "end_ms": 58000, "duration_ms": 58000,"sample_count": 112, "avg_area_px": 20480, "avg_score": 0.912 }],"reframe": {"target_aspect": 0.5625,"crop_px": { "w": 405, "h": 720 },"track_id": "face_01","face_coverage": 0.951,"keyframes": [{ "t_ms": 0, "rect": { "x": 0.31, "y": 0, "w": 0.316, "h": 1 }, "interp": "linear" }]}}}
rectis in normalised 0–1 coordinates relative to the source frame, so it is resolution independent.- Per-frame boxes live in the artifact JSON, not in the dataset row, to keep rows small.
- The keyframes match the crop keyframe format of the Video Render Engine actor, so they can be pasted straight into an edit timeline.
Active speaker detection (op: "asd")
For every face track the actor measures mouth movement over time and correlates it with the audio envelope. A mouth moving while there is sound means that person is talking — no torch, no heavyweight ASD model.
{"segments": [{ "track_id": "face_01", "start_ms": 0, "end_ms": 4000 },{ "track_id": "face_02", "start_ms": 4000, "end_ms": 8000 }],"track_scores": [{ "track_id": "face_01", "speech_score": 0.41, "speaking_ms": 4000 }],"speech_ratio": 1.0}
Add target_aspect to an asd run and the returned crop keyframes follow
whoever is speaking instead of one fixed track (track_id becomes
"active_speaker").
Error handling
{ "ok": false, "reason": "BAD_INPUT", "message": "..." }
reason | Meaning | What to do |
|---|---|---|
BAD_INPUT | missing/unreadable source | check the URL is publicly reachable |
UPSTREAM_BLOCKED | the host refused the download | host the proxy yourself first |
TIMEOUT | analysis exceeded the run timeout | lower sample_fps, analyse a proxy, or shard |
OOM_LIMIT | not enough memory | run with 16 GB |
INTERNAL | unexpected failure | retry; report the run ID |
If no face is visible at all, the run still succeeds and meta.note explains
why no reframe could be produced.
Performance
16 GB run (≈4 vCPU) on a 480p proxy:
| Job | Typical time (10 min source) |
|---|---|
faces at sample_fps: 2 | ~1–2 min |
reframe at sample_fps: 2 | ~1–2 min |
asd at sample_fps: 8, detect_fps: 2 | ~3–5 min |
Cost scales linearly with sample_fps. Fastest recipe: make a 480p proxy, run
faces once, then reuse it with faces_url for asd or a second reframe.
FAQ
Do I need a GPU? No. SCRFD runs on CPU through onnxruntime.
Will the crop jitter? Movement smaller than dead_zone is ignored and the
path is median-filtered (smooth_window), so the virtual camera glides instead
of twitching. Raise smooth_window for an even calmer result.
What if someone turns away? Occluded or back-turned faces cannot be detected; those moments keep the nearest keyframe, so the crop simply holds.
Can I reframe to 1:1 or 4:5? Yes — set target_aspect to any ratio.
Audio-only speaker detection? Use the Speaker Diarization actor; it
needs no visible face and pairs well with asd.