Face Detection, Tracking & Auto Reframe for Video avatar

Face Detection, Tracking & Auto Reframe for Video

Pricing

Pay per usage

Go to Apify Store
Face Detection, Tracking & Auto Reframe for Video

Face Detection, Tracking & Auto Reframe for Video

Detect and track faces in video (SCRFD ONNX), find the active speaker, and get auto-reframe crop keyframes to turn landscape video into 9:16 vertical without cutting faces off.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Andrew Babo

Andrew Babo

Maintained by Community

Actor stats

0

Bookmarked

1.5K

Total users

748

Monthly active users

12 days ago

Last modified

Categories

Share

Detect and track faces in a video, work out who is speaking, and get auto-reframe crop keyframes that turn a landscape recording into a 9:16 vertical clip without cutting anyone's head off.

Runs SCRFD (ONNX) on CPU — no GPU, no torch, no local install. The face model is baked into the image, so nothing is downloaded at run time.

Use it for: vertical clips for short-form video, talking-head reframing, podcast multi-speaker editing, thumbnail selection, face presence analytics, smart cropping in automated video pipelines.

Quick start

{
"op": "reframe",
"source": "https://example.com/podcast.mp4",
"options": { "target_aspect": "9:16" }
}

Feature-detect the build (free, a few seconds, needs no source):

{ "op": "capabilities" }

Analyse a 480p proxy, not the master. The Video & Audio Toolkit actor (op: "proxy") makes one in seconds, and every result is in normalised 0–1 coordinates, so it applies to the full-resolution master unchanged.

Operations

opWhat it returns
facesface detections per sampled frame, grouped into tracks
reframefaces + crop keyframes for a target aspect ratio
asdactive-speaker detection: which tracked face is talking, when
diarizevisual speaker segmentation helper
capabilitiesops, features and vCPUs of this build; needs no source

Options

OptionDefaultNotes
sample_fps2 (8 for asd)analysis frames per second, max 10
detect_fps2 for asdhow often the detector actually runs; boxes are interpolated between
det_size640 (480 for asd)detector input size
batch_size8frames inferred per batch
score_threshold0.5minimum detection confidence
start_sec / duration_sec—analyse one window; timestamps are offset back to the original timeline
track_iou0.3overlap needed to continue a track
track_max_gap_ms700longer disappearance splits the track
min_track_ms400shorter tracks are discarded
include_framesfalseinclude per-frame boxes in the artifact
target_aspect9:16reframe only
track_idmost prominent trackreframe only
smooth_window5median filter; larger = calmer camera
dead_zone0.02ignore movement below 2% of the frame
headroom0.12keep the face slightly above centre
interplinearlinear or hold between keyframes
silence_ratio0.25asd: RMS below this share of peak counts as silence
min_segment_ms600asd: shorter speaking turns are merged away
min_correlation0.05asd: below this a track is treated as never speaking
faces_url—reuse a previous faces result and skip detection entirely

Output

{
"status": "success",
"op": "reframe",
"artifacts": [{ "name": "reframe", "kv_key": "reframe.json" }],
"meta": {
"engine": "scrfd-onnx",
"width": 1280, "height": 720, "duration_sec": 61.4,
"sample_fps": 2, "frame_count": 123, "face_frame_ratio": 0.951,
"track_count": 2,
"tracks": [
{ "id": "face_01", "start_ms": 0, "end_ms": 58000, "duration_ms": 58000,
"sample_count": 112, "avg_area_px": 20480, "avg_score": 0.912 }
],
"reframe": {
"target_aspect": 0.5625,
"crop_px": { "w": 405, "h": 720 },
"track_id": "face_01",
"face_coverage": 0.951,
"keyframes": [
{ "t_ms": 0, "rect": { "x": 0.31, "y": 0, "w": 0.316, "h": 1 }, "interp": "linear" }
]
}
}
}
  • rect is in normalised 0–1 coordinates relative to the source frame, so it is resolution independent.
  • Per-frame boxes live in the artifact JSON, not in the dataset row, to keep rows small.
  • The keyframes match the crop keyframe format of the Video Render Engine actor, so they can be pasted straight into an edit timeline.

Active speaker detection (op: "asd")

For every face track the actor measures mouth movement over time and correlates it with the audio envelope. A mouth moving while there is sound means that person is talking — no torch, no heavyweight ASD model.

{
"segments": [
{ "track_id": "face_01", "start_ms": 0, "end_ms": 4000 },
{ "track_id": "face_02", "start_ms": 4000, "end_ms": 8000 }
],
"track_scores": [
{ "track_id": "face_01", "speech_score": 0.41, "speaking_ms": 4000 }
],
"speech_ratio": 1.0
}

Add target_aspect to an asd run and the returned crop keyframes follow whoever is speaking instead of one fixed track (track_id becomes "active_speaker").

Error handling

{ "ok": false, "reason": "BAD_INPUT", "message": "..." }
reasonMeaningWhat to do
BAD_INPUTmissing/unreadable sourcecheck the URL is publicly reachable
UPSTREAM_BLOCKEDthe host refused the downloadhost the proxy yourself first
TIMEOUTanalysis exceeded the run timeoutlower sample_fps, analyse a proxy, or shard
OOM_LIMITnot enough memoryrun with 16 GB
INTERNALunexpected failureretry; report the run ID

If no face is visible at all, the run still succeeds and meta.note explains why no reframe could be produced.

Performance

16 GB run (≈4 vCPU) on a 480p proxy:

JobTypical time (10 min source)
faces at sample_fps: 2~1–2 min
reframe at sample_fps: 2~1–2 min
asd at sample_fps: 8, detect_fps: 2~3–5 min

Cost scales linearly with sample_fps. Fastest recipe: make a 480p proxy, run faces once, then reuse it with faces_url for asd or a second reframe.

FAQ

Do I need a GPU? No. SCRFD runs on CPU through onnxruntime.

Will the crop jitter? Movement smaller than dead_zone is ignored and the path is median-filtered (smooth_window), so the virtual camera glides instead of twitching. Raise smooth_window for an even calmer result.

What if someone turns away? Occluded or back-turned faces cannot be detected; those moments keep the nearest keyframe, so the crop simply holds.

Can I reframe to 1:1 or 4:5? Yes — set target_aspect to any ratio.

Audio-only speaker detection? Use the Speaker Diarization actor; it needs no visible face and pairs well with asd.