Video Background Remover avatar

Video Background Remover

Pricing

Pay per usage

Go to Apify Store
Video Background Remover

Video Background Remover

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Andrew Babo

Andrew Babo

Maintained by Community

Actor stats

0

Bookmarked

10

Total users

9

Monthly active users

3 days ago

Last modified

Categories

Share

Video Background Remover — person or any object, on CPU

Remove the background from a video and get a green-screen MP4, an alpha-matte MP4 or a demo composite back — with hair- and fur-level edges, the source resolution, frame rate and audio kept.

Two targets in one Actor:

targetWhat it cuts outHow it is chosenModel
person (default)The main person in frameautomaticPP-MattingV2 int8 (human matting)
objectAny single thing: dog, cat, product, car, logo, bag…text prompt (AI draws the box), a box, a point or a painted maskMobileSAM int8 tracking + colour guided-filter edges

Everything runs on CPU — no GPU queue, no per-minute GPU price. One matting process per vCPU, RAM-disk frame pipeline, MPEG-TS slices that stitch frame-exact, and startFrame / endFrame to fan a long clip out across several runs.


🤖 AI Agent / MCP Quickstart

Actor id: andrew_babo/video-background-remover. One call, one dataset row, one video URL.

Person (nothing to select):

{ "videoUrl": "https://example.com/talking-head.mp4", "outputMode": "green" }

Object by description — the vision model finds it (needs visionApiKey, see below):

{ "videoUrl": "https://example.com/park.mp4", "target": "object", "objectPrompt": "the golden retriever",
"visionApiKey": "sk-...", "outputMode": "matte" }

Object by box / point / painted mask — no vision model needed:

{ "videoUrl": "https://example.com/park.mp4", "target": "object", "objectBox": [448, 311, 1221, 902], "selectionFrame": 0 }
{ "videoUrl": "https://example.com/park.mp4", "target": "object", "objectPoint": [947, 595] }
{ "videoUrl": "https://example.com/park.mp4", "target": "object", "objectMaskUrl": "https://example.com/brush.png" }

Check the pick before rendering — add "previewOnly": true: returns previewUrl (frame with the picked object tinted green + its box) and previewMaskUrl in a few seconds, no video.

Read the result: the default dataset has exactly one row.

videoUrl rendered MP4 (null on error)
target, selection person | object; prompt | box | point | mask
seedBox, aiBox object box on the selection frame (pixels)
selectionPreviewUrl JPEG of the selection frame with the picked object tinted (object runs)
objectFrames, lostFrames, sceneCuts, aiCalls, warnings tracking health
renderSeconds, realtimeFactor, totalSeconds speed
errorCode, error only when something the caller can fix went wrong (see Error codes)

Rules of thumb for agents: person for presenters / talking heads / people walking; object for anything else. Prefer objectBox or objectPoint when you already know where the thing is (no API key, deterministic); use objectPrompt when you only have a description. Run previewOnly first when the shot has several similar things. Coordinates are pixels of the source video; fractions 0–1 are accepted too.


Output modes

outputModeYou getUse it for
greensubject on pure #00FF00chroma key in any editor / FFmpeg chromakey
mattegrayscale alpha video (white = subject)luma / track matte, compositing pipelines, transparent WebM via FFmpeg alphamerge
demosubject on a dark studio background with a captionquick visual check, client previews

All outputs are H.264 MP4 (CRF 17, superfast), same size and fps as the source, source audio copied (AAC).


Selecting an object

FieldWhat to sendNeeds vision model
objectPrompt"the red coffee mug on the table"yes
objectBox[x1, y1, x2, y2] around the object on selectionFrameno
objectPoint[x, y] inside the objectno
objectMaskUrlPNG/JPG where the object is painted white (a brush stroke is enough — it snaps to the edges)no

Priority when several are given: mask > box > point > prompt. selectionFrame (default 0) is the frame the selection refers to; tracking runs backwards and forwards from it inside that shot.

How tracking works. The selection becomes a mask on the selection frame. A cheap pre-pass over a reduced copy of the video finds scene cuts and carries the mask with optical flow to the start of every worker's frame range, where MobileSAM re-snaps it. Every full-resolution frame is then segmented by MobileSAM again (box + centre prompt + the flow-warped previous mask as a soft prior) — no frame skipping, no interpolation — and the edge band gets a colour guided filter plus edge colour decontamination so fur and thin parts keep their detail without green/background fringes. The mask is chosen by plausibility gates (contains the previous centre, similar area, IoU > 0.5) and completeness, so a tail or an ear that goes blurry for two frames is picked back up instead of being lost for the rest of the shot.

Scene cuts and lost objects. After a cut the object is re-found with the vision model (prompt, or the reference crop of the selected object when there was no prompt). Without a vision model the tracker holds the last mask for up to 12 frames of occlusion and otherwise renders empty frames after the cut, reporting NO_VISION_MODEL in warnings. Frames without the object are counted in lostFrames and rendered as background only — never a wrong object.


Vision model (only for objectPrompt and re-finding after cuts)

Bring any OpenAI-compatible chat-completions endpoint that accepts images:

FieldDefaultExamples
visionApiKey (secret)—your OpenAI / OpenRouter / Gemini key
visionBaseUrlhttps://api.openai.com/v1https://openrouter.ai/api/v1, https://generativelanguage.googleapis.com/v1beta/openai
visionModelgpt-4.1-minigpt-5-mini, google/gemini-2.5-flash, gemini-2.5-flash
maxAiCalls60hard cost cap per run

A call sends one downscaled frame (long side 1280 px, ~1.2k input tokens) and asks for a JSON box. A run makes 1 call on the selection frame + 1 per scene cut + 1 per worker range that starts after a cut, so a single-shot clip needs exactly one call however long it is; a 2-minute clip with 8 cuts on a 4-vCPU run needs ~10–15. The key is used only inside the run, never logged and never written to the dataset (vision in the row shows base URL and model only). The Actor owner can set MATTE_AI_URL, MATTE_AI_KEY, MATTE_AI_MODEL as Actor environment variables to offer a default; a per-run visionApiKey always wins.


Measured performance

Numbers reported by the Actor itself in its dataset row (renderSeconds, totalSeconds) on real Apify runs, default 16 GB = 4 vCPU (Xeon Ice Lake, AVX-512 VNNI), 1080p 24 fps sources.

TargetClipSetupRenderEnd-to-endNotes
person10 s talking head, 1080p16 GB, one run28.6 s29.0 ssmart crop, 4 workers
person10 s talking head, 1080p4 × 8 GB slices + stitch—22–27 sslices land on different hosts
object10 s dog in a park, 1080p, 1 scene cut at 8.5 s16 GB, one run85.4 s86.2 sbox selection, 203/203 frames of the dog's shot segmented, 37 frames after the cut empty (no vision key); 0.3: 93.4 s
object2 s of the same clip, painted-mask selection16 GB, one run29.1 s29.7 s48/48 frames; 0.3: 32.4 s
objectselection preview only (previewOnly)16 GB—≈ 3.5 spreview + mask PNG

CPU-seconds per 1080p frame on that hardware: person ≈ 0.55 s, object ≈ 1.2 s (MobileSAM encoder ≈ 1.0–1.2 s, i.e. 80–90 % of the object budget on Apify's Ice Lake hosts — 0.4–0.6 s on a current AMX host — depending on the object's aspect ratio, edges ≈ 0.08 s, encode/composite ≈ 0.02 s). Per-stage timings are in cpuSecondsByStage. A long clip scales linearly: 4 workers track ≈ 3 frames/s together, so a 2 min 1080p clip is ≈ 15 min in one run, or ≈ 4 min split over four runs with startFrame / endFrame.

Speed levers, in order: more memory on the run (each 4 GB adds one worker), startFrame / endFrame fan-out across runs (each run gets its own host), x264Preset: "ultrafast" for large frames, maxFrames for tests.


Long videos: fan-out and stitch

Start N runs with consecutive startFrame / endFrame ranges (default sliceFormat: "ts"), download the .ts slices and concatenate:

ffmpeg -f concat -safe 0 -i list.txt -i source.mp4 -map 0:v -map 1:a? -c:v copy -c:a aac -movflags +faststart out.mp4

TS slices concatenate frame-exact with -c copy; MP4 slices can lose a frame at each seam, so only use sliceFormat: "mp4" when a slice must be playable on its own. For object, give every slice the same selection (objectBox/objectPoint/objectMaskUrl + selectionFrame, or the same objectPrompt); the pre-pass carries the selection into each slice's range as long as it lies in the same shot, and the vision model re-finds it otherwise.


Input reference

FieldTypeDefaultNotes
videoUrlstringrequireddirect MP4/MOV/WebM link
targetperson | objectperson
outputModegreen | matte | demogreen
objectPromptstringtext description for the vision model
objectBox[x1,y1,x2,y2]pixels or fractions
objectPoint[x,y]pixels or fractions
objectMaskUrlstringwhite = object
selectionFrameinteger0frame the selection refers to
previewOnlybooleanfalsesegment the selection frame only
visionApiKey / visionBaseUrl / visionModel / visionReasoning / maxAiCallssee above
startFrame / endFrame / sliceFormatfan-out
parallelRuns / parallelMemoryMbytesinteger1 / 8192object only: split the video over N parallel runs of this actor and stitch the slices automatically; cut points are chosen where the object is fully visible. Off by default — only faster than one 16 GB run when your plan can run several machines at once
workersinteger0 = one per vCPU
crf / x264Preset17 / superfastencoder
outputKeystringoutput.mp4key in the run's key-value store; previews use -preview / -selection suffixes
maxFramesinteger0benchmarking aid
disableSmartCropbooleanfalseperson only, debugging
demoLabelstringcaption for demo

Error codes

The run still pushes one dataset row so a caller can read what happened; videoUrl is null and errorCode is one of:

errorCodeMeaningFix
NO_INPUT / BAD_INPUTmissing videoUrl, invalid field, object without any selectioncorrect the input
DOWNLOAD_FAILED / MASK_DOWNLOAD_FAILEDsource or mask URL not reachablemake the URL public / direct
OBJECT_NOT_FOUNDthe box / point / mask / prompt gave no object on selectionFramepick another frame or a tighter selection; use previewOnly
AI_UNAVAILABLEobjectPrompt without a vision modelset visionApiKey or select with box / point / mask
AI_UNAUTHORIZED / AI_PAYMENT_REQUIRED / AI_FAILEDthe vision provider rejected the key, has no credits, or answered garbagecheck key, model id and base URL
AI_BUDGET_EXCEEDEDmaxAiCalls reached before the first object was foundraise the cap
FFMPEG_FAILED / RENDER_FAILEDinfrastructure failure (the run also fails)retry; report with the run id

Successful object runs can still carry warnings: NO_VISION_MODEL (object lost or cut without a way to re-find it), AI_BUDGET_EXCEEDED, AI_FAILED, OBJECT_NEVER_FOUND.


Models and licences

PP-MattingV2 (PaddleSeg, Apache-2.0) exported to static int8 ONNX and run with OpenVINO; MobileSAM (Apache-2.0) image encoder as int8 OpenVINO IR — one static canvas per aspect ratio (long side 1024, short side 512–1024) so a wide or tall object is not paid for as a padded square — and its multi-mask decoder in ONNX Runtime with a matching embedding grid; OpenCV DIS optical flow; FFmpeg / x264 for decode and encode. No frames leave the run except the single downscaled frame per vision call when objectPrompt (or re-finding after a cut) is used with your own key.


Changelog

  • 0.6 — occlusion handling for object: the tracker remembers the object (last whole area, amodal box, occluder negative points), keeps the visible part while something passes in front, and re-acquires the full object afterwards — gated by hue-saturation AND brightness histograms so a same-hue bystander is not absorbed; occluder pixels are cut out by colour pockets and a motion cut that only applies while the object itself stands still. Static scenes render up to 2× faster (unchanged regions are reused instead of recomputed). New opt-in parallelRuns: the run plans cut points where the object is fully visible, calls N parallel runs of this actor and stitches the slices — same masks, shorter wait when the plan allows several machines at once.
  • 0.2 — target: "object": any object by prompt / box / point / painted mask, MobileSAM tracking with scene-cut re-seeding, selection preview, vision provider settings and call cap, machine-readable object metrics and warnings.
  • 0.4 — object pipeline scheduling: the pre-pass no longer blocks the workers (one decode feeds both; worker ranges are released as the chain advances and split by measured per-frame cost, so frames without the object are nearly free), idle workers take the tail of the slowest range at a seeded frame, and the per-frame mask/edge work runs in the object's window only. 10 s dog on Apify 93 s → 85 s (encoder-bound there; −26 % on a 4-core AMX host), a 10 s cat clip −32 % with 9 more frames of the cat and clean seams where 0.3 had picked up a second animal at a range start; masks otherwise unchanged (frame agreement > 0.99).
  • 0.3 — object render 163 s → 93 s on the same 10 s clip and host: aspect-ratio encoder canvases (encoder CPU −52%), mask work confined to the object's window, windowed optical flow, worker ranges balanced against the pre-pass; masks unchanged (frame agreement > 0.99).
  • 0.1 — person matting (PP-MattingV2 int8, smart crop, fan-out slices).