Video Background Remover
Pricing
Pay per usage
Video Background Remover
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Andrew Babo
Maintained by CommunityActor stats
0
Bookmarked
10
Total users
9
Monthly active users
3 days ago
Last modified
Categories
Share
Video Background Remover — person or any object, on CPU
Remove the background from a video and get a green-screen MP4, an alpha-matte MP4 or a demo composite back — with hair- and fur-level edges, the source resolution, frame rate and audio kept.
Two targets in one Actor:
target | What it cuts out | How it is chosen | Model |
|---|---|---|---|
person (default) | The main person in frame | automatic | PP-MattingV2 int8 (human matting) |
object | Any single thing: dog, cat, product, car, logo, bag… | text prompt (AI draws the box), a box, a point or a painted mask | MobileSAM int8 tracking + colour guided-filter edges |
Everything runs on CPU — no GPU queue, no per-minute GPU price. One matting process per vCPU, RAM-disk frame pipeline, MPEG-TS slices that stitch frame-exact, and startFrame / endFrame to fan a long clip out across several runs.
🤖 AI Agent / MCP Quickstart
Actor id: andrew_babo/video-background-remover. One call, one dataset row, one video URL.
Person (nothing to select):
{ "videoUrl": "https://example.com/talking-head.mp4", "outputMode": "green" }
Object by description — the vision model finds it (needs visionApiKey, see below):
{ "videoUrl": "https://example.com/park.mp4", "target": "object", "objectPrompt": "the golden retriever","visionApiKey": "sk-...", "outputMode": "matte" }
Object by box / point / painted mask — no vision model needed:
{ "videoUrl": "https://example.com/park.mp4", "target": "object", "objectBox": [448, 311, 1221, 902], "selectionFrame": 0 }{ "videoUrl": "https://example.com/park.mp4", "target": "object", "objectPoint": [947, 595] }{ "videoUrl": "https://example.com/park.mp4", "target": "object", "objectMaskUrl": "https://example.com/brush.png" }
Check the pick before rendering — add "previewOnly": true: returns previewUrl (frame with the picked object tinted green + its box) and previewMaskUrl in a few seconds, no video.
Read the result: the default dataset has exactly one row.
videoUrl rendered MP4 (null on error)target, selection person | object; prompt | box | point | maskseedBox, aiBox object box on the selection frame (pixels)selectionPreviewUrl JPEG of the selection frame with the picked object tinted (object runs)objectFrames, lostFrames, sceneCuts, aiCalls, warnings tracking healthrenderSeconds, realtimeFactor, totalSeconds speederrorCode, error only when something the caller can fix went wrong (see Error codes)
Rules of thumb for agents: person for presenters / talking heads / people walking; object for anything else. Prefer objectBox or objectPoint when you already know where the thing is (no API key, deterministic); use objectPrompt when you only have a description. Run previewOnly first when the shot has several similar things. Coordinates are pixels of the source video; fractions 0–1 are accepted too.
Output modes
outputMode | You get | Use it for |
|---|---|---|
green | subject on pure #00FF00 | chroma key in any editor / FFmpeg chromakey |
matte | grayscale alpha video (white = subject) | luma / track matte, compositing pipelines, transparent WebM via FFmpeg alphamerge |
demo | subject on a dark studio background with a caption | quick visual check, client previews |
All outputs are H.264 MP4 (CRF 17, superfast), same size and fps as the source, source audio copied (AAC).
Selecting an object
| Field | What to send | Needs vision model |
|---|---|---|
objectPrompt | "the red coffee mug on the table" | yes |
objectBox | [x1, y1, x2, y2] around the object on selectionFrame | no |
objectPoint | [x, y] inside the object | no |
objectMaskUrl | PNG/JPG where the object is painted white (a brush stroke is enough — it snaps to the edges) | no |
Priority when several are given: mask > box > point > prompt. selectionFrame (default 0) is the frame the selection refers to; tracking runs backwards and forwards from it inside that shot.
How tracking works. The selection becomes a mask on the selection frame. A cheap pre-pass over a reduced copy of the video finds scene cuts and carries the mask with optical flow to the start of every worker's frame range, where MobileSAM re-snaps it. Every full-resolution frame is then segmented by MobileSAM again (box + centre prompt + the flow-warped previous mask as a soft prior) — no frame skipping, no interpolation — and the edge band gets a colour guided filter plus edge colour decontamination so fur and thin parts keep their detail without green/background fringes. The mask is chosen by plausibility gates (contains the previous centre, similar area, IoU > 0.5) and completeness, so a tail or an ear that goes blurry for two frames is picked back up instead of being lost for the rest of the shot.
Scene cuts and lost objects. After a cut the object is re-found with the vision model (prompt, or the reference crop of the selected object when there was no prompt). Without a vision model the tracker holds the last mask for up to 12 frames of occlusion and otherwise renders empty frames after the cut, reporting NO_VISION_MODEL in warnings. Frames without the object are counted in lostFrames and rendered as background only — never a wrong object.
Vision model (only for objectPrompt and re-finding after cuts)
Bring any OpenAI-compatible chat-completions endpoint that accepts images:
| Field | Default | Examples |
|---|---|---|
visionApiKey (secret) | — | your OpenAI / OpenRouter / Gemini key |
visionBaseUrl | https://api.openai.com/v1 | https://openrouter.ai/api/v1, https://generativelanguage.googleapis.com/v1beta/openai |
visionModel | gpt-4.1-mini | gpt-5-mini, google/gemini-2.5-flash, gemini-2.5-flash |
maxAiCalls | 60 | hard cost cap per run |
A call sends one downscaled frame (long side 1280 px, ~1.2k input tokens) and asks for a JSON box. A run makes 1 call on the selection frame + 1 per scene cut + 1 per worker range that starts after a cut, so a single-shot clip needs exactly one call however long it is; a 2-minute clip with 8 cuts on a 4-vCPU run needs ~10–15. The key is used only inside the run, never logged and never written to the dataset (vision in the row shows base URL and model only). The Actor owner can set MATTE_AI_URL, MATTE_AI_KEY, MATTE_AI_MODEL as Actor environment variables to offer a default; a per-run visionApiKey always wins.
Measured performance
Numbers reported by the Actor itself in its dataset row (renderSeconds, totalSeconds) on real Apify runs, default 16 GB = 4 vCPU (Xeon Ice Lake, AVX-512 VNNI), 1080p 24 fps sources.
| Target | Clip | Setup | Render | End-to-end | Notes |
|---|---|---|---|---|---|
| person | 10 s talking head, 1080p | 16 GB, one run | 28.6 s | 29.0 s | smart crop, 4 workers |
| person | 10 s talking head, 1080p | 4 × 8 GB slices + stitch | — | 22–27 s | slices land on different hosts |
| object | 10 s dog in a park, 1080p, 1 scene cut at 8.5 s | 16 GB, one run | 85.4 s | 86.2 s | box selection, 203/203 frames of the dog's shot segmented, 37 frames after the cut empty (no vision key); 0.3: 93.4 s |
| object | 2 s of the same clip, painted-mask selection | 16 GB, one run | 29.1 s | 29.7 s | 48/48 frames; 0.3: 32.4 s |
| object | selection preview only (previewOnly) | 16 GB | — | ≈ 3.5 s | preview + mask PNG |
CPU-seconds per 1080p frame on that hardware: person ≈ 0.55 s, object ≈ 1.2 s (MobileSAM encoder ≈ 1.0–1.2 s, i.e. 80–90 % of the object budget on Apify's Ice Lake hosts — 0.4–0.6 s on a current AMX host — depending on the object's aspect ratio, edges ≈ 0.08 s, encode/composite ≈ 0.02 s). Per-stage timings are in cpuSecondsByStage. A long clip scales linearly: 4 workers track ≈ 3 frames/s together, so a 2 min 1080p clip is ≈ 15 min in one run, or ≈ 4 min split over four runs with startFrame / endFrame.
Speed levers, in order: more memory on the run (each 4 GB adds one worker), startFrame / endFrame fan-out across runs (each run gets its own host), x264Preset: "ultrafast" for large frames, maxFrames for tests.
Long videos: fan-out and stitch
Start N runs with consecutive startFrame / endFrame ranges (default sliceFormat: "ts"), download the .ts slices and concatenate:
ffmpeg -f concat -safe 0 -i list.txt -i source.mp4 -map 0:v -map 1:a? -c:v copy -c:a aac -movflags +faststart out.mp4
TS slices concatenate frame-exact with -c copy; MP4 slices can lose a frame at each seam, so only use sliceFormat: "mp4" when a slice must be playable on its own. For object, give every slice the same selection (objectBox/objectPoint/objectMaskUrl + selectionFrame, or the same objectPrompt); the pre-pass carries the selection into each slice's range as long as it lies in the same shot, and the vision model re-finds it otherwise.
Input reference
| Field | Type | Default | Notes |
|---|---|---|---|
videoUrl | string | required | direct MP4/MOV/WebM link |
target | person | object | person | |
outputMode | green | matte | demo | green | |
objectPrompt | string | text description for the vision model | |
objectBox | [x1,y1,x2,y2] | pixels or fractions | |
objectPoint | [x,y] | pixels or fractions | |
objectMaskUrl | string | white = object | |
selectionFrame | integer | 0 | frame the selection refers to |
previewOnly | boolean | false | segment the selection frame only |
visionApiKey / visionBaseUrl / visionModel / visionReasoning / maxAiCalls | see above | ||
startFrame / endFrame / sliceFormat | fan-out | ||
parallelRuns / parallelMemoryMbytes | integer | 1 / 8192 | object only: split the video over N parallel runs of this actor and stitch the slices automatically; cut points are chosen where the object is fully visible. Off by default — only faster than one 16 GB run when your plan can run several machines at once |
workers | integer | 0 = one per vCPU | |
crf / x264Preset | 17 / superfast | encoder | |
outputKey | string | output.mp4 | key in the run's key-value store; previews use -preview / -selection suffixes |
maxFrames | integer | 0 | benchmarking aid |
disableSmartCrop | boolean | false | person only, debugging |
demoLabel | string | caption for demo |
Error codes
The run still pushes one dataset row so a caller can read what happened; videoUrl is null and errorCode is one of:
errorCode | Meaning | Fix |
|---|---|---|
NO_INPUT / BAD_INPUT | missing videoUrl, invalid field, object without any selection | correct the input |
DOWNLOAD_FAILED / MASK_DOWNLOAD_FAILED | source or mask URL not reachable | make the URL public / direct |
OBJECT_NOT_FOUND | the box / point / mask / prompt gave no object on selectionFrame | pick another frame or a tighter selection; use previewOnly |
AI_UNAVAILABLE | objectPrompt without a vision model | set visionApiKey or select with box / point / mask |
AI_UNAUTHORIZED / AI_PAYMENT_REQUIRED / AI_FAILED | the vision provider rejected the key, has no credits, or answered garbage | check key, model id and base URL |
AI_BUDGET_EXCEEDED | maxAiCalls reached before the first object was found | raise the cap |
FFMPEG_FAILED / RENDER_FAILED | infrastructure failure (the run also fails) | retry; report with the run id |
Successful object runs can still carry warnings: NO_VISION_MODEL (object lost or cut without a way to re-find it), AI_BUDGET_EXCEEDED, AI_FAILED, OBJECT_NEVER_FOUND.
Models and licences
PP-MattingV2 (PaddleSeg, Apache-2.0) exported to static int8 ONNX and run with OpenVINO; MobileSAM (Apache-2.0) image encoder as int8 OpenVINO IR — one static canvas per aspect ratio (long side 1024, short side 512–1024) so a wide or tall object is not paid for as a padded square — and its multi-mask decoder in ONNX Runtime with a matching embedding grid; OpenCV DIS optical flow; FFmpeg / x264 for decode and encode. No frames leave the run except the single downscaled frame per vision call when objectPrompt (or re-finding after a cut) is used with your own key.
Changelog
- 0.6 — occlusion handling for
object: the tracker remembers the object (last whole area, amodal box, occluder negative points), keeps the visible part while something passes in front, and re-acquires the full object afterwards — gated by hue-saturation AND brightness histograms so a same-hue bystander is not absorbed; occluder pixels are cut out by colour pockets and a motion cut that only applies while the object itself stands still. Static scenes render up to 2× faster (unchanged regions are reused instead of recomputed). New opt-inparallelRuns: the run plans cut points where the object is fully visible, calls N parallel runs of this actor and stitches the slices — same masks, shorter wait when the plan allows several machines at once. - 0.2 —
target: "object": any object by prompt / box / point / painted mask, MobileSAM tracking with scene-cut re-seeding, selection preview, vision provider settings and call cap, machine-readable object metrics and warnings. - 0.4 — object pipeline scheduling: the pre-pass no longer blocks the workers (one decode feeds both; worker ranges are released as the chain advances and split by measured per-frame cost, so frames without the object are nearly free), idle workers take the tail of the slowest range at a seeded frame, and the per-frame mask/edge work runs in the object's window only. 10 s dog on Apify 93 s → 85 s (encoder-bound there; −26 % on a 4-core AMX host), a 10 s cat clip −32 % with 9 more frames of the cat and clean seams where 0.3 had picked up a second animal at a range start; masks otherwise unchanged (frame agreement > 0.99).
- 0.3 — object render 163 s → 93 s on the same 10 s clip and host: aspect-ratio encoder canvases (encoder CPU −52%), mask work confined to the object's window, windowed optical flow, worker ranges balanced against the pre-pass; masks unchanged (frame agreement > 0.99).
- 0.1 — person matting (PP-MattingV2 int8, smart crop, fan-out slices).