Video Background Remover
Pricing
Pay per usage
Go to Apify Store
Video Background Remover
Pricing
Pay per usage
Direct link to the source video (MP4/MOV/WebM, H.264/H.265/VP9). Must be downloadable without login. Output keeps the source resolution, frame rate and audio.
person: the main person in frame (human matting model, no selection needed, fastest). object: any single thing — a dog, a cat, a product, a car, a logo — chosen with objectPrompt, objectBox, objectPoint or objectMaskUrl and tracked through the video.
green: subject on pure green (#00FF00) for chroma keying. matte: grayscale alpha matte video (white = subject) to use as a luma/track matte. demo: subject composited on a dark studio background for a quick visual check.
target=object only. Plain-text description of the thing to cut out, e.g. "the golden retriever", "the red coffee mug on the table", "the white car". A vision model draws the box on the selection frame and again after every scene cut. Needs visionApiKey (or the actor owner's MATTE_AI_* environment). Ignored when objectBox / objectPoint / objectMaskUrl is given, except as a hint for re-finding the object after cuts.
target=object only. Rectangle around the object on the selection frame, in pixels of the source video (or fractions 0-1 of width/height). Example: [448, 311, 1221, 902]. No vision model needed.
target=object only. One pixel inside the object on the selection frame (pixels, or fractions 0-1). Example: [947, 595]. No vision model needed.
target=object only. URL of a PNG/JPG the same size as the video (or any size, it is scaled) where the object is painted white / opaque and the rest is black / transparent — a brush stroke over the object is enough, it is snapped to the object edges. Takes priority over objectBox, objectPoint and objectPrompt.
target=object only. 0-based frame index the box / point / mask / prompt refers to. Pick a frame where the whole object is visible; tracking runs backwards and forwards from it inside that shot.
target=object only. Segment just the selection frame and return previewUrl (frame with the picked object tinted green + box) and previewMaskUrl. Takes a few seconds; use it to confirm the right object was picked before rendering.
Only for objectPrompt (and for re-finding a lost object after scene cuts). Key for an OpenAI-compatible chat-completions endpoint that accepts images: OpenAI, OpenRouter, Google Gemini (OpenAI-compatible endpoint), etc. Stored encrypted, never logged. A run makes 1 call per scene cut plus 1 per worker start after a cut (typically 1–10 small calls, ~1.2k tokens each).
Base URL of the OpenAI-compatible API. Examples: https://api.openai.com/v1 (default), https://openrouter.ai/api/v1, https://generativelanguage.googleapis.com/v1beta/openai.
Model id that accepts images and returns JSON. Examples: gpt-4.1-mini (default, OpenAI), gpt-5-mini, google/gemini-2.5-flash (OpenRouter), gemini-2.5-flash (Gemini endpoint).
Hard cap on vision model requests in one run (cost guard). When reached, the object is simply not re-found after further cuts and the row carries the AI_BUDGET_EXCEEDED warning.
First frame to render (0-based). Use with endFrame to split one long video across several runs and stitch the slices afterwards.
Render frames up to, but not including, this index. Leave empty for the rest of the video.
Parallel matting processes inside this run. 0 = one per available vCPU (recommended).
Split the video over this many parallel runs of the actor and stitch the slices - same masks, much shorter wait. Off by default (1 = one run); only worth it when your plan can run several machines at once (total memory above 16 GB). Only for target=object without startFrame/endFrame.
Memory (MB) for each parallel run; 8192 gives 4 vCPUs. Only used when parallelRuns applies.
Lower is better quality and bigger files. 17 is visually lossless.
Encoder speed/size trade-off. superfast keeps encoding under 10% of the render time.
Only with startFrame/endFrame. ts (default) slices stitch frame-exact with ffmpeg -f concat -c copy; mp4 slices are playable on their own but can lose a frame at each seam when concatenated.
Key of the rendered video in the run's key-value store. Previews and the selection snapshot use the same name with -preview / -selection suffixes.
target=person only. Always run the model on the whole frame instead of a window around the person. Slower; only for debugging.
Stop after this many frames (benchmarking aid). 0 = whole video.
Caption drawn on the demo composite background.