Video Render Engine - Remotion & HTML/CSS to MP4 avatar

Video Render Engine - Remotion & HTML/CSS to MP4

Pricing

Pay per usage

Go to Apify Store
Video Render Engine - Remotion & HTML/CSS to MP4

Video Render Engine - Remotion & HTML/CSS to MP4

Render Remotion / React / HTML / CSS / SVG / Canvas compositions to MP4, WebM, MOV or GIF in seconds. Parallel frame sharding across cores and containers, seamless stitching, voiceover and ducked music mixed in. 900 frames of 1080x1920 in 13-28 s. Webhook and standby mode.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

Andrew Babo

Andrew Babo

Maintained by Community

Actor stats

0

Bookmarked

63

Total users

12

Monthly active users

6 days ago

Last modified

Share

Hyper Video Render Engine — Remotion & HTML/CSS to MP4 at high speed

Turn a Remotion project (React, HTML, CSS, SVG, Canvas, WebGL) into a finished MP4 / WebM / MOV / GIF in seconds, on demand, through an API — without keeping a render farm alive.

The engine splits the composition into frame shards, renders them in parallel across many Chromium instances, and stitches the pieces with Remotion's own seamless concat: no re-encode, no duplicated or dropped frames at the seams, no drift between picture and sound. Voiceover and background music are mixed in one final FFmpeg pass, so the browser never has to deal with audio.


Why it is fast

StageWhat most render setups doWhat this engine does
WorkspaceWrite frames to the container diskEverything lives on the RAM disk (/dev/shm) — no disk I/O at all
ParallelismOne browser, more tabsOne browser per shard, so the work actually uses every CPU core
Chunk formatMP4 pieces that need a new headerMPEG-TS pieces that concatenate with a stream copy in ~50 ms
AudioRendered inside the browserMuxed at the end with FFmpeg — under a second, always in sync
Cold startDownload Chromium and the bundle on every runChromium is baked into the image; bundles are cached by content hash
Warm callsNew container every timeStandby mode keeps browsers and unpacked bundles warm between requests

On Apify you get one CPU core per 4 GB of memory, so memory is the speed dial:

MemoryCoresParallel renderersUse it for
8 GB22cheapest, short clips
16 GB (default)44balanced, or the orchestrator of a distributed run
32 GB88fastest single container
mode: "distributed"orchestrator + N workers16–64long videos and heavy templates

Measured performance

Real runs on Apify, 1080x1920 at 30 fps, h264 CRF 18, chromiumGl: "angle" — no cherry-picking, these are the numbers the Actor reports in its own output row.

VideoSetupRenderTotalThroughput
30 s (900 frames), light template32 GB, one container12.8 s13.4 s67 fps
30 s (900 frames), heavy template¹32 GB, one container35.2 s51.9 s17 fps
30 s (900 frames), heavy template¹16 GB + 6 x 8 GB workers26.6 s28.3 s32 fps
30 s, heavy + voiceover + ducked music16 GB + 6 x 8 GB workers39.6 s44.7 s20 fps
2 min (3600 frames), light template32 GB, one container59.1 s60.1 s60 fps
2 min (3600 frames), light template16 GB + 6 x 8 GB workers39.9 s40.7 s89 fps
10 s (300 frames), transparent VP9²16 GB, one container29.4 s37.2 s10 fps
10 s (300 frames), transparent VP9 + music²16 GB, one container28.8 s38.5 s10 fps

¹ The heavy template is deliberately brutal: full-screen backdrop-filter blur, an animated conic gradient, a 120-particle canvas and animated SVG dashes on every single frame. Ordinary marketing and social templates behave like the light one.

² Transparency costs frame rate everywhere: lossless PNG frames and an alpha-capable intermediate are unavoidable. The point of the number is what it replaces — the same clip encoded the standard way spent 6.7 minutes in libvpx alone. See Transparent output.

Every output above was verified with ffprobe: exactly 900 / 3600 frames, 30.000 s, audio track the same length as the picture — no dropped, duplicated or drifting frames at the shard seams. The transparent runs were decoded frame by frame: 94% of every frame is fully transparent, alpha intact at the shard seams.

Cost is the memory-time Apify bills: roughly $0.10–0.13 per 30 s render at the 64 GB tier, under $0.01 for a light 30 s clip in a single container.


Quick start

Run it with no input at all to render the built-in demo composition and see the output format.

1. Build your Remotion bundle

npx remotion bundle # produces the "build" folder (index.html + JS + your public/ assets)
cd build && zip -r ../bundle.zip .

Upload bundle.zip anywhere with a direct download link (S3, R2, GitHub release, Apify key-value store).

2. Run the Actor

{
"bundleUrl": "https://example.com/bundle.zip",
"compositionId": "MainVideo",
"inputProps": { "title": "Hello world", "brandColor": "#ff6a00" },
"audioUrl": "https://example.com/voiceover.mp3",
"bgmUrl": "https://example.com/music.mp3",
"bgmVolume": 0.15,
"videoConfig": {
"width": 1080,
"height": 1920,
"fps": 30,
"durationInFrames": 900,
"format": "mp4",
"crf": 18
},
"sharding": { "mode": "auto" },
"webhookUrl": "https://example.com/api/render-webhook"
}

3. Collect the result

The finished file lands in the run's key-value store under OUTPUT and the dataset gets one report row:

{
"status": "SUCCEEDED",
"videoUrl": "https://api.apify.com/v2/key-value-stores/<STORE_ID>/records/OUTPUT.mp4",
"compositionId": "MainVideo",
"width": 1080,
"height": 1920,
"fps": 30,
"totalFrames": 900,
"durationSeconds": 30,
"renderTimeMs": 28329,
"effectiveSpeed": "31.8 fps",
"shards": 16,
"localShards": 4,
"workerRuns": 6,
"shardingMode": "distributed",
"fileSizeBytes": 24313344,
"computeUnits": 0.428,
"costUsd": 0.107,
"warnings": [],
"errorCode": null
}

If a webhookUrl is set, that exact JSON is POSTed to it the moment the video is ready.


Three ways to supply the bundle

FieldUse it when
bundleUrlYou have a .zip / .tar.gz of the bundle on any HTTP host. Downloaded once, then cached by content hash.
serveUrlThe bundle is already hosted (S3 website, Cloudflare R2, npx remotion deploy). Nothing is downloaded.
bundleKeyValueStoreKeyAnother Actor produced the bundle and left it in a key-value store.

Leave all three empty and the built-in demo composition is rendered — useful for a smoke test.


Input reference

Composition

FieldDefaultMeaning
compositionIdfirst composition foundWhich composition to render
inputProps{}Props passed to your React component
envVariables{}Values available as process.env inside the composition

videoConfig

FieldDefaultMeaning
width, heightfrom the compositionOutput resolution
fpsfrom the compositionFrame rate
durationInFramesfrom the compositionLength in frames
frameRangefull video[from, to] to render a slice only
formatmp4mp4, webm, mov, mkv, gif
codecfrom the formath264, h265, vp8, vp9, av1, prores, gif
crfcodec defaultQuality, lower is better (18 is visually lossless for H.264)
presetfasterx264 speed/size trade-off
pixelFormatyuv420pyuva420p + codec vp8, vp9 or prores gives a transparent background. PNG frames and the ProRes 4444 profile are switched on automatically; H.264/H.265 cannot carry alpha, so the request falls back to opaque
proResProfilehq (4444 when transparent)4444-xq, 4444, hq, standard, light, proxy
videoBitrate, audioBitrate— / 192kFixed bitrates instead of CRF
audioCodeccontainer defaultaac, mp3, opus, pcm-16
scale1Render at a multiple of the composition size
mutedfalseDrop the composition's own audio
imageFormat, jpegQualityjpeg, 90Intermediate frame format

Transparent output

{ "compositionId": "Overlay", "videoConfig": { "format": "webm", "codec": "vp9", "pixelFormat": "yuva420p" } }

Transparent WebM takes a detour that keeps it fast. Handed straight to libvpx at its stock speed setting, an alpha channel encodes at well under one frame per second — a four second 1080x1920 clip took 6.7 minutes in testing. So the shards are rendered as ProRes 4444 instead, converted to VP8/VP9 in parallel with libvpx tuned for speed, and stitched with a stream copy: the same clip now finishes in seconds. ProRes 4444 (format: "mov") keeps alpha natively and needs no conversion.

Two things to know: transparent jobs always render in one container (the intermediate frames are too heavy to ship between machines), and most players and ffprobe report an alpha WebM as yuv420p — decode it with

ffmpeg -c:v libvpx-vp9
or drop it on a web page to see the transparency.

sharding

FieldDefaultMeaning
workersone per CPU coreParallel renderers in local mode; number of extra containers in distributed mode
concurrencyPerWorker1Tabs inside each renderer (raise only for very light compositions)
modeautolocal = one container, distributed = many containers, auto = switch at the threshold
distributedThresholdFrames4000When auto fans out to several containers
workerMemoryMbytes8192Memory (and therefore cores) per worker container — 8 GB buys 2 cores
masterRenderstrueThe orchestrator renders its own share instead of idling while the workers work

Audio

FieldDefaultMeaning
audioUrl—Voiceover or main audio track
audioVolume1Level of that track
bgmUrl—Background music
bgmVolume0.15Level of the music bed
bgmDuckingtrueMusic dips automatically while the voice speaks
bgmLooptrueShort music beds repeat to the end of the video

Both files are muxed onto the finished video with -c:v copy, so the picture is never re-encoded and the audio cannot drift.

Delivery and advanced

FieldDefaultMeaning
outputKeyOUTPUTName of the record in the key-value store
webhookUrl, webhookHeaders—Where to POST the report when the render finishes
remotionLicenseKey—Your Remotion company licence key, if you need one
chromiumGlangleGraphics backend. angle measured 8x faster than swangle here; if a machine refuses it the render falls back automatically and says so in warnings
shardTimeoutSecs600A silent shard is abandoned and retried on a fresh browser
failOnRenderErrorfalseTurn on if your pipeline watches run status instead of the output row
verbosefalseFull Remotion renderer log

Standby mode — near-zero cold start

Enable Standby on the Actor and the container stays alive with warm Chromium instances and cached bundles. Then render over plain HTTP:

curl -X POST "https://<your-actor>.apify.actor?token=<APIFY_TOKEN>" \
-H 'Content-Type: application/json' \
-d '{"bundleUrl":"https://example.com/bundle.zip","compositionId":"MainVideo","inputProps":{"title":"Hi"}}'

The response is the same report JSON. A GET on the same URL returns a readiness check. Warm requests skip the bundle download, the browser launch and the composition resolution — typically a third of the total time on short videos.


Distributed mode — for long videos

With "sharding": { "mode": "distributed" } (or automatically past distributedThresholdFrames), the run becomes an orchestrator: it starts worker runs of the same Actor, each rendering its own frame range with its own CPUs, hands the finished chunks back through its own key-value store, and the orchestrator stitches everything. A worker that crashes or times out is retried once on its own, so one bad shard never fails the whole job.

The orchestrator does not sit idle: it claims one shard per core it owns and renders alongside the workers, so a 16 GB orchestrator is four extra renderers rather than a supervisor.

Two things to know before turning it on:

  • A worker container costs about 9 seconds to boot. Below roughly 900 frames of a light template, a single 32 GB container finishes before the workers have even started.
  • Cost scales with total memory-time, so a fan-out is worth it when the render is long or heavy enough that wall-clock time matters more than a couple of cents.

The whole render is reported as one number: computeUnits and costUsd in the output row already include every worker run, and workerRuns / localShards show how the work was split.


What renders correctly

  • Modern CSS: backdrop-filter, conic and radial gradients, CSS variables, filters, mix-blend-mode
  • Web fonts — Chromium waits for document.fonts.ready, so there is no flash of unstyled text on the first frames
  • SVG (including strokeDashoffset path animation), Canvas 2D and WebGL
  • Remotion spring() / interpolate() physics, identical across shard boundaries
  • Images and videos embedded in the bundle, plus any remote asset your composition fetches
  • Liberation, DejaVu, Noto (including CJK) and colour emoji fonts are installed in the image

Error codes

errorCodeWhat happenedWhat to do
NO_BUNDLENo source givenSet bundleUrl, serveUrl or bundleKeyValueStoreKey
BUNDLE_DOWNLOAD_FAILEDThe archive could not be fetchedCheck that the link is public and direct
BUNDLE_FORMAT / BUNDLE_INVALIDNot a Remotion bundle archiveZip the folder that contains index.html
BUNDLE_TOO_LARGEBigger than maxBundleSizeMbRaise the limit or trim the assets
COMPOSITION_NOT_FOUNDNo composition with that IDCheck the ID in your Root file
SHARD_FAILEDA frame range failed three timesUsually a runtime error in the composition; run with verbose: true
COMBINE_FAILEDStitching produced an empty fileReport it — include the run ID
AUDIO_DOWNLOAD_FAILED / AUDIO_MUX_FAILEDAudio could not be fetched or mixedCheck the audio links
WORKER_FAILEDA distributed worker run did not succeedLook at the worker run log, or use mode: "local"

A failed render still finishes as SUCCEEDED with an explanatory row, so scheduled pipelines stay readable. Set failOnRenderError: true to flip that.


Licensing note

Remotion itself is free for individuals, non-profits and companies with up to three people. Larger companies need a Remotion company licence — put your key in remotionLicenseKey. This Actor does not include a Remotion licence.


🤖 AI Agent / MCP Skill

Role & Trigger — Use this Actor whenever the task is "render / produce / export a video" from a React, HTML or Remotion template, or to re-render the same template with different text, colours or data. Do not use it to edit an existing video file, to scrape videos, or to generate footage with AI.

Parameter Rules

  1. A source is required unless the user wants the demo: bundleUrl (zip link), serveUrl (hosted bundle) or bundleKeyValueStoreKey. Never invent a bundle URL — ask for it.
  2. compositionId must match a composition in that bundle. If you do not know it, omit it: the first composition is rendered and its ID is reported back.
  3. Put all dynamic content in inputProps. Never patch text by rebuilding the bundle.
  4. Vertical social video: {"width":1080,"height":1920,"fps":30}. Landscape: {"width":1920,"height":1080,"fps":30}. Transparency needs format: "webm" with pixelFormat: "yuva420p".
  5. Leave sharding alone unless the user asks for speed. For speed, raise the run's memory first (each 4 GB adds a core); only use mode: "distributed" above ~2 minutes of video.
  6. For batch jobs from one template, enable Standby and POST each job — the bundle is unpacked once and reused.

Output Interpretation

  • status: "SUCCEEDED" → the video is at videoUrl; report durationSeconds, renderTimeMs and effectiveSpeed if asked about performance.
  • status: "FAILED" → read errorCode and error and fix the input; do not retry the same input unchanged.
  • warnings is never fatal: it lists things that were ignored (for example audio on a container that cannot carry it).
  • Every run writes exactly one report row plus a RUN_SUMMARY record with the full timing breakdown (bundleMs, renderMs, combineMs, audioMs, uploadMs).