Video Render Engine: Timeline JSON to MP4
Pricing
Pay per usage
Video Render Engine: Timeline JSON to MP4
Render a JSON edit timeline into an MP4: layered clips, crop and pan keyframes, captions, transitions and audio mixing, rendered headlessly and joined with ffmpeg. Scales across machines.
Pricing
Pay per usage
Rating
0.0
(0)
Developer
Andrew Babo
Maintained by CommunityActor stats
0
Bookmarked
792
Total users
418
Monthly active users
an hour ago
Last modified
Categories
Share
Video Render Engine — Timeline JSON to MP4
Send a JSON edit timeline, get back a finished MP4. Layered clips, crop and pan keyframes, captions, transitions and audio mixing are rendered frame by frame in a headless browser, then encoded and joined with ffmpeg — so the output matches what a browser preview shows, pixel for pixel.
Use it for: automated short-form video, subtitle burn-in, vertical 9:16 reframing at scale, templated social clips, batch rendering from a CMS or an AI pipeline.
- Declarative JSON in, MP4 out — no timeline software, no GPU, no local ffmpeg
- Crop / pan / zoom keyframes with linear or hold interpolation
- Burned-in captions with word timing
- Audio mixing: voice gain, source gain, background music, loudness normalisation
- Long renders are split into shards and rendered in parallel, then joined losslessly
Quick start
{"mode": "editplan","source_url": "https://example.com/master.mp4","edit_plan": {"canvas": { "width": 1080, "height": 1920, "fps": 30 },"tracks": [{"type": "media","clips": [{"source": "https://example.com/master.mp4","start_ms": 0,"end_ms": 61400,"layout": "single-center","crop": {"keyframes": [{ "t_ms": 0, "rect": { "x": 0.31, "y": 0, "w": 0.316, "h": 1 }, "interp": "linear" },{ "t_ms": 61400, "rect": { "x": 0.36, "y": 0, "w": 0.316, "h": 1 }, "interp": "linear" }]}}]}]}}
Feature-detect the build (free, a few seconds, needs no plan):
{ "mode": "capabilities" }
Input
| Field | Notes |
|---|---|
mode | editplan (default) or capabilities |
edit_plan | the timeline document, inline |
edit_plan_url | URL of the timeline JSON, when it is too big to inline |
source_url | shorthand when the plan has exactly one source |
source_map | { "<sourceId>": "https://…" } for multi-source plans |
video | { width, height, fps, crf, preset } — defaults come from the plan canvas |
audio | { voice_gain_db, source_gain, bgm_url, bgm_gain_db, loudnorm } |
options | { workers, threads, profile, preset, crf, min_shard_sec, captions, maxAssetBytes } |
output | { signed_upload_url, fallback_kv_key } |
callback | webhook called with the result JSON when the run finishes |
Timeline basics
canvas— output size and frame rate; everything else is expressed relative to it.tracks[]— layered top to bottom; media, caption and overlay tracks.clips[]—source,start_ms,end_ms,layout, and optionalcrop,transform,opacity,transition,filters.crop.keyframes[]—rectin normalised 0–1 coordinates, so the same plan renders at any resolution. This is exactly the format the Face Detection & Auto Reframe actor produces, so auto-reframe output can be pasted in directly.
Output
{"status": "success","artifacts": [{ "name": "video", "kv_key": "output.mp4", "url": "https://api.apify.com/v2/key-value-stores/.../output.mp4", "bytes": 18442310 }],"meta": {"width": 1080, "height": 1920, "fps": 30,"duration_sec": 61.4,"frame_count": 1842,"shards": 4,"render_sec": 96.3}}
Set output.signed_upload_url to have the MP4 PUT straight into your own
storage bucket instead of staying on Apify.
Error handling
{ "ok": false, "reason": "BAD_INPUT", "message": "..." }
reason | Meaning | What to do |
|---|---|---|
BAD_INPUT | invalid plan, unknown source id, bad keyframes | validate the plan; check every source resolves |
UPSTREAM_BLOCKED | a source URL could not be fetched | host the media somewhere publicly reachable |
TIMEOUT | render exceeded the run timeout | raise timeoutSecs, or lower resolution / fps |
OOM_LIMIT | not enough memory | run with 16 GB |
INTERNAL | unexpected failure | retry; report the run ID |
Performance
16 GB run (≈4 vCPU), 1080×1920 at 30 fps:
| Output length | Typical time |
|---|---|
| 15 s | ~25–40 s |
| 60 s | ~1.5–3 min |
| 5 min | ~8–15 min |
Rendering is split into shards across the available vCPUs (options.workers)
and the shards are concatenated with a stream copy, so there is no second
encode. 16 GB and a run timeout of 2 hours are the recommended settings.
Pipeline example
- Video Downloader — fetch the source MP4 from a page URL.
- Video & Audio Toolkit — build a 480p proxy and a 16 kHz mono audio track.
- Speech to Text (Whisper) — get word timestamps for the captions.
- Face Detection & Auto Reframe — get 9:16 crop keyframes that follow the speaker.
- Video Render Engine — assemble the plan and render the final vertical MP4.
FAQ
Do I need ffmpeg or an editor installed? No. Everything happens in the actor.
Why a headless browser? So the rendered frames use the same drawing rules as a browser-based preview — one implementation instead of two that drift apart.
Can I burn in subtitles? Yes — add a caption track with word timings, or
enable options.captions.
Can I add background music? Yes — audio.bgm_url plus bgm_gain_db, with
optional loudnorm for consistent loudness.
Is the output web-ready? Yes — H.264/AAC MP4 with faststart.
Building from source
src/generated/ and src/vendor/ are generated; do not edit them by hand.
Regenerate from the repository root with:
$bun scripts/build-actor-frames.mjs