AI Viral Clip Generator – Long Video to Shorts
Pricing
from $50.00 / 1,000 rendered clip minutes
AI Viral Clip Generator – Long Video to Shorts
Turn podcasts, interviews, webinars, and long videos into distinct, non-overlapping 9:16 clips. AI selects strong moments, follows the active speaker, adds or translates subtitles in 10 languages, validates every MP4, and returns SRT, VTT, metadata, and ZIP.
Pricing
from $50.00 / 1,000 rendered clip minutes
Rating
0.0
(0)
Developer
François Fernandez
Maintained by CommunityActor stats
0
Bookmarked
5
Total users
2
Monthly active users
10 days ago
Last modified
Categories
Share
Turn one podcast, interview or webinar into several distinct vertical clips — each one cut on speech boundaries, reframed 9:16 around the active speaker, and subtitled or translated.
A 60-minute podcast → 3 clips, 720×1280, no translation: about $1.95. Most of that is the source, not the clips: transcription and analysis are billed per source minute, rendering per delivered clip minute. A 10-minute source with 2 short clips costs about $0.40. Full breakdown in Pricing.
Your first run, in four stepssvgsvg
- Input tab → Video file → Upload new files. Alternatively, paste one publicly reachable direct media-file URL under Direct video URL. Do not fill both source fields.
- Tick I own this content or am authorised to transform it.
- For a first run, use these recommended settings: 3 clips, 30–60 seconds, 9:16 at 720×1280, face tracking with centre-crop fallback, original language,
maxCostUsd: 5. - Press Start.
The restorable example intentionally contains no media. Add your own video before pressing Start. Starting with no video performs a free readiness check and therefore produces no clips.
▶ Try it now — use the recommended first-run settings
Choose the exact subtitle resultsvgsvg
These options are independent. burnSubtitles controls visible text inside the MP4. generateSrt and generateVtt control separate subtitle files.
Desired result**outputLanguageoutputLanguageModeburnSubtitlesgenerateSrt**** / ******generateVtt | ||||
|---|---|---|---|---|
| Video only, no subtitle artifacts | empty | original | false | both false |
| Original-language subtitles | empty | original | as desired | as desired |
| Translated subtitles only | choose a target | translations | as desired | as desired |
| Original + translation | choose a target | both | as desired | as desired |
In both mode, two subtitle levels in the video are intentional: the translation is the primary lower track and the original is a smaller secondary track above it. Choose translations, not both, when you want French-only, Spanish-only, or another single translated language.
Example — French translated subtitles, burned in, plus an SRT file:
undefined
undefined
svg
undefined
{
svg
"outputLanguage": "FR",
"outputLanguageMode": "translations",
"burnSubtitles": true,
"generateSrt": true,
"generateVtt": false
}
svg
Example — clean video, no subtitle artifacts at all:
undefined
undefined
svg
undefined
{
svg
"outputLanguageMode": "original",
"burnSubtitles": false,
"generateSrt": false,
"generateVtt": false
}
svg
If subtitles are already burned into the source image, the Actor cannot remove them. Burning a new track on top can look like duplicate subtitles. Embedded text or bitmap subtitle streams are not used as transcription input: this version transcribes the audio.
Inputssvgsvg
| FieldMeaning | |
|---|---|
videoFile | One real file upload. In JSON it is an array containing the URL created by Apify's upload widget. Do not add characters to that URL manually. |
videoUrl | One publicly reachable direct media-file URL. It must return the video bytes, not an HTML page. |
confirmContentRights | Required. Processing does not start until this is true. |
clipCount | 1 to 10; default 3. It is a target, not permission to duplicate material. |
targetClipDuration | 15-30, 30-60, or 60-90 seconds. The Actor may adjust a boundary to preserve a complete sentence. |
platform | TikTok, Instagram Reels, YouTube Shorts, or Generic vertical. This affects suggested metadata only. JSON values are case-sensitive. |
sourceLanguage | Leave empty for automatic detection. If set through JSON, use a language code such as en, fr, es, de, it, pt, ar, zh, ja, or ko. |
outputLanguage | Optional translation target: EN-US, ES, FR, DE, IT, PT-BR, AR, ZH-HANS, JA, or KO. |
outputLanguageMode | original, translations, or both. A translated mode requires outputLanguage. |
aspectRatio | 9:16. Other ratios are refused rather than silently rendered incorrectly. |
resolution | 720x1280 or 1080x1920. HD has a higher render price. |
cropMode | face_tracking, center_crop, or blurred_background. |
fallbackCropMode | center_crop or blurred_background, used when tracking has no reliable subject. |
burnSubtitles | Whether subtitles are visibly rendered into the MP4. |
subtitleStyle | Clean, Bold, or Minimal; relevant only when subtitles are burned. |
generateSrt, generateVtt | Whether separate subtitle files are produced. |
createZip | Whether all produced artifacts are also packaged into ai-viral-clips.zip. |
maxCostUsd | Internal cost ceiling checked before any paid provider call. Minimum $1. |
minimumViralScore | Optional editorial-score threshold. A high value can reduce the number of delivered clips. |
avoidOverlaps | Keep enabled to prevent overlapping or reused source passages. |
maxMediaMinutes | Refuse longer sources. Default 90 minutes; maximum 180. |
The form supplies the allowed values. When using JSON, copy their spelling and capitalisation exactly. For example, use "platform": "TikTok", not "tiktok".
How the clips remain differentsvgsvg
The analysis model proposes candidate passages from the timed transcript. The Actor then validates the candidates independently of the model:
- source intervals cannot significantly overlap;
- textual duplicates and near-duplicates are rejected;
- neighbouring slices of the same exchange are rejected;
- delivered clips normally keep a 20-second editorial gap;
- cuts are moved to word, phrase and sentence boundaries where possible;
- if the model concentrates all proposals in one section, a deterministic pass searches other complete transcript regions.
clipCount: 3 means "find up to three honest clips," not "produce three files at any cost." A short or repetitive source may correctly return only one or two clips. Reducing the target duration may reveal more distinct passages; turning off overlap protection is not recommended.
What the viral score meanssvgsvg
The 0–100 score is an editorial estimate based on hook strength, standalone quality, clarity, informative value, curiosity, emotional intensity, pacing and conclusion, with penalties for context dependency and artificial cuts. The full breakdown is returned with every clip.
It is not a performance prediction or a promise that a clip will go viral.
Framing and active-speaker trackingsvgsvg
| ModeBehaviourBest for | ||
|---|---|---|
face_tracking | Detects persistent faces, compares visible speech with the audio, follows the likely active speaker, and smooths the virtual camera. | Interviews, podcasts, panels and talking-head footage. |
center_crop | Uses a fixed vertical crop through the centre of the landscape image. | A subject that remains centred, or a completely static result. |
blurred_background | Preserves the full landscape frame over a blurred vertical background. | Slides, demonstrations, wide group shots or footage where cropping would remove important content. |
Face tracking does not simply choose the largest face. YuNet detects faces and Light-ASD helps associate visible speech with the audio. The track is then median-filtered, dead-zoned, velocity-limited and smoothed. A speaker change must remain dominant before the crop switches, and scene cuts are reacquired without a slow pan.
When active-speaker evidence is unavailable, the Actor may retain a stable visual face track. If no face is reliable, it uses fallbackCropMode. During B-roll or voice-over, there may be no visible person who corresponds to the voice; in that situation no system can infer the narrator from lip movement.
Profiles, covered mouths, heavy backlight, small faces, overlapping speech and crowded scenes remain difficult. For footage where every part of the wide frame matters, choose blurred_background instead of tracking.
"Stabilised face tracking" means stabilised crop coordinates. It does not apply a general camera-shake stabiliser to the source video.
Languages and translationsvgsvg
Leave Input language empty unless automatic detection is wrong. The Actor normalises common two-letter and three-letter source codes before translation; you do not need to know the translation provider's internal code format.
Validated output languages:
| LanguageCode | |
|---|---|
| English (US) | EN-US |
| Spanish | ES |
| French | FR |
| German | DE |
| Italian | IT |
| Portuguese (Brazil) | PT-BR |
| Arabic | AR |
| Simplified Chinese | ZH-HANS |
| Japanese | JA |
| Korean | KO |
Input detection covers more languages, but only these ten have been validated end to end for translation, wrapping, direction and rendering. Translation is machine-produced and marked machine-translated-unreviewed. When source and target are the same language, including EN to EN-US, no translation is performed or charged.
Subtitle behavioursvgsvg
All ten output languages have validated language profiles. English, French, Spanish, German, Italian, Portuguese and Arabic avoid awkward line endings on frequent grammatical function words. Simplified Chinese and Japanese use punctuation-aware kinsoku boundaries. Korean prefers spaces and safely falls back to Hangul-syllable boundaries. Arabic uses right-to-left shaping through libass. Unicode grapheme measurement prevents internal cuts in combining marks and emoji sequences.
Subtitles use at most two lines and remain inside a phone-safe region. The reading-speed target is 23 characters per second, with script-specific limits where appropriate. When fast speech makes readability and synchronisation incompatible, synchronisation wins and the subtitle track is labelled needs_manual_review; text is not silently removed or desynchronised.
There is no word-by-word karaoke or dynamic highlight style. The transcription provider does not guarantee the word timings required for an honest result.
Outputs and statusessvgsvg
One Dataset row is created per selected clip. It includes the source interval, duration, score and breakdown, hook, title, description, hashtags, selectionMethod, cropMethod, language codes, compliance statuses, warnings, and links to files that actually exist.
The run's key-value store can contain:
- one MP4 and one metadata JSON per delivered clip;
- SRT and VTT tracks only when their toggles are enabled;
- original and translated tracks with language suffixes when both are requested;
manifest.json, including counters, billing reconciliation, warnings, artifact sizes and SHA-256 checksums;ai-viral-clips.zipwhencreateZipis enabled.
An absent artifact is represented by null, never by an empty fake link.
| StatusMeaning | |
|---|---|
success | Every selected clip was rendered and delivered. |
partial_success | At least one clip was delivered and at least one failed. Inspect each row. |
failed | No clip could be delivered. Inspect error and the logs. |
needs_manual_review | A subtitle track exists but violates a linguistic reading constraint, usually because speech is too fast. |
A subtitle warning does not mean the MP4 failed to decode. Technical media validation and linguistic subtitle review are separate facts.
Media validationsvgsvg
Every rendered clip is probed again before delivery. Validation checks include the required audio and video streams, dimensions, pixel format, duration, decodable beginning and end, and timestamp consistency. This catches damaged files that can report a plausible duration yet fail when their final frames are decoded.
The MP4 output uses H.264 video, AAC audio and yuv420p for broad social and mobile compatibility. Rendering re-encodes the cut so boundaries remain frame-accurate instead of drifting to the nearest source keyframe.
Pricingsvgsvg
What a real run costssvgsvg
| SourceOutputTotal | ||
|---|---|---|
| 10-minute recording | 2 clips, 30 s, 720p | about $0.40 |
| 20-minute interview | 3 clips, 60 s, 720p | about $0.75 |
| 60-minute podcast | 3 clips, 60 s, 720p | about $1.95 |
| 60-minute podcast | 5 clips, 60 s, 720p | about $2.05 |
| 60-minute podcast | 5 clips, 60 s, 1080p, translated | about $2.35 |
Transcription and analysis are charged per source minute, so the length of what you upload drives most of the cost. Rendering is charged per delivered clip, with every started minute rounded up independently. Source length still drives most of the total cost.
If you are on a free Apify plan, start with a 10–20 minute source.
▶ Start with a 10-minute file — about $0.40
Event tablesvgsvg
Pay per event. A started minute is a billed minute.
| EventPriceCharged when | ||
|---|---|---|
apify-actor-start | $0.00005 | Automatically charged by Apify when the Actor starts; never charged manually by the code. |
transcribed-minute | $0.02 | Per source minute, after a usable transcript is produced. |
analysed-minute | $0.01 | Per source minute when the model selects passages. |
fallback-analysed-minute | $0.01 | Per source minute when deterministic selection is used instead. |
translated-minute | $0.03 | Per delivered clip minute, only when translation is produced. |
rendered-clip-minute | $0.05 | Per validated clip minute at 720 × 1280. |
rendered-hd-clip-minute | $0.08 | Per validated clip minute at 1080 × 1920. |
Exactly one of the two analysis event types is used in a run. For each delivered clip, exactly one of the two render event types is used. Failed renders, rejected candidates and unnecessary translations are not charged.
Example for a one-hour source producing five 30–60 second clips:
- 720 × 1280, no translation: about $2.05;
- 720 × 1280, translated: about $2.14–$2.20;
- 1080 × 1920, no translation: about $2.20;
- 1080 × 1920, translated: about $2.29–$2.35.
Cost ceilingssvgsvg
maxCostUsd is the Actor's preflight ceiling. Apify's separate Maximum cost per run setting must also be high enough. If the Actor estimate exceeds maxCostUsd, the run stops with cost_limit_exceeded before a paid provider is contacted. If the Apify charge ceiling is exhausted later, the run reports billing_limit_reached and keeps already produced artifacts.
A complete JSON examplesvgsvg
For a direct file URL, three distinct French-subtitled clips:
undefined
undefined
svg
undefined
{
svg
"videoUrl": "https://example.com/interview.mp4",
"confirmContentRights": true,
"clipCount": 3,
"targetClipDuration": "15-30",
"platform": "TikTok",
"sourceLanguage": "en",
"outputLanguage": "FR",
"outputLanguageMode": "translations",
"aspectRatio": "9:16",
"resolution": "720x1280",
"cropMode": "face_tracking",
"fallbackCropMode": "center_crop",
"burnSubtitles": true,
"subtitleStyle": "Clean",
"generateSrt": true,
"generateVtt": true,
"createZip": true,
"maxCostUsd": 5,
"minimumViralScore": 0,
"avoidOverlaps": true,
"maxMediaMinutes": 90
}
svg
The videoFile value generated by the Console is normally an array containing one Apify storage URL. Use the upload widget instead of typing that structure by hand whenever possible.
Before you commit a long sourcesvgsvg
Read this before running a 90-minute source. Nothing here is a defect; these are the deliberate boundaries of an automatic first edit.
Important: this is an automatic first edit, not a human editorial or linguistic review. Watch every delivered clip before publishing it.
Sourcessvgsvg
- One uploaded MP4, MOV, MKV or WebM, or one direct URL to such a file.
- Platform page URLs — YouTube, TikTok, Instagram and similar — are refused. Upload a file you own, or provide a direct media-file URL you control.
- Source limits are 500 MB, 90 minutes by default and 180 minutes maximum.
- The source must contain both video and an audible speech track. Weak, distorted, overlapping or music-covered speech can reduce transcription, selection and active-speaker accuracy.
What is not guaranteedsvgsvg
- The requested clip count. The Actor returns fewer clips when the source does not contain enough distinct, complete passages.
- The viral score is an editorial estimate, not a performance prediction.
- Active-speaker tracking can fail on profiles, small or obscured faces, overlapping speakers, heavy backlight and crowded scenes.
- "Follow the main face" stabilises the virtual crop around a subject. It does not remove camera shake already present in the source footage.
- Automatic selection, transcription, metadata and translation all require human review before publication. Translation is machine-produced and unreviewed.
What this Actor does not dosvgsvg
- It does not publish to social networks. It produces files and metadata; you review and publish them yourself.
- Only vertical 9:16 output is rendered in this version, at 720 × 1280 or 1080 × 1920.
- Subtitle streams already embedded in the source are ignored; the audio is transcribed instead. Text already burned into the image cannot be removed.
- Subtitles are never compulsory: burned subtitles, SRT files and VTT files are three independent options.
- There is no automatic B-roll insertion, voice cloning, face modification, AI-generated imagery or word-by-word karaoke.
Troubleshootingsvgsvg
| Symptom or resultExplanation and action | |
|---|---|
| The run succeeds in a few seconds but returns no clips | No source was supplied, so the free readiness check ran. Upload a video and start a new run. |
| The upload field is missing | The selected Actor build does not expose this project's .actor/input_schema.json. Select the deployed build/version that contains the complete project. |
validation_error | Read the message and correct the named input. JSON enum values are case-sensitive. |
| A direct URL is refused | It is probably a platform page, private page or redirect rather than a direct media file. Upload the video instead. |
| Fewer clips than requested | The source lacks enough distinct passages under the current duration/score constraints. Try a shorter duration or a longer source. |
| Two subtitle levels appear | outputLanguageMode is both, or subtitles were already burned into the source. Choose translations for translated-only output. |
| No visible subtitles, but SRT/VTT files exist | This is expected when burnSubtitles is false but generateSrt or generateVtt is true. |
| The wrong person is framed during a conversation | Active-speaker evidence may be ambiguous. Profiles, overlapping speech and covered mouths are difficult. Try center_crop or blurred_background for that source. |
| The image shakes | Face tracking filters crop jitter but cannot remove shake present in the original camera footage. |
cost_limit_exceeded | Increase maxCostUsd, or reduce source length, clip count, clip duration or resolution. Nothing was charged. |
billing_limit_reached | Increase Apify's Maximum cost per run and start a new run. Already produced files remain available. |
configuration_error | A required provider is not configured on the Actor. Nothing was charged; report it to the Actor owner. |
Subtitle needs_manual_review | Speech was too fast for the configured reading limit. Review or edit the external subtitle before publication. |
Permissions and securitysvg
This Actor currently requires Full permissions because Apify's native file uploader stores uploaded media in a private Key-value store and provides the Actor with an Apify storage URL. The Actor uses that access only to fetch the uploaded source video required for the run.
The Actor does not intentionally enumerate or read unrelated Actors, datasets, Key-value stores, schedules, or other account data. If you use Direct video URL, the source must be a publicly reachable direct media-file URL.
Privacy and third-party processingsvgsvg
The media is downloaded into the run's container and temporary working files are deleted when the run ends. Audio is sent to the transcription provider; the timed transcript is sent to the selection provider; clip text is sent to the translation provider only when translation is requested. Face detection, active-speaker analysis, cropping, subtitle rendering and media validation run inside the Actor container.
Output files remain in the run's Apify storage according to your Apify storage and retention settings.
Copyright and content rightssvgsvg
You must own the source or be authorised to transform it. Confirmation is required before any download or paid processing.
The Actor does not download from YouTube, TikTok, Instagram or similar platform pages, and does not bypass authentication, access controls or DRM. Download or obtain the source lawfully, then upload the file.
Other Lumaxys Actorssvgsvg
This Actor is one stage of a larger media workflow. If you need a different stage, these cover it — you only pay for the processing you actually use.
- Video & Audio to Subtitles – Transcribe & Translatesvgsvg — create subtitle files from video or audio without producing short clips.
- Professional Subtitle QC – SRT/VTT Compliance Checkersvgsvg — check or safely correct existing SRT/VTT files.
- Professional Subtitle Translator & Video Captionersvgsvg — transcribe, translate, quality-check and burn subtitles into one complete source video.
- Subtitle Converter & Resync — SRT/VTT/ASSsvgsvg — convert subtitle formats or adjust their timing.
▶ Try AI Viral Clip Generator — upload a video, tick the rights box, press Start.
API callsvgsvg
undefined
undefined
svg
undefined
curl -X POST "https://api.apify.com/v2/acts/lumaxys~ai-viral-clip-generator/runs?token=$APIFY_TOKEN" \
svg
-H "Content-Type: application/json" \
-d '{
"videoUrl": "https://example.com/interview.mp4",
"confirmContentRights": true,
"clipCount": 3,
"targetClipDuration": "30-60",
"platform": "TikTok",
"resolution": "720x1280",
"cropMode": "face_tracking",
"outputLanguageMode": "original",
"maxCostUsd": 5
}'
svg
Dataset rows and file URLs are available in the run output. The full artifact inventory, validation results and billing reconciliation are recorded in manifest.json.
Built by Lumaxys.