AI Viral Clip Generator – Long Video to Shorts avatar

AI Viral Clip Generator – Long Video to Shorts

Pricing

from $50.00 / 1,000 rendered clip minutes

Go to Apify Store
AI Viral Clip Generator – Long Video to Shorts

AI Viral Clip Generator – Long Video to Shorts

Turn podcasts, interviews, webinars, and long videos into distinct, non-overlapping 9:16 clips. AI selects strong moments, follows the active speaker, adds or translates subtitles in 10 languages, validates every MP4, and returns SRT, VTT, metadata, and ZIP.

Pricing

from $50.00 / 1,000 rendered clip minutes

Rating

0.0

(0)

Developer

François Fernandez

François Fernandez

Maintained by Community

Actor stats

0

Bookmarked

5

Total users

2

Monthly active users

10 days ago

Last modified

Share

Turn one podcast, interview or webinar into several distinct vertical clips — each one cut on speech boundaries, reframed 9:16 around the active speaker, and subtitled or translated.

A 60-minute podcast → 3 clips, 720×1280, no translation: about $1.95. Most of that is the source, not the clips: transcription and analysis are billed per source minute, rendering per delivered clip minute. A 10-minute source with 2 short clips costs about $0.40. Full breakdown in Pricing.


Your first run, in four stepssvgsvg

  1. Input tab → Video file → Upload new files. Alternatively, paste one publicly reachable direct media-file URL under Direct video URL. Do not fill both source fields.
  2. Tick I own this content or am authorised to transform it.
  3. For a first run, use these recommended settings: 3 clips, 30–60 seconds, 9:16 at 720×1280, face tracking with centre-crop fallback, original language, maxCostUsd: 5.
  4. Press Start.

The restorable example intentionally contains no media. Add your own video before pressing Start. Starting with no video performs a free readiness check and therefore produces no clips.

▶ Try it now — use the recommended first-run settings


Choose the exact subtitle resultsvgsvg

These options are independent. burnSubtitles controls visible text inside the MP4. generateSrt and generateVtt control separate subtitle files.

Desired result**outputLanguageoutputLanguageModeburnSubtitlesgenerateSrt**** / ******generateVtt
Video only, no subtitle artifactsemptyoriginalfalseboth false
Original-language subtitlesemptyoriginalas desiredas desired
Translated subtitles onlychoose a targettranslationsas desiredas desired
Original + translationchoose a targetbothas desiredas desired

In both mode, two subtitle levels in the video are intentional: the translation is the primary lower track and the original is a smaller secondary track above it. Choose translations, not both, when you want French-only, Spanish-only, or another single translated language.

Example — French translated subtitles, burned in, plus an SRT file:

undefined
undefined

svg

undefined
{

svg

"outputLanguage": "FR",

"outputLanguageMode": "translations",

"burnSubtitles": true,

"generateSrt": true,

"generateVtt": false

}

svg

Example — clean video, no subtitle artifacts at all:

undefined
undefined

svg

undefined
{

svg

"outputLanguageMode": "original",

"burnSubtitles": false,

"generateSrt": false,

"generateVtt": false

}

svg

If subtitles are already burned into the source image, the Actor cannot remove them. Burning a new track on top can look like duplicate subtitles. Embedded text or bitmap subtitle streams are not used as transcription input: this version transcribes the audio.


Inputssvgsvg

FieldMeaning
videoFileOne real file upload. In JSON it is an array containing the URL created by Apify's upload widget. Do not add characters to that URL manually.
videoUrlOne publicly reachable direct media-file URL. It must return the video bytes, not an HTML page.
confirmContentRightsRequired. Processing does not start until this is true.
clipCount1 to 10; default 3. It is a target, not permission to duplicate material.
targetClipDuration15-3030-60, or 60-90 seconds. The Actor may adjust a boundary to preserve a complete sentence.
platformTikTokInstagram ReelsYouTube Shorts, or Generic vertical. This affects suggested metadata only. JSON values are case-sensitive.
sourceLanguageLeave empty for automatic detection. If set through JSON, use a language code such as enfresdeitptarzhja, or ko.
outputLanguageOptional translation target: EN-USESFRDEITPT-BRARZH-HANSJA, or KO.
outputLanguageModeoriginaltranslations, or both. A translated mode requires outputLanguage.
aspectRatio9:16. Other ratios are refused rather than silently rendered incorrectly.
resolution720x1280 or 1080x1920. HD has a higher render price.
cropModeface_trackingcenter_crop, or blurred_background.
fallbackCropModecenter_crop or blurred_background, used when tracking has no reliable subject.
burnSubtitlesWhether subtitles are visibly rendered into the MP4.
subtitleStyleCleanBold, or Minimal; relevant only when subtitles are burned.
generateSrtgenerateVttWhether separate subtitle files are produced.
createZipWhether all produced artifacts are also packaged into ai-viral-clips.zip.
maxCostUsdInternal cost ceiling checked before any paid provider call. Minimum $1.
minimumViralScoreOptional editorial-score threshold. A high value can reduce the number of delivered clips.
avoidOverlapsKeep enabled to prevent overlapping or reused source passages.
maxMediaMinutesRefuse longer sources. Default 90 minutes; maximum 180.

The form supplies the allowed values. When using JSON, copy their spelling and capitalisation exactly. For example, use "platform": "TikTok", not "tiktok".


How the clips remain differentsvgsvg

The analysis model proposes candidate passages from the timed transcript. The Actor then validates the candidates independently of the model:

  • source intervals cannot significantly overlap;
  • textual duplicates and near-duplicates are rejected;
  • neighbouring slices of the same exchange are rejected;
  • delivered clips normally keep a 20-second editorial gap;
  • cuts are moved to word, phrase and sentence boundaries where possible;
  • if the model concentrates all proposals in one section, a deterministic pass searches other complete transcript regions.

clipCount: 3 means "find up to three honest clips," not "produce three files at any cost." A short or repetitive source may correctly return only one or two clips. Reducing the target duration may reveal more distinct passages; turning off overlap protection is not recommended.


What the viral score meanssvgsvg

The 0–100 score is an editorial estimate based on hook strength, standalone quality, clarity, informative value, curiosity, emotional intensity, pacing and conclusion, with penalties for context dependency and artificial cuts. The full breakdown is returned with every clip.

It is not a performance prediction or a promise that a clip will go viral.


Framing and active-speaker trackingsvgsvg

ModeBehaviourBest for
face_trackingDetects persistent faces, compares visible speech with the audio, follows the likely active speaker, and smooths the virtual camera.Interviews, podcasts, panels and talking-head footage.
center_cropUses a fixed vertical crop through the centre of the landscape image.A subject that remains centred, or a completely static result.
blurred_backgroundPreserves the full landscape frame over a blurred vertical background.Slides, demonstrations, wide group shots or footage where cropping would remove important content.

Face tracking does not simply choose the largest face. YuNet detects faces and Light-ASD helps associate visible speech with the audio. The track is then median-filtered, dead-zoned, velocity-limited and smoothed. A speaker change must remain dominant before the crop switches, and scene cuts are reacquired without a slow pan.

When active-speaker evidence is unavailable, the Actor may retain a stable visual face track. If no face is reliable, it uses fallbackCropMode. During B-roll or voice-over, there may be no visible person who corresponds to the voice; in that situation no system can infer the narrator from lip movement.

Profiles, covered mouths, heavy backlight, small faces, overlapping speech and crowded scenes remain difficult. For footage where every part of the wide frame matters, choose blurred_background instead of tracking.

"Stabilised face tracking" means stabilised crop coordinates. It does not apply a general camera-shake stabiliser to the source video.


Languages and translationsvgsvg

Leave Input language empty unless automatic detection is wrong. The Actor normalises common two-letter and three-letter source codes before translation; you do not need to know the translation provider's internal code format.

Validated output languages:

LanguageCode
English (US)EN-US
SpanishES
FrenchFR
GermanDE
ItalianIT
Portuguese (Brazil)PT-BR
ArabicAR
Simplified ChineseZH-HANS
JapaneseJA
KoreanKO

Input detection covers more languages, but only these ten have been validated end to end for translation, wrapping, direction and rendering. Translation is machine-produced and marked machine-translated-unreviewed. When source and target are the same language, including EN to EN-US, no translation is performed or charged.


Subtitle behavioursvgsvg

All ten output languages have validated language profiles. English, French, Spanish, German, Italian, Portuguese and Arabic avoid awkward line endings on frequent grammatical function words. Simplified Chinese and Japanese use punctuation-aware kinsoku boundaries. Korean prefers spaces and safely falls back to Hangul-syllable boundaries. Arabic uses right-to-left shaping through libass. Unicode grapheme measurement prevents internal cuts in combining marks and emoji sequences.

Subtitles use at most two lines and remain inside a phone-safe region. The reading-speed target is 23 characters per second, with script-specific limits where appropriate. When fast speech makes readability and synchronisation incompatible, synchronisation wins and the subtitle track is labelled needs_manual_review; text is not silently removed or desynchronised.

There is no word-by-word karaoke or dynamic highlight style. The transcription provider does not guarantee the word timings required for an honest result.


Outputs and statusessvgsvg

One Dataset row is created per selected clip. It includes the source interval, duration, score and breakdown, hook, title, description, hashtags, selectionMethodcropMethod, language codes, compliance statuses, warnings, and links to files that actually exist.

The run's key-value store can contain:

  • one MP4 and one metadata JSON per delivered clip;
  • SRT and VTT tracks only when their toggles are enabled;
  • original and translated tracks with language suffixes when both are requested;
  • manifest.json, including counters, billing reconciliation, warnings, artifact sizes and SHA-256 checksums;
  • ai-viral-clips.zip when createZip is enabled.

An absent artifact is represented by null, never by an empty fake link.

StatusMeaning
successEvery selected clip was rendered and delivered.
partial_successAt least one clip was delivered and at least one failed. Inspect each row.
failedNo clip could be delivered. Inspect error and the logs.
needs_manual_reviewA subtitle track exists but violates a linguistic reading constraint, usually because speech is too fast.

A subtitle warning does not mean the MP4 failed to decode. Technical media validation and linguistic subtitle review are separate facts.


Media validationsvgsvg

Every rendered clip is probed again before delivery. Validation checks include the required audio and video streams, dimensions, pixel format, duration, decodable beginning and end, and timestamp consistency. This catches damaged files that can report a plausible duration yet fail when their final frames are decoded.

The MP4 output uses H.264 video, AAC audio and yuv420p for broad social and mobile compatibility. Rendering re-encodes the cut so boundaries remain frame-accurate instead of drifting to the nearest source keyframe.


Pricingsvgsvg

What a real run costssvgsvg

SourceOutputTotal
10-minute recording2 clips, 30 s, 720pabout $0.40
20-minute interview3 clips, 60 s, 720pabout $0.75
60-minute podcast3 clips, 60 s, 720pabout $1.95
60-minute podcast5 clips, 60 s, 720pabout $2.05
60-minute podcast5 clips, 60 s, 1080p, translatedabout $2.35

Transcription and analysis are charged per source minute, so the length of what you upload drives most of the cost. Rendering is charged per delivered clip, with every started minute rounded up independently. Source length still drives most of the total cost.

If you are on a free Apify plan, start with a 10–20 minute source.

▶ Start with a 10-minute file — about $0.40

Event tablesvgsvg

Pay per event. A started minute is a billed minute.

EventPriceCharged when
apify-actor-start$0.00005Automatically charged by Apify when the Actor starts; never charged manually by the code.
transcribed-minute$0.02Per source minute, after a usable transcript is produced.
analysed-minute$0.01Per source minute when the model selects passages.
fallback-analysed-minute$0.01Per source minute when deterministic selection is used instead.
translated-minute$0.03Per delivered clip minute, only when translation is produced.
rendered-clip-minute$0.05Per validated clip minute at 720 × 1280.
rendered-hd-clip-minute$0.08Per validated clip minute at 1080 × 1920.

Exactly one of the two analysis event types is used in a run. For each delivered clip, exactly one of the two render event types is used. Failed renders, rejected candidates and unnecessary translations are not charged.

Example for a one-hour source producing five 30–60 second clips:

  • 720 × 1280, no translation: about $2.05;
  • 720 × 1280, translated: about $2.14–$2.20;
  • 1080 × 1920, no translation: about $2.20;
  • 1080 × 1920, translated: about $2.29–$2.35.

Cost ceilingssvgsvg

maxCostUsd is the Actor's preflight ceiling. Apify's separate Maximum cost per run setting must also be high enough. If the Actor estimate exceeds maxCostUsd, the run stops with cost_limit_exceeded before a paid provider is contacted. If the Apify charge ceiling is exhausted later, the run reports billing_limit_reached and keeps already produced artifacts.


A complete JSON examplesvgsvg

For a direct file URL, three distinct French-subtitled clips:

undefined
undefined

svg

undefined
{

svg

"videoUrl": "https://example.com/interview.mp4",

"confirmContentRights": true,

"clipCount": 3,

"targetClipDuration": "15-30",

"platform": "TikTok",

"sourceLanguage": "en",

"outputLanguage": "FR",

"outputLanguageMode": "translations",

"aspectRatio": "9:16",

"resolution": "720x1280",

"cropMode": "face_tracking",

"fallbackCropMode": "center_crop",

"burnSubtitles": true,

"subtitleStyle": "Clean",

"generateSrt": true,

"generateVtt": true,

"createZip": true,

"maxCostUsd": 5,

"minimumViralScore": 0,

"avoidOverlaps": true,

"maxMediaMinutes": 90

}

svg

The videoFile value generated by the Console is normally an array containing one Apify storage URL. Use the upload widget instead of typing that structure by hand whenever possible.


Before you commit a long sourcesvgsvg

Read this before running a 90-minute source. Nothing here is a defect; these are the deliberate boundaries of an automatic first edit.

Important: this is an automatic first edit, not a human editorial or linguistic review. Watch every delivered clip before publishing it.

Sourcessvgsvg

  • One uploaded MP4, MOV, MKV or WebM, or one direct URL to such a file.
  • Platform page URLs — YouTube, TikTok, Instagram and similar — are refused. Upload a file you own, or provide a direct media-file URL you control.
  • Source limits are 500 MB, 90 minutes by default and 180 minutes maximum.
  • The source must contain both video and an audible speech track. Weak, distorted, overlapping or music-covered speech can reduce transcription, selection and active-speaker accuracy.

What is not guaranteedsvgsvg

  • The requested clip count. The Actor returns fewer clips when the source does not contain enough distinct, complete passages.
  • The viral score is an editorial estimate, not a performance prediction.
  • Active-speaker tracking can fail on profiles, small or obscured faces, overlapping speakers, heavy backlight and crowded scenes.
  • "Follow the main face" stabilises the virtual crop around a subject. It does not remove camera shake already present in the source footage.
  • Automatic selection, transcription, metadata and translation all require human review before publication. Translation is machine-produced and unreviewed.

What this Actor does not dosvgsvg

  • It does not publish to social networks. It produces files and metadata; you review and publish them yourself.
  • Only vertical 9:16 output is rendered in this version, at 720 × 1280 or 1080 × 1920.
  • Subtitle streams already embedded in the source are ignored; the audio is transcribed instead. Text already burned into the image cannot be removed.
  • Subtitles are never compulsory: burned subtitles, SRT files and VTT files are three independent options.
  • There is no automatic B-roll insertion, voice cloning, face modification, AI-generated imagery or word-by-word karaoke.

Troubleshootingsvgsvg

Symptom or resultExplanation and action
The run succeeds in a few seconds but returns no clipsNo source was supplied, so the free readiness check ran. Upload a video and start a new run.
The upload field is missingThe selected Actor build does not expose this project's .actor/input_schema.json. Select the deployed build/version that contains the complete project.
validation_errorRead the message and correct the named input. JSON enum values are case-sensitive.
A direct URL is refusedIt is probably a platform page, private page or redirect rather than a direct media file. Upload the video instead.
Fewer clips than requestedThe source lacks enough distinct passages under the current duration/score constraints. Try a shorter duration or a longer source.
Two subtitle levels appearoutputLanguageMode is both, or subtitles were already burned into the source. Choose translations for translated-only output.
No visible subtitles, but SRT/VTT files existThis is expected when burnSubtitles is false but generateSrt or generateVtt is true.
The wrong person is framed during a conversationActive-speaker evidence may be ambiguous. Profiles, overlapping speech and covered mouths are difficult. Try center_crop or blurred_background for that source.
The image shakesFace tracking filters crop jitter but cannot remove shake present in the original camera footage.
cost_limit_exceededIncrease maxCostUsd, or reduce source length, clip count, clip duration or resolution. Nothing was charged.
billing_limit_reachedIncrease Apify's Maximum cost per run and start a new run. Already produced files remain available.
configuration_errorA required provider is not configured on the Actor. Nothing was charged; report it to the Actor owner.
Subtitle needs_manual_reviewSpeech was too fast for the configured reading limit. Review or edit the external subtitle before publication.

Permissions and securitysvg

This Actor currently requires Full permissions because Apify's native file uploader stores uploaded media in a private Key-value store and provides the Actor with an Apify storage URL. The Actor uses that access only to fetch the uploaded source video required for the run.

The Actor does not intentionally enumerate or read unrelated Actors, datasets, Key-value stores, schedules, or other account data. If you use Direct video URL, the source must be a publicly reachable direct media-file URL.


Privacy and third-party processingsvgsvg

The media is downloaded into the run's container and temporary working files are deleted when the run ends. Audio is sent to the transcription provider; the timed transcript is sent to the selection provider; clip text is sent to the translation provider only when translation is requested. Face detection, active-speaker analysis, cropping, subtitle rendering and media validation run inside the Actor container.

Output files remain in the run's Apify storage according to your Apify storage and retention settings.


You must own the source or be authorised to transform it. Confirmation is required before any download or paid processing.

The Actor does not download from YouTube, TikTok, Instagram or similar platform pages, and does not bypass authentication, access controls or DRM. Download or obtain the source lawfully, then upload the file.


Other Lumaxys Actorssvgsvg

This Actor is one stage of a larger media workflow. If you need a different stage, these cover it — you only pay for the processing you actually use.


▶ Try AI Viral Clip Generator — upload a video, tick the rights box, press Start.


API callsvgsvg

undefined
undefined

svg

undefined
curl -X POST "https://api.apify.com/v2/acts/lumaxys~ai-viral-clip-generator/runs?token=$APIFY_TOKEN" \

svg

-H "Content-Type: application/json" \

-d '{

"videoUrl": "https://example.com/interview.mp4",

"confirmContentRights": true,

"clipCount": 3,

"targetClipDuration": "30-60",

"platform": "TikTok",

"resolution": "720x1280",

"cropMode": "face_tracking",

"outputLanguageMode": "original",

"maxCostUsd": 5

}'

svg

Dataset rows and file URLs are available in the run output. The full artifact inventory, validation results and billing reconciliation are recorded in manifest.json.


Built by Lumaxys.