Bulk FFmpeg Media Processor – 100 URLs/Run
Pricing
from $10.00 / 1,000 media transformeds
Bulk FFmpeg Media Processor – 100 URLs/Run
Process up to 100 direct video or audio file URLs in one run. Probe metadata, extract 16 kHz speech audio, create a thumbnail, remux, or build a lightweight analysis proxy, with one status row per file.
Pricing
from $10.00 / 1,000 media transformeds
Rating
0.0
(0)
Developer
Vadim Bezrukov
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
Process up to 100 independent video or audio URLs in one Actor run. Probe, extract speech audio, create thumbnails, remux, or build lightweight analysis proxies with per-file status and no local FFmpeg server.
Paste direct https:// links to files you are allowed to process. The run writes one Dataset row per file. A successful file also leaves an artifact in the run Key-Value Store. A file that fails stays on its own row and does not cancel the others.
The prefilled input probes one short CC0 flower video published by MDN. You should see one SUCCESS row and a probe JSON record. If that sample URL is withdrawn, replace it with a file you can fetch. Run the same operation again when the next batch of URLs arrives: a schedule, an n8n/Make step, or a fresh dataset from the Actor upstream of this one.
This build is not on the Apify Store yet. Event prices below are the private staging tariff measured on 2026-09-24. Update history is in CHANGELOG.md.
Batch example
{"operation": "audio16k","items": [{"id": "a", "source": "https://example.com/a.mp4"},{"id": "b", "source": "https://example.com/b.mp4"}]}
id is optional. When you omit it, the row id is item-0000, item-0001, and so on. Send items or a dataset, not both.
Operations
| Operation | What you get | Options |
|---|---|---|
probe | ffprobe metadata stored as JSON: duration, container, codecs, dimensions, fps, bitrate | none |
audio16k | mono 16 kHz WAV for speech pipelines | format must be wav when set |
thumbnail | one JPEG or WebP frame | atSeconds or atPercent (default 10), imageFormat jpeg or webp |
remux | same streams in a new container, no re-encode | container mp4 (default) or mkv |
proxy | low-resolution MP4, width capped at maxWidth (160-640, default 640) and checked again after encode | maxWidth |
There is one operation per run. Unknown option names are rejected for that file.
Dataset input
Point the run at a dataset from a previous Actor when the media URL lives on each row:
{"operation": "probe","datasetId": "aaaaaaaaaaaaaaaaa","sourceField": "videoUrl","idField": "clipId"}
Pass the dataset id in top-level datasetId. In an integration, set that field to {{resource.defaultDatasetId}} so the platform substitutes the id before the run starts. A raw payload.resource.defaultDatasetId is not read. sourceField is the top-level column that contains the file URL. The run reads at most maxItems rows (default 100). A row with an empty URL becomes INVALID_INPUT and the other rows continue. Under limited permissions, read access is requested for datasetId only.
This mode reads a dataset in your Apify account. It does not read another customer's storage. kvs://record-name reads a binary record from this run's own Key-Value Store only.
Output row
Field names are snake_case. The values below are illustrative, not a recorded customer run.
{"source": "bulk-ffmpeg-media-processor","source_id": "a","source_url": "https://example.com/a.mp4","source_index": 0,"id": "a","status": "SUCCESS","operation": "audio16k","input": {"bytes": 1200, "duration_seconds": 8.0, "container": "mp4"},"output": {"kvs_key": "artifacts-a-0000-audio16k.wav","url": "https://api.apify.com/v2/key-value-stores/STORE/records/artifacts-a-0000-audio16k.wav","bytes": 256000,"sha256": "…","content_type": "audio/wav"},"output_url": "https://api.apify.com/v2/key-value-stores/STORE/records/artifacts-a-0000-audio16k.wav","processing_ms": 900,"error": null,"error_code": null,"billed_units": 1,"scraped_at": "2026-09-23T12:00:00Z","fingerprint": "…","schema_version": 1}
output_url repeats output.url so the table view can link to the file. Key-Value Store keys cannot contain slashes, so the prefix is artifacts- rather than a folder path. fingerprint covers the semantic fields and skips the observation time, which makes a later comparison a hash check. scraped_at is when this run observed the file.
RUN_SUMMARY in the Key-Value Store has the overall status, per-status counts, the FFmpeg version line, charged event counts, and nextAction. Read that before you fetch a large dataset. BILLING_RECEIPT is the charge ledger for the run.
Pricing
These event prices come from private staging runs on 2026-09-24. They are the tariff on this Actor. The Actor is not on the Apify Store yet.
| Event | When it is charged | Price |
|---|---|---|
apify-actor-start | Actor start, once per GB of memory (minimum one) | $0.00005 |
media-probed | probe metadata stored | $0.003 |
media-transformed | one started minute of a stored audio, thumbnail, remux or proxy file | $0.01 |
A transform bills one event per started minute of the input, with a minimum of one. An 8 second file is one event ($0.01). A 10 minute file is ten events ($0.10). billed_units on the row is that count. A file with no duration is not transformed and is not charged. A failed file, a rejected URL, a retry and a missing artifact are not charged. The artifact is stored before the event is charged. Set maxTotalChargeUsd high enough for the minutes you expect. When the limit cannot cover the next file, that file is BUDGET_EXCEEDED and files that already succeeded stay in the dataset.
These prices are the event charges. This Actor does not add a separate platform-usage line on top of them. Compute, transfer and storage still cost the publisher; they are not a second bill to you unless the pricing page says pay per event plus usage.
The smallest maxTotalChargeUsd the platform accepts for this Actor is $0.0031. A one-file probe at 1024 MB is $0.00305 in events, and at 2048 MB it is $0.00310 because the start event is billed once per GB. Both fit under a $0.0031 cap. A cap of $0.00305 is rejected before the run starts. A 100-file transform of one-minute files is about $1.00005 in events at 1024 MB.
Limits
- 1-100 files per run (
maxItems) - concurrency 1-4, and the operation may use fewer workers (
audio16kandproxyuse at most 2) - 256 MiB per file by default, 1 GiB maximum
- 2 GiB total download by default
- media longer than
maxDurationSeconds(default 1 hour, maximum 2 hours) isTOO_LARGE - download ports are 80 and 443
- one operation per run
- no ZIP of the outputs
Memory defaults to 1024 MB and cannot be raised above 2048 MB. The run timeout default is 1 hour.
API, n8n and Make
After the Actor is in your account, start it with the Apify run API and the JSON above. In n8n or Make, use the Apify node or an HTTP POST to the run endpoint, wait for the run, then read the dataset items. Keep rows where status is SUCCESS and pass output.url to the next step.
Create a schedule in the Apify Console for the dataset input above when an upstream run refreshes that dataset. Keep the same sourceField. A daily schedule is the usual second run. Do not schedule the prefilled flower sample. It is only the first-run check. In an integration, set datasetId to {{resource.defaultDatasetId}}. The platform substitutes the id before the run, and that top-level field is the one this Actor can read. A raw payload.resource.defaultDatasetId is not read. Retry ids whose status is not SUCCESS, so stored files are not charged again.
Ask an agent: "Probe these direct media URLs and return one status row per file." The direct MCP endpoint is https://mcp.apify.com?tools=automa-flow/bulk-ffmpeg-media-processor. The agent should call this for direct file URLs and a bounded batch. It should not call it for YouTube, TikTok or other page URLs. Execution uses the caller's Apify account. Read RUN_SUMMARY first. If status is PARTIAL or ALL_FAILED, inspect error.code before retrying. Retry the failed ids only, so successful files are not charged again.
Anonymous MCP discovery and a live tools list are not verified for this unpublished build.
Failure semantics
SUCCESS means the expected artifact exists, is non-empty, passed an output check, and was stored. Thumbnails must match JPEG or WebP bytes. Speech audio must be mono 16 kHz. Remux and proxy files must be readable by ffprobe, and a proxy must not be wider than maxWidth. Any other status has output null and an error object. An HTML page, even with HTTP 200, is SOURCE_FAILED (HTML_BODY), not a successful media file. A 404 is SOURCE_NOT_FOUND. A private or metadata address is INVALID_INPUT (SSRF_REJECTED) when it is caught before the download, or SOURCE_BLOCKED when a redirect points there. A timeout is TIMEOUT. Those are different outcomes: an unreachable host is not the same as a file that was fetched and contains no matching media.
| Status | Meaning |
|---|---|
SUCCESS | Artifact stored |
INVALID_INPUT | Bad id, duplicate id, unknown option, or blocked URL shape |
SOURCE_NOT_FOUND | HTTP 404 or missing kvs:// record |
SOURCE_BLOCKED | HTTP 401/403, or a redirect to a blocked address |
SOURCE_FAILED | Other HTTP errors, DNS failure, HTML body, network error |
TOO_LARGE | Byte cap or duration cap |
UNSUPPORTED_MEDIA | ffprobe could not read a container or streams |
PROCESSING_FAILED | ffmpeg failed or the output failed validation |
TIMEOUT | Download or ffmpeg exceeded its timeout |
BUDGET_EXCEEDED | maxTotalChargeUsd cannot cover another success |
The run finishes successfully when item failures are the only problem. It fails the whole run when FFmpeg is missing or built with --enable-nonfree, when the Key-Value Store or dataset cannot be written, or when a charge call returns an uncertain result. A restarted process does not replay charges for a delivery that already started.
Security
Media URLs are downloaded with a normal HTTP client and only then passed to FFmpeg as local files. FFmpeg does not receive the URL, a shell string, or a custom filter graph.
The client allows http and https on ports 80 and 443. It rejects other schemes, embedded passwords, localhost, link-local and metadata names, private and reserved IP ranges, IPv6 loopback, unique-local and link-local addresses, and non-canonical IP spellings. DNS is checked before the request and again when the connection opens, including after every redirect. The body is streamed and cut at the byte cap even when Content-Length is missing or too small. Query values are redacted in logs. The client does not send cookies or an Authorization header.
You supply the media. This Actor does not extract videos from web pages, does not run yt-dlp, and does not accept raw FFmpeg arguments.
Licensing and your media
You need the rights or permission to process the files you submit.
The container installs Debian's unmodified ffmpeg package from the Actor base image and refuses a build that contains --enable-nonfree. Debian's FFmpeg is typically GPL-2 or later because optional GPL components such as libx264 are enabled. The image keeps Debian's copyright file and the installed package version. Source for that binary is the corresponding Debian source package (apt-get source ffmpeg on the same Debian release). This Actor does not add its own FFmpeg patches.
MVP encoders are the native WAV and JPEG paths, stream copy for remux, WebP when libwebp is present, and libx264 or MPEG-4 for the proxy. MP3 and Opus are not written in this version. The measured image is Debian ffmpeg 7.1.5-0+deb13u1 with --enable-gpl and without --enable-nonfree. The owner accepted that build on 2026-09-24. Codec patents are not claimed to be exhausted. This section is the license note shipped with the Actor.
FFmpeg is a trademark of the FFmpeg project. This Actor is not affiliated with, endorsed by, or sponsored by FFmpeg. The name is used to describe command-line compatibility.
Technical details
Python 3.12 on the Apify Python image. Downloads use httpx with a pinned DNS backend. FFmpeg and ffprobe run as argument arrays against temporary files, with a timeout and a kill on expiry. Temporary files are removed after each item is stored or fails. There is no browser, no proxy product, and no external database.
The startup log records the FFmpeg version line and whether the build enabled GPL components.