# Bulk FFmpeg Media Processor – 100 URLs/Run (`automa-flow/bulk-ffmpeg-media-processor`) Actor

Process up to 100 direct video or audio file URLs in one run. Probe metadata, extract 16 kHz speech audio, create a thumbnail, remux, or build a lightweight analysis proxy, with one status row per file.

- **URL**: https://apify.com/automa-flow/bulk-ffmpeg-media-processor.md
- **Developed by:** [Vadim Bezrukov](https://apify.com/automa-flow) (community)
- **Categories:** Automation, Developer tools, Videos
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $10.00 / 1,000 media transformeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk FFmpeg Media Processor – 100 URLs/Run

Process up to 100 independent video or audio URLs in one Actor run. Probe, extract speech audio, create thumbnails, remux, or build lightweight analysis proxies with per-file status and no local FFmpeg server.

Paste direct `https://` links to files you are allowed to process. The run writes one Dataset row per file. A successful file also leaves an artifact in the run Key-Value Store. A file that fails stays on its own row and does not cancel the others.

The prefilled input probes one short CC0 flower video published by MDN. You should see one `SUCCESS` row and a probe JSON record. If that sample URL is withdrawn, replace it with a file you can fetch. Run the same operation again when the next batch of URLs arrives: a schedule, an n8n/Make step, or a fresh dataset from the Actor upstream of this one.

This build is not on the Apify Store yet. Event prices below are the private staging tariff measured on 2026-09-24. Update history is in [CHANGELOG.md](CHANGELOG.md).

### Batch example

```json
{
  "operation": "audio16k",
  "items": [
    {"id": "a", "source": "https://example.com/a.mp4"},
    {"id": "b", "source": "https://example.com/b.mp4"}
  ]
}
```

`id` is optional. When you omit it, the row id is `item-0000`, `item-0001`, and so on. Send `items` or a dataset, not both.

### Operations

| Operation | What you get | Options |
| --- | --- | --- |
| `probe` | ffprobe metadata stored as JSON: duration, container, codecs, dimensions, fps, bitrate | none |
| `audio16k` | mono 16 kHz WAV for speech pipelines | `format` must be `wav` when set |
| `thumbnail` | one JPEG or WebP frame | `atSeconds` or `atPercent` (default 10), `imageFormat` `jpeg` or `webp` |
| `remux` | same streams in a new container, no re-encode | `container` `mp4` (default) or `mkv` |
| `proxy` | low-resolution MP4, width capped at `maxWidth` (160-640, default 640) and checked again after encode | `maxWidth` |

There is one operation per run. Unknown option names are rejected for that file.

### Dataset input

Point the run at a dataset from a previous Actor when the media URL lives on each row:

```json
{
  "operation": "probe",
  "datasetId": "aaaaaaaaaaaaaaaaa",
  "sourceField": "videoUrl",
  "idField": "clipId"
}
```

Pass the dataset id in top-level `datasetId`. In an integration, set that field to `{{resource.defaultDatasetId}}` so the platform substitutes the id before the run starts. A raw `payload.resource.defaultDatasetId` is not read. `sourceField` is the top-level column that contains the file URL. The run reads at most `maxItems` rows (default 100). A row with an empty URL becomes `INVALID_INPUT` and the other rows continue. Under limited permissions, read access is requested for `datasetId` only.

This mode reads a dataset in your Apify account. It does not read another customer's storage. `kvs://record-name` reads a binary record from this run's own Key-Value Store only.

### Output row

Field names are snake\_case. The values below are illustrative, not a recorded customer run.

```json
{
  "source": "bulk-ffmpeg-media-processor",
  "source_id": "a",
  "source_url": "https://example.com/a.mp4",
  "source_index": 0,
  "id": "a",
  "status": "SUCCESS",
  "operation": "audio16k",
  "input": {"bytes": 1200, "duration_seconds": 8.0, "container": "mp4"},
  "output": {
    "kvs_key": "artifacts-a-0000-audio16k.wav",
    "url": "https://api.apify.com/v2/key-value-stores/STORE/records/artifacts-a-0000-audio16k.wav",
    "bytes": 256000,
    "sha256": "…",
    "content_type": "audio/wav"
  },
  "output_url": "https://api.apify.com/v2/key-value-stores/STORE/records/artifacts-a-0000-audio16k.wav",
  "processing_ms": 900,
  "error": null,
  "error_code": null,
  "billed_units": 1,
  "scraped_at": "2026-09-23T12:00:00Z",
  "fingerprint": "…",
  "schema_version": 1
}
```

`output_url` repeats `output.url` so the table view can link to the file. Key-Value Store keys cannot contain slashes, so the prefix is `artifacts-` rather than a folder path. `fingerprint` covers the semantic fields and skips the observation time, which makes a later comparison a hash check. `scraped_at` is when this run observed the file.

`RUN_SUMMARY` in the Key-Value Store has the overall status, per-status counts, the FFmpeg version line, charged event counts, and `nextAction`. Read that before you fetch a large dataset. `BILLING_RECEIPT` is the charge ledger for the run.

### Pricing

These event prices come from private staging runs on 2026-09-24. They are the tariff on this Actor. The Actor is not on the Apify Store yet.

| Event | When it is charged | Price |
| --- | --- | --- |
| `apify-actor-start` | Actor start, once per GB of memory (minimum one) | $0.00005 |
| `media-probed` | probe metadata stored | $0.003 |
| `media-transformed` | one started minute of a stored audio, thumbnail, remux or proxy file | $0.01 |

A transform bills one event per started minute of the input, with a minimum of one. An 8 second file is one event ($0.01). A 10 minute file is ten events ($0.10). `billed_units` on the row is that count. A file with no duration is not transformed and is not charged. A failed file, a rejected URL, a retry and a missing artifact are not charged. The artifact is stored before the event is charged. Set `maxTotalChargeUsd` high enough for the minutes you expect. When the limit cannot cover the next file, that file is `BUDGET_EXCEEDED` and files that already succeeded stay in the dataset.

These prices are the event charges. This Actor does not add a separate platform-usage line on top of them. Compute, transfer and storage still cost the publisher; they are not a second bill to you unless the pricing page says pay per event plus usage.

The smallest `maxTotalChargeUsd` the platform accepts for this Actor is $0.0031. A one-file probe at 1024 MB is $0.00305 in events, and at 2048 MB it is $0.00310 because the start event is billed once per GB. Both fit under a $0.0031 cap. A cap of $0.00305 is rejected before the run starts. A 100-file transform of one-minute files is about $1.00005 in events at 1024 MB.

### Limits

- 1-100 files per run (`maxItems`)
- concurrency 1-4, and the operation may use fewer workers (`audio16k` and `proxy` use at most 2)
- 256 MiB per file by default, 1 GiB maximum
- 2 GiB total download by default
- media longer than `maxDurationSeconds` (default 1 hour, maximum 2 hours) is `TOO_LARGE`
- download ports are 80 and 443
- one operation per run
- no ZIP of the outputs

Memory defaults to 1024 MB and cannot be raised above 2048 MB. The run timeout default is 1 hour.

### API, n8n and Make

After the Actor is in your account, start it with the Apify run API and the JSON above. In n8n or Make, use the Apify node or an HTTP POST to the run endpoint, wait for the run, then read the dataset items. Keep rows where `status` is `SUCCESS` and pass `output.url` to the next step.

Create a schedule in the Apify Console for the dataset input above when an upstream run refreshes that dataset. Keep the same `sourceField`. A daily schedule is the usual second run. Do not schedule the prefilled flower sample. It is only the first-run check. In an integration, set `datasetId` to `{{resource.defaultDatasetId}}`. The platform substitutes the id before the run, and that top-level field is the one this Actor can read. A raw `payload.resource.defaultDatasetId` is not read. Retry ids whose `status` is not `SUCCESS`, so stored files are not charged again.

Ask an agent: "Probe these direct media URLs and return one status row per file." The direct MCP endpoint is <https://mcp.apify.com?tools=automa-flow/bulk-ffmpeg-media-processor>. The agent should call this for direct file URLs and a bounded batch. It should not call it for YouTube, TikTok or other page URLs. Execution uses the caller's Apify account. Read `RUN_SUMMARY` first. If `status` is `PARTIAL` or `ALL_FAILED`, inspect `error.code` before retrying. Retry the failed ids only, so successful files are not charged again.

Anonymous MCP discovery and a live tools list are not verified for this unpublished build.

### Failure semantics

`SUCCESS` means the expected artifact exists, is non-empty, passed an output check, and was stored. Thumbnails must match JPEG or WebP bytes. Speech audio must be mono 16 kHz. Remux and proxy files must be readable by ffprobe, and a proxy must not be wider than `maxWidth`. Any other status has `output` null and an `error` object. An HTML page, even with HTTP 200, is `SOURCE_FAILED` (`HTML_BODY`), not a successful media file. A 404 is `SOURCE_NOT_FOUND`. A private or metadata address is `INVALID_INPUT` (`SSRF_REJECTED`) when it is caught before the download, or `SOURCE_BLOCKED` when a redirect points there. A timeout is `TIMEOUT`. Those are different outcomes: an unreachable host is not the same as a file that was fetched and contains no matching media.

| Status | Meaning |
| --- | --- |
| `SUCCESS` | Artifact stored |
| `INVALID_INPUT` | Bad id, duplicate id, unknown option, or blocked URL shape |
| `SOURCE_NOT_FOUND` | HTTP 404 or missing `kvs://` record |
| `SOURCE_BLOCKED` | HTTP 401/403, or a redirect to a blocked address |
| `SOURCE_FAILED` | Other HTTP errors, DNS failure, HTML body, network error |
| `TOO_LARGE` | Byte cap or duration cap |
| `UNSUPPORTED_MEDIA` | ffprobe could not read a container or streams |
| `PROCESSING_FAILED` | ffmpeg failed or the output failed validation |
| `TIMEOUT` | Download or ffmpeg exceeded its timeout |
| `BUDGET_EXCEEDED` | `maxTotalChargeUsd` cannot cover another success |

The run finishes successfully when item failures are the only problem. It fails the whole run when FFmpeg is missing or built with `--enable-nonfree`, when the Key-Value Store or dataset cannot be written, or when a charge call returns an uncertain result. A restarted process does not replay charges for a delivery that already started.

### Security

Media URLs are downloaded with a normal HTTP client and only then passed to FFmpeg as local files. FFmpeg does not receive the URL, a shell string, or a custom filter graph.

The client allows `http` and `https` on ports 80 and 443. It rejects other schemes, embedded passwords, localhost, link-local and metadata names, private and reserved IP ranges, IPv6 loopback, unique-local and link-local addresses, and non-canonical IP spellings. DNS is checked before the request and again when the connection opens, including after every redirect. The body is streamed and cut at the byte cap even when `Content-Length` is missing or too small. Query values are redacted in logs. The client does not send cookies or an `Authorization` header.

You supply the media. This Actor does not extract videos from web pages, does not run yt-dlp, and does not accept raw FFmpeg arguments.

### Licensing and your media

You need the rights or permission to process the files you submit.

The container installs Debian's unmodified `ffmpeg` package from the Actor base image and refuses a build that contains `--enable-nonfree`. Debian's FFmpeg is typically GPL-2 or later because optional GPL components such as libx264 are enabled. The image keeps Debian's copyright file and the installed package version. Source for that binary is the corresponding Debian source package (`apt-get source ffmpeg` on the same Debian release). This Actor does not add its own FFmpeg patches.

MVP encoders are the native WAV and JPEG paths, stream copy for remux, WebP when libwebp is present, and libx264 or MPEG-4 for the proxy. MP3 and Opus are not written in this version. The measured image is Debian `ffmpeg 7.1.5-0+deb13u1` with `--enable-gpl` and without `--enable-nonfree`. The owner accepted that build on 2026-09-24. Codec patents are not claimed to be exhausted. This section is the license note shipped with the Actor.

FFmpeg is a trademark of the FFmpeg project. This Actor is not affiliated with, endorsed by, or sponsored by FFmpeg. The name is used to describe command-line compatibility.

### Technical details

Python 3.12 on the Apify Python image. Downloads use `httpx` with a pinned DNS backend. FFmpeg and ffprobe run as argument arrays against temporary files, with a timeout and a kill on expiry. Temporary files are removed after each item is stored or fails. There is no browser, no proxy product, and no external database.

The startup log records the FFmpeg version line and whether the build enabled GPL components.

# Changelog

This Actor's version history is a separate document: https://apify.com/automa-flow/bulk-ffmpeg-media-processor/changelog.md

# Actor input Schema

## `operation` (type: `string`):

probe reads metadata. audio16k writes 16 kHz mono WAV. thumbnail saves one JPEG or WebP frame. remux copies streams into MP4 or MKV. proxy builds a low-resolution H.264 or MPEG-4 MP4.

## `items` (type: `array`):

1-100 direct media files. Use this or a dataset, not both. id is optional. source is an http(s) file URL or kvs://record in this run's Key-Value Store.

## `datasetId` (type: `string`):

Apify dataset to read instead of items. The run requests read access to this field. In an integration, set it to {{resource.defaultDatasetId}}. A nested resource id is not read.

## `sourceField` (type: `string`):

Top-level field on each dataset row that holds the media URL. Required in dataset mode.

## `idField` (type: `string`):

Optional top-level field used as the item id. Rows without it get item-0000 style ids.

## `resource` (type: `object`):

Optional webhook object. This Actor does not read defaultDatasetId. Put the id in datasetId.

## `payload` (type: `object`):

Optional webhook wrapper. A nested resource.defaultDatasetId is not read. Put the id in datasetId.

## `concurrency` (type: `integer`):

Upper bound on files processed at once (1-4). The operation may use fewer workers. Default 2.

## `maxItems` (type: `integer`):

Maximum files to accept from items or a dataset. 1-100. Default 100.

## `maxInputBytesPerItem` (type: `integer`):

Reject a file whose Content-Length or streamed body exceeds this size. Default 256 MiB. Maximum 1 GiB.

## `maxTotalInputBytes` (type: `integer`):

Stop accepting further files once the downloaded bytes in this run would pass this total. Default 2 GiB. Must be at least the per-file cap.

## `maxDurationSeconds` (type: `integer`):

Reject media whose probed duration is longer than this. Default 3600. Maximum 7200.

## `downloadTimeoutSeconds` (type: `integer`):

Total time allowed to download one file, including redirects. Default 120.

## `processingTimeoutSeconds` (type: `integer`):

Maximum time ffprobe or ffmpeg may run for one file. The process is killed when this expires. Default 600.

## Actor input object example

```json
{
  "operation": "probe",
  "items": [
    {
      "id": "flower",
      "source": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4"
    }
  ],
  "concurrency": 2,
  "maxItems": 100,
  "maxInputBytesPerItem": 268435456,
  "maxTotalInputBytes": 2147483648,
  "maxDurationSeconds": 3600,
  "downloadTimeoutSeconds": 120,
  "processingTimeoutSeconds": 600
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

## `billingReceipt` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "operation": "probe",
    "items": [
        {
            "id": "flower",
            "source": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4"
        }
    ],
    "concurrency": 2
};

// Run the Actor and wait for it to finish
const run = await client.actor("automa-flow/bulk-ffmpeg-media-processor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "operation": "probe",
    "items": [{
            "id": "flower",
            "source": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
        }],
    "concurrency": 2,
}

# Run the Actor and wait for it to finish
run = client.actor("automa-flow/bulk-ffmpeg-media-processor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "operation": "probe",
  "items": [
    {
      "id": "flower",
      "source": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4"
    }
  ],
  "concurrency": 2
}' |
apify call automa-flow/bulk-ffmpeg-media-processor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automa-flow/bulk-ffmpeg-media-processor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/xx3kXcNjdzma9Rodb/builds/x904lASjEfWsaSoFe/openapi.json
