# Video Reuse Monitor (`human_intelligence/video-reuse-monitor`) Actor

Detect reused video clips in supplied videos or Apify datasets. Compare originals, locate matching segments, track repeat observations, and export visual evidence reports. Supports direct video URLs; does not search the entire web.

- **URL**: https://apify.com/human\_intelligence/video-reuse-monitor.md
- **Developed by:** [human intelligence](https://apify.com/human_intelligence) (community)
- **Stats:** 2 total users, 1 monthly users, 66.7% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $50.00 / 1,000 video minute processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Video Reuse Monitor — Clip Matching & Usage Tracking

Connect existing ad-scraper datasets to your original video library. Find full-video reuse and overlapping clips, group observed ads by reference asset, and track newly observed placements across runs.

**This Actor analyzes supplied media. It does not search the whole internet or scrape ad libraries itself.** It processes real video files with FFmpeg, visual perceptual fingerprints and temporal matching; no paid AI API, language-model key, GPU service or face recognition is needed.

### What you get

- Exact-file matches, even for static files.
- Perceptual matches for re-encoding/resizing and supported same-speed excerpts.
- Start/end timestamps in both your reference and the inspected video.
- Reference-linked families: multiple ads using an original share its asset-family ID. A montage may match several originals; different originals are not falsely collapsed into one family.
- First observed / still observed / observed again and conservative not-observed records.
- Optional supplied usage-window checks, with active-record and unverified-activity flags separated.
- An HTML report with paired comparison frames, JSON and spreadsheet-safe CSV.
- Persistent fingerprint reuse and bounded history when you supply a named state store.

### Quick start: real videos

Set **Mode** to `monitor`, add references, then add direct comparison URLs or an existing dataset ID.

```json
{
  "mode": "monitor",
  "references": [
    {
      "id": "asset-01",
      "name": "Summer product demo",
      "url": "https://your-public-media-host.example/original.mp4",
      "usageEnd": "2026-12-31"
    }
  ],
  "videos": [
    {
      "id": "ad-01",
      "url": "https://your-public-media-host.example/ad.mp4",
      "pageUrl": "https://www.facebook.com/ads/library/?id=YOUR_AD_ID",
      "advertiser": "Your client",
      "platform": "meta",
      "isActive": true
    }
  ],
  "stateStoreName": "vrm-my-project",
  "monitorId": "client-01"
}
```

Replace example media URLs with your actual direct downloadable MP4/MOV/WebM URLs. **An Instagram/TikTok/YouTube page URL is not a video file.** HLS/DASH manifests and authenticated downloads are not supported. A hosting link returning HTML rather than the actual file will fail explicitly.

IDs must be stable. Reference IDs must be unique. One comparison video can match multiple references.

#### Existing scraper datasets

Set `datasetId` to an existing dataset you can access. The Actor reads it; it does not start or charge for another scraper.

Supported automatic layouts include:

- Meta: `snapshot.videos[].video_hd_url` (or `video_sd_url`) and `snapshot.cards[]`.
- Generic: `videoUrl`, `downloadUrl`, `download_url`, `video_hd_url`, `video_sd_url`.
- Downloaded TikTok media: `mediaUrls[]` and common direct download fields.

For another layout, set `datasetMediaField` to a dotted path such as `snapshot.videos.0.video_hd_url` and optionally `datasetIdField` to your stable ID field. A custom media field can contain a URL or a list of URLs.

For input datasets and named history storage, configure the Actor to allow that storage access. **Limited permissions may prevent reading other datasets or named stores.** Direct URLs without cross-run history need no external storage access.

Rows without usable direct video URLs are disclosed in the summary. Dataset row/video caps are disclosed; incomplete datasets never generate a missing-video inference.

### Genuine self-test

Run with:

```json
{"mode":"self-test"}
```

This generates and encodes real MP4 files, then uses the same decoding and matching pipeline to inspect an identical copy, a resized/re-encoded copy, a montage containing an excerpt, and unrelated footage. It does **not** return canned detection results. `SUMMARY.selfTestPassed` must be `true`.

Self-test does not charge the custom processing event. Platform compute/storage or a platform-configured actor-start fee can still apply. Do not publish a self-test task as the product's actual customer workflow.

### Outputs

| Output | Contents |
|---|---|
| Default dataset | Match rows, explicit failures, optional unmatched rows and not-observed records |
| `REPORT.html` | Standalone evidence report; first 100 match rows can have comparison frames |
| `RESULTS.json` | Full structured results |
| `RESULTS.csv` | CSV with timestamp segments and protection against formula injection |
| `SUMMARY` | Coverage, counts, limits, failures, cache usage and processing-event count |

Match rows contain `referenceId`, `candidateId`, `familyId`, `confidence`, `method`, `similarity`, `matchedSeconds`, coverage ratios and `segments` with start/end positions in both videos.

#### Interpret matches correctly

- `exact_sha256`: identical file bytes. `high` evidence for file identity.
- `temporal_perceptual`: visually consistent frame fingerprints across time. Conservative `high` or `review` tier.
- `similarity` measures fingerprint agreement. **It is not a calibrated probability.**
- No match means no match passed these thresholds; it does not prove originality.
- Shared stock footage, templates and recurring logo animations can legitimately match; a match does not establish ownership.
- Match boundaries are approximate (normally within roughly one sample interval for supported edits), not frame-accurate forensic timestamps.

### History and repeat runs

Use the same `stateStoreName`, `monitorId`, reference IDs and candidate IDs for repeat runs. A named key-value store persists content fingerprints and sightings. No state-store name means a one-off cloud comparison without cross-run history.

**Run one job at a time per monitor.** Key-value stores do not provide transactional locking; overlapping schedules for the same monitor are unsupported. Use different monitor IDs or storage names for independent projects.

Files are downloaded again so changed content behind a URL can be detected. A cached content fingerprint skips decoding, not network transfer. Cache identity includes the algorithm and sampling rate. New metadata or usage dates are evaluated again; stored fingerprints do not preserve stale contract decisions.

Histories retain up to 10,000 recently observed reference/video pairs. Each monitor tracks up to 5,000 cached fingerprints; old entries are evicted. Apify storage and network costs remain applicable.

`first_observed` is the first sighting by this monitor, not the platform's upload date. `observed_again` means it returned after not being observed in an earlier completed input snapshot. `not_observed_in_current_input` means absent from supplied input, **not that an ad ended**. Failed, truncated or unsupported inputs suppress missing inference.

### Usage-window flags

Dates are optional and supplied by the user, inclusive `YYYY-MM-DD`, compared using the UTC observation date. An outside-window match is flagged only at high match confidence. The flag distinguishes active input records, inactive records and unknown activity.

The Actor does not independently verify ad delivery, rights ownership, license exceptions, continuous advertising, infringements or amounts owed. A media URL remaining reachable is not proof that an advertisement is running.

### Pricing for developers

The implemented custom event is **`video-minute-processed`**.

- One event per started minute of each newly fingerprinted file (references and candidates), after a successful decode.
- Identical file content is fingerprinted/charged once within the run; valid stored fingerprints are not charged again.
- Failed downloads/decodes do not charge this event.
- Fingerprinting is the priced operation. A successfully fingerprinted video can have no matches or hit the bounded comparison limit.
- Match rows do not create additional processing charges.
- Budget availability is checked before decoding each new file. When the requested processing units will not fit, the run saves partial results and stops.
- Disable the default automatic `apify-default-dataset-item` charge; using it would charge customers for each evidence row. The Actor rejects that conflicting pricing configuration.
- The Actor's price is configured in Apify Console, not in source code. No claim of a validated profitable price is included.

A fixed actor-start fee is optional, platform-managed, and separate from this event. It is not included in the processing count. Start fees may scale with memory; a processing-only price is clearer for this product.

### Limits and appropriate scope

This first version is visual matching for short videos (default 180 seconds; hard cap 300), approximately unchanged playback speed. It is not a generic semantic similarity model or a face/person identifier.

Heavy cropping, mirror flips, different speeds, large overlays, occlusion, very short clips, low-information slides and severe edits may be missed or need review. Audio is not matched in version 1. Automated cross-platform discovery and private/authenticated videos are outside the scope.

The comparison index skips saturated low-information buckets and bounds frame comparisons. This can miss difficult/repetitive matches rather than spending unbounded CPU. Check summary/failure records instead of assuming complete coverage.

Media downloads validate public DNS addresses and redirect targets; private networks and non-HTTP schemes are blocked. Only downloaded local files enter FFmpeg, with protocol restrictions, duration/file/resolution limits and timeouts. This is a bounded media processor, not a general URL fetcher.

### Development

Python 3.12, FFmpeg/ffprobe with H.264 support, dependencies in `requirements.txt`.

```sh
python main.py --input example-self-test.json --output local-output
```

Run `python -m pytest tests` after installing `pytest` and `jsonschema`. Use `python tools/package.py` to regenerate the five-file upload variant and release ZIP. The simplified upload `main.py` bundles the same auditable `vrm/` source; no separate detection implementation is used.

### Release status

Version 1.0 implements the described scoped workflow. Local generated-video, state, billing, adapter and package checks are recorded in `VALIDATION.md`. It has not been deployed or validated against your live ad library/account as part of this delivery. Broad commercial accuracy requires a representative real-media benchmark.

# Actor input Schema

## `mode` (type: `string`):

Choose monitor for real videos or self-test for generated test clips.

## `references` (type: `array`):

Original videos with a direct media URL and optional stable ID.

## `videos` (type: `array`):

Comparison videos with a direct media URL and optional stable ID.

## `datasetId` (type: `string`):

Optional existing dataset containing downloadable video URLs.

## `stateStoreName` (type: `string`):

Optional named storage for history and fingerprint reuse.

## `monitorId` (type: `string`):

Stable identifier for this monitoring project.

## Actor input object example

```json
{
  "mode": "self-test",
  "references": [],
  "videos": [],
  "datasetId": "",
  "stateStoreName": "",
  "monitorId": "default"
}
```

# Actor output Schema

## `matches` (type: `string`):

No description

## `report` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "self-test"
};

// Run the Actor and wait for it to finish
const run = await client.actor("human_intelligence/video-reuse-monitor").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "mode": "self-test" }

# Run the Actor and wait for it to finish
run = client.actor("human_intelligence/video-reuse-monitor").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "self-test"
}' |
apify call human_intelligence/video-reuse-monitor --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,human_intelligence/video-reuse-monitor"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/lYeyCAdSadfUmN9Jt/builds/eNcilg1oZBBzbnz2E/openapi.json
