# 🧪 YouTube, Instagram & TikTok Video Evidence Compiler (`thenetaji/youtube-instagram-tiktok-video-evidence-compiler`) Actor

Compile public YouTube, Instagram, and TikTok videos into timestamped phrase evidence or ordered tutorial drafts.

- **URL**: https://apify.com/thenetaji/youtube-instagram-tiktok-video-evidence-compiler.md
- **Developed by:** [The Netaji](https://apify.com/thenetaji) (community)
- **Categories:** Videos, Automation, Education
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $8.00 / 1,000 analyzed videos

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## YouTube, Instagram and TikTok Video Evidence Compiler

Compile public YouTube, Instagram, and TikTok videos into timestamped transcript evidence or an ordered tutorial draft.

### Accepted input

`videoUrls` accepts public YouTube, Instagram, and TikTok post URLs, YouTube video IDs, and direct public media URLs. Select `videoEvidence` to locate exact spoken phrases or `tutorialToSop` to preserve spoken tutorial segments as draft steps.

```json
{
  "scraperType": "videoEvidence",
  "videoUrls": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://www.instagram.com/reel/example/"
  ],
  "evidenceTerms": ["turn off power", "remove the cover"],
  "language": "en",
  "detectScenes": true,
  "sceneThreshold": 27,
  "includeTranscriptText": true,
  "maxItems": 20
}
```

Matching is literal after case and punctuation normalization.

### Analysis modes

`videoEvidence` returns the spoken opening found within the first three seconds, exact selected phrase matches, word and segment timestamps, and the scene indexes overlapping each piece of evidence.

`tutorialToSop` returns one ordered draft step per usable spoken transcript segment, up to `maxSopSteps`. Each step keeps the spoken instruction, timestamps, transcript segment ID, and overlapping scene indexes. The draft preserves spoken evidence; it is not a verified procedure.

### Response fields

Each row represents one submitted video:

- `video_source` contains the platform, public post URL, and resolved public metadata.
- `analysis_status` is `complete`, `partial`, `not_found`, or `source_error`.
- `media` contains observed duration, size, format, dimensions, and audio/video availability.
- `transcript` contains language and coverage fields, plus text, segments, and words when enabled.
- `scenes` contains detected cut timings and motion labels.
- `video_evidence` contains the opening hook, exact matches, and the timestamped transcript timeline.
- `sop_draft` contains the ordered spoken steps in tutorial mode.
- `analysis` and `provenance` record which stages completed.

```json
{
  "record_type": "video_evidence",
  "video_source": {
    "platform": "youtube",
    "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "id": "dQw4w9WgXcQ",
    "title": "Example repair guide",
    "author": "Fix Lab"
  },
  "analysis_status": "complete",
  "video_evidence": {
    "hook": {
      "start": 0,
      "end": 2.4,
      "text": "First, turn off power."
    },
    "exact_matches": [
      {
        "phrase": "turn off power",
        "start": 0.7,
        "end": 2.4,
        "quote": "turn off power.",
        "segment_ids": [0],
        "scene_indexes": [0, 1]
      }
    ]
  },
  "analysis": {
    "resolve": { "status": "ok" },
    "transcribe": { "status": "ok" },
    "scenes": { "status": "ok" },
    "compile": { "status": "ok" }
  }
}
```

### Behaviour on partial results

A video remains in the dataset when one selected stage fails. For example, successful scene detection with failed transcription produces `partial`, empty transcript evidence, and the scene timeline. An invalid source or a post with no playable media produces `source_error`; a resolved post with no record produces `not_found`.

Media analysis is limited to two minutes and 100 MB per video. Scene labels describe cuts and broad motion; they are not a substitute for visual interpretation.

Direct media URLs must be publicly reachable.

# Actor input Schema

## `scraperType` (type: `string`):

Choose the dataset for this run, then fill in the section for that mode below.

## `videoUrls` (type: `array`):

Public YouTube, Instagram, or TikTok post URLs, YouTube video IDs, or direct media URLs. Add one per line.

## `maxItems` (type: `integer`):

Maximum video rows to save across the run. Set 0 for no limit.

## `evidenceTerms` (type: `array`):

Optional exact phrases to locate in the transcript. Matching ignores case and punctuation.

## `maxSopSteps` (type: `integer`):

Maximum spoken transcript segments to preserve as ordered SOP draft steps.

## `language` (type: `string`):

Optional two-letter language code for transcription. Leave empty to detect the language.

## `viewerRegion` (type: `string`):

Optional two-letter region used for YouTube results.

## `proxyTier` (type: `string`):

Optional setting for accessing the selected video media.

## `detectScenes` (type: `boolean`):

Detect visual cuts and align transcript evidence to the scenes it overlaps.

## `sceneThreshold` (type: `number`):

Cut-detection threshold from 5 to 100. Lower values detect subtler changes.

## `includeTranscriptText` (type: `boolean`):

Include transcript text, segments, and word timings in each result row.

## Actor input object example

```json
{
  "scraperType": "videoEvidence",
  "videoUrls": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "maxItems": 20,
  "evidenceTerms": [
    "turn off power",
    "remove the cover"
  ],
  "maxSopSteps": 25,
  "language": "en",
  "viewerRegion": "US",
  "proxyTier": "none",
  "detectScenes": true,
  "sceneThreshold": 27,
  "includeTranscriptText": true
}
```

# Actor output Schema

## `dataset` (type: `string`):

All records scraped by this run

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "scraperType": "videoEvidence",
    "videoUrls": [
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
    ],
    "maxItems": 20,
    "maxSopSteps": 25,
    "viewerRegion": "US",
    "proxyTier": "none",
    "detectScenes": true,
    "sceneThreshold": 27,
    "includeTranscriptText": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("thenetaji/youtube-instagram-tiktok-video-evidence-compiler").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "scraperType": "videoEvidence",
    "videoUrls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ"],
    "maxItems": 20,
    "maxSopSteps": 25,
    "viewerRegion": "US",
    "proxyTier": "none",
    "detectScenes": True,
    "sceneThreshold": 27,
    "includeTranscriptText": True,
}

# Run the Actor and wait for it to finish
run = client.actor("thenetaji/youtube-instagram-tiktok-video-evidence-compiler").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "scraperType": "videoEvidence",
  "videoUrls": [
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "maxItems": 20,
  "maxSopSteps": 25,
  "viewerRegion": "US",
  "proxyTier": "none",
  "detectScenes": true,
  "sceneThreshold": 27,
  "includeTranscriptText": true
}' |
apify call thenetaji/youtube-instagram-tiktok-video-evidence-compiler --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,thenetaji/youtube-instagram-tiktok-video-evidence-compiler"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/pF9VZEpEO6nWhpdrl/builds/Qe7KRMZDwLEW5qr9M/openapi.json
