𝕏 (twitter) Transcript
Pricing
from $0.2592 / transcript
𝕏 (twitter) Transcript
X Twitter Transcript turns one public X video post into structured text for news monitoring, quote capture, and social research. Output includes language detection, ordered timestamps, post metadata, and optional translation into 133 languages. The base transcript rate is $0.2592.
Pricing
from $0.2592 / transcript
Rating
0.0
(0)
Developer
AgentX
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
3 days ago
Last modified
Categories
Share
X Twitter Transcript - Timestamped X Speech-to-Text API
X Twitter Transcript converts one public X post containing video into detected-language text, ordered timestamped segments, source metadata, and an optional translated transcript. It processes audible speech from the media; it does not promise to return X caption tracks or existing subtitle files.
- One public X Post URL containing video is accepted per run.
- Platform precheck confirms the resolved extractor before the formal download begins.
- Every successful result follows the 22-field Dataset contract used by the Actor.
- Requested translation is either complete and segment-aligned or absent; partial translation is not published as success.
Run X Twitter Transcript · Open the API page
Start with a short public video that contains clear speech. Confirm the output before scheduling large, long, translated, or business-critical workloads.
Why Choose This API
A narrow contract for one known X video post
Use this Actor when you already know the public X video that you want to transcribe. The input is intentionally small: one video_url and one optional translate target. The Actor validates that the resolved media belongs to the X extraction path before the formal download begins.
The result combines source context and spoken content in one record:
| Capability | What is returned |
|---|---|
| Speech recognition | Detected language, combined transcript text, and ordered start / end / text segments |
| Video context | Title, description, author, account identifiers, duration, publication time, thumbnail, categories, and tags when exposed |
| Engagement context | Views, likes, dislikes, shares, and comments when the source exposes those values |
| Optional translation | A translated text and segment list that preserves the original segment timestamps |
This is not an account-feed crawler, multi-post processor, live recorder, native-caption service, speaker-diarization system, summary generator, or video downloader. A successful run creates one Dataset item. Missing source metadata remains null, zero, or empty according to the Dataset schema; the Actor does not invent values.
The transcript comes from speech recognition over the downloaded media. Names, accents, overlapping speakers, music, poor audio, and specialist terminology can reduce accuracy. Review quotations and high-stakes uses against the source video.
Quick Start Guide
Run in Apify Console
- Open X Twitter Transcript on Apify.
- Paste one public X Post URL in Video URL.
- Leave Translate empty for the source-language transcript, or choose one target language.
- Click Start and open the default Dataset after the run succeeds.
The schema example is a real public input:
{"video_url": "https://x.com/captainamerica/status/719944021058060289","translate": "spanish"}
The same URL is used in this guide's API and output examples so that the scenario remains testable. Public videos can later be removed, restricted, made private, age-gated, or changed by their owners. Replace an unavailable example with another public X video containing audible speech.
What success means
A completed production run must download a media file with an audio track, detect speech, publish one Dataset item, and charge the transcript event. An HTTP response, title lookup, or metadata-only extraction is not enough to prove transcript success.
Input Parameters
Input configuration
| Input | Type | Required | Description |
|---|---|---|---|
video_url | string | Yes | One public X video URL to download and transcribe. |
translate | string | No | One target language from the schema's 133 selectable values. Leave empty to skip translation. |
Only one video is processed per run. Profile URLs, search pages, feed URLs, Posts without video, private media, uploads, cookies, credentials, and lists of URLs are outside this Actor's public input contract. The downloader is configured not to process playlists.
X documents how to copy a permanent Post URL in its official URL guide. The schema accepts a public Post URL when the source makes video media accessible. A URL may still fail because of removal, access restrictions, regional availability, temporary platform changes, or the absence of an audio track.
Translation is attempted only after the source transcript succeeds. If any segment cannot be translated after the available attempts, the Actor leaves translation empty and does not charge the translation event. It never presents the source text as if it were a successful translation.
Output Data Schema
One 22-field Dataset item
The default Dataset contains one record for a successful video. The fields are:
| Group | Fields | Notes |
|---|---|---|
| Processing | processor, processed_at | Actor URL and processing timestamp |
| Identity | platform, title, description, thumbnail, published_at | Source-provided video identity |
| Author | author, author_id, author_url | Channel or uploader context when available |
| Media | duration, audio_title, audio_artist | Source-reported values; duration can be 0 when metadata is absent |
| Engagement | view_count, like_count, shares_count, dislike_count, comment_count | Nullable source metrics |
| Labels | categories, tags | Source classifications when exposed |
| Speech | transcript, translation | Timestamped source transcript and optional translated version |
Abbreviated output from the production-checked scenario:
{"processor": "https://apify.com/agentx/x-twitter-transcript?fpr=aiagentapi","processed_at": "2026-07-21T13:30:00+00:00","platform": "Twitter","title": "Captain America - Are you sure you made the right choice?","author": "Captain America","author_id": null,"author_url": "https://twitter.com/CaptainAmerica","duration": 3,"like_count": 2,"categories": [],"tags": [],"transcript": {"language": "English","text": "Jose, Team Cap, welcome to the team!","segments": [{"start": "00:00:00.350","end": "00:00:03.178","text": "Jose, Team Cap, welcome to the team!"}]},"translation": {"language": "Spanish","text": "¡Jose, equipo Cap, bienvenido al equipo!","segments": [{"start": "00:00:00.350","end": "00:00:03.178","text": "¡Jose, equipo Cap, bienvenido al equipo!"}]}}
Counts and source text are snapshots, not permanent facts. The Dataset does not add a source media file, caption track, SRT/VTT file, speaker labels, word-level timing, confidence score, summary, or legal verification.
Integration Examples
REST API
The synchronous endpoint returns Dataset items directly:
curl -L "https://api.apify.com/v2/actors/agentx~x-twitter-transcript/run-sync-get-dataset-items" \-H "Authorization: Bearer $APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"video_url": "https://x.com/captainamerica/status/719944021058060289","translate": "spanish"}'
For long videos or workflows that must not hold one HTTP connection open, start an asynchronous run and retrieve its Dataset afterward. Apify documents that synchronous Dataset responses can time out after 300 seconds while the Actor run continues.
Python client
from apify_client import ApifyClientclient = ApifyClient("YOUR_APIFY_TOKEN")run = client.actor("agentx/x-twitter-transcript").call(run_input={"video_url": "https://x.com/captainamerica/status/719944021058060289","translate": "spanish",})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item["transcript"]["text"])
MCP for AI clients
Configure the Apify MCP server with the Actor-scoped tool URL:
{"mcpServers": {"apify": {"url": "https://mcp.apify.com?tools=agentx/x-twitter-transcript","headers": {"Authorization": "Bearer YOUR_APIFY_TOKEN"}}}}
Call agentx/x-twitter-transcript with the same input fields. See the Apify MCP documentation and the generated Actor API page.
Pricing & Cost Calculator
X Twitter Transcript uses pay-per-event pricing. The repository metadata is authoritative for the build being audited; the public Store's “from” price reflects the lowest transcript tier and may lag a deployment.
| Event | Current price |
|---|---|
| Actor start | $0.001 per charged start event; the 8192 MB run configuration charges eight start events |
| Actor usage | $0.00001 per usage unit; total depends on runtime resources |
| Transcript - Free | $0.28800 |
| Transcript - Bronze | $0.27840 |
| Transcript - Silver | $0.26880 |
| Transcript - Gold, Platinum, Diamond | $0.25920 |
| Translation - Free | $0.10000 |
| Translation - Bronze | $0.09667 |
| Translation - Silver | $0.09333 |
| Translation - Gold, Platinum, Diamond | $0.09000 |
A Free-tier run with one transcript, no translation, and the fixed 8192 MB configuration starts at $0.288 + (8 × $0.001) = $0.296, plus usage events. With translation, it starts at $0.288 + $0.10 + $0.008 = $0.396, plus usage events. Failed work may still consume start and usage events; the transcript event is tied to successful Dataset publication.
Check the live pricing page before scheduling a large workload. Estimate with representative media because speech density, duration, network behavior, and translation length affect runtime.
Use Cases & Applications
Search and retrieval over one video
Store transcript.text for whole-video search or index the timestamped segments as smaller retrieval units. A downstream application can display the matching text and its approximate video interval. The Actor does not create embeddings or answer questions; those are downstream responsibilities.
Editorial and research review
Researchers, journalists, analysts, and content teams can inspect spoken claims alongside source title, author, publication time, and engagement fields. Metadata availability varies, and ASR can be wrong, so cite the original video and manually verify consequential quotations.
Accessibility drafts and content repurposing
The segment list can seed caption editing, show notes, summaries, articles, or social copy. It is a draft transcript, not a certified accessibility artifact. Human review remains appropriate before publication.
Multilingual review
Request one translation target when a team needs to inspect the spoken content in another language. The translation preserves segment time ranges, which helps compare source and translated text. For multiple target languages, run the Actor separately for each target and account for each translation event.
Automation boundaries
This Actor is suited to a known video URL. Use an X discovery workflow to identify video Posts first, then invoke one transcription run per chosen URL. Do not treat a successful transcription as permission to republish the underlying video or transcript.
FAQ
Does the Actor return the X Post text or on-screen text?
No. The verified path downloads the Post's video audio and generates speech text. It does not return the Post body as the transcript and does not perform OCR on text displayed in the video.
What kind of X URL is supported?
Use the permanent URL of one public Post containing downloadable video with an audio track. Protected, deleted, suspended, restricted, audio-only Space, and live content is not promised.
Can it process an entire X account or feed?
No. The public input accepts one video URL, and multi-post processing is disabled. Use a discovery Actor first and submit individual video URLs in separate runs.
What happens if the video has no speech?
The run fails without a transcript Dataset item. Silent clips, music-only media, missing audio, or speech that cannot be detected do not satisfy the successful output contract.
Why are some metadata fields null?
X or the extraction response did not expose them for that video. The schema intentionally permits nullable values instead of fabricating counts, dates, audio attribution, or author details.
Are timestamps word-accurate?
No. They are segment time ranges formatted as HH:MM:SS.mmm. They are useful for navigation and downstream processing but are not guaranteed word-level alignments.
How are long videos handled?
The Actor uses a fixed 8192 MB configuration. Media up to 3600 seconds is transcribed directly. Longer media is processed sequentially in 900-second core chunks with 15 seconds of context on each available side; only one temporary chunk is retained at a time. Runtime and cost still grow with duration and speech density, so use asynchronous runs for long inputs.
Is translation always charged when requested?
No. Translation is charged only when the complete translated structure is produced. If any segment ultimately fails, the output keeps translation empty and the translation event is not charged.
SEO Keywords & Search Terms
People may describe this task as converting a X video to text, generating a X transcript with timestamps, extracting spoken words from a X URL, or preparing video speech for search and RAG. Those phrases describe the same narrow workflow: one public video enters, and one structured transcript record leaves.
The Actor is not positioned as a X subtitle downloader, channel scraper, bulk transcript endpoint, audio downloader, video summarizer, or guaranteed verbatim transcription service. Keeping those boundaries explicit helps users choose the right tool and prevents misleading search claims.
For programmatic discovery, use the stable Actor name agentx/x-twitter-transcript. For human evaluation, inspect the Input, Output, Pricing, Reviews, and Issues tabs on the live Store page.
Trust & Certifications
The strongest available evidence is the public contract: one schema-constrained Post URL, a platform extractor check, a 22-field Dataset schema, ordered timestamped segments, and a complete-or-absent translation object. These properties can be inspected before integrating and validated from each completed run.
No Store rating, user count, monthly activity, bookmark total, or one-off test result is presented as evidence of future availability or transcription accuracy. X delivery and speech recognition remain variable.
No independent accuracy certification, SLA, legal-transcript status, GDPR certification, CCPA certification, or guaranteed uptime is claimed here. Apify supplies the Actor execution and storage platform; source availability and extraction behavior remain external dependencies.
Legal & Compliance
Process only content you are authorized to use. Public accessibility does not remove copyright, privacy, contractual, publicity, data-protection, or platform-policy obligations. X's current Terms of Service prohibit crawling or scraping without prior written consent and restrict automated access to published interfaces unless X specifically allows otherwise. Use this Actor only where you have the required permission or another lawful basis that applies to your situation.
The Actor returns source metadata and generated speech text. It does not grant a license to the video, verify ownership, determine fair use, provide legal advice, or certify the accuracy of quotations. Avoid submitting private credentials or confidential URLs because they are not part of the supported input contract.
If a source owner deletes or restricts a video, future runs can fail even when an older run succeeded. Retain the source URL and processing time when provenance matters.
Related Tools
Related AgentX Actors
- TikTok Transcript handles one public TikTok video.
- YouTube Transcript handles one public YouTube video.
- Video Transcript handles uploads and supported public video URLs.
Enrich with broader discovery
- All Video Scraper provides broader video metadata and media workflows.
- All Jobs Scraper supplies cross-platform job data for content or labor-market research.
- All Shopping Scraper supplies cross-market product data for commerce research.
Choose a discovery Actor when you do not yet have a specific video URL. Keep transcription runs one video at a time.
Support & Community
When reporting a problem, include the public video URL, run ID, expected result, observed error, and whether translation was enabled. Do not post private tokens or credentials.
Schema, code behavior, pricing metadata, the example URL, and official documentation links were last checked on July 21, 2026. X availability, terms, and Store statistics are volatile.