# Social URL to LLM Text — Xiaohongshu, Douyin, YouTube, X (`loongnian714/social-url-to-llm-text`) Actor

Turn a Xiaohongshu, Douyin, YouTube or X link into text an LLM can read: transcript, on-screen text, image descriptions, caption and metadata.

- **URL**: https://apify.com/loongnian714/social-url-to-llm-text.md
- **Developed by:** [Programming with Jack Chew](https://apify.com/loongnian714) (community)
- **Categories:**
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $20.00 / 1,000 link digesteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Social URL to LLM Text

Paste a Xiaohongshu, Douyin, YouTube or X link and get back text a language
model can actually read: **transcript, on-screen text, image descriptions,
caption and metadata**.

Official website: **[linkdigest.dev](https://linkdigest.dev)** — powered by
LinkDigest, which also offers this as an MCP server and a REST API.

### The problem this solves

Fetching a social link yourself returns nothing useful. Share URLs are tokenised
(`xsec_token`, `app_code_link`), the content lives in video and images rather
than HTML, and the server sends an app-download shell or a login wall instead of
the post. There is nothing in the page to parse.

```
$ curl -s 'https://xhslink.com/o/1WiQ1QI6Uc0' | grep -o '<title>.*</title>'
<title>小红书</title>
```

That is the whole page. No caption, no images, no text.

This Actor returns the post instead:

```json
{
  "platform": "douyin",
  "author": "逸蒙的赛博空间",
  "title": "这只丑萌AI竟然月售100万美金！",
  "posted_at": "2026-08-27",
  "key_points": [
    "该AI产品名为 Tolen，主打「记住你」的核心功能",
    "商业模式为订阅制，提供免费体验，付费解锁更长时间的聊天服务"
  ],
  "transcript": [{ "t": 0, "text": "谁能想到，就这么一个丑萌小玩意儿…" }],
  "transcript_source": "asr",
  "images": [{ "description": "…", "ocr": "…" }],
  "ocr_text": ["…"],
  "credits": 7
}
```

### Platforms

Every row below was checked against a real post, not a documentation page.

| Platform | Status |
|---|---|
| Xiaohongshu 小红书 | Works — including image notes, with every image described and OCR'd |
| Douyin 抖音 | Works — video transcribed |
| YouTube | Works — native captions where available, ASR otherwise |
| X | Works |
| Web pages / articles | Works |
| TikTok | **Not supported.** TikTok blocks our servers' IP range |
| Bilibili | **Not supported.** Returns HTTP 412 to our servers |
| Instagram, Facebook | Not supported |

The two unsupported rows are listed on purpose. Finding out after you have wired
something in is worse than knowing now.

### Input

| Field | Type | Notes |
|---|---|---|
| `urls` | array of strings | Required. Share links work as-is. |
| `format` | `json` or `markdown` | `json` gives separate fields; `markdown` gives one ready-to-paste document per link. |
| `maxItems` | integer | Optional ceiling so a long list cannot cost more than expected. |

### Output

One dataset item per link. In `json` mode:

`platform`, `author`, `title`, `posted_at`, `caption`, `transcript` (`{t, text}`),
`ocr_text`, `images` (`{description, ocr}`), `key_points`, `source_url`,
`transcript_source` (`native_captions` | `asr` | `none`), `degraded`, `cached`,
`credits`.

A link that cannot be read still produces an item, carrying `url` and `error`.
Nothing is silently dropped — a short dataset is the worst way to discover that
three of your fifty links failed.

### Pricing

| Event | Price |
|---|---|
| `digest-completed` | $0.02 per link successfully read |
| `extra-credit` | $0.01 per credit beyond the first |

Credits track the two things that actually cost money: images to describe and
minutes of media to transcribe. One base credit covers a typical post and its
first six images; each further six images adds one, and each minute of audio or
video adds two.

| Example | Price |
|---|---|
| A web article or short text post | $0.02 |
| A Xiaohongshu note with 17 images | $0.04 |
| A 2.5-minute Douyin video | $0.08 |
| A 19-minute YouTube video | $0.40 |

**A link that fails is never charged.** Neither is a run that fails because of a
problem on our side.

### How long it takes

Real measurements, not estimates:

- Already-digested link: about **1 second** (anything anyone has run before is
  cached)
- Xiaohongshu note with images: **1–2 minutes**
- YouTube video with captions: about **2.5 minutes**

Links are processed four at a time.

### Notes

- The heavy work runs on our servers, not on Apify. That is deliberate: the
  platforms above refuse datacenter IP ranges, so fetching from inside a cloud
  Actor is exactly what does not work.
- Chinese text is returned as-is, not translated.

### Also available at linkdigest.dev

The same service runs at **[linkdigest.dev](https://linkdigest.dev)**, where you
can use it three other ways:

- **MCP server** — inside Claude Code or Cursor, so an agent can read a link
  mid-conversation:
  ```
  claude mcp add --transport http linkdigest \
    https://linkdigest.dev/mcp \
    --header "Authorization: Bearer ld_live_..."
  ```
- **REST API** — `POST https://linkdigest.dev/api/v1/digest`
- **Web app** — paste a link and read the result in the browser

Three digests free, no card. Questions or a link that should work but doesn't:
[linkdigest.dev](https://linkdigest.dev).

# Actor input Schema

## `urls` (type: `array`):

Xiaohongshu, Douyin, YouTube, X or ordinary web page links. Share links (xhslink.com, v.douyin.com) work as-is.

## `format` (type: `string`):

json gives every field separately. markdown gives one ready-to-paste document per link.

## `maxItems` (type: `integer`):

Optional ceiling, so a long list cannot cost more than you expect.

## Actor input object example

```json
{
  "urls": [
    "https://v.douyin.com/h4tziS8FoxI/"
  ],
  "format": "json"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://v.douyin.com/h4tziS8FoxI/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("loongnian714/social-url-to-llm-text").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://v.douyin.com/h4tziS8FoxI/"] }

# Run the Actor and wait for it to finish
run = client.actor("loongnian714/social-url-to-llm-text").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://v.douyin.com/h4tziS8FoxI/"
  ]
}' |
apify call loongnian714/social-url-to-llm-text --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,loongnian714/social-url-to-llm-text"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gE3hbjVmU6ulF4nqh/builds/KcYSGkibI38m0tXX1/openapi.json
