# Link Preview API - Open Graph, Twitter Card & oEmbed Metadata (`neverempty/link-preview-api`) Actor

For chat apps, CMS editors and AI agents that unfurl links: send URLs, get title, description, image, site name, favicon, canonical URL and language from Open Graph, Twitter Card, oEmbed and the HTML head. About a second per URL. Respects robots.txt; URLs that cannot be previewed are free.

- **URL**: https://apify.com/neverempty/link-preview-api.md
- **Developed by:** [NeverEmpty](https://apify.com/neverempty) (community)
- **Categories:** Developer tools, Automation, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.40 / 1,000 preview returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Link Preview API - Open Graph, Twitter Card & oEmbed Metadata

Give it a URL, get back the link preview: **title, description, image, site name, favicon, canonical URL, language and type** in one clean JSON row, read from the page's **Open Graph** tags, **Twitter Card** tags, **oEmbed** endpoint and HTML `<head>`.

Built for software that unfurls links: chat apps and bots (Slack, Discord, Telegram, Teams), CMS and newsletter editors, bookmarking and read-later apps, CRMs, AI agents and RAG pipelines that need a title and thumbnail for a URL, and SEO tools that check `og:image` and `twitter:card` tags.

- **One URL in about a second.** Only the `<head>` of each page is read; the download stops at `</head>`, so a 1.3 MB YouTube page costs the same memory as a small blog post. A run with one URL finished in 2.5 to 3.2 seconds on Apify, including start-up (2026-09-24).
- **Only real values.** A page with no `og:image` has `image: null`. The favicon comes from the page's own `<link rel="icon">`; `/favicon.ico` is not guessed. Nothing is filled in with made-up defaults.
- **Nothing is charged for a URL that could not be previewed.** A page that does not exist, refuses automated reading, needs a sign-in or has no title, description or image comes back as a free row that says why.
- **Respects each site's robots.txt** (RFC 9309) for every URL, every redirect and every oEmbed endpoint, and never signs in, solves check pages or rotates proxies to get around a refusal.
- **Any character set.** The `charset` is taken from the response header, a byte-order mark or the `<meta charset>` in the page itself, and the page is decoded with that character set instead of being forced into UTF-8 (tested on kakaku.com, which declares Shift\_JIS only in `<meta>`).
- **Monitor mode** returns only URLs whose preview changed since the last run, with the previous title, description and image. Each change is confirmed by a second read, so a rotating image is not reported as a change.

### Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `urls` | array of strings | (empty: the example `https://github.com/apify/crawlee` is used) | Page URLs, 1 to 1,000 per run. A URL without `http://` or `https://` is read as `https://`. The same URL given twice (also when only the `#fragment` differs) is read and charged once. |
| `acceptLanguage` | string | (not sent) | Optional `Accept-Language` header, for sites that show a different language version by language, for example `en`, `ja-JP` or `de-DE,de;q=0.9`. Empty means the version the site shows by default. |
| `includeOembed` | boolean | `true` | When the page links to its own oEmbed JSON endpoint in its `<head>` (on 2026-09-24: YouTube and Spotify among the 23 pages read), also read it and add the `oembed` object. |
| `onlyChanges` | boolean | `false` | Monitor mode: return only URLs whose title, description, image, site name, canonical URL or type changed since the last run with the same `watchName`, plus URLs new to the watch. |
| `watchName` | string | (none) | Name of the remembered state (letters, digits, `.`, `-`, `_`; up to 40). Setting it fills `changeType` and the previous-value columns. With monitor mode on and no name, `default` is used. |
| `resetMonitoringState` | boolean | `false` | Forget what this watch remembered, so every URL is returned as a first check. |
| `maxConcurrency` | integer | `4` | URLs read in parallel, 1 to 8. |
| `requestTimeoutSecs` | integer | `20` | Time limit for one request, 5 to 60 seconds. HTTP 429, 500, 502, 503 and 504 are asked again up to two more times; a site that does not answer in time is asked once more. |

```json
{
  "urls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ", "github.com/apify/crawlee"]
}
```

### Output

One row per URL. A real row from 2026-09-24 (long text shortened here):

```json
{
  "status": "ok",
  "position": 1,
  "url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "finalUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "httpStatus": 200,
  "contentType": "text/html; charset=utf-8",
  "redirects": [],
  "title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
  "description": "The official video for “Never Gonna Give You Up” by Rick Astley. ...",
  "image": "https://i.ytimg.com/vi/dQw4w9WgXcQ/maxresdefault.jpg",
  "imageWidth": 1280,
  "imageHeight": 720,
  "imageAlt": null,
  "siteName": "YouTube",
  "favicon": "https://www.youtube.com/s/desktop/02b72088/img/favicon_32x32.png",
  "appleTouchIcon": null,
  "canonicalUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "ogUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "language": "en",
  "languageSource": "html",
  "type": "video.other",
  "twitterCard": "summary_large_image",
  "twitterSite": "@youtube",
  "author": "Rick Astley",
  "publishedTime": null,
  "modifiedTime": null,
  "themeColor": "rgba(255, 255, 255, 0.98)",
  "video": "https://www.youtube.com/embed/dQw4w9WgXcQ",
  "oembed": {
    "type": "video",
    "title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
    "authorName": "Rick Astley",
    "authorUrl": "https://www.youtube.com/@RickAstleyYT",
    "providerName": "YouTube",
    "providerUrl": "https://www.youtube.com/",
    "thumbnailUrl": "https://i.ytimg.com/vi/dQw4w9WgXcQ/hqdefault.jpg",
    "thumbnailWidth": 480,
    "thumbnailHeight": 360,
    "html": "<iframe width=\"200\" height=\"113\" src=\"https://www....",
    "width": 200,
    "height": 113
  },
  "oembedEndpoint": "https://www.youtube.com/oembed?format=json&url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DdQw4w9WgXcQ",
  "jsonLdTypes": ["VideoObject"],
  "robotsMeta": null,
  "sources": ["openGraph", "oembed"],
  "openGraph": { "title": "...", "type": "video.other", "...": "..." },
  "twitter": { "card": "summary_large_image", "site": "@youtube", "...": "..." },
  "changeType": null,
  "changedFields": null,
  "previousTitle": null,
  "previousDescription": null,
  "previousImage": null,
  "previousCheckedAt": null,
  "charset": "utf-8",
  "headBytesRead": 716367,
  "headComplete": true,
  "watchName": null,
  "note": null,
  "fetchedAt": "2026-09-24T13:07:11.920Z"
}
```

Where each value comes from:

| Column | Taken from (first one present) |
|---|---|
| `title` | `og:title`, `twitter:title`, `<title>`, JSON-LD `headline`, oEmbed `title` |
| `description` | `og:description`, `twitter:description`, `<meta name="description">` |
| `image` | `og:image:secure_url` / `og:image`, `twitter:image`, `<link rel="image_src">`, oEmbed `thumbnail_url`; for a URL that is itself an image, the URL |
| `imageWidth`, `imageHeight`, `imageAlt` | `og:image:width`, `og:image:height`, `og:image:alt` (or `twitter:image:alt`) as the page states them; the image is not downloaded |
| `siteName` | `og:site_name`, `<meta name="application-name">`, oEmbed `provider_name` |
| `favicon`, `appleTouchIcon` | `<link rel="icon">` (the size closest to 32 px), `<link rel="apple-touch-icon">` |
| `canonicalUrl`, `ogUrl` | `<link rel="canonical">`, `og:url` |
| `language`, `languageSource` | `<html lang>`, the `Content-Language` header, `og:locale` |
| `type` | `og:type` |
| `author`, `publishedTime`, `modifiedTime` | `<meta name="author">`, `article:author`, `article:published_time`, `article:modified_time`, `og:updated_time`, JSON-LD, oEmbed `author_name` |
| `openGraph`, `twitter` | every `og:*` and `twitter:*` tag, as written |
| `sources` | which of `openGraph`, `twitterCard`, `html`, `jsonLd` supplied the title, description, image or site name, plus `oembed` when the oEmbed endpoint was read, or `url` when the URL is itself an image |

Relative URLs are resolved against the final URL (and `<base href>`), so `image`, `favicon` and `canonicalUrl` are always absolute.

#### Rows that are not charged

| `status` | Meaning |
|---|---|
| `robots-disallowed` | The site's robots.txt does not allow automated reading of this address; it was not requested. |
| `robots-unreachable` | The site's robots.txt answered 5xx or could not be reached; following RFC 9309 the page was not requested. |
| `login-required` | HTTP 401, or the page redirects to a sign-in page. |
| `blocked` | The site refused this reader (HTTP 403, 429 after retries, 451) or showed a CAPTCHA or browser-check page (with any status; a check page is never asked again). A person with a browser may still see the page. |
| `not-found` | HTTP 404 or 410. |
| `http-error` | Another HTTP answer that is not a page (for example 400 or 405), or a broken redirect. |
| `unreachable` | The domain does not resolve, or the site did not answer. |
| `unreadable` | 5xx after retries, an answer that kept stopping before the end of the page's `<head>`, or more than 8 redirects. |
| `no-metadata` | The page was read but has no title, description or image. |
| `not-html` | The address is a PDF, JSON, video or other file that is not a web page or an image. |
| `bad-input` | Not an http(s) URL, a URL with a user name or password, or an address in a private network. |
| `no-change` | Monitor mode: nothing changed (one summary row), or a change could not be confirmed by a second read. |
| `budget-reached` | The run reached the maximum total charge you set; the rest was not requested. |

These rows keep the same columns as a preview row, with `note` explaining the reason and `httpStatus`, `finalUrl` and `redirects` filled in when known.

### Pricing

Pay per event:

- **Preview returned**: charged for each row with `status: "ok"`. A URL that is itself an image (Content-Type `image/*`) is an `ok` row with `image` set to that URL and the other preview values (title, description, site name and so on) null.
- **Run start**: charged once per run that returns at least one preview. In monitor mode it is charged once per run that read and compared at least one URL, whether or not anything changed. It is not charged when no URL could be previewed or the input could not be used.

A run whose maximum total charge has no room for the run start fee plus one preview requests nothing and is charged nothing. The current prices are on the Pricing tab.

### Calling it from code

Synchronous call that returns the rows directly (replace `<YOUR_TOKEN>`):

```bash
curl -X POST "https://api.apify.com/v2/acts/neverempty~link-preview-api/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://github.com/apify/crawlee"]}'
```

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('neverempty/link-preview-api').call({ urls: ['https://github.com/apify/crawlee'] });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].title, items[0].image);
```

### Monitor mode

Run the same list on a schedule with `onlyChanges: true` and a `watchName`. The first run returns every URL (`changeType: "first-check"`). Later runs return only URLs whose title, description, image, site name, canonical URL or type changed (`changeType: "changed"`, with `changedFields`, `previousTitle`, `previousDescription`, `previousImage` and `previousCheckedAt`), and URLs added to the list (`"new"`).

- A change is reported only when a second read a moment later shows the same new value. A value that differs between the two reads (for example an image that rotates on every visit) is not reported and is not compared for that URL again.
- Image URLs are compared without their query string, so signed or cache-busting CDN parameters do not count as a change.
- Use a different `watchName` for each list you track on its own schedule, and avoid two overlapping schedules with the same `watchName` (the remembered state is saved after each run and Apify has no atomic update).

### Measured on 2026-09-24

40 commonly shared URLs (news sites, GitHub, YouTube, Wikipedia, Spotify, Vimeo, PyPI, social networks, Japanese, French and German sites) in one run from Apify: **23 previews**, 6 not requested because robots.txt disallows them (the big social networks), 7 refused or check pages, 2 not found, 2 sites that did not answer. The run took 54 seconds with 4 in parallel; peak memory 176 MB of 256 MB.

### Limits

- JavaScript is not run. Sites that write their tags only in the browser (single-page apps without server rendering) may return fewer values; the row says what the served HTML contains.
- Pages behind a sign-in, and sites whose robots.txt disallows automated reading, are not read. Most large social networks disallow it.
- Values are what the site serves to a reader in a US data center; a site may serve different content by country or language.
- This is an independent tool and is not affiliated with any website it reads.

### Support

Questions, a URL that returns something unexpected, or a missing tag: open an issue in the Issues tab with the URL and the run ID.

# Actor input Schema

## `urls` (type: `array`):

Page URLs to turn into link previews, one per line (1 to 1,000 per run). A URL without http:// or https:// is read as https://. The same URL given twice (also when only the #fragment differs) is read and charged once. If this is empty, the example URL https://github.com/apify/crawlee is used.

## `acceptLanguage` (type: `string`):

Optional. Sent as the Accept-Language header, for sites that show a different language version (and different title and description) by language, for example "en", "ja-JP" or "de-DE,de;q=0.9". Leave empty to get the version the site shows by default (nothing is sent). The language column reports what the page itself declares.

## `includeOembed` (type: `boolean`):

When a page links to its own oEmbed JSON endpoint in its <head> (for example YouTube and Spotify), also read that endpoint and add the oembed object (type, provider, author, thumbnail, embed HTML, width, height). It is one extra small request for those pages only. The endpoint's robots.txt is respected too.

## `onlyChanges` (type: `boolean`):

Return only URLs whose title, description, image, site name, canonical URL or type changed since the last run with the same watch name (plus URLs new to the watch). The first run returns every URL as the starting point. A change is confirmed by reading the page a second time a moment later; a value that differs between those two reads (for example a rotating image) is not reported and is ignored for that URL from then on. Image URLs are compared without their query string. A run in which nothing changed returns a free row saying so and charges only the run start fee.

## `watchName` (type: `string`):

Name of the remembered state used to compare runs (letters, digits, dot, dash, underscore; up to 40). Setting it (or turning on monitor mode) fills changeType and the previous-value columns. Use a different name for each list of URLs you track on its own schedule. With monitor mode on and no name, the name "default" is used.

## `resetMonitoringState` (type: `boolean`):

Start this watch over: forget the remembered previews before this run, so every URL is returned as a first check.

## `maxConcurrency` (type: `integer`):

How many URLs are read in parallel (1 to 8). Each page is read only up to the end of its <head>, so memory stays small.

## `requestTimeoutSecs` (type: `integer`):

How long to wait for one site to answer (5 to 60 seconds). HTTP 429, 500, 502, 503 and 504 are asked again up to two more times; a site that does not answer in time is asked once more.

## Actor input object example

```json
{
  "urls": [
    "https://github.com/apify/crawlee",
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ],
  "includeOembed": true,
  "onlyChanges": false,
  "resetMonitoringState": false,
  "maxConcurrency": 4,
  "requestTimeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

One row per URL: title, description, image (with width, height and alt), site name, favicon, canonical URL, language, type, Twitter Card, author, dates, oEmbed and the raw Open Graph and Twitter tags, plus the final URL after redirects. In monitor mode, changeType, changedFields and the previous title, description and image. A URL that robots.txt does not allow, that needs a sign-in, that shows a check page, that does not exist or that has no preview comes back as a free row that says why.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://github.com/apify/crawlee",
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("neverempty/link-preview-api").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://github.com/apify/crawlee",
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("neverempty/link-preview-api").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://github.com/apify/crawlee",
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ"
  ]
}' |
apify call neverempty/link-preview-api --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,neverempty/link-preview-api"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/VDNbNmVFWHFrW07x3/builds/eZplP3GZiDR8JyUYH/openapi.json
