# Link Preview Metadata: Open Graph, Twitter Card, Favicon (`jtpalms/link-preview-metadata`) Actor

Get what a link unfurler shows for any URL: title, description, image, site name, every Open Graph and Twitter card tag, best favicon, theme color, language, author, publish date, JSON-LD types, oEmbed URL, H1 and word count. For app link previews and AI agents. USD 1 per 1,000 URLs.

- **URL**: https://apify.com/jtpalms/link-preview-metadata.md
- **Developed by:** [JT Palms](https://apify.com/jtpalms) (community)
- **Categories:** Developer tools, AI, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 url previeweds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Link Preview Metadata: Open Graph, Twitter card and favicon for any URL

Give it URLs and get back what a link unfurler (Slack, iMessage, X, LinkedIn, an AI agent) would show: title, description, preview image, site name and favicon, plus every Open Graph and Twitter card tag, canonical URL, language, author, published and modified dates, JSON-LD types, oEmbed endpoint, first H1 and a word count.

**USD 1 per 1,000 URLs.** Malformed URLs and pages that fail to load are free.

### What people use it for

- **Link previews in your app.** Show rich cards for links your users paste, without running a headless browser or fighting with each site's quirks.
- **AI agents and RAG.** Give an LLM a clean summary of a URL (title, description, type, author, date, language, word count) before deciding whether to read the full page.
- **Social card QA.** Check that every page has `og:title`, `og:description`, `og:image` and a Twitter card before a launch.
- **Content audits.** Pull titles, descriptions, canonicals, authors, publish dates, H1s and word counts for a list of articles into one spreadsheet.
- **Favicon and brand lookup.** Get the best available icon for a domain, from `<link rel="icon">`, `apple-touch-icon` or the web app manifest.

### Sample output

A real news article, trimmed:

```json
{
  "input": "https://www.bbc.com/news/articles/cqx2z898p6p4o",
  "ok": true,
  "url": "https://www.bbc.com/news/articles/cqx2z898p6p4o",
  "statusCode": 200,
  "title": "Utility firms under fire over Ormskirk pavement repairs",
  "description": "Residents say streets have been left with temporary-looking repairs instead of matching stone flags.",
  "image": "https://ichef.bbci.co.uk/news/1024/branded_news/c2eb/live/c6c68a30-baa0-11f1-afa8-43b99a243ced.jpg",
  "type": "article",
  "canonicalUrl": "https://www.bbc.com/news/articles/cqx2z898p6p4o",
  "openGraph": {
    "og:title": "Utility firms under fire over Ormskirk pavement repairs",
    "og:type": "article",
    "og:image": "https://ichef.bbci.co.uk/news/1024/branded_news/c2eb/live/c6c68a30-baa0-11f1-afa8-43b99a243ced.jpg",
    "og:image:width": "1024",
    "og:image:height": "576"
  },
  "twitterCard": { "twitter:card": "summary_large_image", "twitter:title": "Utility firms under fire over Ormskirk pavement repairs" },
  "article": { "article:modified_time": "2026-09-27T19:20:10.231Z" },
  "language": "en-GB",
  "author": "Paul Faulkner",
  "publishedTime": "2026-09-27T19:20:10.231Z",
  "modifiedTime": "2026-09-27T19:20:10.231Z",
  "themeColor": "#ffffff",
  "jsonLdTypes": ["ReportageNewsArticle"],
  "oembedUrl": null,
  "h1": "Utility firms under fire over pavement repairs",
  "wordCount": 865,
  "favicon": "https://static.files.bbci.co.uk/bbcdotcom/web/20260922-091849-a8520cb5b4-web-3.22.0-3/android-chrome-512x512.png",
  "faviconSizes": "512x512",
  "faviconSource": "manifest"
}
```

A YouTube video adds the video tags and the oEmbed endpoint:

```json
{
  "title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
  "image": "https://i.ytimg.com/vi/dQw4w9WgXcQ/maxresdefault.jpg",
  "siteName": "YouTube",
  "type": "video.other",
  "openGraph": { "og:video:url": "https://www.youtube.com/embed/dQw4w9WgXcQ", "og:video:width": "1280", "og:video:height": "720" },
  "jsonLdTypes": ["VideoObject"],
  "publishedTime": "2009-10-24T23:57:33-07:00",
  "oembedUrl": "https://www.youtube.com/oembed?format=json&url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DdQw4w9WgXcQ",
  "favicon": "https://www.gstatic.com/youtube/img/branding/favicon/favicon_192x192_v2.png"
}
```

A page with nothing but a title (`https://example.com`) returns `"title": "Example Domain"`, empty `openGraph` and `twitterCard`, and `null` for the image and description, so missing tags are easy to spot.

| Field | What it is |
|---|---|
| `title`, `description`, `image`, `siteName`, `type` | The preview fields an unfurler shows: Open Graph first, then Twitter card, then the plain HTML `<title>` and meta description. `image` is made absolute. |
| `htmlTitle`, `metaDescription` | The plain HTML title and meta description, for SEO comparisons. |
| `openGraph`, `twitterCard`, `article` | Every `og:*`, `twitter:*` and `article:*` tag as found. Repeated tags (several `og:image`) become arrays. |
| `canonicalUrl`, `language`, `author`, `themeColor` | From `<link rel="canonical">`, `<html lang>`, meta author (or article:author or JSON-LD) and `theme-color`. |
| `publishedTime`, `modifiedTime` | From `article:published_time` / `modified_time`, `og:updated_time` or JSON-LD dates. |
| `jsonLdTypes` | Schema.org types in the page's JSON-LD, such as `NewsArticle`, `Product`, `Recipe`, `Organization`. |
| `oembedUrl` | The oEmbed endpoint the page advertises, if any. |
| `favicon`, `faviconSizes`, `faviconSource`, `icons` | The best icon (the largest declared, from link tags or the manifest; `/favicon.ico` only if it really exists), plus every icon found. |
| `h1`, `wordCount` | The first H1 and an estimate of the visible words (scripts and styles excluded). |
| `url`, `statusCode`, `redirected`, `contentType` | The final URL after redirects and its response. |
| `ok`, `errorType`, `error` | `ok` is `false` for pages that could not be loaded, with `DNS`, `TLS`, `TIMEOUT`, `HTTP_ERROR` and so on. Those rows are free. |

### How to use it

1. Paste URLs into **URLs**, one per line, or send a JSON array through the API.
2. Optional: set **Fetch as** to Chrome on Android to see the mobile page.
3. Click **Start**. Export CSV, JSON or Excel, or call the actor from your app through the Apify API and read the dataset.

For an app, call it synchronously through the API (`run-sync-get-dataset-items`) with a handful of URLs and cache the results.

### Pricing

| What | Price |
|---|---|
| URL previewed (a page or file that loaded with 2xx) | USD 0.001 (USD 1 per 1,000) |
| Malformed URL, DNS or TLS error, timeout, 4xx or 5xx | Free |
| Duplicate URLs in the list | Fetched and billed once |

Set a maximum cost per run in the run options; the actor stops cleanly when it is reached.

### Limits

- It reads the HTML the server sends, like Slack and Facebook do, and does not run JavaScript. Sites that add their tags with JavaScript only show what is in the raw HTML.
- Up to 5 MB of HTML per page. Image dimensions are not measured; use `og:image:width` and `og:image:height` when the site provides them.
- Some sites block data-center traffic. Use a proxy if you see 403 or 429 for sites that work in your browser.

### Related actors

- [Website Screenshot API: Full Page, Mobile, PDF](https://apify.com/JTPalms/website-screenshot): Screenshot any list of URLs as full-page or viewport PNG, JPEG, WebP or PDF, on desktop, laptop, tablet or mobile.
- [Website Tech Stack Detector (Wappalyzer Alternative)](https://apify.com/JTPalms/tech-stack-detector): Find the technology behind any website in bulk: CMS, ecommerce platform, analytics, frameworks, CDN, hosting, payment, email and...

### FAQ

**Why is the title different from what I see in the browser tab?** Link previews use `og:title` first. The browser tab uses `<title>`, which is in `htmlTitle`.

**Does it fetch the oEmbed data?** It returns the endpoint URL. Call it yourself if you need the embed HTML.

**Is it legal?** It requests public URLs that you supply, the same way a browser or a chat app's link preview does, and reads the page's own metadata. It does not log in and collects no personal data beyond what the page publishes about itself (for example the article's author line).

**Something missing or wrong?** Open an issue with the URL.

# Actor input Schema

## `urls` (type: `array`):

One URL per line. A URL without http:// or https:// is fetched as https://. Duplicates are fetched once. Malformed URLs and pages that fail to load are reported free.

## `userAgent` (type: `string`):

Which browser the request presents as. Chrome sends realistic browser headers. Custom sends the user agent text below, for example your own bot name.

## `customUserAgent` (type: `string`):

Only used when Fetch as is Custom, for example MyAppPreview/1.0 (+https://example.com).

## `fetchManifest` (type: `boolean`):

Also read the site's web app manifest (one small extra request) so the best favicon can come from its icon list, which often has the largest icons.

## `maxConcurrency` (type: `integer`):

How many URLs to fetch at the same time.

## `timeoutSecs` (type: `integer`):

How long to wait for one page before reporting it as timed out.

## `proxyConfiguration` (type: `object`):

Optional. Only needed if a site blocks requests from data centers.

## Actor input object example

```json
{
  "urls": [
    "https://github.com/apify/crawlee",
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://example.com"
  ],
  "userAgent": "chrome-desktop",
  "fetchManifest": true,
  "maxConcurrency": 10,
  "timeoutSecs": 20
}
```

# Actor output Schema

## `results` (type: `string`):

All output rows in the default dataset (JSON, CSV, Excel via the format parameter).

## `summary` (type: `string`):

Counts and per-input status for this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://github.com/apify/crawlee",
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://example.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("jtpalms/link-preview-metadata").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://github.com/apify/crawlee",
        "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
        "https://example.com",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("jtpalms/link-preview-metadata").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://github.com/apify/crawlee",
    "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
    "https://example.com"
  ]
}' |
apify call jtpalms/link-preview-metadata --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jtpalms/link-preview-metadata"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/n0t5pcRcc5gdxxCJG/builds/pZf3Z8fug1VMytjXH/openapi.json
