Link Preview API - Open Graph, Twitter Card & oEmbed Metadata avatar

Link Preview API - Open Graph, Twitter Card & oEmbed Metadata

Pricing

from $2.40 / 1,000 preview returneds

Go to Apify Store
Link Preview API - Open Graph, Twitter Card & oEmbed Metadata

Link Preview API - Open Graph, Twitter Card & oEmbed Metadata

For chat apps, CMS editors and AI agents that unfurl links: send URLs, get title, description, image, site name, favicon, canonical URL and language from Open Graph, Twitter Card, oEmbed and the HTML head. About a second per URL. Respects robots.txt; URLs that cannot be previewed are free.

Pricing

from $2.40 / 1,000 preview returneds

Rating

0.0

(0)

Developer

NeverEmpty

NeverEmpty

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Give it a URL, get back the link preview: title, description, image, site name, favicon, canonical URL, language and type in one clean JSON row, read from the page's Open Graph tags, Twitter Card tags, oEmbed endpoint and HTML <head>.

Built for software that unfurls links: chat apps and bots (Slack, Discord, Telegram, Teams), CMS and newsletter editors, bookmarking and read-later apps, CRMs, AI agents and RAG pipelines that need a title and thumbnail for a URL, and SEO tools that check og:image and twitter:card tags.

  • One URL in about a second. Only the <head> of each page is read; the download stops at </head>, so a 1.3 MB YouTube page costs the same memory as a small blog post. A run with one URL finished in 2.5 to 3.2 seconds on Apify, including start-up (2026-09-24).
  • Only real values. A page with no og:image has image: null. The favicon comes from the page's own <link rel="icon">; /favicon.ico is not guessed. Nothing is filled in with made-up defaults.
  • Nothing is charged for a URL that could not be previewed. A page that does not exist, refuses automated reading, needs a sign-in or has no title, description or image comes back as a free row that says why.
  • Respects each site's robots.txt (RFC 9309) for every URL, every redirect and every oEmbed endpoint, and never signs in, solves check pages or rotates proxies to get around a refusal.
  • Any character set. The charset is taken from the response header, a byte-order mark or the <meta charset> in the page itself, and the page is decoded with that character set instead of being forced into UTF-8 (tested on kakaku.com, which declares Shift_JIS only in <meta>).
  • Monitor mode returns only URLs whose preview changed since the last run, with the previous title, description and image. Each change is confirmed by a second read, so a rotating image is not reported as a change.

Input

FieldTypeDefaultWhat it does
urlsarray of strings(empty: the example https://github.com/apify/crawlee is used)Page URLs, 1 to 1,000 per run. A URL without http:// or https:// is read as https://. The same URL given twice (also when only the #fragment differs) is read and charged once.
acceptLanguagestring(not sent)Optional Accept-Language header, for sites that show a different language version by language, for example en, ja-JP or de-DE,de;q=0.9. Empty means the version the site shows by default.
includeOembedbooleantrueWhen the page links to its own oEmbed JSON endpoint in its <head> (on 2026-09-24: YouTube and Spotify among the 23 pages read), also read it and add the oembed object.
onlyChangesbooleanfalseMonitor mode: return only URLs whose title, description, image, site name, canonical URL or type changed since the last run with the same watchName, plus URLs new to the watch.
watchNamestring(none)Name of the remembered state (letters, digits, ., -, _; up to 40). Setting it fills changeType and the previous-value columns. With monitor mode on and no name, default is used.
resetMonitoringStatebooleanfalseForget what this watch remembered, so every URL is returned as a first check.
maxConcurrencyinteger4URLs read in parallel, 1 to 8.
requestTimeoutSecsinteger20Time limit for one request, 5 to 60 seconds. HTTP 429, 500, 502, 503 and 504 are asked again up to two more times; a site that does not answer in time is asked once more.
{
"urls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ", "github.com/apify/crawlee"]
}

Output

One row per URL. A real row from 2026-09-24 (long text shortened here):

{
"status": "ok",
"position": 1,
"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"finalUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"httpStatus": 200,
"contentType": "text/html; charset=utf-8",
"redirects": [],
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"description": "The official video for “Never Gonna Give You Up” by Rick Astley. ...",
"image": "https://i.ytimg.com/vi/dQw4w9WgXcQ/maxresdefault.jpg",
"imageWidth": 1280,
"imageHeight": 720,
"imageAlt": null,
"siteName": "YouTube",
"favicon": "https://www.youtube.com/s/desktop/02b72088/img/favicon_32x32.png",
"appleTouchIcon": null,
"canonicalUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"ogUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
"language": "en",
"languageSource": "html",
"type": "video.other",
"twitterCard": "summary_large_image",
"twitterSite": "@youtube",
"author": "Rick Astley",
"publishedTime": null,
"modifiedTime": null,
"themeColor": "rgba(255, 255, 255, 0.98)",
"video": "https://www.youtube.com/embed/dQw4w9WgXcQ",
"oembed": {
"type": "video",
"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)",
"authorName": "Rick Astley",
"authorUrl": "https://www.youtube.com/@RickAstleyYT",
"providerName": "YouTube",
"providerUrl": "https://www.youtube.com/",
"thumbnailUrl": "https://i.ytimg.com/vi/dQw4w9WgXcQ/hqdefault.jpg",
"thumbnailWidth": 480,
"thumbnailHeight": 360,
"html": "<iframe width=\"200\" height=\"113\" src=\"https://www....",
"width": 200,
"height": 113
},
"oembedEndpoint": "https://www.youtube.com/oembed?format=json&url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DdQw4w9WgXcQ",
"jsonLdTypes": ["VideoObject"],
"robotsMeta": null,
"sources": ["openGraph", "oembed"],
"openGraph": { "title": "...", "type": "video.other", "...": "..." },
"twitter": { "card": "summary_large_image", "site": "@youtube", "...": "..." },
"changeType": null,
"changedFields": null,
"previousTitle": null,
"previousDescription": null,
"previousImage": null,
"previousCheckedAt": null,
"charset": "utf-8",
"headBytesRead": 716367,
"headComplete": true,
"watchName": null,
"note": null,
"fetchedAt": "2026-09-24T13:07:11.920Z"
}

Where each value comes from:

ColumnTaken from (first one present)
titleog:title, twitter:title, <title>, JSON-LD headline, oEmbed title
descriptionog:description, twitter:description, <meta name="description">
imageog:image:secure_url / og:image, twitter:image, <link rel="image_src">, oEmbed thumbnail_url; for a URL that is itself an image, the URL
imageWidth, imageHeight, imageAltog:image:width, og:image:height, og:image:alt (or twitter:image:alt) as the page states them; the image is not downloaded
siteNameog:site_name, <meta name="application-name">, oEmbed provider_name
favicon, appleTouchIcon<link rel="icon"> (the size closest to 32 px), <link rel="apple-touch-icon">
canonicalUrl, ogUrl<link rel="canonical">, og:url
language, languageSource<html lang>, the Content-Language header, og:locale
typeog:type
author, publishedTime, modifiedTime<meta name="author">, article:author, article:published_time, article:modified_time, og:updated_time, JSON-LD, oEmbed author_name
openGraph, twitterevery og:* and twitter:* tag, as written
sourceswhich of openGraph, twitterCard, html, jsonLd supplied the title, description, image or site name, plus oembed when the oEmbed endpoint was read, or url when the URL is itself an image

Relative URLs are resolved against the final URL (and <base href>), so image, favicon and canonicalUrl are always absolute.

Rows that are not charged

statusMeaning
robots-disallowedThe site's robots.txt does not allow automated reading of this address; it was not requested.
robots-unreachableThe site's robots.txt answered 5xx or could not be reached; following RFC 9309 the page was not requested.
login-requiredHTTP 401, or the page redirects to a sign-in page.
blockedThe site refused this reader (HTTP 403, 429 after retries, 451) or showed a CAPTCHA or browser-check page (with any status; a check page is never asked again). A person with a browser may still see the page.
not-foundHTTP 404 or 410.
http-errorAnother HTTP answer that is not a page (for example 400 or 405), or a broken redirect.
unreachableThe domain does not resolve, or the site did not answer.
unreadable5xx after retries, an answer that kept stopping before the end of the page's <head>, or more than 8 redirects.
no-metadataThe page was read but has no title, description or image.
not-htmlThe address is a PDF, JSON, video or other file that is not a web page or an image.
bad-inputNot an http(s) URL, a URL with a user name or password, or an address in a private network.
no-changeMonitor mode: nothing changed (one summary row), or a change could not be confirmed by a second read.
budget-reachedThe run reached the maximum total charge you set; the rest was not requested.

These rows keep the same columns as a preview row, with note explaining the reason and httpStatus, finalUrl and redirects filled in when known.

Pricing

Pay per event:

  • Preview returned: charged for each row with status: "ok". A URL that is itself an image (Content-Type image/*) is an ok row with image set to that URL and the other preview values (title, description, site name and so on) null.
  • Run start: charged once per run that returns at least one preview. In monitor mode it is charged once per run that read and compared at least one URL, whether or not anything changed. It is not charged when no URL could be previewed or the input could not be used.

A run whose maximum total charge has no room for the run start fee plus one preview requests nothing and is charged nothing. The current prices are on the Pricing tab.

Calling it from code

Synchronous call that returns the rows directly (replace <YOUR_TOKEN>):

curl -X POST "https://api.apify.com/v2/acts/neverempty~link-preview-api/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \
-H "Content-Type: application/json" \
-d '{"urls": ["https://github.com/apify/crawlee"]}'
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('neverempty/link-preview-api').call({ urls: ['https://github.com/apify/crawlee'] });
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].title, items[0].image);

Monitor mode

Run the same list on a schedule with onlyChanges: true and a watchName. The first run returns every URL (changeType: "first-check"). Later runs return only URLs whose title, description, image, site name, canonical URL or type changed (changeType: "changed", with changedFields, previousTitle, previousDescription, previousImage and previousCheckedAt), and URLs added to the list ("new").

  • A change is reported only when a second read a moment later shows the same new value. A value that differs between the two reads (for example an image that rotates on every visit) is not reported and is not compared for that URL again.
  • Image URLs are compared without their query string, so signed or cache-busting CDN parameters do not count as a change.
  • Use a different watchName for each list you track on its own schedule, and avoid two overlapping schedules with the same watchName (the remembered state is saved after each run and Apify has no atomic update).

Measured on 2026-09-24

40 commonly shared URLs (news sites, GitHub, YouTube, Wikipedia, Spotify, Vimeo, PyPI, social networks, Japanese, French and German sites) in one run from Apify: 23 previews, 6 not requested because robots.txt disallows them (the big social networks), 7 refused or check pages, 2 not found, 2 sites that did not answer. The run took 54 seconds with 4 in parallel; peak memory 176 MB of 256 MB.

Limits

  • JavaScript is not run. Sites that write their tags only in the browser (single-page apps without server rendering) may return fewer values; the row says what the served HTML contains.
  • Pages behind a sign-in, and sites whose robots.txt disallows automated reading, are not read. Most large social networks disallow it.
  • Values are what the site serves to a reader in a US data center; a site may serve different content by country or language.
  • This is an independent tool and is not affiliated with any website it reads.

Support

Questions, a URL that returns something unexpected, or a missing tag: open an issue in the Issues tab with the URL and the run ID.