Link Preview API - Open Graph, Twitter Card & oEmbed Metadata
Pricing
from $2.40 / 1,000 preview returneds
Link Preview API - Open Graph, Twitter Card & oEmbed Metadata
For chat apps, CMS editors and AI agents that unfurl links: send URLs, get title, description, image, site name, favicon, canonical URL and language from Open Graph, Twitter Card, oEmbed and the HTML head. About a second per URL. Respects robots.txt; URLs that cannot be previewed are free.
Pricing
from $2.40 / 1,000 preview returneds
Rating
0.0
(0)
Developer
NeverEmpty
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Give it a URL, get back the link preview: title, description, image, site name, favicon, canonical URL, language and type in one clean JSON row, read from the page's Open Graph tags, Twitter Card tags, oEmbed endpoint and HTML <head>.
Built for software that unfurls links: chat apps and bots (Slack, Discord, Telegram, Teams), CMS and newsletter editors, bookmarking and read-later apps, CRMs, AI agents and RAG pipelines that need a title and thumbnail for a URL, and SEO tools that check og:image and twitter:card tags.
- One URL in about a second. Only the
<head>of each page is read; the download stops at</head>, so a 1.3 MB YouTube page costs the same memory as a small blog post. A run with one URL finished in 2.5 to 3.2 seconds on Apify, including start-up (2026-09-24). - Only real values. A page with no
og:imagehasimage: null. The favicon comes from the page's own<link rel="icon">;/favicon.icois not guessed. Nothing is filled in with made-up defaults. - Nothing is charged for a URL that could not be previewed. A page that does not exist, refuses automated reading, needs a sign-in or has no title, description or image comes back as a free row that says why.
- Respects each site's robots.txt (RFC 9309) for every URL, every redirect and every oEmbed endpoint, and never signs in, solves check pages or rotates proxies to get around a refusal.
- Any character set. The
charsetis taken from the response header, a byte-order mark or the<meta charset>in the page itself, and the page is decoded with that character set instead of being forced into UTF-8 (tested on kakaku.com, which declares Shift_JIS only in<meta>). - Monitor mode returns only URLs whose preview changed since the last run, with the previous title, description and image. Each change is confirmed by a second read, so a rotating image is not reported as a change.
Input
| Field | Type | Default | What it does |
|---|---|---|---|
urls | array of strings | (empty: the example https://github.com/apify/crawlee is used) | Page URLs, 1 to 1,000 per run. A URL without http:// or https:// is read as https://. The same URL given twice (also when only the #fragment differs) is read and charged once. |
acceptLanguage | string | (not sent) | Optional Accept-Language header, for sites that show a different language version by language, for example en, ja-JP or de-DE,de;q=0.9. Empty means the version the site shows by default. |
includeOembed | boolean | true | When the page links to its own oEmbed JSON endpoint in its <head> (on 2026-09-24: YouTube and Spotify among the 23 pages read), also read it and add the oembed object. |
onlyChanges | boolean | false | Monitor mode: return only URLs whose title, description, image, site name, canonical URL or type changed since the last run with the same watchName, plus URLs new to the watch. |
watchName | string | (none) | Name of the remembered state (letters, digits, ., -, _; up to 40). Setting it fills changeType and the previous-value columns. With monitor mode on and no name, default is used. |
resetMonitoringState | boolean | false | Forget what this watch remembered, so every URL is returned as a first check. |
maxConcurrency | integer | 4 | URLs read in parallel, 1 to 8. |
requestTimeoutSecs | integer | 20 | Time limit for one request, 5 to 60 seconds. HTTP 429, 500, 502, 503 and 504 are asked again up to two more times; a site that does not answer in time is asked once more. |
{"urls": ["https://www.youtube.com/watch?v=dQw4w9WgXcQ", "github.com/apify/crawlee"]}
Output
One row per URL. A real row from 2026-09-24 (long text shortened here):
{"status": "ok","position": 1,"url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","finalUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","httpStatus": 200,"contentType": "text/html; charset=utf-8","redirects": [],"title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","description": "The official video for “Never Gonna Give You Up” by Rick Astley. ...","image": "https://i.ytimg.com/vi/dQw4w9WgXcQ/maxresdefault.jpg","imageWidth": 1280,"imageHeight": 720,"imageAlt": null,"siteName": "YouTube","favicon": "https://www.youtube.com/s/desktop/02b72088/img/favicon_32x32.png","appleTouchIcon": null,"canonicalUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","ogUrl": "https://www.youtube.com/watch?v=dQw4w9WgXcQ","language": "en","languageSource": "html","type": "video.other","twitterCard": "summary_large_image","twitterSite": "@youtube","author": "Rick Astley","publishedTime": null,"modifiedTime": null,"themeColor": "rgba(255, 255, 255, 0.98)","video": "https://www.youtube.com/embed/dQw4w9WgXcQ","oembed": {"type": "video","title": "Rick Astley - Never Gonna Give You Up (Official Video) (4K Remaster)","authorName": "Rick Astley","authorUrl": "https://www.youtube.com/@RickAstleyYT","providerName": "YouTube","providerUrl": "https://www.youtube.com/","thumbnailUrl": "https://i.ytimg.com/vi/dQw4w9WgXcQ/hqdefault.jpg","thumbnailWidth": 480,"thumbnailHeight": 360,"html": "<iframe width=\"200\" height=\"113\" src=\"https://www....","width": 200,"height": 113},"oembedEndpoint": "https://www.youtube.com/oembed?format=json&url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DdQw4w9WgXcQ","jsonLdTypes": ["VideoObject"],"robotsMeta": null,"sources": ["openGraph", "oembed"],"openGraph": { "title": "...", "type": "video.other", "...": "..." },"twitter": { "card": "summary_large_image", "site": "@youtube", "...": "..." },"changeType": null,"changedFields": null,"previousTitle": null,"previousDescription": null,"previousImage": null,"previousCheckedAt": null,"charset": "utf-8","headBytesRead": 716367,"headComplete": true,"watchName": null,"note": null,"fetchedAt": "2026-09-24T13:07:11.920Z"}
Where each value comes from:
| Column | Taken from (first one present) |
|---|---|
title | og:title, twitter:title, <title>, JSON-LD headline, oEmbed title |
description | og:description, twitter:description, <meta name="description"> |
image | og:image:secure_url / og:image, twitter:image, <link rel="image_src">, oEmbed thumbnail_url; for a URL that is itself an image, the URL |
imageWidth, imageHeight, imageAlt | og:image:width, og:image:height, og:image:alt (or twitter:image:alt) as the page states them; the image is not downloaded |
siteName | og:site_name, <meta name="application-name">, oEmbed provider_name |
favicon, appleTouchIcon | <link rel="icon"> (the size closest to 32 px), <link rel="apple-touch-icon"> |
canonicalUrl, ogUrl | <link rel="canonical">, og:url |
language, languageSource | <html lang>, the Content-Language header, og:locale |
type | og:type |
author, publishedTime, modifiedTime | <meta name="author">, article:author, article:published_time, article:modified_time, og:updated_time, JSON-LD, oEmbed author_name |
openGraph, twitter | every og:* and twitter:* tag, as written |
sources | which of openGraph, twitterCard, html, jsonLd supplied the title, description, image or site name, plus oembed when the oEmbed endpoint was read, or url when the URL is itself an image |
Relative URLs are resolved against the final URL (and <base href>), so image, favicon and canonicalUrl are always absolute.
Rows that are not charged
status | Meaning |
|---|---|
robots-disallowed | The site's robots.txt does not allow automated reading of this address; it was not requested. |
robots-unreachable | The site's robots.txt answered 5xx or could not be reached; following RFC 9309 the page was not requested. |
login-required | HTTP 401, or the page redirects to a sign-in page. |
blocked | The site refused this reader (HTTP 403, 429 after retries, 451) or showed a CAPTCHA or browser-check page (with any status; a check page is never asked again). A person with a browser may still see the page. |
not-found | HTTP 404 or 410. |
http-error | Another HTTP answer that is not a page (for example 400 or 405), or a broken redirect. |
unreachable | The domain does not resolve, or the site did not answer. |
unreadable | 5xx after retries, an answer that kept stopping before the end of the page's <head>, or more than 8 redirects. |
no-metadata | The page was read but has no title, description or image. |
not-html | The address is a PDF, JSON, video or other file that is not a web page or an image. |
bad-input | Not an http(s) URL, a URL with a user name or password, or an address in a private network. |
no-change | Monitor mode: nothing changed (one summary row), or a change could not be confirmed by a second read. |
budget-reached | The run reached the maximum total charge you set; the rest was not requested. |
These rows keep the same columns as a preview row, with note explaining the reason and httpStatus, finalUrl and redirects filled in when known.
Pricing
Pay per event:
- Preview returned: charged for each row with
status: "ok". A URL that is itself an image (Content-Typeimage/*) is anokrow withimageset to that URL and the other preview values (title, description, site name and so on) null. - Run start: charged once per run that returns at least one preview. In monitor mode it is charged once per run that read and compared at least one URL, whether or not anything changed. It is not charged when no URL could be previewed or the input could not be used.
A run whose maximum total charge has no room for the run start fee plus one preview requests nothing and is charged nothing. The current prices are on the Pricing tab.
Calling it from code
Synchronous call that returns the rows directly (replace <YOUR_TOKEN>):
curl -X POST "https://api.apify.com/v2/acts/neverempty~link-preview-api/run-sync-get-dataset-items?token=<YOUR_TOKEN>" \-H "Content-Type: application/json" \-d '{"urls": ["https://github.com/apify/crawlee"]}'
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('neverempty/link-preview-api').call({ urls: ['https://github.com/apify/crawlee'] });const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items[0].title, items[0].image);
Monitor mode
Run the same list on a schedule with onlyChanges: true and a watchName. The first run returns every URL (changeType: "first-check"). Later runs return only URLs whose title, description, image, site name, canonical URL or type changed (changeType: "changed", with changedFields, previousTitle, previousDescription, previousImage and previousCheckedAt), and URLs added to the list ("new").
- A change is reported only when a second read a moment later shows the same new value. A value that differs between the two reads (for example an image that rotates on every visit) is not reported and is not compared for that URL again.
- Image URLs are compared without their query string, so signed or cache-busting CDN parameters do not count as a change.
- Use a different
watchNamefor each list you track on its own schedule, and avoid two overlapping schedules with the samewatchName(the remembered state is saved after each run and Apify has no atomic update).
Measured on 2026-09-24
40 commonly shared URLs (news sites, GitHub, YouTube, Wikipedia, Spotify, Vimeo, PyPI, social networks, Japanese, French and German sites) in one run from Apify: 23 previews, 6 not requested because robots.txt disallows them (the big social networks), 7 refused or check pages, 2 not found, 2 sites that did not answer. The run took 54 seconds with 4 in parallel; peak memory 176 MB of 256 MB.
Limits
- JavaScript is not run. Sites that write their tags only in the browser (single-page apps without server rendering) may return fewer values; the row says what the served HTML contains.
- Pages behind a sign-in, and sites whose robots.txt disallows automated reading, are not read. Most large social networks disallow it.
- Values are what the site serves to a reader in a US data center; a site may serve different content by country or language.
- This is an independent tool and is not affiliated with any website it reads.
Support
Questions, a URL that returns something unexpected, or a missing tag: open an issue in the Issues tab with the URL and the run ID.