URL Metadata & Open Graph Extractor (Bulk) avatar

URL Metadata & Open Graph Extractor (Bulk)

Pricing

from $4.00 / 1,000 page reads

Go to Apify Store
URL Metadata & Open Graph Extractor (Bulk)

URL Metadata & Open Graph Extractor (Bulk)

Paste any list of URLs and get for each page: title, meta description, Open Graph, Twitter card, favicon, canonical, language, author, dates, keywords, structured data types and the link preview a chat app would show. Respects robots.txt.

Pricing

from $4.00 / 1,000 page reads

Rating

0.0

(0)

Developer

Bruno Petrelli

Bruno Petrelli

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

12 hours ago

Last modified

Share

Paste a list of URLs and get one clean row of metadata per page:

  • Basics: title, meta description, canonical URL, language, first H1, word count.
  • Open Graph: og:title, og:description, og:image, og:type, og:url, site name.
  • Twitter / X card: card type, title, description, image, @site.
  • Favicon: the largest icon the page declares (or /favicon.ico, flagged as a guess).
  • Article data: author, published and modified dates (as ISO dates), keywords.
  • Indexing: meta robots, noindex, hreflang alternates, JSON-LD structured data types.
  • preview: the title, description and image a chat app or social network would show when the link is shared (Open Graph first, then Twitter, then the plain tags).
  • HTTP: status code, final URL after redirects and the redirect chain.

Respects robots.txt of every site. No browser, no proxies; the User-Agent says url-metadata. In our test, 7 URLs took 3 seconds.

Use it for

  • Link previews in your app, newsletter or CMS without running your own fetcher.
  • Content and SEO inventories: titles, descriptions and canonicals of hundreds of URLs in one table.
  • Social sharing QA: find pages with no og:image or a missing description before a campaign.
  • Enrich lists of URLs (leads, bookmarks, backlinks, search results) with a title, description, image and icon.
  • AI agents and RAG pipelines: a compact, structured summary of any page, callable through the Apify MCP server.

Input

{
"urls": ["https://userpilot.com/", "stripe.com/blog"]
}

Output

One item per URL (real result for userpilot.com on 2026-09-25, lists shortened):

{
"input": "https://userpilot.com/",
"url": "https://userpilot.com/",
"finalUrl": null,
"statusCode": 200,
"redirectChain": [],
"contentType": "text/html; charset=utf-8",
"title": "Product Growth Platform | Userpilot",
"metaDescription": "Userpilot is the AI-powered, no-code product growth platform that helps teams onboard users, drive feature adoption, and boost retention across web and mobile.",
"canonical": "https://userpilot.com/",
"lang": "en-US",
"h1": "You ship the feature. Userpilot AI gets it adopted.",
"modifiedTime": "2026-08-06T14:16:54.000Z",
"siteName": "Userpilot",
"ogTitle": "Product Growth Platform | Userpilot",
"ogImage": "https://userpilot-website-assets.s3.us-west-2.amazonaws.com/wp-content/uploads/2026/05/11144757/UserpilotPreview-size-2026.png",
"ogType": "website",
"twitterCard": "summary_large_image",
"twitterSite": "@teamuserpilot",
"favicon": "https://userpilot-website-assets.s3.us-west-2.amazonaws.com/wp-content/uploads/2025/08/20164814/cropped-favicon-dark-512x512-1-192x192.png",
"faviconGuessed": false,
"noindex": false,
"structuredDataTypes": ["Organization", "WebSite", "SoftwareApplication"],
"wordCount": 1413,
"preview": {
"title": "Product Growth Platform | Userpilot",
"description": "Userpilot is the AI-powered, no-code product growth platform that helps teams onboard users, drive feature adoption, and boost retention across web and mobile.",
"image": "https://userpilot-website-assets.s3.us-west-2.amazonaws.com/wp-content/uploads/2026/05/11144757/UserpilotPreview-size-2026.png"
},
"status": "ok",
"checkedAt": "2026-09-25T15:33:53.574Z"
}

status is ok, error (4xx, 5xx or unreachable, with the reason in error), not_html (a PDF, an image...), blocked_by_robots or invalid.

Pricing

Pay per event: one event per page read (status: ok). Errors, non-HTML URLs, robots-blocked URLs, invalid inputs and duplicates are saved free. Set a maximum cost for the run and the Actor stops cleanly when it is reached.

Good to know

  • The Actor reads the HTML the server sends. Tags that a page adds with JavaScript after loading are not seen (most sites put their meta tags in the HTML for exactly this reason: social networks read it the same way).
  • Dates are converted to ISO format when they can be parsed, and kept as written otherwise.
  • Need a whole website instead of a list of URLs? Use SEO Audit & Broken Link Checker by the same author.

Support

A tag that should be read and isn't? Open an issue on the Actor's Issues tab with the URL. Issues get an answer within a few days.