SEO Audit: Meta Tags, Open Graph & JSON-LD Checker
Pricing
$2.00 / 1,000 page auditeds
SEO Audit: Meta Tags, Open Graph & JSON-LD Checker
On-page SEO audit for any list of URLs: title and meta description lengths, canonical, robots, hreflang, share-preview meta tags, headings, image alt text and JSON-LD errors. One row per page.
Pricing
$2.00 / 1,000 page auditeds
Rating
0.0
(0)
Developer
James Power
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
4 days ago
Last modified
Categories
Share
Give this Actor a list of web page addresses and it returns one tidy row per page with the facts an on-page SEO check needs: the title and meta description (with their lengths and a pass/fail against common limits), the canonical tag, robots meta and X-Robots-Tag, hreflang, the SEO meta tags in one simple name-to-content map built from named allowlists (the core names, the named Open Graph and article keys, and share-preview names — see What is never returned for the exact lists), the headings outline, image alt-text coverage, and every schema.org JSON-LD block parsed and checked for obvious errors. Each row ends with a short, plain-English issues list such as "Title too long: 72 chars (maximum 60)" or "JSON-LD error: Block 2 (Product): missing required property offers or review or aggregateRating".
It uses plain HTTP requests only: no browser, no proxies, no login, no cookies and no API key.
Price: $0.002 per page audited
You pay $0.002 for each page that is fetched and audited successfully ($2 per 1,000 pages). One charge buys one complete row for one page: all the fields listed below plus the issues list. There is no start fee and no monthly rental.
- 100 pages cost 100 x $0.002 = $0.20.
- 10,000 pages cost 10,000 x $0.002 = $20.00.
Set a maximum charge per run in the run options if you want a hard budget: the Actor stops cleanly before it would go over it and never charges more.
Every address you send gets a row. Failed pages, pages blocked by robots.txt or by the site, non-HTML files, invalid addresses, duplicates, addresses refused for safety, and entries the run did not get to (because of your charge cap or the run's time limit) all still get a row that says why, and those rows are free. A blank line, a line with nothing on it, and an entry beyond your maxUrls cap each get their own free row too, so every line you pasted can be accounted for.
Who it is for
SEO consultants and agencies checking client pages before and after a release; developers and QA teams who want a quick metadata and structured-data regression check on a list of URLs; content teams fixing titles, descriptions and share previews in bulk; and anyone building a spreadsheet or dashboard of on-page SEO facts (export as CSV, Excel or JSON, or call it from the Apify API).
What each row contains
| Area | Fields |
|---|---|
| Fetch | url, inputIndex, finalUrl, status, error, httpStatus, redirectCount, redirects (every hop with its status code), contentType, responseBytes, responseTruncated, responseTimeMs |
| Title and description | title, titleLength, titleStatus, metaDescription, metaDescriptionLength, metaDescriptionStatus (ok, too_short, too_long or missing) |
| Indexing | canonical, canonicalStatus (self, other, missing), robotsMeta, xRobotsTag, indexable, nofollow, lang, hasViewport |
| International | hreflang (code and URL), with checks for invalid codes, duplicates and a missing self-reference |
| Meta tags | metaTags: the SEO meta tags in the returned allowlists (the core names, the named Open Graph and article keys, and share-preview names) as name to content; metaTagsDropped (tags outside those lists); removedForPrivacy |
| Headings | h1, h1Count, headingCounts, headingsOutline (H1-H6 in page order, first 100) |
| Images | imageCount, imagesWithAlt, imagesEmptyAlt, imagesMissingAlt, imageAltCoveragePercent |
| Structured data | jsonLdTypes, jsonLdErrorCount, jsonLd (block count, types, errors, warnings, the allowlisted parsed blocks), microdataTypes |
| Source and licence | attribution: the source address and the notice to carry if you republish that page's content |
| Summary | wordCount, issueCount, issues, checkedAt |
The Console shows an SEO overview table and a Share-preview and meta tags table.
JSON-LD checks. Invalid JSON, a missing @context or @type, properties Google needs for common rich results (Product, BreadcrumbList, FAQPage, Event, Recipe, VideoObject, JobPosting, Review), offers without a price, and dates that are not ISO 8601. Every nested entity is visited, so a Product inside an ItemList is checked like any other; recommended but optional properties are listed as warnings, not errors. The checks read the page's own data, so a value the returned block does not include is still validated and reported; the returned block itself keeps only the allowlisted types and properties named under What is never returned. These are quick sanity checks, not a replacement for Google's Rich Results Test.
The run summary. Every run also writes an OUTPUT record to the run's key-value store: how many addresses were requested, audited, charged and left unprocessed, why the run stopped if it stopped early, and the limits it ran with. The dataset holds the rows; OUTPUT explains the run.
Input
| Field | What it does | Default |
|---|---|---|
urls | The pages to audit, one per line. example.com/page becomes https://example.com/page. | 3 example pages |
maxUrls | Safety cap on how many unique pages are attempted. Every other entry still gets an indexed free not_processed row. | 1000 |
maxConcurrency | Pages fetched at once across all sites. | 10 |
maxConcurrencyPerHost | Pages fetched at once from the same site (1 or 2). | 2 |
requestTimeoutSecs | Timeout per request. | 20 |
titleMinLength / titleMaxLength | Title limits used for pass/fail. | 30 / 60 |
descriptionMinLength / descriptionMaxLength | Meta description limits. | 70 / 160 |
includeJsonLdBlocks | Include the full parsed JSON-LD blocks (turn off for smaller rows). | on |
includeHeadingsOutline | Include the H1-H6 outline. | on |
Example input:
{"urls": ["https://www.gov.uk/","https://docs.python.org/3/","https://www.sqlite.org/index.html"],"titleMaxLength": 60,"descriptionMaxLength": 160}
Output example
A real row from our own test run on 28 September 2026 (the long headings outline is left out here to keep the example short). The page is the Python documentation home page, content copyright Python Software Foundation, licensed under the PSF License Version 2; the row's own attribution field repeats that notice for a reuser.
{"url": "https://docs.python.org/3/","inputIndex": 6,"finalUrl": "https://docs.python.org/3/","status": "ok","error": null,"httpStatus": 200,"redirectCount": 0,"redirects": [],"contentType": "text/html","responseBytes": 19641,"responseTruncated": false,"responseTimeMs": 302,"title": "3.14.7 Documentation","titleLength": 20,"titleStatus": "too_short","metaDescription": "The official Python documentation.","metaDescriptionLength": 34,"metaDescriptionStatus": "too_short","canonical": "https://docs.python.org/3/index.html","canonicalStatus": "other","robotsMeta": null,"xRobotsTag": null,"indexable": true,"nofollow": false,"lang": "en","hasViewport": true,"hreflangCount": 0,"hreflang": [],"metaTags": {"viewport": "width=device-width, initial-scale=1.0","og:title": "Python 3.14 documentation","og:type": "website","og:url": "https://docs.python.org/3/","og:site_name": "Python documentation","og:description": "The official Python documentation.","og:image": "https://docs.python.org/3/_static/og-image.png","description": "The official Python documentation.","og:image:width": "200","og:image:height": "200","theme-color": "#3776ab"},"metaTagsDropped": 1,"removedForPrivacy": 0,"h1Count": 1,"h1": ["Python 3.14.7 documentation"],"headingCounts": {"h1": 1,"h2": 0,"h3": 8,"h4": 0,"h5": 0,"h6": 0},"imageCount": 3,"imagesMissingAlt": 0,"imagesEmptyAlt": 0,"imagesWithAlt": 3,"imageAltCoveragePercent": 100.0,"jsonLdTypes": [],"jsonLdErrorCount": 0,"jsonLd": {"blockCount": 0,"types": [],"errors": [],"warnings": [],"removedForPrivacy": 0,"blocks": []},"microdataTypes": [],"wordCount": 463,"issueCount": 5,"issues": ["Title too short: 20 chars (minimum 30)","Meta description too short: 34 chars (minimum 70)","Canonical points to a different URL: https://docs.python.org/3/index.html","No share-card meta tag (a name ending in \":card\")","No structured data (JSON-LD or microdata)"],"checkedAt": "2026-09-28T22:40:13Z"}
Pages that could not be audited still get a row, free of charge, with status set to failed, blocked_by_robots, blocked_by_site, not_html, private_address, invalid_url, duplicate, empty_input or not_processed and a plain error such as HTTP 404, robots.txt disallows /private/ for this crawler or Not an HTML page (content-type application/pdf).
Rows arrive in the order pages finish. Use inputIndex to sort them back into input order; every row carries one, including the free explanation rows for blank, invalid, duplicate and unprocessed entries.
Politeness and limits (please read)
- It respects robots.txt. Before fetching any page of a site it reads that site's robots.txt and skips disallowed pages (free rows). If robots.txt cannot be read because of a server error, the site is skipped, as the robots.txt standard (RFC 9309) requires. A
Crawl-delayis honoured in full, and while it applies, pages from that site are fetched one at a time rather than two at once; sites asking for more than 30 seconds between requests are skipped rather than fetched faster. A long Crawl-delay slows down a big list. A site with a 5-second Crawl-delay limits itself to about 12 pages a minute (measured, local run, 28 September 2026); a 10-second delay is about 6 a minute (estimate, same formula). Spread a big list across several sites, or expect a single slow site to take a while. - It backs off at once. At most 2 requests at a time per site. A site that answers 403 or 429 anywhere — robots.txt or a page, and whatever hostname or scheme the next request would have used — is never retried and gets no further requests in that run; its remaining pages become free
blocked_by_siterows. Only server errors and network errors are retried (at most twice). The User-Agent is honest and fixed (ApifySeoAudit). It does not use proxies or environment proxy settings, does not store or send cookies, and does not try to get around bot protection. - It never fetches a private or internal address, and it never dials one by accident. Every outbound request — the page, its robots.txt, and every redirect hop of both — has its destination address checked twice: when the request is prepared, and again at the moment the connection is opened. The Actor connects only to the address it checked, so a hostname whose DNS answer changes to a private one between the two checks is refused rather than dialled, and a name that answers with any non-public address is refused outright. Loopback, private, link-local (including the
169.254.169.254cloud metadata address), carrier-grade NAT, multicast, reserved and tunnelling addresses are refused. A typed address is refused too: an IP literal, or a shorthand, hex, octal or decimal form of one, is not an address this Actor audits. Refused rows are free (private_address,invalid_url). - It does not log in as anyone. A URL carrying an embedded user name and password is refused before any request is made — in the input and on any redirect — and none of it is kept in the row or the error text. The Actor never sends an Authorization header, never reads netrc files, and has no login of its own.
- It reads the HTML the server sends. It does not run JavaScript, so tags a site adds only with JavaScript are not seen.
- It audits the pages you list. It does not crawl a whole site or check links. Pages over 3 MB are cut at 3 MB (
responseTruncated: true). If two of your addresses end on the same page (a redirect or a trailing-slash alias), the page is audited and charged once and the other address gets a freeduplicaterow. - It stops cleanly before the run's time limit. The default limit is 3,600 seconds (1 hour), enough for a large list even with a slow site or two. About 30 seconds before the deadline it stops starting new pages and cancels anything still in flight, then still writes the summary and delivers every row it has, plus a free
not_processedrow for each address it did not finish. The status message says how many were left over, and none of them is charged. - You are charged once for a page, after its complete row has been saved. Failures, robots and site blocks, non-HTML files, invalid and duplicate entries, addresses refused for safety, and anything left unprocessed are never charged.
Your data and the site's data
The page content in each row (titles, descriptions, meta tags, headings, JSON-LD) belongs to the site it came from, and each row carries its source address in url and finalUrl. Each audited row also carries an attribution object: the source address plus the notice a reuser should carry. For the documentation sites used in our own examples that is the Open Government Licence v3.0 for GOV.UK, the PSF License Version 2 for the Python documentation, and the public-domain dedication for SQLite. For any other site it tells you to check that site's own terms and licence before you republish its content. Please do: this Actor reads a site's published rules before it fetches anything, but permission to read a page is not a licence to reuse it. The terms of our own test targets were read before any content was fetched, and are recorded in SOURCES.md beside this README.
What is never returned
Person-related data is removed on purpose. A value reaches your row only when it sits under a key on one of the named allowlists below — the output is built from those lists, not by deleting known-bad names from everything else — so a key nobody here has thought of cannot carry a person's name into your dataset. The same rules clean the issues list, the warnings and the error text as well as the fields.
The meta-tag map is built from these lists and nothing else:
- Core meta names, the returned set:
canonical,description,googlebot,keywords,referrer,robots,theme-color,viewport. - Open Graph keys, the returned set:
og:audio,og:audio:secure_url,og:audio:type,og:audio:url,og:description,og:image,og:image:alt,og:image:height,og:image:secure_url,og:image:type,og:image:url,og:image:width,og:locale,og:locale:alternate,og:site_name,og:title,og:type,og:url,og:video,og:video:duration,og:video:height,og:video:secure_url,og:video:type,og:video:url,og:video:width. - Article keys, the returned set:
article:expiration_time,article:modified_time,article:published_time,article:section,article:tag. - Share-preview names, the returned set:
alt,card,description,expirationtime,height,image,locale,modifiedtime,publishedtime,title,type,url,width.
The share-preview names are accepted under any namespace prefix except the reserved og:, fb: and profile: ones, and the NAME is what is allowlisted: a share-preview key must be one of those words under whatever prefix the site uses, and a name-part key such as first_name, last_name or username is out by construction whichever prefix it arrives with (the three folded time names match their published / modified / expiration time spellings). The name comes back under the fixed preview: prefix whatever prefix the page wrote — a tag named something:card is returned as preview:card — so no part of a returned key is ever text the page chose. The fb: and profile: namespaces are never returned, and neither is any key whose name is not on a list — a CMS's own tag, generator, copyright, a site-invented tag: all left out, and metaTagsDropped counts them. On top of the lists, a person-role name is removed wherever it appears, however the page spells it (camelCase, _/-/:-separated, or with a dc./citation_-style prefix): such a tag is taken out whole whatever its value and counted in removedForPrivacy; the paired share-preview label/data tags that WordPress SEO plugins add ("Written by" plus a name) go the same way, and so does a handle value beginning with @.
The returned JSON-LD blocks are rebuilt from these two lists as well:
- JSON-LD types, the returned set:
Article,BlogPosting,BreadcrumbList,Event,FAQPage,JobPosting,LocalBusiness,NewsArticle,Organization,Product,Recipe,Review,VideoObject,WebSite. - JSON-LD properties, the returned set:
acceptedAnswer,aggregateRating,alternativeHeadline,articleBody,caption,comment,contentUrl,dateCreated,dateModified,datePosted,datePublished,description,embedUrl,endDate,gtin,gtin12,gtin13,gtin14,gtin8,headline,highPrice,identifier,image,isbn,issn,itemListElement,itemReviewed,location,logo,lowPrice,mainEntity,mainEntityOfPage,mpn,name,offers,price,priceCurrency,priceSpecification,priceValidUntil,productID,ratingValue,review,reviewBody,reviewCount,reviewRating,serialNumber,sku,startDate,suggestedAnswer,text,thumbnailUrl,title,uploadDate,url,validFrom,validThrough.
A node is returned only when every one of its types is on the JSON-LD type list; a node of any other type — Person above all — is dropped whole with every field on it, and an untyped node is dropped too, including a bare reference or a node whose @id looks like a person. A property is returned only when it is on the property list and its value is a text, number or true/false: a list or an object is a nested entity, and nested entities are never returned as content — the property itself comes back as true, the fact that the page carries one — so nothing inside a Product's offers, a Question's answers, an about node or an image node reaches you. An @id is reduced to true the same way (its value is never returned), and a returned block's @context is only ever https://schema.org. hiringOrganization is deliberately not on the property list — it is a recruiter/employer field — even though the JobPosting check still validates it on the page's own data.
Contact details are dropped or scrubbed everywhere. E-mail addresses and phone numbers are removed from the free text of a returned value whichever key it sits in — an articleBody, reviewBody, comment or caption, the returned meta tag values — while identifier fields such as an ISBN, GTIN or SKU keep their digits. Contact properties (contactPoint, email, telephone) are never returned, and a whole Person node never is either, with person name parts (givenName, familyName, jobTitle and similar) going with it.
metaTagsDroppedcounts the tags left out because they are not on the returned lists, andremovedForPrivacycounts the person-data items taken out (a person-role tag or property, a Person node, a name part or a contact detail), so you can see both filters working.
FAQ
Why is a page "blocked_by_robots"? The site's robots.txt disallows that path for crawlers like this one. We never bypass it. The row is free.
Why is a page "blocked_by_site"? The site answered 403 or 429 to this crawler, so it is left alone for the rest of the run. The row is free.
Why does my page show no JSON-LD when I can see it in the browser? The site probably adds it with JavaScript after the page loads. This Actor reads the HTML as served.
Why is a tag missing from metaTags? The map is built from the named allowlists listed under What is never returned: the core names, the named Open Graph and article keys, and the share-preview names. Anything not on those lists — a CMS's own tag, a site-invented tag, a name-part key such as first_name — is left out, and the row's metaTagsDropped counts how many. A person-role tag is never returned either, and is counted in removedForPrivacy instead. That is deliberate: it is what stops a key nobody expected from carrying a person's name into your dataset.
Why did some addresses come back "not_processed"? The run reached your maximum charge per run or its time limit before starting them. They were never fetched and never charged; run them again in a new run.
What if I paste a blank line, or send an empty list? Nothing breaks and nothing is charged: a blank line gets its own free invalid_url row, and an empty list completes with a single free empty_input row. The summary says what happened.
Can I change the title and description limits? Yes. Set titleMinLength, titleMaxLength, descriptionMinLength and descriptionMaxLength.
How fast is it? In our local test on 28 September 2026, 30 pages across 8 sites took under 4 seconds with the default settings. Speed depends on the sites and on their robots.txt rules.
Can I run it on a schedule? Yes. Use Apify Schedules to rerun the same list daily or weekly and compare the issues over time.