Bulk URL Status, Redirect & Indexability Checker avatar

Bulk URL Status, Redirect & Indexability Checker

Pricing

Pay per event

Go to Apify Store
Bulk URL Status, Redirect & Indexability Checker

Bulk URL Status, Redirect & Indexability Checker

Check thousands of URLs: HTTP status, full redirect chain, loops, response time, and SEO indexability (noindex, X-Robots-Tag, canonical, robots.txt for Googlebot, title, H1). Find broken links and bad redirects after migrations. Pay per URL.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Mohamed T.

Mohamed T.

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 hours ago

Last modified

Share

Check thousands of URLs in minutes: HTTP status, the full redirect chain and whether each page can actually be indexed by Google. It's built for SEO migrations, broken-link audits, backlink checks and QA of marketing links.

What does Bulk URL Status Checker do?

For every URL it records:

  • Status: first status code, final status code, response time, content type and size, server, HTTPS and HSTS.
  • Redirects: every hop (url → status → location), redirect count, loop detection, cross-domain redirects, a max-hop limit.
  • Indexability (HTML pages): <title> and length, meta description, H1 count, meta robots / googlebot, X-Robots-Tag header, canonical URL and whether it points to itself, hreflang count, robots.txt verdict for Googlebot, and an overall indexable flag.
  • Errors in plain words: HTTP 404, DNS resolution failed, timeout, TLS/SSL error, redirect loop…

Transient failures (429, 5xx, timeouts) are retried with exponential backoff, and requests to each host are spaced out politely.

Why use it?

  • Site migrations: verify that every old URL 301-redirects to the right new page in one hop.
  • Broken-link audits: feed it URLs from a crawl, a sitemap (see our Sitemap URL Extractor) or a CMS export.
  • Indexability QA: catch stray noindex, wrong canonicals or robots.txt blocks before they cost traffic.
  • Backlink and affiliate checks: confirm that links pointing to you (or that you pay for) still resolve.
  • Monitoring: schedule it daily on your key URLs and alert on anything that isn't ok.

How to use it

  1. Paste URLs into URLs (one per line; example.com/page works too).
  2. Click Start. Results appear live. Filter by ok = false to see problems.
  3. Export to CSV/Excel/JSON, or connect a Slack/email integration for scheduled checks.

Input example

{
"urls": ["https://apify.com", "http://apify.com/pricing", "https://apify.com/this-page-does-not-exist"],
"checkIndexability": true,
"maxRedirects": 10
}

Output example

{
"url": "http://apify.com/pricing",
"ok": true,
"statusCode": 301,
"finalUrl": "https://apify.com/pricing",
"finalStatusCode": 200,
"redirectCount": 1,
"redirectChain": [{"url": "http://apify.com/pricing", "statusCode": 301, "location": "https://apify.com/pricing"}],
"redirectLoop": false,
"responseTimeMs": 14,
"contentType": "text/html",
"hsts": true,
"httpsFinal": true,
"title": "Apify pricing - flexible plan + pay as you go · Apify",
"h1Count": 1,
"metaRobots": "index,follow",
"canonicalUrl": "https://apify.com/pricing",
"canonicalIsSelf": true,
"noindex": false,
"robotsTxtAllowsGooglebot": true,
"indexable": true,
"error": null
}

The Output tab has a Status view and an SEO view. You can download the dataset in various formats such as JSON, CSV, Excel, XML or HTML.

How much does it cost?

You pay per URL checked. Broken URLs are results too, so they count. URLs skipped because robots.txt disallows them aren't charged. See the Pricing tab.

Tips

  • Checking your own site? Lower Delay between requests to the same host and raise Max concurrency.
  • indexable is false when the final status isn't 200, a noindex directive is present, robots.txt blocks Googlebot or the canonical points elsewhere.
  • Response times are measured from Apify's servers, not from your visitors' locations.

FAQ

Does it render JavaScript? No. It checks what servers return over HTTP, which is what matters for status codes, redirects and most indexability signals. Canonicals or robots meta tags injected by JavaScript won't be seen.

Why is a URL reported as blocked by robots.txt? By default the checker respects robots.txt like a well-behaved crawler. For sites you own or are authorized to audit, switch off Respect robots.txt.

Something looks wrong? Open an issue on the Issues tab with the URL.