Bulk URL Status Checker avatar

Bulk URL Status Checker

Pricing

$0.50 / 1,000 url checkeds

Go to Apify Store
Bulk URL Status Checker

Bulk URL Status Checker

Check thousands of URLs for broken links: status code, full redirect chain, response time, content type and size. Chains straight off the Sitemap URL Extractor.

Pricing

$0.50 / 1,000 url checkeds

Rating

0.0

(0)

Developer

Ali Alsaidi

Ali Alsaidi

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Give it a list of URLs and get back what each one actually does — the status code, every hop of the redirect chain, how long it took, the content type and the size.

What does Bulk URL Status Checker do?

  • Finds broken links at scale. Built to run over tens of thousands of URLs in one go without hammering anyone.
  • Reports the whole redirect chain, hop by hop, with a loop guard. Most checkers give you a final status code; the chain is what tells you whether a migration kept its redirects, where a loop is, and which links cost users two extra round trips.
  • Reads URLs straight from another run. Point it at a dataset ID — for example the output of the Sitemap URL Extractor — instead of pasting a list.
  • Falls back from HEAD to GET automatically when a server refuses HEAD with 405, 501 or 403, and tells you which method answered.
  • Polite by default. Optional minimum gap per host, backoff on 429 and 5xx that honours Retry-After, and robots.txt respected.

What data can you get?

FieldWhat it means
urlThe URL you asked for
finalUrlWhere it ended up after redirects
statusHTTP status of the final response
oktrue for a 2xx final status
redirectsNumber of hops
redirectChainEvery hop: url, status, location
responseTimeMsWhole check, including redirects and retries
contentTypeContent-Type of the final response
sizeBytesSize from Content-Length, or measured if you ask
serverServer header, when the site sends one
methodHEAD or GET — which one actually answered
attemptsHow many tries it took
errorError type, or null

How to use it

  1. Create a free Apify account.
  2. Open the Actor and click Try for free.
  3. Paste your URLs into URLs, or leave that empty and put a Dataset ID from a previous run instead.
  4. Set Maximum URLs to cap the run, then click Start.
  5. Open the Dataset tab and download as JSON, CSV, Excel or XML, or pull it through the Apify API.

Input

FieldTypeDefaultWhat it does
urlsarray—URLs to check. Leave empty if you use datasetId
datasetIdstring—Read URLs from another run's dataset instead
datasetFieldstringurlWhich field of that dataset holds the URL
maxUrlsinteger10000Stop after this many URLs. A hard ceiling on run cost
followRedirectsbooleantrueFollow each redirect and report the chain
measureSizebooleanfalseDownload the body to measure size when the server sends no Content-Length
concurrencyinteger10How many URLs to check at once
minHostIntervalSecsinteger0Smallest gap between two requests to the same host
requestTimeoutSecsinteger20Per-request timeout
maxRetriesinteger2Retries with backoff on timeouts and 429/5xx
maxRedirectsinteger10Hops before calling it a redirect loop
respectRobotsTxtbooleantrueSkip URLs that robots.txt disallows
{
"urls": ["https://apify.com", "https://apify.com/store"],
"maxUrls": 1000,
"followRedirects": true
}

To chain it onto a sitemap run instead:

{ "urls": [], "datasetId": "PASTE_DATASET_ID", "datasetField": "url" }

Output

One item per URL. This is a real item from a run, a URL that redirected twice:

{
"url": "https://httpbin.org/redirect/2",
"finalUrl": "https://httpbin.org/get",
"status": 200,
"ok": true,
"redirects": 2,
"redirectChain": [
{ "url": "https://httpbin.org/redirect/2", "status": 302, "location": "/relative-redirect/1" },
{ "url": "https://httpbin.org/relative-redirect/1", "status": 302, "location": "/get" },
{ "url": "https://httpbin.org/get", "status": 200, "location": null }
],
"responseTimeMs": 3754,
"contentType": "application/json",
"sizeBytes": 383,
"server": "gunicorn/19.9.0",
"method": "HEAD",
"attempts": 1,
"error": null
}

Every run also writes a RUN_SUMMARY record to the key-value store — urlsChecked, a byStatus breakdown (2xx, 3xx, 4xx, 5xx), averageResponseMs, requests, retries and blockedByRobotsTxt — so you can see the shape of a big run without reading the dataset:

{ "urlsChecked": 4, "byStatus": { "2xx": 3, "4xx": 1 },
"averageResponseMs": 1708, "requests": 10, "retries": 0, "blockedByRobotsTxt": 0 }

How much does it cost?

$0.50 per 1,000 URLs checked ($0.0005 each). 10,000 URLs = $5. Apify platform usage — compute, storage, transfer — is included in that price, not billed on top.

A URL skipped for robots.txt is not charged, and maxUrls is a hard ceiling on what a run can cost.

Integrations

Connect it to webhooks, Make, Zapier, Google Sheets, Slack, Airbyte or GitHub from the Integrations tab — for example, run it weekly and get a Slack message when the 4xx count goes up.

Call it from your own code with the Apify API:

from apify_client import ApifyClient
client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("palgenius/bulk-url-status-checker").call(
run_input={"urls": ["https://apify.com", "https://apify.com/store"], "maxUrls": 1000}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
if not item["ok"]:
print(item["url"], item["status"], item["error"])

AI agents can call it through the Apify MCP server at https://mcp.apify.com, e.g. https://mcp.apify.com?tools=palgenius/bulk-url-status-checker.

Error items and limits

A URL that could not be checked still comes back as an item, with ok: false and the reason in error — never a silent gap:

{ "url": "https://example.com/private", "finalUrl": "https://example.com/private",
"status": null, "ok": false, "redirects": 0, "redirectChain": [],
"responseTimeMs": null, "error": "blocked by robots.txt" }

Typical error values: blocked by robots.txt (skipped, and not charged), a timeout, a DNS failure, or a redirect loop past maxRedirects.

Honest limits:

  • A status code is not a promise. Some sites return 200 with an error page, and some return 403 to any non-browser client, including this one. server and contentType help you spot both.
  • HEAD can lie. A few servers answer HEAD differently from GET. That is why a refusal triggers a GET, and why method is reported.
  • Timing is indicative. responseTimeMs covers the whole check from one cloud region. It is not a substitute for real user monitoring.
  • Size comes from the server unless you set measureSize, and servers are not always truthful about Content-Length.

FAQ

Is this legal? Yes. It requests public URLs the same way any browser or link checker does, obeys robots.txt by default, never logs in, and collects no personal data.

Can I check a whole website? Run the Sitemap URL Extractor first, then paste its dataset ID here.

Will it overload a small server? Set minHostIntervalSecs to put a gap between requests to the same host, and lower concurrency.

Why is sizeBytes null? The server sent no Content-Length. Turn on measureSize to download the body and measure it.

Does it find redirect loops? Yes — it stops at maxRedirects and reports the chain it followed.

Do I pay for URLs that are skipped? No. Only URLs actually checked are charged.

You might also like

  • Sitemap URL Extractor — get every page URL a site publishes, then feed its dataset ID straight into this Actor.
  • Tech Stack Detector — find the platform, CMS and frameworks behind any domain, with evidence.

Feedback

Hit a site it reports wrongly, or want a field added? Open an issue on the Actor's Issues tab and include the URL. Fixes ship within days.