Bulk URL Status Checker
Pricing
$0.50 / 1,000 url checkeds
Bulk URL Status Checker
Check thousands of URLs for broken links: status code, full redirect chain, response time, content type and size. Chains straight off the Sitemap URL Extractor.
Pricing
$0.50 / 1,000 url checkeds
Rating
0.0
(0)
Developer
Ali Alsaidi
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Give it a list of URLs and get back what each one actually does — the status code, every hop of the redirect chain, how long it took, the content type and the size.
What does Bulk URL Status Checker do?
- Finds broken links at scale. Built to run over tens of thousands of URLs in one go without hammering anyone.
- Reports the whole redirect chain, hop by hop, with a loop guard. Most checkers give you a final status code; the chain is what tells you whether a migration kept its redirects, where a loop is, and which links cost users two extra round trips.
- Reads URLs straight from another run. Point it at a dataset ID — for example the output of the Sitemap URL Extractor — instead of pasting a list.
- Falls back from HEAD to GET automatically when a server refuses HEAD with 405, 501 or 403, and tells you which method answered.
- Polite by default. Optional minimum gap per host, backoff on 429 and 5xx that honours
Retry-After, androbots.txtrespected.
What data can you get?
| Field | What it means |
|---|---|
url | The URL you asked for |
finalUrl | Where it ended up after redirects |
status | HTTP status of the final response |
ok | true for a 2xx final status |
redirects | Number of hops |
redirectChain | Every hop: url, status, location |
responseTimeMs | Whole check, including redirects and retries |
contentType | Content-Type of the final response |
sizeBytes | Size from Content-Length, or measured if you ask |
server | Server header, when the site sends one |
method | HEAD or GET — which one actually answered |
attempts | How many tries it took |
error | Error type, or null |
How to use it
- Create a free Apify account.
- Open the Actor and click Try for free.
- Paste your URLs into URLs, or leave that empty and put a Dataset ID from a previous run instead.
- Set Maximum URLs to cap the run, then click Start.
- Open the Dataset tab and download as JSON, CSV, Excel or XML, or pull it through the Apify API.
Input
| Field | Type | Default | What it does |
|---|---|---|---|
urls | array | — | URLs to check. Leave empty if you use datasetId |
datasetId | string | — | Read URLs from another run's dataset instead |
datasetField | string | url | Which field of that dataset holds the URL |
maxUrls | integer | 10000 | Stop after this many URLs. A hard ceiling on run cost |
followRedirects | boolean | true | Follow each redirect and report the chain |
measureSize | boolean | false | Download the body to measure size when the server sends no Content-Length |
concurrency | integer | 10 | How many URLs to check at once |
minHostIntervalSecs | integer | 0 | Smallest gap between two requests to the same host |
requestTimeoutSecs | integer | 20 | Per-request timeout |
maxRetries | integer | 2 | Retries with backoff on timeouts and 429/5xx |
maxRedirects | integer | 10 | Hops before calling it a redirect loop |
respectRobotsTxt | boolean | true | Skip URLs that robots.txt disallows |
{"urls": ["https://apify.com", "https://apify.com/store"],"maxUrls": 1000,"followRedirects": true}
To chain it onto a sitemap run instead:
{ "urls": [], "datasetId": "PASTE_DATASET_ID", "datasetField": "url" }
Output
One item per URL. This is a real item from a run, a URL that redirected twice:
{"url": "https://httpbin.org/redirect/2","finalUrl": "https://httpbin.org/get","status": 200,"ok": true,"redirects": 2,"redirectChain": [{ "url": "https://httpbin.org/redirect/2", "status": 302, "location": "/relative-redirect/1" },{ "url": "https://httpbin.org/relative-redirect/1", "status": 302, "location": "/get" },{ "url": "https://httpbin.org/get", "status": 200, "location": null }],"responseTimeMs": 3754,"contentType": "application/json","sizeBytes": 383,"server": "gunicorn/19.9.0","method": "HEAD","attempts": 1,"error": null}
Every run also writes a RUN_SUMMARY record to the key-value store — urlsChecked, a
byStatus breakdown (2xx, 3xx, 4xx, 5xx), averageResponseMs, requests, retries
and blockedByRobotsTxt — so you can see the shape of a big run without reading the dataset:
{ "urlsChecked": 4, "byStatus": { "2xx": 3, "4xx": 1 },"averageResponseMs": 1708, "requests": 10, "retries": 0, "blockedByRobotsTxt": 0 }
How much does it cost?
$0.50 per 1,000 URLs checked ($0.0005 each). 10,000 URLs = $5. Apify platform usage — compute, storage, transfer — is included in that price, not billed on top.
A URL skipped for robots.txt is not charged, and maxUrls is a hard ceiling on what a
run can cost.
Integrations
Connect it to webhooks, Make, Zapier, Google Sheets, Slack, Airbyte or GitHub from the Integrations tab — for example, run it weekly and get a Slack message when the 4xx count goes up.
Call it from your own code with the Apify API:
from apify_client import ApifyClientclient = ApifyClient("<YOUR_API_TOKEN>")run = client.actor("palgenius/bulk-url-status-checker").call(run_input={"urls": ["https://apify.com", "https://apify.com/store"], "maxUrls": 1000})for item in client.dataset(run["defaultDatasetId"]).iterate_items():if not item["ok"]:print(item["url"], item["status"], item["error"])
AI agents can call it through the Apify MCP server at https://mcp.apify.com, e.g.
https://mcp.apify.com?tools=palgenius/bulk-url-status-checker.
Error items and limits
A URL that could not be checked still comes back as an item, with ok: false and the reason
in error — never a silent gap:
{ "url": "https://example.com/private", "finalUrl": "https://example.com/private","status": null, "ok": false, "redirects": 0, "redirectChain": [],"responseTimeMs": null, "error": "blocked by robots.txt" }
Typical error values: blocked by robots.txt (skipped, and not charged), a timeout, a DNS
failure, or a redirect loop past maxRedirects.
Honest limits:
- A status code is not a promise. Some sites return 200 with an error page, and some
return 403 to any non-browser client, including this one.
serverandcontentTypehelp you spot both. - HEAD can lie. A few servers answer HEAD differently from GET. That is why a refusal
triggers a GET, and why
methodis reported. - Timing is indicative.
responseTimeMscovers the whole check from one cloud region. It is not a substitute for real user monitoring. - Size comes from the server unless you set
measureSize, and servers are not always truthful aboutContent-Length.
FAQ
Is this legal? Yes. It requests public URLs the same way any browser or link checker
does, obeys robots.txt by default, never logs in, and collects no personal data.
Can I check a whole website? Run the Sitemap URL Extractor first, then paste its dataset ID here.
Will it overload a small server? Set minHostIntervalSecs to put a gap between requests
to the same host, and lower concurrency.
Why is sizeBytes null? The server sent no Content-Length. Turn on measureSize to
download the body and measure it.
Does it find redirect loops? Yes — it stops at maxRedirects and reports the chain it
followed.
Do I pay for URLs that are skipped? No. Only URLs actually checked are charged.
You might also like
- Sitemap URL Extractor — get every page URL a site publishes, then feed its dataset ID straight into this Actor.
- Tech Stack Detector — find the platform, CMS and frameworks behind any domain, with evidence.
Feedback
Hit a site it reports wrongly, or want a field added? Open an issue on the Actor's Issues tab and include the URL. Fixes ship within days.