API Origin Health Report - is this API alive, and paid? avatar

API Origin Health Report - is this API alive, and paid?

Pricing

from $20.00 / 1,000 origin reports

Go to Apify Store
API Origin Health Report - is this API alive, and paid?

API Origin Health Report - is this API alive, and paid?

Bulk due-diligence on API origins. One row per URL: a liveness verdict from a closed set (LIVE, PAUSED, GATED, CHALLENGED, DEGRADED, DEAD, UNKNOWN), a confidence, coded reasons, the origin reached after redirects, the RFC 9309 robots disposition, any payment rail on the wire, catalogue surfaces.

Pricing

from $20.00 / 1,000 origin reports

Rating

0.0

(0)

Developer

Stephen Psaradellis

Stephen Psaradellis

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

21 days ago

Last modified

Share

API Origin Health Report

Is this API alive, and does it actually take money? Give it a list of API origins. Get back one row each: a verdict, the evidence behind it, and the sources that evidence came from.

Built for the moment before you commit: you are choosing between four vendors, inheriting a service catalogue nobody has audited, or checking whether the 200 integrations in your .env are all still there.

Use case

Your team inherited a service catalogue with 300 API integrations in it and no one knows which are still real. Or you are down to four vendors and want to know, before the call, which of them actually answers, which sits behind a login, and which quietly went away. Paste the list in, run it in report mode, and sort the output by verdict: DEAD and GATED rows are the conversation you were going to have anyway, with the evidence already attached.

What you get

One dataset row per origin:

fieldwhat it is
verdictLIVE, PAUSED, GATED, CHALLENGED, DEGRADED, DEAD or UNKNOWN. A closed set - never a score out of 100
confidencehow much of the evidence the verdict rests on
reasonscoded reasons, each naming the source that produced it
origin_measuredthe scheme and host actually reached, after redirects - often not the one you submitted
final_urlwhere the primary GET ended up
http_statusthe status of that GET
robotsallowed, disallowed or unknown, matched under RFC 9309
paymentany payment rail the origin advertises on the wire (a 402 with an accepts array, an x402 well-known), or null
surfacesfive catalogue paths - /openapi.json, /api/v1/services, /api/v1/discover, /discover, /.well-known/x402 - each with its status and counts of JSON rows, priced rows and payee-naming rows. null in liveness mode
charged_eventwhich of the two events this row billed
warrantywhat the row does and does not claim

A RUN_SUMMARY record in the key-value store carries the counts for the run: submitted, measured, skipped by robots, stranded over the cap, the charged events and the verdict histogram.

The two modes

Full report - robots, the primary GET, up to 3 of the health sidecars a service may publish (/health, /healthz, /api/health, /api/v1/health, /status, /.well-known/health) and the 5 catalogue surfaces. About 10 requests per origin. Use it on a shortlist.

Liveness only - robots and the primary GET. About 2 requests. Use it to sweep a long list and then re-run the interesting rows as full reports.

Sample output

Two rows, exactly as the dataset holds them - one in each mode, so you can see what the dearer read buys you. The first is https://api.github.com in report mode; the second is https://httpbin.org in liveness mode, from platform run OFivUcrHgKl1bFuSA on 2026-09-07:

[
{
"schema": "exp24.report/1",
"url": "https://api.github.com",
"origin_measured": "https://api.github.com",
"checked_at": "2026-09-07T05:43:30+00:00",
"mode": "report",
"verdict": "LIVE",
"confidence": "partial",
"http_status": 200,
"robots": "unknown",
"final_url": "https://api.github.com/",
"payment": null,
"reasons": [
{
"code": "http_ok",
"detail": "HTTP 200 in 127ms",
"source": "https://api.github.com"
}
],
"surfaces": {
"/openapi.json": {
"status": 404,
"error": null,
"json": true,
"rows": 1,
"priced_rows": 0,
"payee_rows": 0
},
"/api/v1/services": {
"status": 404,
"error": null,
"json": true,
"rows": 1,
"priced_rows": 0,
"payee_rows": 0
},
"/api/v1/discover": {
"status": 404,
"error": null,
"json": true,
"rows": 1,
"priced_rows": 0,
"payee_rows": 0
},
"/discover": {
"status": 404,
"error": null,
"json": true,
"rows": 1,
"priced_rows": 0,
"payee_rows": 0
},
"/.well-known/x402": {
"status": 404,
"error": null,
"json": true,
"rows": 1,
"priced_rows": 0,
"payee_rows": 0
}
},
"charged_event": "origin-report",
"warranty": "This report describes only what was measured at checked_at, from the sources named in every reason. It warrants nothing about future availability and is not advice."
},
{
"schema": "exp24.liveness/1",
"url": "https://httpbin.org",
"origin_measured": "https://httpbin.org/",
"checked_at": "2026-09-07T02:56:46+00:00",
"mode": "liveness",
"verdict": "LIVE",
"confidence": "partial",
"http_status": 200,
"robots": "allowed",
"final_url": "https://httpbin.org/",
"payment": null,
"reasons": [
{
"code": "http_ok",
"detail": "HTTP 200 in 93ms",
"source": "https://httpbin.org"
}
],
"surfaces": null,
"charged_event": "origin-liveness",
"warranty": "This report describes only what was measured at checked_at, from the sources named in every reason. It warrants nothing about future availability and is not advice."
}
]

How to run it

  1. Open the Actor in Apify Console, paste your origins into Origins - one per line, a bare host is read as https - and press Start.
  2. Sweep first if the list is long: run it in liveness mode at $0.0004 an origin, then re-run the rows you care about in report mode.
  3. Read the dataset in Console, sort by verdict, export it as JSON or CSV, or fetch it from the API with the run's defaultDatasetId.

Starting the Actor with no input at all measures nothing and costs nothing. That is deliberate: an Actor with a default origin list would bill whoever merely pressed Start.

Sample input

{
"origins": [
"https://api.apis.guru",
"https://api.github.com",
"https://httpbin.org"
],
"mode": "report",
"respectRobots": true,
"maxOrigins": 100,
"concurrency": 8
}

Input

  • origins - the URLs to measure, one per line. Duplicates are collapsed and charged once. No default, on purpose.
  • mode - report (about 10 requests per origin) or liveness (about 2).
  • respectRobots - on by default. A URL the origin disallows for this user agent is reported as skipped, not requested, and not charged.
  • maxOrigins - hard cap on how many origins one run measures, after de-duplication. Default 100, ceiling 5000. This is also your bill ceiling.
  • concurrency - how many origins are measured at once.

Pricing

Pay per event, with no charge for starting the Actor and no charge for platform usage:

eventpricewhen
origin-report$0.02one origin measured in full
origin-liveness$0.0004one origin measured in liveness mode

You are not charged for a URL skipped by robots.txt, a URL stranded beyond maxOrigins, an entry that is not a parseable URL, or a duplicate. An origin that does not resolve is billed at the liveness price, not the report price

  • the report's other nine requests never happened, so you do not pay for them.

What a run costs:

runeventsprice
the sample input above: 3 origins, report3 report$0.06
100 origins swept for liveness100 liveness$0.04
100 origins, full report100 report$2.00
the ceiling: 5,000 origins, full report5,000 report$100.00

Use it from code

The API tab on this page has ready-made snippets for Node.js, Python and curl, including run-sync-get-dataset-items, which starts a run and returns the rows in one call.

Use it from Claude or ChatGPT (Apify MCP)

This Actor is a tool your assistant can call directly, over the Model Context Protocol. Apify hosts the server at https://mcp.apify.com; adding ?tools= pins this Actor as one named tool rather than loading the whole store.

Claude Desktop or Claude Code - add a custom connector with the URL:

https://mcp.apify.com?tools=vital_tuxedo/api-origin-health-report

Any other MCP client (Cursor, VS Code, ChatGPT developer mode, your own agent) - add the server to its config:

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com?tools=vital_tuxedo/api-origin-health-report",
"headers": { "Authorization": "Bearer <APIFY_TOKEN>" }
}
}
}

Drop the headers block to use OAuth instead: the browser opens once so you can sign in to Apify and approve access. Then ask in plain language:

Here are the 40 API origins on our vendor shortlist - which are dead, gated, or paused? Give me the DEAD ones with their evidence.

Your assistant starts the run, waits for it and reads the dataset back. The per-event prices above are the whole cost - going through MCP adds nothing.

No Apify account at all? This Actor is flagged eligible for agentic payments, so an agent holding USDC on Base can buy a prepaid Apify token at https://agi.apify.com over x402 and run it with no account and no signup. The instrument behind this Actor sells the same report on its own x402 route, so the rail is not a bolt-on.

FAQ

Is this uptime monitoring? No, and it will not invent uptime. One read at one moment is one read at one moment; there is no rolling availability figure here because none was measured. If you need alerting, a status page or an SLA, buy one of those. This is the instrument you point at 500 origins when you need to know what they are doing right now, with the evidence attached.

Will it grade or rank my vendors? No. You get verdicts from a closed set, coded reasons, and the source behind each one. What that means for your vendor choice is yours to decide.

Does it ignore robots.txt? No. With Respect robots.txt on - the default - a URL the origin disallows for this user agent is reported as skipped, no request is made to it, and you are not charged for it. The match is RFC 9309, so * and $ mean what the publisher meant by them, rather than the startswith comparison the Python standard library still does - which silently passes exactly the paths a site wrote down to forbid.

What if robots.txt is unreadable? It is recorded as unknown, never allowed, and it does not block the read. A partial answer is reported as partial: an origin the input cap stranded is reported as stranded, not guessed.

Why is one of my origins UNKNOWN rather than DEAD? Because the evidence did not settle it. A DNS miss, a TLS failure and a connection reset are different facts, and reasons names which one happened and where it came from.

Which mode should I pay for? Sweep in liveness at $0.0004 to find the rows worth looking at, then re-run those in report at $0.02. The two modes are exclusive per origin - a row bills one event or the other, never both.

Is this a re-implementation of something? The opposite. It runs liveness.py, the instrument behind a working paid API that has been serving this exact report - same verdict set, same reason codes, same catalogue sweep - on its own metered route. That file is vendored here byte for byte and pinned to its source by a test, so the two cannot drift.

Running locally

python -m src.main

Input is read from storage/key_value_stores/default/INPUT.json; the dataset lands under storage/datasets/default/.