Job Freshness Validator - kill ghost jobs in your dataset
Pricing
$0.75 / 1,000 posting checkeds
Job Freshness Validator - kill ghost jobs in your dataset
Validate a list of job postings: detect dead links (404/410), expired pages, dead redirects, blocked pages and duplicates. Input: array of jobs with a URL field. Output: per-row verdict + stale-percentage summary.
$0.75 per 1,000 postings checked, no monthly fee.
You already have the jobs. Which ones are still real?
Job data goes stale fast: roles get filled, boards 404 quietly, and reposts pile up with tracking-parameter URLs that look like new listings. Paste your jobs as JSON or CSV rows, and every posting comes back labeled live, dead, expired, redirected, or duplicate - with the reason, per row. You also get a deduplicated result set and a freshness summary (stale % by status).
Built for job boards, aggregators, and recruiting-ops teams sitting on scraped or purchased inventory. It validates the data you own; it is not a scraper.
Input
{"jobs": [{"url": "https://example.com/jobs/123", "title": "Data Engineer", "company": "Acme"}],"urlField": "url","concurrency": 10}
Every job object needs a posting URL (field name configurable with urlField).
All other fields pass through to the output untouched.
Output
One dataset item per input row:
| field | meaning |
|---|---|
validity | live, dead_404, expired_content, redirect_dead, duplicate, blocked, invalid_row, or unknown |
reason | why (HTTP code, marker phrase found, redirect target) |
canonical_url | dedup key: tracking params (utm, gh_src, lever-source) stripped |
duplicate_of | canonical URL of the earlier row, when a duplicate |
Plus an OUTPUT key-value record: counts by status, stale count and stale
percentage, unique URL and duplicate counts.
Validity statuses
- live - HTTP 200, no expiry markers
- dead_404 - server says 404/410
- expired_content - page renders but says the role is filled/closed (Greenhouse, Lever and generic markers)
- redirect_dead - posting URL now redirects up to a list/careers root
- duplicate - same canonical URL as an earlier row (not re-fetched)
- blocked - 403/429 (bot protection or rate limit); reported, never bypassed
- invalid_row - missing/invalid URL
- unknown - timeout or 5xx; safe default, try again later
Compliance posture
It fetches only the URLs you provide, once per unique URL, with a declared user
agent. No logins, no personal data, no proxy rotation, no anti-bot bypass. Blocked
pages are reported as blocked, never circumvented. Each source's terms still need
your review before production use.
Honest limits
unknownis a safe default, not a failure to decide: timeouts and 5xx are transient by nature.- Content-expiry detection relies on marker phrases; coverage spans Greenhouse/Lever/generic markers today and grows per ATS.
- Sites behind aggressive bot protection will return
blocked- reported honestly, not routed around.