Job Freshness Validator - kill ghost jobs in your dataset avatar

Job Freshness Validator - kill ghost jobs in your dataset

Pricing

$0.75 / 1,000 posting checkeds

Go to Apify Store
Job Freshness Validator - kill ghost jobs in your dataset

Job Freshness Validator - kill ghost jobs in your dataset

Validate a list of job postings: detect dead links (404/410), expired pages, dead redirects, blocked pages and duplicates. Input: array of jobs with a URL field. Output: per-row verdict + stale-percentage summary.

Pricing

$0.75 / 1,000 posting checkeds

Rating

0.0

(0)

Developer

muazah

muazah

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

$0.75 per 1,000 postings checked, no monthly fee.

You already have the jobs. Which ones are still real?

Job data goes stale fast: roles get filled, boards 404 quietly, and reposts pile up with tracking-parameter URLs that look like new listings. Paste your jobs as JSON or CSV rows, and every posting comes back labeled live, dead, expired, redirected, or duplicate - with the reason, per row. You also get a deduplicated result set and a freshness summary (stale % by status).

Built for job boards, aggregators, and recruiting-ops teams sitting on scraped or purchased inventory. It validates the data you own; it is not a scraper.

Input

{
"jobs": [
{"url": "https://example.com/jobs/123", "title": "Data Engineer", "company": "Acme"}
],
"urlField": "url",
"concurrency": 10
}

Every job object needs a posting URL (field name configurable with urlField). All other fields pass through to the output untouched.

Output

One dataset item per input row:

fieldmeaning
validitylive, dead_404, expired_content, redirect_dead, duplicate, blocked, invalid_row, or unknown
reasonwhy (HTTP code, marker phrase found, redirect target)
canonical_urldedup key: tracking params (utm, gh_src, lever-source) stripped
duplicate_ofcanonical URL of the earlier row, when a duplicate

Plus an OUTPUT key-value record: counts by status, stale count and stale percentage, unique URL and duplicate counts.

Validity statuses

  • live - HTTP 200, no expiry markers
  • dead_404 - server says 404/410
  • expired_content - page renders but says the role is filled/closed (Greenhouse, Lever and generic markers)
  • redirect_dead - posting URL now redirects up to a list/careers root
  • duplicate - same canonical URL as an earlier row (not re-fetched)
  • blocked - 403/429 (bot protection or rate limit); reported, never bypassed
  • invalid_row - missing/invalid URL
  • unknown - timeout or 5xx; safe default, try again later

Compliance posture

It fetches only the URLs you provide, once per unique URL, with a declared user agent. No logins, no personal data, no proxy rotation, no anti-bot bypass. Blocked pages are reported as blocked, never circumvented. Each source's terms still need your review before production use.

Honest limits

  • unknown is a safe default, not a failure to decide: timeouts and 5xx are transient by nature.
  • Content-expiry detection relies on marker phrases; coverage spans Greenhouse/Lever/generic markers today and grows per ATS.
  • Sites behind aggressive bot protection will return blocked - reported honestly, not routed around.