Employment News Scraper avatar

Employment News Scraper

Pricing

from $2.99 / 1,000 employment news notices

Go to Apify Store
Employment News Scraper

Employment News Scraper

Scrape job notifications from Employment News (employmentnews.gov.in), India's official weekly employment gazette published by the Ministry of Information and Broadcasting. Extract notification titles, categories, application deadlines, and eligibility for government job seekers.

Pricing

from $2.99 / 1,000 employment news notices

Rating

0.0

(0)

Developer

Jobs API

Jobs API

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

11 days ago

Last modified

Share

Employment News Jobs Search Scraper

This Apify Actor reads the public Employment News jobs table and joins each matching notification to its official advertisement PDF. It emits complete, source-grounded notification records rather than listing-only rows.

Public input

{
"query": "recruitment",
"maxItems": 3
}

query is matched against the public organisation, post, and appointment-method columns. maxItems is bounded to 1–50. The Actor uses ordinary direct access and exposes no proxy, stealth, fingerprint, retry, concurrency, fixture, or temporary context controls.

Source workflow

  1. Read https://employmentnews.gov.in/NewEmp/AllJobs.aspx?k=All.
  2. Read the official Web Advertisement index at https://employmentnews.gov.in/NewEmp/MoreContentS.aspx?n=WebAdvertisement.
  3. Match each selected organisation to its official PDF.
  4. Fetch the PDF directly, record its hash, size, headers, page count, text-layer status, and source receipt, then emit the record.

Image-only official PDFs are reported as pdfTextExtractionStatus: "IMAGE_ONLY"; the Actor does not invent OCR text or substitute another source.

Output

Each dataset record contains the five source table columns, normalized date objects, matching issue and advertisement metadata, official PDF integrity and extraction metadata, direct request receipts, stable URLs, and quality/provenance fields. The schema in .actor/dataset_schema.json is strict at the top level and accepts only fields produced by the Actor.

The run also stores OUTPUT_SUMMARY, REQUEST_RECEIPTS, RUN_DIAGNOSTICS, and OUTPUT in the default key-value store.

status reports the Apify process outcome (SUCCEEDED or FAILED); resultStatus separately reports data completeness (COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED). Only complete notifications that pass the first-party provenance/quality gate and whose awaited Dataset writes resolve are counted. If a later failure occurs after rows were stored, the platform run remains SUCCEEDED with resultStatus: "LIMITED". Fatal zero-row runtime or storage failures call Actor.fail() and report FAILED / FAILED. Diagnostics and request receipts are KVS-only and never become Dataset rows.

Zero matching rows are DEFERRED unless the source itself provides a readable empty-results statement with a successful first-party receipt for every requested search. An empty parse alone is not positive NO_DATA evidence.

Access policy

An explicit HTTP 401/403/429/451 response or verification/access-denial challenge from the official Employment News HTTPS domain stops the route and discards all staged records; the platform status is SUCCEEDED and resultStatus is SKIPPED. A denial/challenge on a non-first-party redirect, HTTP 407, other unverified responses, and browser/controller/transport failures are DEFERRED, not source blocks. If real rows were already stored before a later error, they are preserved as SUCCEEDED / LIMITED. The Actor does not switch proxies, identities, browser fingerprints, or alternate routes to work around a restriction.

Local validation

npm install
npm test
npm run lint
apify validate-schema
apify run --purge --input-file ./INPUT.json
npm run validate:dataset -- storage

The dataset validator requires successful source receipts, a matched official PDF, direct-source provenance, and more than 20 meaningful populated fields per emitted record.