Employment News Scraper
Pricing
from $2.99 / 1,000 employment news notices
Employment News Scraper
Scrape job notifications from Employment News (employmentnews.gov.in), India's official weekly employment gazette published by the Ministry of Information and Broadcasting. Extract notification titles, categories, application deadlines, and eligibility for government job seekers.
Pricing
from $2.99 / 1,000 employment news notices
Rating
0.0
(0)
Developer
Jobs API
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
11 days ago
Last modified
Categories
Share
Employment News Jobs Search Scraper
This Apify Actor reads the public Employment News jobs table and joins each matching notification to its official advertisement PDF. It emits complete, source-grounded notification records rather than listing-only rows.
Public input
{"query": "recruitment","maxItems": 3}
query is matched against the public organisation, post, and appointment-method columns. maxItems is bounded to 1–50. The Actor uses ordinary direct access and exposes no proxy, stealth, fingerprint, retry, concurrency, fixture, or temporary context controls.
Source workflow
- Read
https://employmentnews.gov.in/NewEmp/AllJobs.aspx?k=All. - Read the official Web Advertisement index at
https://employmentnews.gov.in/NewEmp/MoreContentS.aspx?n=WebAdvertisement. - Match each selected organisation to its official PDF.
- Fetch the PDF directly, record its hash, size, headers, page count, text-layer status, and source receipt, then emit the record.
Image-only official PDFs are reported as pdfTextExtractionStatus: "IMAGE_ONLY"; the Actor does not invent OCR text or substitute another source.
Output
Each dataset record contains the five source table columns, normalized date objects, matching issue and advertisement metadata, official PDF integrity and extraction metadata, direct request receipts, stable URLs, and quality/provenance fields. The schema in .actor/dataset_schema.json is strict at the top level and accepts only fields produced by the Actor.
The run also stores OUTPUT_SUMMARY, REQUEST_RECEIPTS, RUN_DIAGNOSTICS, and OUTPUT in the default key-value store.
status reports the Apify process outcome (SUCCEEDED or FAILED); resultStatus separately reports data completeness (COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED). Only complete notifications that pass the first-party provenance/quality gate and whose awaited Dataset writes resolve are counted. If a later failure occurs after rows were stored, the platform run remains SUCCEEDED with resultStatus: "LIMITED". Fatal zero-row runtime or storage failures call Actor.fail() and report FAILED / FAILED. Diagnostics and request receipts are KVS-only and never become Dataset rows.
Zero matching rows are DEFERRED unless the source itself provides a readable empty-results statement with a successful first-party receipt for every requested search. An empty parse alone is not positive NO_DATA evidence.
Access policy
An explicit HTTP 401/403/429/451 response or verification/access-denial challenge from the official Employment News HTTPS domain stops the route and discards all staged records; the platform status is SUCCEEDED and resultStatus is SKIPPED. A denial/challenge on a non-first-party redirect, HTTP 407, other unverified responses, and browser/controller/transport failures are DEFERRED, not source blocks. If real rows were already stored before a later error, they are preserved as SUCCEEDED / LIMITED. The Actor does not switch proxies, identities, browser fingerprints, or alternate routes to work around a restriction.
Local validation
npm installnpm testnpm run lintapify validate-schemaapify run --purge --input-file ./INPUT.jsonnpm run validate:dataset -- storage
The dataset validator requires successful source receipts, a matched official PDF, direct-source provenance, and more than 20 meaningful populated fields per emitted record.