PNet Search Scraper
Pricing
from $2.99 / 1,000 pnet job records
PNet Search Scraper
Scrape job listings from PNet.co.za, South Africa's leading job board. Extract job titles, companies, locations, salary ranges, and descriptions for South African recruitment.
Pricing
from $2.99 / 1,000 pnet job records
Rating
0.0
(0)
Developer
Jobs API
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
11 days ago
Last modified
Categories
Share
PNet Jobs Search Scraper
This Actor retrieves complete public PNet South Africa job postings from the official www.pnet.co.za listing and detail pages. It follows the public search redirect, parses server-rendered listing cards, enriches each selected job from its official detail page, and writes only detail-verified records.
Input modes
search: one query and location, using PNet's official search route.searchMultiple: several queries with stable job-ID deduplication.single: one exact official PNet detail URL.multiple: several exact official PNet detail URLs.startUrls: mixed official listing/search URLs and detail URLs.jobUrlandjobUrls: compatibility aliases for direct modes.
All URLs must use HTTPS and the PNet host. Search and detail requests are bounded by page, candidate, detail, retry, timeout, and total-request limits.
Output
Job records include the PNet identifier and reference number, title, company, listing snippet, location, structured salary facts, contract/work type, publication and closing dates, employment metadata, skills and requirements where published, company links, all available structured locations/coordinates, complete description text and HTML, headings, sections, bullets, links, source evidence, canonical identity checks, request receipts, and a list/count of meaningful populated source fields. Complete job rows must contain more than 20 meaningful PNet source fields. Published email addresses and telephone numbers are redacted from output text and structured fields.
RUN_SUMMARY, OUTPUT_SUMMARY, RUN_DIAGNOSTICS, RUN_METADATA, and RUN_HEALTH are stored in the default key-value store. Incomplete, expired, mismatched, blocked, or otherwise unverifiable pages are reported as structured diagnostics and never emitted as job rows. HTTP 401, 403, 429, or 451 and visible source-authored access challenges stop requests immediately, discard buffered job rows, and produce SKIPPED with the matching receipt and evidence. Generic request or tool failures that do not prove a source block are classified as DEFERRED.
Transport policy
The implementation uses ordinary native HTTPS requests and Cheerio in local or Apify cloud runs. It does not use a browser, proxy, fingerprint injection, CAPTCHA/WAF bypass, login, or third-party mirror. Dataset writes are buffered until validation succeeds.
Local verification
npm ci --ignore-scripts --no-audit --no-fundnpm run checknpm testnpx --yes apify-cli validate-schema .actor/input_schema.jsonNew-Item -ItemType Directory -Path storage-pnet-validation$env:APIFY_LOCAL_STORAGE_DIR = 'storage-pnet-validation'npx --yes apify-cli run --resurrect --input-file INPUT.jsonnpm run validate -- storage-pnet-validation
Before running the Actor, confirm the selected storage directory does not already exist and inspect both root and actor-local legacy apify_storage directories; preserve their contents separately rather than letting the CLI rename or replace them. Do not use --purge or write/modify dataset files manually. The exact detail URLs in the input samples can expire on the source; the negative input verifies diagnostics for an official-shaped missing URL.