PNet Search Scraper avatar

PNet Search Scraper

Pricing

from $2.99 / 1,000 pnet job records

Go to Apify Store
PNet Search Scraper

PNet Search Scraper

Scrape job listings from PNet.co.za, South Africa's leading job board. Extract job titles, companies, locations, salary ranges, and descriptions for South African recruitment.

Pricing

from $2.99 / 1,000 pnet job records

Rating

0.0

(0)

Developer

Jobs API

Jobs API

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

11 days ago

Last modified

Categories

Share

PNet Jobs Search Scraper

This Actor retrieves complete public PNet South Africa job postings from the official www.pnet.co.za listing and detail pages. It follows the public search redirect, parses server-rendered listing cards, enriches each selected job from its official detail page, and writes only detail-verified records.

Input modes

  • search: one query and location, using PNet's official search route.
  • searchMultiple: several queries with stable job-ID deduplication.
  • single: one exact official PNet detail URL.
  • multiple: several exact official PNet detail URLs.
  • startUrls: mixed official listing/search URLs and detail URLs.
  • jobUrl and jobUrls: compatibility aliases for direct modes.

All URLs must use HTTPS and the PNet host. Search and detail requests are bounded by page, candidate, detail, retry, timeout, and total-request limits.

Output

Job records include the PNet identifier and reference number, title, company, listing snippet, location, structured salary facts, contract/work type, publication and closing dates, employment metadata, skills and requirements where published, company links, all available structured locations/coordinates, complete description text and HTML, headings, sections, bullets, links, source evidence, canonical identity checks, request receipts, and a list/count of meaningful populated source fields. Complete job rows must contain more than 20 meaningful PNet source fields. Published email addresses and telephone numbers are redacted from output text and structured fields.

RUN_SUMMARY, OUTPUT_SUMMARY, RUN_DIAGNOSTICS, RUN_METADATA, and RUN_HEALTH are stored in the default key-value store. Incomplete, expired, mismatched, blocked, or otherwise unverifiable pages are reported as structured diagnostics and never emitted as job rows. HTTP 401, 403, 429, or 451 and visible source-authored access challenges stop requests immediately, discard buffered job rows, and produce SKIPPED with the matching receipt and evidence. Generic request or tool failures that do not prove a source block are classified as DEFERRED.

Transport policy

The implementation uses ordinary native HTTPS requests and Cheerio in local or Apify cloud runs. It does not use a browser, proxy, fingerprint injection, CAPTCHA/WAF bypass, login, or third-party mirror. Dataset writes are buffered until validation succeeds.

Local verification

npm ci --ignore-scripts --no-audit --no-fund
npm run check
npm test
npx --yes apify-cli validate-schema .actor/input_schema.json
New-Item -ItemType Directory -Path storage-pnet-validation
$env:APIFY_LOCAL_STORAGE_DIR = 'storage-pnet-validation'
npx --yes apify-cli run --resurrect --input-file INPUT.json
npm run validate -- storage-pnet-validation

Before running the Actor, confirm the selected storage directory does not already exist and inspect both root and actor-local legacy apify_storage directories; preserve their contents separately rather than letting the CLI rename or replace them. Do not use --purge or write/modify dataset files manually. The exact detail URLs in the input samples can expire on the source; the negative input verifies diagnostics for an official-shaped missing URL.