# Pnet Jobs Search Scraper (`alexist/pnet-jobs-search-scraper`) Actor

Scrape job search results from PNet.co.za with precision. This scraper collects 30+ fields per listing including titles, company details, salary data, locations, skills, and job rankings — perfect for recruiters, job aggregators, and labor market researchers in South Africa.

- **URL**: https://apify.com/alexist/pnet-jobs-search-scraper.md
- **Developed by:** [Alex](https://apify.com/alexist) (community)
- **Categories:** Jobs, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are a software tools running on the Apify platform, for all kinds of web data extraction and automation use cases.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

In JavaScript/TypeScript projects, use official [JavaScript/TypeScript client](https://docs.apify.com/api/client/js/docs.md):

```bash
npm install apify-client
```

In Python projects, use official [Python client library](https://docs.apify.com/api/client/python/docs.md):

```bash
pip install apify-client
```

In shell scripts, use [Apify CLI](https://docs.apify.com/cli/docs.md):

````bash
# MacOS / Linux
curl -fsSL https://apify.com/install-cli.sh | bash
# Windows
irm https://apify.com/install-cli.ps1 | iex
```bash

In AI frameworks, you might use the [Apify MCP server](https://docs.apify.com/integrations/mcp.md).

If your project is in a different language, use the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).


# README

## PNet Jobs Search Scraper: Extract South African Job Listings at Scale

---

### What Is PNet.co.za?

PNet.co.za is South Africa's largest online job board, connecting millions of job seekers with employers across all industries and experience levels. It hosts thousands of active listings ranging from entry-level positions to executive roles. Manually collecting and organizing this data is tedious and error-prone — the **PNet Jobs Search Scraper** automates the extraction of search results, delivering clean, structured job data for analysis and integration.

---

### Overview

The **PNet Jobs Search Scraper** extracts job listings from PNet.co.za search result pages, converting search queries into machine-readable datasets. It is designed for:

- **Recruiters** monitoring job market trends and competitor activity in South Africa
- **HR professionals** benchmarking salaries and skill requirements across sectors
- **Job aggregators** feeding PNet listings into multi-platform job boards
- **Labor researchers** analyzing employment patterns and hiring demand
- **Career coaches** tracking industry-specific opportunities

Key strengths include support for targeted search URLs (by industry, location, keyword), failure tolerance via `ignore_url_failures`, and configurable scraping volume per URL.

---

### Input Format

The scraper accepts a JSON configuration object to control what and how much to scrape:

```json
{
  "urls": [
    "https://www.pnet.co.za/jobs/design?searchOrigin=Homepage_top-search"
  ],
  "ignore_url_failures": true,
  "max_items_per_url": 20
}
````

| Field | Description | Example |
|---|---|---|
| `urls` | Array of PNet search result page URLs to scrape | `https://www.pnet.co.za/jobs/design`, `https://www.pnet.co.za/jobs/marketing?location=johannesburg` |
| `ignore_url_failures` | Boolean flag: if `true`, scraper continues if a URL fails; if `false`, entire run stops | `true` or `false` |
| `max_items_per_url` | Maximum number of job listings extracted per URL (1–100+) | `20`, `50`, `100` |

**Tips:**

- Use filtered search URLs (e.g., `https://www.pnet.co.za/jobs/[category]?location=[city]`) to target specific job types or regions
- Set `max_items_per_url` based on your needs: `20` for quick samples, `50–100` for comprehensive datasets
- Enable `ignore_url_failures: true` for bulk runs to avoid interruptions if a search page is temporarily unavailable

***

### Output Format

**Sample output**

```json
{
  "id": 4232678,
  "title": "VJ 18979 - Design Draughtsman (Mining Equipment) - North West",
  "labels": [],
  "url": "/jobs--VJ-18979-Design-Draughtsman-Mining-Equipment-North-West-North-West-Professional-Career-Services-Gauteng--4232678-inline.html?rltr=1_1_25_seorl_m_0_0_0_0_0_0",
  "company_id": 12590,
  "company_name": "Professional Career Services - Gauteng",
  "company_url": "https://www.pnet.co.za/cmp/en/professional-career-services-gauteng-12590/jobs",
  "company_logo_url": "https://www.pnet.co.za/upload_za/logo/P/logoProfessional-Career-Services-Gauteng-12590ZEN.gif",
  "date_posted": "2026-07-14T13:07:19+02:00",
  "location": "North West",
  "is_anonymous": false,
  "salary": "R35,000 – R40,000 p/m",
  "post_code": null,
  "partnership": {
    "is_partnership_job": false,
    "show_partnership_label": false,
    "is_backfilled": false,
    "source_site_friendly_name": "",
    "is_cross_posted": false
  },
  "unified_salary": null,
  "work_from_home": "0",
  "meta_data": {
    "position_on_page": 1,
    "position_absolute": 1
  },
  "harmonised_id": "7E598A48-3557-4CE8-9D78-D13D83CA0EF9",
  "job_posting_sequence": null,
  "period_posted_date": null,
  "publish_from_date": null,
  "publish_to_date": null,
  "has_future_posting": false,
  "fingerprint_count": null,
  "section": "main",
  "top_labels": [],
  "skills": [],
  "text_snippet": "Our client is a specialist <strong>design</strong> company for the mining sector * Detailed <strong>design</strong> and draughting of mineral processing plants and mining equipment, including plant layouts, process flow integration, conveyors, chutes and screening equipment * Minimum 5+ years' experience in <strong>design</strong> draughting * Strong plant layout <strong>design</strong> experience * Experience <strong>designing</strong> conveyors, chutes, screens, mining equipment",
  "cv_to_job_score": null,
  "is_highlighted": false,
  "is_sponsored": false,
  "travel_time": null,
  "unified_travel_time": null,
  "is_top_job": false,
  "is_traffic_from_partner": false,
  "from_url": "https://www.pnet.co.za/jobs/design?searchOrigin=Homepage_top-search"
}
```

Each scraped job listing returns 30+ structured fields:

#### Core Identification

| Field | Meaning |
|---|---|
| `id` | Unique job listing identifier within PNet |
| `title` | Job title as displayed on the search result |
| `url` | Direct link to the full job posting |
| `labels` | Category tags assigned to the listing (e.g., "Full Time", "Permanent") |
| `top_labels` | High-priority category labels highlighting key job attributes |

#### Company Information

| Field | Meaning |
|---|---|
| `company_id` | Unique identifier for the employer |
| `company_name` | Name of the hiring company |
| `company_url` | Company's profile or website URL on PNet |
| `company_logo_url` | URL to the company's logo image |
| `partnership` | Whether the company is a PNet partner or premium advertiser |
| `is_traffic_from_partner` | Boolean flag: `true` if job traffic originated from a partner source |

#### Location & Logistics

| Field | Meaning |
|---|---|
| `location` | Primary job location (city, province, or region) |
| `post_code` | Postal code or zip code of the work location |
| `work_from_home` | Boolean: `true` if remote/WFH is available |
| `travel_time` | Estimated commute time from user location (if applicable) |
| `unified_travel_time` | Standardized travel time value for filtering |

#### Compensation & Timeline

| Field | Meaning |
|---|---|
| `salary` | Salary or salary range as displayed (raw format) |
| `unified_salary` | Standardized salary value for easy filtering and comparison |
| `date_posted` | Date the listing was published on PNet |
| `period_posted_date` | Generalized posting date (e.g., "1 week ago") |
| `publish_from_date` | Date when the job listing becomes visible |
| `publish_to_date` | Closing date for applications |
| `has_future_posting` | Boolean: `true` if job is scheduled to post in the future |

#### Job Attributes & Flags

| Field | Meaning |
|---|---|
| `is_anonymous` | Boolean: `true` if company details are hidden |
| `is_highlighted` | Boolean: `true` if listing has paid highlighting/premium visibility |
| `is_sponsored` | Boolean: `true` if job appears as a sponsored/promoted result |
| `is_top_job` | Boolean: `true` if job is featured in top positions |
| `cv_to_job_score` | Relevance score comparing user CV to job requirements (0–100) |

#### Candidate & Content

| Field | Meaning |
|---|---|
| `skills` | Array of required or preferred skills for the role |
| `text_snippet` | Brief excerpt from the job description |
| `section` | Job category section (e.g., "IT", "Sales", "Engineering") |
| `fingerprint_count` | Number of times this job has been viewed/interacted with |

#### Technical & Metadata

| Field | Meaning |
|---|---|
| `meta_data` | Additional structured metadata about the listing |
| `harmonised_id` | Standardized identifier used across aggregator platforms |
| `job_posting_sequence` | Internal sequence number for the job posting |

***

### How to Use

1. **Identify search URLs** — Navigate to PNet.co.za and perform a job search. Copy the URL from the search results page (e.g., `https://www.pnet.co.za/jobs/design`). You can filter by category, location, keyword, or any combination.

2. **Build your configuration** — Add one or more search URLs to the `urls` array. Adjust `max_items_per_url` based on your data volume needs:
   - `20–30`: Quick samples or testing
   - `50–100`: Comprehensive datasets
   - `100+`: Full market analysis

3. **Set failure handling** — Keep `ignore_url_failures: true` to ensure the scraper continues even if one URL times out or changes.

4. **Run the scraper** — Initiate the actor and monitor progress in the logs.

5. **Export and analyze** — Download results as JSON, CSV, or Excel for use in spreadsheets, databases, or BI tools.

**Common best practices:**

- Use specific search URLs (e.g., by job title or location) to gather targeted datasets
- For salary analysis, rely on the `unified_salary` field for consistent comparisons
- Filter listings using `is_top_job`, `is_highlighted`, or `is_sponsored` to segment premium postings
- Combine `skills` with `text_snippet` for quick content review without opening each listing

***

### Use Cases & Business Value

- **Recruitment analytics:** Track hiring trends by industry, location, and salary band in South Africa
- **Job aggregation:** Feed PNet listings into multi-site job boards or specialized portals
- **Salary benchmarking:** Compare compensation across roles, companies, and regions using unified salary data
- **Market research:** Analyze skill demand, in-demand roles, and competitor hiring activity
- **Career insights:** Help job seekers understand market competitiveness and skill gaps
- **HR intelligence:** Monitor hiring velocity and talent competition for your industry

By automating data collection, the scraper eliminates manual work and reveals patterns that guide strategic hiring, recruitment planning, and compensation strategy.

***

### Conclusion

The **PNet Jobs Search Scraper** is a powerful tool for anyone who needs structured South African job market data. With 30+ fields per listing and flexible URL targeting, it transforms unstructured search results into actionable insights. Whether you're a recruiter, researcher, or platform operator, this scraper accelerates your workflow and unlocks competitive advantage through data-driven job market intelligence.

# Actor input Schema

## `urls` (type: `array`):

Add the URLs of the Jobs list urls you want to scrape. You can paste URLs one by one, or use the Bulk edit section to add a prepared list.

## `ignore_url_failures` (type: `boolean`):

If true, the scraper will continue running even if some URLs fail to be scraped.

## `max_items_per_url` (type: `integer`):

The maximum number of items to scrape per URL.

## Actor input object example

```json
{
  "urls": [
    "https://www.pnet.co.za/jobs/design?searchOrigin=Homepage_top-search"
  ],
  "ignore_url_failures": true,
  "max_items_per_url": 20
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.pnet.co.za/jobs/design?searchOrigin=Homepage_top-search"
    ],
    "ignore_url_failures": true,
    "max_items_per_url": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("alexist/pnet-jobs-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://www.pnet.co.za/jobs/design?searchOrigin=Homepage_top-search"],
    "ignore_url_failures": True,
    "max_items_per_url": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("alexist/pnet-jobs-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.pnet.co.za/jobs/design?searchOrigin=Homepage_top-search"
  ],
  "ignore_url_failures": true,
  "max_items_per_url": 20
}' |
apify call alexist/pnet-jobs-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=alexist/pnet-jobs-search-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

```json
{
    "openapi": "3.0.1",
    "info": {
        "title": "Pnet Jobs Search Scraper",
        "description": "Scrape job search results from PNet.co.za with precision. This scraper collects 30+ fields per listing including titles, company details, salary data, locations, skills, and job rankings — perfect for recruiters, job aggregators, and labor market researchers in South Africa.",
        "version": "0.0",
        "x-build-id": "d63hjkyfIlgWl7ail"
    },
    "servers": [
        {
            "url": "https://api.apify.com/v2"
        }
    ],
    "paths": {
        "/acts/alexist~pnet-jobs-search-scraper/run-sync-get-dataset-items": {
            "post": {
                "operationId": "run-sync-get-dataset-items-alexist-pnet-jobs-search-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for its completion, and returns Actor's dataset items in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        },
        "/acts/alexist~pnet-jobs-search-scraper/runs": {
            "post": {
                "operationId": "runs-sync-alexist-pnet-jobs-search-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor and returns information about the initiated run in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK",
                        "content": {
                            "application/json": {
                                "schema": {
                                    "$ref": "#/components/schemas/runsResponseSchema"
                                }
                            }
                        }
                    }
                }
            }
        },
        "/acts/alexist~pnet-jobs-search-scraper/run-sync": {
            "post": {
                "operationId": "run-sync-alexist-pnet-jobs-search-scraper",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for completion, and returns the OUTPUT from Key-value store in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        }
    },
    "components": {
        "schemas": {
            "inputSchema": {
                "type": "object",
                "properties": {
                    "urls": {
                        "title": "URLs of the Jobs list urls to scrape",
                        "type": "array",
                        "description": "Add the URLs of the Jobs list urls you want to scrape. You can paste URLs one by one, or use the Bulk edit section to add a prepared list.",
                        "items": {
                            "type": "string"
                        }
                    },
                    "ignore_url_failures": {
                        "title": "Continue running even if some URLs fail to be scraped",
                        "type": "boolean",
                        "description": "If true, the scraper will continue running even if some URLs fail to be scraped."
                    },
                    "max_items_per_url": {
                        "title": "Max items per URL",
                        "type": "integer",
                        "description": "The maximum number of items to scrape per URL."
                    }
                }
            },
            "runsResponseSchema": {
                "type": "object",
                "properties": {
                    "data": {
                        "type": "object",
                        "properties": {
                            "id": {
                                "type": "string"
                            },
                            "actId": {
                                "type": "string"
                            },
                            "userId": {
                                "type": "string"
                            },
                            "startedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "finishedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "status": {
                                "type": "string",
                                "example": "READY"
                            },
                            "meta": {
                                "type": "object",
                                "properties": {
                                    "origin": {
                                        "type": "string",
                                        "example": "API"
                                    },
                                    "userAgent": {
                                        "type": "string"
                                    }
                                }
                            },
                            "stats": {
                                "type": "object",
                                "properties": {
                                    "inputBodyLen": {
                                        "type": "integer",
                                        "example": 2000
                                    },
                                    "rebootCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "restartCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "resurrectCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "computeUnits": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "options": {
                                "type": "object",
                                "properties": {
                                    "build": {
                                        "type": "string",
                                        "example": "latest"
                                    },
                                    "timeoutSecs": {
                                        "type": "integer",
                                        "example": 300
                                    },
                                    "memoryMbytes": {
                                        "type": "integer",
                                        "example": 1024
                                    },
                                    "diskMbytes": {
                                        "type": "integer",
                                        "example": 2048
                                    }
                                }
                            },
                            "buildId": {
                                "type": "string"
                            },
                            "defaultKeyValueStoreId": {
                                "type": "string"
                            },
                            "defaultDatasetId": {
                                "type": "string"
                            },
                            "defaultRequestQueueId": {
                                "type": "string"
                            },
                            "buildNumber": {
                                "type": "string",
                                "example": "1.0.0"
                            },
                            "containerUrl": {
                                "type": "string"
                            },
                            "usage": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "integer",
                                        "example": 1
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "usageTotalUsd": {
                                "type": "number",
                                "example": 0.00005
                            },
                            "usageUsd": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "number",
                                        "example": 0.00005
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}
```
