# Phenom Jobs Scraper — RBC, CVS Health, BMO & 140+ Career Sites (`yugenox/phenom-jobs-scraper`) Actor

Every open job from any Phenom-powered careers site: RBC, BMO, Manulife, CVS Health, Cigna, Mastercard, DHL, Lowe's, UPS and 140+ more. Paste a URL or type a company. Complete past 10,000 jobs, AI skills, salary, full descriptions, hiring stats, only-new mode.

- **URL**: https://apify.com/yugenox/phenom-jobs-scraper.md
- **Developed by:** [Yugenox Corp](https://apify.com/yugenox) (community)
- **Categories:** Jobs, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 jobs

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Phenom Career Sites Jobs Scraper

Get **every open job** from any company careers site built on Phenom. That includes RBC, BMO, Manulife, CVS Health, The Cigna Group, Mastercard, Thermo Fisher, DHL, Lowe's, UPS, Humana, PNC, U.S. Bank, Kohl's, Dollar Tree, TJX, Five Below, Wawa, Abbott and GE HealthCare, plus big health systems like Trinity Health, Sutter Health and Duke Health. About 150 employers (260,000+ open jobs) are in the built-in list, and any other Phenom site works from its URL.

Paste a careers URL or type a company name. Add keywords, locations and filters if you want them. You get clean, structured job data: title, location with coordinates, category, job type, remote/hybrid, posting date, AI-extracted skills, apply link and, optionally, the full description with salary.

### Why this scraper

- **Complete results, even for the largest employers.** Career sites stop showing results after about 10,000 jobs, and paging through them skips or repeats around 5% of jobs. This scraper splits each search into slices of 500 jobs or fewer by category, job family, job type and location. Every job is collected exactly once. CVS Health: all 19,936 of 19,936 jobs in one run.
- **Works with any Phenom careers site.** No fixed list of boards. Paste the search page, a category page, a single job page or just the domain. Filters already in the URL are kept.
- **Company names work too.** Type "RBC", "CVS Health" or "Mastercard" and the built-in list finds the site.
- **Richer data than generic job scrapers.** You get AI-extracted skills, the site's AI job summary, salary range, compensation grade, job family, career level, business unit, legal employer, closing date, all locations with latitude/longitude, and more.
- **Hiring stats.** Optional rows count open jobs by country, state, city, category, job type, work arrangement and business unit, for hiring-trend or competitor tracking.
- **Only-new mode.** Schedule it daily and get only the jobs you haven't seen before.
- **Contact details are hidden.** Recruiter names and emails are never included, and email addresses and phone numbers inside descriptions are masked.

### Input

| Field | What it does |
|---|---|
| **Career site URLs** | Any page of a Phenom careers site, e.g. `https://jobs.rbc.com/ca/en/search-results`, `https://jobs.cvshealth.com/us/en/search-results?keywords=pharmacist`, `https://jobs.rbc.com/ca/en/c/technology-analytics-research-jobs`, `jobs.bmo.com`. |
| **Companies** | Company names, e.g. `RBC`, `BMO`, `Manulife`, `CVS Health`, `Cigna`, `Mastercard`, `Thermo Fisher`, `DHL`, `Lowe's`, `UPS`. |
| **Keywords** | Each keyword is its own search, the same as the site's search box. Leave empty to get every job. Keywords with no sites or companies search the largest employers in the built-in list. |
| **Locations** | `Canada`, `Ontario`, `ON`, `Toronto`, `Texas`, `USA`, or `City, State` to be exact. Matched to each site's own spelling. |
| **Job categories / Employment type / Work arrangement / Posted within** | Optional filters. |
| **Full job details** | Adds the full description, salary, closing date, job family and more. Makes one extra request per job. |
| **Max jobs / Max jobs per site** | Limits. Leave empty to get everything. |
| **Only new jobs** | Skips jobs returned in earlier runs. |
| **Hiring stats** | Adds job-count rows per country, state, city, category, and more. |
| **Advanced: any site filter** | `{"jobFamily": ["Audit"]}`, `{"businessUnit": ["Aetna"]}` and so on. |

#### Example inputs

All software jobs in the US at Mastercard and RBC, with full details:

```json
{
  "companies": ["Mastercard", "RBC"],
  "keywords": ["software"],
  "locations": ["USA"],
  "includeDetails": true
}
```

Every nursing job at three health systems, posted in the last 7 days, as a daily feed:

```json
{
  "companies": ["Trinity Health", "Sutter Health", "Duke Health"],
  "categories": ["Nursing"],
  "postedWithinDays": 7,
  "onlyNew": true
}
```

### Output

One row per job. A real row (RBC, with full details), shortened:

```json
{
  "title": "Financial Analyst",
  "company": "Royal Bank of Canada",
  "hiringEntity": "RBC Capital Markets, LLC",
  "location": "Minneapolis, Minnesota, United States of America",
  "city": "Minneapolis",
  "state": "Minnesota",
  "country": "United States of America",
  "latitude": 44.97997,
  "longitude": -93.26384,
  "employmentType": "Full time",
  "category": "Finance | Accounting",
  "subCategory": "Finance, Accounting and Treasury",
  "postedDate": "2026-09-03",
  "postedDaysAgo": 21,
  "summary": "We are looking for a Financial Analyst to provide critical financial insights that support strategic decision-making…",
  "skills": ["financial modeling", "variance analysis", "financial reporting", "microsoft excel", "budgeting"],
  "url": "https://jobs.rbc.com/ca/en/job/R-0000184407/financial-analyst",
  "applyUrl": "https://rbc.wd3.myworkdayjobs.com/RBCGLOBAL1/job/…/Financial-Analyst_R-0000184407-1/apply",
  "jobId": "R-0000184407",
  "descriptionText": "What is the opportunity?\n\nResponsible for providing the business with financial information…",
  "aiSummary": "The Financial Analyst at RBC will be responsible for delivering financial information to support strategic business decisions…",
  "salaryMin": 50000,
  "salaryMax": 85000,
  "salaryCurrency": "USD",
  "salaryPeriod": "YEAR",
  "salarySource": "description",
  "payType": "Salaried",
  "closingDate": "2026-10-30",
  "jobFamily": "Financial Performance Management",
  "jobProfile": "Financial Performance Management Analyst",
  "careerLevel": "Professional",
  "compensationGrade": "GG09",
  "workerType": "Employee",
  "detailsIncluded": true,
  "careerSite": "jobs.rbc.com",
  "scrapedAt": "2026-09-24T10:55:45.849Z"
}
```

Fields a site doesn't publish are `null`. Salary comes from the site's own pay fields when it has them, otherwise from the pay range written in the description (`salarySource` tells you which).

When **Hiring stats** is on, you also get rows like `{"recordType": "facet", "company": "CVS Health", "field": "state", "value": "Texas", "count": 1511}`.

The run's key-value store has a `SUMMARY` record for each site. It shows how many jobs matched, how many were collected, and the coverage.

### Reliability and billing

- **The run always finishes with what it found.** A site that is down, blocked or not built on Phenom is named in the log and the status message while the other sites continue. If the run's time limit is near, it stops early, saves everything collected and reports `Partial — …`. If your spending limit is reached it stops with `Stopped at your spending limit`. Bad input gives a how-to message and 0 rows.
- **You only pay for rows you receive.** Each job row is billed once it is written to the dataset; a job whose details could not be fetched is delivered as a list row and is not billed as a detailed one; hiring-stats rows have their own (much lower) price; a run that delivers nothing costs nothing.
- **Nothing to configure.** No account, cookie or API key: the scraper reads each site's public configuration on every run, so per-site changes heal themselves. Apify's datacenter proxy is enough for every verified site; a request a firewall refuses is retried on a fresh IP and then through Apify's unblocking proxy automatically.
- **Coverage is reported.** The `SUMMARY` record in the run's key-value store lists, per site and query, how many jobs matched and how many were collected.

### Use cases

- **Job boards and aggregators.** Pull full, de-duplicated job feeds from large enterprise employers.
- **Recruiting and sales intelligence.** See who is hiring for what, where, and how fast.
- **Labour-market research.** Track open roles, skills demand, pay ranges and remote share by company, region or category.
- **Job seekers.** Get a daily feed of new postings at the companies you care about.

### FAQ

**Which sites does it work with?** Career sites built on the Phenom platform. The page address usually looks like `…/us/en/search-results`, `…/ca/en/search-results` or `…/global/en/search-results`. If a URL isn't a Phenom site, the run tells you and moves on without charging you.

**How fast is it?** A list of 1,000 jobs takes about 5 seconds. CVS Health's roughly 20,000 jobs take under 2 minutes on Apify (most of that is saving the rows), and five employers with 70,000 jobs between them took 37 seconds to collect. With full details it runs at about 10 jobs per second per site, and several sites run in parallel.

**Why does "Toronto" work when the site spells it "TORONTO"?** Locations, categories and job types are matched to each site's own values, ignoring case and accents.

**Do I need a proxy?** No. The default Apify proxy is used. If a site's firewall blocks a request, that request is retried on a new IP and then through Apify's unblocking proxy automatically.

**A site was "blocked on datacenter IPs". What do I change?** A few career sites (mostly ones behind a strict firewall) refuse Apify's standard datacenter IPs. For those, the scraper needs Apify's unblocker or residential proxy on your plan. If your plan has neither, the other sites still deliver, you are not charged for the blocked site, and the status message names each blocked site. To get them too, enable residential proxies (or the unblocker proxy) in Apify Console → Proxy and run again.

**Is anyone's personal data included?** No. Recruiter and hiring-manager fields are never output, and emails and phone numbers inside job descriptions are masked.

**Is it legal to scrape Phenom career sites?** This scraper only collects job postings that employers publish on their own public career sites, the same pages anyone can open in a browser without logging in. Collecting public data like this is generally allowed, but how you use it is up to you: follow the data protection laws that apply to you (such as GDPR, PIPEDA or CCPA) and the terms of the sites you collect from. If you are unsure whether your use case is allowed, check with a lawyer. Apify's guide [Is web scraping legal?](https://blog.apify.com/is-web-scraping-legal/) is a good starting point.

**Does it access any private data?** No. It never logs in and uses no accounts, cookies or API keys. It reads only what each career site shows to every visitor: the public job listings, job pages and the site's own job counts. Internal fields such as recruiter or hiring-manager names and emails are dropped before anything is saved, and contact details inside descriptions are masked.

# Actor input Schema

## `startUrls` (type: `array`):

Any page of a Phenom-powered careers site: the job search page (filters in the URL are kept), a category page (…/c/technology-jobs), a single job page, or just the domain (jobs.rbc.com). Examples: https://jobs.rbc.com/ca/en/search-results, https://jobs.cvshealth.com/us/en/search-results, https://careers.mastercard.com/us/en/search-results.

## `companies` (type: `array`):

Company names, resolved through a built-in list of verified Phenom career sites (e.g. RBC, BMO, Manulife, CVS Health, Cigna, Mastercard, Thermo Fisher, DHL, Lowe's, UPS, Humana, PNC, Trinity Health, Sutter Health). A careers URL works here too.

## `keywords` (type: `array`):

Job title or keywords, searched on each site exactly like its own search box. Each keyword is a separate search; results are de-duplicated. Leave empty for every open job. Keywords with no sites/companies search the largest employers in the built-in list.

## `maxItems` (type: `integer`):

Maximum number of jobs in the whole run. Leave empty to get every matching job (tens of thousands for the biggest employers).

## `maxItemsPerSite` (type: `integer`):

Cap per career site, so one big employer does not use up the whole run.

## `includeDetails` (type: `boolean`):

Adds the full description (text + HTML), AI-written summary, salary range, closing date, job family, career level, business unit, legal employer, worker type and schedule. One extra request per job, so it is slower; leave off for a fast list with title, location, category, skills, dates and links.

## `locations` (type: `array`):

Countries, states/provinces or cities: Canada, Ontario, ON, Toronto, Texas, USA. Use “City, State” to be exact (e.g. “New York, New York”). Matched to each site's own location filters, whatever their spelling (TORONTO, Montréal…). Several locations are combined with OR.

## `categories` (type: `array`):

The site's job categories, e.g. Technology, Nursing, Finance, Operations, Pharmacy. Partial names work (“Technology” matches “Technology | Analytics | Research”).

## `employmentTypes` (type: `array`):

Matched to each site's own job types.

## `workplaceTypes` (type: `array`):

Remote, hybrid or on-site. Uses the site's own remote filter where it has one, otherwise what each job states.

## `postedWithinDays` (type: `integer`):

Only jobs posted in the last N days.

## `facetFilters` (type: `object`):

Filter on any field the site indexes, as {"field": \["value", …]} — e.g. {"jobFamily": \["Audit"]}, {"businessUnit": \["Aetna"]}, {"remote": \["Hybrid"]}. Values are matched case-insensitively.

## `onlyNew` (type: `boolean`):

Remembers every job this actor returned to you and skips it next time. Schedule it daily for a new-jobs feed; you only pay for new jobs.

## `onlyNewStateKey` (type: `string`):

Use a different name per saved task/schedule to keep their “already seen” lists separate.

## `includeFacets` (type: `boolean`):

Also output how many matching jobs each site has per country, state, city, category, job type, work arrangement, job family and business unit (rows with recordType "facet"). Useful for hiring-trend and competitor tracking.

## `maxSites` (type: `integer`):

With keywords but no sites or companies, this many of the largest employers in the built-in list are searched (only those hiring in your locations' countries).

## `maxConcurrency` (type: `integer`):

Requests in flight across all sites.

## `proxyConfiguration` (type: `object`):

Apify's default datacenter proxy works for the sites we have tested and is the cheapest. When a site's firewall refuses a request, that request is retried on a fresh IP and then through Apify's unblocking proxy automatically.

## `_noUnblocker` (type: `boolean`):

Internal (canaries): behave exactly like an account without the UNBLOCKER proxy group.

## `_noResidential` (type: `boolean`):

Internal (canaries): behave exactly like an account without the RESIDENTIAL proxy group.

## Actor input object example

```json
{
  "startUrls": [
    "https://jobs.rbc.com/ca/en/search-results"
  ],
  "keywords": [
    "data analyst"
  ],
  "maxItems": 50,
  "includeDetails": false,
  "onlyNew": false,
  "onlyNewStateKey": "default",
  "includeFacets": false,
  "maxSites": 60,
  "maxConcurrency": 16,
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "_noUnblocker": false,
  "_noResidential": false
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped job postings.

## `summary` (type: `string`):

Per site: jobs matching, jobs collected, coverage.

## `run` (type: `string`):

Status and statistics for this run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://jobs.rbc.com/ca/en/search-results"
    ],
    "keywords": [
        "data analyst"
    ],
    "maxItems": 50,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("yugenox/phenom-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://jobs.rbc.com/ca/en/search-results"],
    "keywords": ["data analyst"],
    "maxItems": 50,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("yugenox/phenom-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://jobs.rbc.com/ca/en/search-results"
  ],
  "keywords": [
    "data analyst"
  ],
  "maxItems": 50,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call yugenox/phenom-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,yugenox/phenom-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/TTfZPVMYgwomMaCBN/builds/HsL4cxBrSR3Zlli47/openapi.json
