# HigherEdJobs Scraper: Faculty, Staff & Admin Jobs (`santamaria-automations/higheredjobs-scraper`) Actor

Scrape HigherEdJobs.com listings for faculty, administrative and staff roles at US universities and colleges. Returns title, institution, location, salary, category, badges and, with details on, the full description, apply URL, employer website, requisition id and rank.

- **URL**: https://apify.com/santamaria-automations/higheredjobs-scraper.md
- **Developed by:** [NanoScrape](https://apify.com/santamaria-automations) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 actor starts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## HigherEdJobs Scraper: Faculty, Staff and Administrative Jobs in US Higher Education

Extract job listings from [HigherEdJobs.com](https://www.higheredjobs.com), the largest US higher-education job board. Faculty, administrative, executive and staff roles at universities, colleges and community colleges, returned as structured data. No login and no browser needed.

### What data can you extract?

Every job includes:

- **Position:** title, category, position type (Full-Time, Adjunct/Part-Time), employment type, department, detected academic rank, tenure-track flag, credentials required (PhD, MD, JD, EdD, Master's and more)
- **Institution:** name, location (city, state, country), remote flag, institution website, HigherEdJobs institution profile URL
- **Pay and timing:** salary text (trimmed to the pay sentence) with parsed min, max and period (see Salary below), posted date and time, application deadline (`deadline_at` as a `YYYY-MM-DD` date when the posting gives a future date, `deadline_text` with the raw text, `deadline_open_until_filled`)
- **Listing badges:** Military-friendly, Inclusive Workplace, Dual Career, Priority
- **With `includeJobDetails`:** full description (text and HTML), final apply URL, the institution's own requisition id, industry tags
- **Meta:** search query, results page, scrape time, job URL

Most fields appear twice: the original name and a canonical name (`institution_name` and `company_name`, `job_url` and `source_url`, `application_url` and `apply_url`, `posted_at` and `posted_at_datetime`).

### Salary fields

Only dollar-prefixed amounts are parsed, and `45K` means 45,000. `salary_period` is `annual`, `monthly` or `hourly`. It comes from the pay text (hour, month, year, annual); when no period is named and every amount is 15,000 or more, it is `annual`. Smaller amounts with no period leave it null. Non-standard pay (per semester, per course, per credit hour, per load, per week or day) keeps the amount in `salary_text`, leaves `salary_min` and `salary_max` null and puts the unit in `salary_period_raw` (`semester`, `course`, `credit`, ...). `salary_period_raw` also keeps the older vocabulary (`year`, `month`, `hour`) for standard periods.

### Dates, remote and institution type

- Search result pages show only a relative posted date ("Posted 3 days ago"), so `posted_at_datetime` is date-only (midnight UTC) until the detail page adds the exact time (`includeJobDetails`).
- `remote_option` is `onsite`, `hybrid` or `remote`. Postings with no remote or hybrid signal are `onsite`, because the site files them under Location Bound.
- `institution_type` is exact when you use the `institutionType` filter. Otherwise it is inferred from the institution name and is null when unclear, because the site does not print it on job pages.
- `industry` is one constant tag set by the site, `dual_career` is a rare badge and `priority` is a paid upgrade shown on a small share of jobs. Treat them as informational.

### Pricing

Pay per event, no subscription:

- **Start:** $0.001 per run
- **Search result:** $0.003 per job returned
- **Detail result:** $0.005 per job where the description was parsed (only when `includeJobDetails` is on)
- **Institution details:** $0.005 per unique institution (only when `includeCompanyDetails` is on)

Example: 100 jobs without details cost $0.301 ($0.001 + $0.30). With details, 100 jobs cost $0.80 ($0.001 + $0.30 + $0.50). You are only charged for jobs that were written to your dataset.

### Input

```json
{
  "searchQueries": ["assistant professor computer science"],
  "sortBy": "newest",
  "maxResultsPerQuery": 20,
  "includeJobDetails": true,
  "location": "California",
  "category": "faculty",
  "institutionType": "four_year"
}
```

| Field | Notes |
|-------|-------|
| `searchQueries` | Keywords. Each runs as its own search. Defaults to `professor` when empty. |
| `searchUrls` | HigherEdJobs search result URLs. Pagination is added for you. |
| `directUrls` | Single `details.cfm?JobCode=` job URLs. |
| `companyUrls` | Institution pages (`InstitutionProfile.cfm?ProfileID=N`). Returns that institution's jobs plus `company_about`, `company_social_urls` and `company_profile_url`. |
| `startUrls` | Alias. Any mix of the three URL kinds, routed by URL shape. |
| `postedWithinDays` | Keep jobs posted in the last N days. The site has no date filter, so it is applied to the posted date and, with `sortBy: newest`, paging stops early. |
| `positionType` | `any`, `full_time`, `adjunct_part_time` (site filter). |
| `remoteOnly` | Only Online/Remote positions (site filter). |
| `titleOnly` | Keyword must match the job title (site filter). |
| `maxResultsPerQuery` | Jobs per keyword or URL. Default 20. |
| `maxResults` | Total cap. Empty or 0 means per-query cap times the number of queries. |
| `sortBy` | `newest` (default), `priority`, `institution`, `location`, `title`, `category`. |
| `includeJobDetails` | Fetch each detail page. Default on. |
| `includeCompanyDetails` | Fetch each institution profile once and add `company_website`, `company_about`, `company_careers_url`, `company_headquarters`, `company_social_urls`. Default off. |
| `location` | State name or code, or `City, ST` (for example `Boston, MA`). |
| `category` | `any`, `faculty`, `administrative`, `executive`, `staff`. |
| `institutionType` | `any`, `four_year`, `two_year`, `k12`, `other`. |
| `maxJobsPerCompany` | Cap per institution. Default 25. |
| `maxConcurrency` | Detail pages fetched in parallel. Default 10, maximum 20. |

Institution details are read from the most recent archived copy of the institution profile page (the live profile pages sit behind a bot check that cannot be passed without a full browser), so mission text and links can be a few months old. Job rows are never delayed by this: company enrichment is time-boxed, jobs are always delivered, and an institution is only billed when its profile was read.

### Sample output

```json
{
  "job_id": "179572646",
  "title": "Lab Technician (Temp) - Genetics, Molecular and Cellular Biology",
  "institution_name": "Kennesaw State University",
  "location": "Kennesaw, GA",
  "category": "Laboratory and Research",
  "position_type": "Full-Time",
  "posted_at": "Posted Today",
  "posted_at_datetime": "2026-09-29T19:47:25Z",
  "employer_job_id": "304188",
  "company_website": "https://www.kennesaw.edu/",
  "military_friendly": true,
  "job_url": "https://www.higheredjobs.com/search/details.cfm?JobCode=179572646"
}
```

### Use with AI agents (MCP)

Connect this actor to any MCP-compatible AI client: Claude Desktop, Claude.ai, Cursor, VS Code, LangChain, LlamaIndex, or custom agents.

Configure the MCP server with this actor preconfigured at `mcp.apify.com?tools=santamaria-automations/higheredjobs-scraper`.

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=santamaria-automations/higheredjobs-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}
```

Example prompt: "Use higheredjobs-scraper to find the 30 newest assistant professor jobs in biology with salary and apply URL."

See the [auto-generated API tab](https://console.apify.com/actors/3SrNxGGn5JdY5bPyf/info/api) for language-specific examples (cURL, JS, Python, .NET, Ruby, PHP).

### Tips

- Raise `maxResultsPerQuery` for bigger pulls. The scraper pages until the cap is reached or results run out.
- Postings often say "commensurate with experience" and list no salary, so salary fields are empty for many jobs.
- Turn `includeJobDetails` off for a faster, cheaper listing-only pull.
- Use `company_website` with the email scraper and contact extractor below to find hiring contacts.

### Related actors

- [Website Job Extractor](https://apify.com/santamaria-automations/website-job-extractor): jobs from any company career page.
- [Website Email Scraper](https://apify.com/santamaria-automations/website-email-scraper): find emails on the institution websites in `company_website`.
- [Website Contact Extractor](https://apify.com/santamaria-automations/website-contact-extractor): names, roles and phones from those sites.
- [Indeed Scraper](https://apify.com/santamaria-automations/indeed-http-scraper): Indeed listings across many countries.
- [Glassdoor Scraper](https://apify.com/santamaria-automations/glassdoor-scraper): jobs with salary estimates and company ratings.
- [USAJobs Federal Scraper](https://apify.com/santamaria-automations/usajobs-federal-scraper): US federal government jobs.

### Support

- Questions or problems? Email contact@nanoscrape.com, we usually reply within 6 hours. You can also open a ticket in the [Issues tab](https://console.apify.com/actors/3SrNxGGn5JdY5bPyf/info/issues).
- Need a field that is missing? Open a ticket with the field name and a HigherEdJobs URL that shows it.
- Found a bug? Include the run URL so we can reproduce it.

# Actor input Schema

## `searchQueries` (type: `array`):

One or more job search terms, for example 'professor', 'assistant professor computer science', 'director of admissions'. Each query runs as its own search. Defaults to 'professor' when empty.

## `searchUrls` (type: `array`):

HigherEdJobs search result URLs (advanced\_action.cfm). Pagination is added automatically.

## `directUrls` (type: `array`):

Single job URLs (details.cfm?JobCode=...). Each is fetched directly.

## `companyUrls` (type: `array`):

Institution pages (InstitutionProfile.cfm?ProfileID=N). Returns all jobs of that institution plus its about text and social links.

## `startUrls` (type: `array`):

Backwards-compatible alias. Any mix of the URLs above, routed by shape to searchUrls, directUrls or companyUrls.

## `maxResultsPerQuery` (type: `integer`):

Maximum jobs to return per keyword or start URL.

## `maxResults` (type: `integer`):

Hard cap on total jobs across all keywords and URLs. Leave empty or 0 to use max results per query times the number of queries.

## `sortBy` (type: `string`):

Result order, using the site's own sort options. Newest posted first is the default.

## `postedWithinDays` (type: `integer`):

Keep only jobs posted in the last N days. The site has no date filter, so this is applied to the posted date (day precision). With newest-first sorting the scraper stops paging once a page has nothing in range.

## `includeJobDetails` (type: `boolean`):

Fetch each job's detail page for the full description, exact posting time, apply URL, employer website, requisition id, department, rank, credentials and more. Billed as a separate 'job-detail-result' event for each job where a description was parsed.

## `includeCompanyDetails` (type: `boolean`):

Fetch each institution's HigherEdJobs profile page once and add company\_website, company\_about, company\_careers\_url, company\_headquarters and company\_social\_urls to every job from that institution. Billed as a 'company-detail-result' event ($0.005) once per unique institution.

## `location` (type: `string`):

State name or code, or 'City, ST', for example 'California', 'TX' or 'Boston, MA'. Leave empty for nationwide.

## `category` (type: `string`):

Broad category filter. 'any' = no filter. Staff maps to the site's Admin group.

## `institutionType` (type: `string`):

Institution type filter. 'any' = no filter. K-12 and Other both map to the site's 'Outside Higher Education'.

## `positionType` (type: `string`):

Server-side filter. Full-time or Adjunct / Part-time.

## `remoteOnly` (type: `boolean`):

Only Online/Remote positions (site filter).

## `titleOnly` (type: `boolean`):

Match the keyword in the job title only (site filter), for higher precision.

## `maxJobsPerCompany` (type: `integer`):

Keep at most N jobs per institution. Useful when one large university dominates the results.

## `maxConcurrency` (type: `integer`):

Number of detail pages fetched in parallel. Each request uses its own proxy session, so a higher value is a safe speedup. Default 10, maximum 20.

## Actor input object example

```json
{
  "searchQueries": [
    "professor"
  ],
  "maxResultsPerQuery": 20,
  "sortBy": "newest",
  "includeJobDetails": true,
  "includeCompanyDetails": false,
  "location": "",
  "category": "any",
  "institutionType": "any",
  "positionType": "any",
  "remoteOnly": false,
  "titleOnly": false,
  "maxJobsPerCompany": 25,
  "maxConcurrency": 10
}
```

# Actor output Schema

## `jobListings` (type: `string`):

Dataset containing all scraped HigherEdJobs listings. Each record represents one job posting deduplicated across queries by job\_id.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQueries": [
        "professor"
    ],
    "sortBy": "newest"
};

// Run the Actor and wait for it to finish
const run = await client.actor("santamaria-automations/higheredjobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQueries": ["professor"],
    "sortBy": "newest",
}

# Run the Actor and wait for it to finish
run = client.actor("santamaria-automations/higheredjobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQueries": [
    "professor"
  ],
  "sortBy": "newest"
}' |
apify call santamaria-automations/higheredjobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,santamaria-automations/higheredjobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3SrNxGGn5JdY5bPyf/builds/OmwK0XQ0SMIUqhwO6/openapi.json
