# Glassdoor Jobs Scraper (`mlg14/glassdoor-jobs-scraper`) Actor

Scrape public Glassdoor jobs by keyword and location. Export titles, employers, ratings, pay, posting age, full descriptions, and job links with useful search filters.

- **URL**: https://apify.com/mlg14/glassdoor-jobs-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** Jobs, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Glassdoor Jobs Scraper

Scrape Glassdoor jobs by keyword and location, then export public job listings to JSON, CSV, or Excel. This Glassdoor jobs scraper provides a search data API alternative for collecting titles, employers, locations, pay information, ratings, and full job descriptions in one dataset.

The scraper accepts a familiar job search: enter a title or skill, choose a location, and set a result limit. Optional filters narrow the results by posting age, remote work, Easy Apply, employer rating, distance, employer size, and sort order. Each result has a stable job ID so recurring runs can skip listings already collected.

### What data can you extract from Glassdoor?

Each dataset item is a flat record. Fields that are absent from a particular public listing are returned as `null`. The same column names appear in JSON, CSV, and spreadsheet exports.

| Field | Description | Example |
| --- | --- | --- |
| `id` | Public job listing ID, stored as a string | `1010191586777` |
| `title` | Job title shown in search results | `Web Analytics Analyst` |
| `url` | Direct job link using the listing ID | `https://www.glassdoor.com/job-listing/j?jl=1010191586777` |
| `seoUrl` | Descriptive job link supplied by the site | A job listing URL |
| `ageInDays` | Listing age reported in search results | `82` |
| `discoveredAt` | Date returned in the job detail record | `2026-09-06T00:00:00` |
| `rating` | Overall employer rating when present | `3.7` |
| `easyApply` | Whether the listing offers Easy Apply | `false` |
| `isSponsored` | Whether the listing is marked sponsored | `false` |
| `employerId` | Public employer ID | `2937` |
| `employerName` | Employer display name | A public employer name |
| `employerUrl` | Employer overview link | An overview URL |
| `employerLogoUrl` | Public employer logo link | An image URL |
| `employerSize` | Employer size range, when listed | `10000+ Employees` |
| `employerWebsite` | Employer website, when listed | A website URL |
| `industry` | Employer industry, when listed | `Other Retail Stores` |
| `locationId` | Public location ID of the job | `1132348` |
| `locationName` | Displayed job location | `New York, NY` |
| `locationType` | Location type code | `C` |
| `countryId` | Public country ID | `1` |
| `payCurrency` | Currency code for pay values | `USD` |
| `payPeriod` | Pay period, such as annual or hourly | `ANNUAL` |
| `payMin` | Lower pay value from the listing data | `68200` |
| `payMedian` | Middle pay value from the listing data | `78100` |
| `payMax` | Upper pay value from the listing data | `88000` |
| `salarySource` | Whether pay is estimated or employer supplied | `EMPLOYER_PROVIDED` |
| `jobType` | Job type, when listed | `Full-time` |
| `remoteWorkType` | Remote arrangement, when listed | `null` |
| `skills` | Public skill labels attached to the job | An array of skill labels |
| `description` | Full job description in HTML | `<p>Role details...</p>` |
| `descriptionText` | Plain text version or listing excerpt | `Role details...` |
| `searchKeywords` | Keywords used for this search | `data analyst` |
| `searchLocation` | Resolved location used for this search | `New York, NY` |

The pay values retain the site's own currency, period, and source label. They are not converted across currencies or periods. A low, middle, and high value may represent an estimated distribution rather than a guaranteed offer; inspect `salarySource` before comparing compensation. `description` preserves the listing's HTML structure, while `descriptionText` is easier to read in a table or search index.

### How to scrape Glassdoor jobs

1. Enter a job title, skill, or phrase in `keywords` and a city, state, or country in `location`.
2. Set `limit` to the maximum number of unique jobs you want. The default is 100 and the input maximum is 1,000.
3. Add filters if you need a narrower result set. For example, use `daysOld` and `sortBy` to focus on recent postings, or select Easy Apply and remote work.
4. Run the actor. It resolves the location, collects search pages, and fetches each selected job's public details when `includeDescription` is enabled.
5. Open the dataset to download JSON, CSV, or Excel. Keep the `id` values if you plan to exclude previously collected jobs on later runs.

The scraper saves one item per unique job ID within a run. Search results are returned in the site's selected order. A sponsored listing can appear among ordinary results; use `isSponsored` if you need to distinguish it.

### Input

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `keywords` | string | Required | Title, skill, or search phrase. |
| `location` | string | Required | City, state, or country. The first matching public location is selected if there is no exact label match. |
| `daysOld` | integer | Any age | Maximum reported posting age in days. |
| `easyApply` | boolean | `false` | Restrict results to Easy Apply listings. |
| `remoteWorkType` | boolean | `false` | Apply the remote work filter. |
| `minRating` | number | No minimum | Minimum employer rating from 0 to 5. |
| `radius` | string | `"25"` | Distance in miles: 0, 5, 10, 15, 25, 50, or 100. |
| `employerSizes` | string | Any size | Site size category from 1 to 5. |
| `sortBy` | string | `relevant_desc` | Most relevant or `date_desc` for most recent. |
| `limit` | integer | `100` | Maximum unique results; allowed range is 1 to 1,000. |
| `includeDescription` | boolean | `true` | Fetch full descriptions and additional detail fields. |
| `excludeJobIds` | array | Empty | Public job IDs to omit, useful for recurring collection. |
| `urlParam` | key/value array | Empty | Additional supported search filters, such as an industry ID. |
| `proxyConfiguration` | object | Platform proxy | Optional proxy configuration. |

For example, this input collects up to 60 recent analyst jobs in the selected area:

```json
{
  "keywords": "data analyst",
  "location": "New York",
  "daysOld": 30,
  "sortBy": "date_desc",
  "radius": "25",
  "limit": 60,
  "includeDescription": true
}
```

Use a more specific location label when a name exists in multiple places. A city and state is usually clearer than a city name alone. If the lookup has no public match, the run reports an error instead of returning results for an unrelated location.

### Output example

This is a selected set of fields from a real 35-job run using `data analyst` and `New York`. The full record also contains the employer and description fields listed above; the long HTML description is omitted here for readability.

```json
{
  "id": "1010191586777",
  "title": "Web Analytics Analyst",
  "url": "https://www.glassdoor.com/job-listing/j?jl=1010191586777",
  "ageInDays": 82,
  "discoveredAt": "2026-09-06T00:00:00",
  "rating": 3.7,
  "easyApply": false,
  "isSponsored": false,
  "locationId": 1132348,
  "locationName": "New York, NY",
  "payCurrency": "USD",
  "payPeriod": "ANNUAL",
  "payMin": 68200,
  "payMedian": 78100,
  "payMax": 88000,
  "salarySource": "EMPLOYER_PROVIDED",
  "jobType": "Full-time",
  "searchKeywords": "data analyst",
  "searchLocation": "New York, NY"
}
```

In that run, all 35 jobs had an ID, title, URL, employer name, location name, full description, and plain text description. The result count is a measurement of that search at that time, not a promise that another search will return the same number of jobs.

### Use cases

- **Job market tracking:** Save periodic datasets for a role and location, then compare new IDs across runs to see newly surfaced listings.
- **Salary research:** Compare pay values for a single role and location, separating estimated values from employer supplied values with `salarySource`.
- **Hiring activity analysis:** Count listings by employer, location, job type, and search date. Keep sponsored listings identifiable with `isSponsored`.
- **Recruiting research:** Find public roles that match a skill, posting age, and location. Use the direct job URL to inspect the current listing.
- **Career search:** Build a focused list of openings, with descriptions and job links available in one export.
- **Regional comparisons:** Run the same keywords in several locations and compare the returned jobs while preserving `searchLocation` on each row.

For repeated searches, save the `id` column from the last dataset and pass those IDs through `excludeJobIds`. The exclusion applies to output. It does not cause the site to return replacement jobs beyond the search pages that would otherwise be reachable, so set a sufficient `limit` and use narrower searches when you need a large set of new results.

### How much does it cost to scrape Glassdoor jobs?

The current event price is **$1.00 per 1,000 saved jobs**, or **$0.001 per result**. Platform usage is included in that result price. The actor charges for records it saves, subject to the run's spending limit.

| Saved jobs | Result charge |
| ---: | ---: |
| 35 | $0.035 |
| 100 | $0.10 |
| 1,000 | $1.00 |

These examples are based on the number of dataset items, not the number of search pages or detail requests. If a search only has 48 available jobs and `limit` is 100, the dataset contains at most 48 jobs and the result charge is at most $0.048. If the run budget stops output early, the dataset can contain fewer items than `limit`.

### Tips for best results

Search with a specific role and a clear location. Broad terms can return many loosely related jobs; combining a precise keyword with `daysOld`, `radius`, and `sortBy` can make the dataset easier to review. The remote work and Easy Apply switches can reduce the available count substantially, especially when used together with a short posting window or a high minimum rating.

Start with a moderate `limit` to inspect relevance and field coverage. Then increase it for routine collection. The search service returns jobs in pages of about 30. The scraper follows the cursor supplied for the next page, deduplicates job IDs, and stops at your result limit or when the service no longer supplies a cursor. A page 6 request returned 30 jobs during validation, so collection is not confined to the first page.

Leave `includeDescription` enabled when you need full text, employer size, website, industry, job type, or skill labels. Turn it off for a faster listing-only run. With it off, `description` and some detail-only fields are empty; `descriptionText` may still contain a short excerpt from the search result.

Review `searchLocation` after a run. A short location name can match more than one place. The scraper prefers an exact label match and otherwise takes the first suggestion. For recurring jobs in the same area, use a distinct city and region label to reduce ambiguity.

### Limits

Only public job data is collected. A listing can disappear, expire, or change between the search request and its detail request. If a detail request fails, the scraper still saves the listing with the fields available from search; the full description and detail-only fields may then be empty. Some listings never provide pay, a rating, a logo, a job type, or a remote arrangement. Empty values are represented by `null`.

The site's `ageInDays` and detail record date can differ. Treat `ageInDays` as the search result's reported age and `discoveredAt` as the date exposed in the detail record. Neither field is recalculated by the scraper. When exact publication timing matters, check the current listing directly.

The result limit is an output cap, not a guarantee of that many matches. Filters can leave fewer results, and pagination depends on cursors returned by the public search service. The accepted input maximum is 1,000 jobs. Searches with more matches may require narrower keywords, locations, or filters across separate runs; each run deduplicates only within its own search.

The standard job search page presented a challenge during source checks. Public location, search, and detail endpoints served the validated run. Site behavior can change, including access controls, field availability, ordering, and pagination depth. The actor does not require an account or collect private profile data.

### Use with automated workflows

The input and output are ordinary structured JSON, so a scheduled workflow can run the same search repeatedly and compare job IDs. Two example requests are: “Collect 100 recent analyst jobs in New York and export their pay fields,” and “Find remote designer jobs posted within seven days, omitting IDs saved last week.” A workflow can then read the default dataset, filter rows by the fields it needs, and send the data to a database or spreadsheet.

Keep `description` when your downstream process needs the original formatting. Use `descriptionText` when it needs plain text. Preserve `salarySource`, `payCurrency`, and `payPeriod` beside the pay values so a downstream comparison does not mix currencies, time periods, or estimated and employer supplied data.

### FAQ

#### Is scraping public Glassdoor jobs legal?

Rules vary by location and use. Work with public listings, review the site's terms, respect applicable privacy and data protection law, and use the dataset for a legitimate purpose. This actor does not sign into an account or collect private candidate information.

#### Do I need to configure a proxy?

A proxy is configured by default. Most users can leave `proxyConfiguration` untouched. If a run encounters an access block, the fetch layer can rotate its session and try the available proxy routes. Access is still dependent on the site at run time.

#### How fast is a run?

Speed depends on the number of jobs, the number of detail requests, and site response times. The validated 35-job run completed in roughly 72 seconds on the remote platform. Listing-only runs can require fewer requests; larger searches with full descriptions take longer.

#### Can I schedule and monitor searches?

Yes. Schedule the actor with saved input, then inspect each run's status and dataset. Keeping the same keywords and location makes comparisons easier. Pass prior IDs in `excludeJobIds` if you only want to save jobs that were not in your earlier dataset.

#### Can I export to a spreadsheet?

Yes. The default dataset supports JSON, CSV, and Excel downloads. Flat columns make it straightforward to sort by pay, location, rating, date, or employer. Long descriptions may need wider cells or a separate text field in spreadsheet views.

#### Why is a field empty?

The source may not provide it for that listing, or the detail request may have been unavailable. Pay, remote arrangement, industry, and employer details are particularly variable. If `includeDescription` is disabled, full descriptions and other detail-only fields are intentionally empty.

#### Why did I receive fewer jobs than `limit`?

`limit` is a maximum. Narrow filters, exclusions, duplicate IDs, or the end of available pagination can all reduce the output count. Broaden a filter or search another location if you need more matching jobs.

### Integrations

Use the actor's run endpoint to start a search and the default dataset endpoint to retrieve results. Scheduling, webhooks, and ordinary data exports support recurring collection and delivery to databases, spreadsheets, or internal reporting systems. Stable job IDs let a downstream process merge successive datasets without treating the same listing as a new job each time.

### Support

Open an issue on the Issues tab; we reply within 24h and add fields on request.

# Actor input Schema

## `keywords` (type: `string`):

Job title, skill, or search phrase.

## `location` (type: `string`):

City, state, or country to search. The first matching public location is used.

## `daysOld` (type: `integer`):

Only jobs posted within this many days. Leave empty for any age.

## `easyApply` (type: `boolean`):

Return jobs marked Easy Apply.

## `remoteWorkType` (type: `boolean`):

Apply the remote work filter.

## `minRating` (type: `number`):

Filter by employer rating from 0 to 5.

## `radius` (type: `string`):

Distance in miles from the selected location.

## `employerSizes` (type: `string`):

Filter by the site employer size category.

## `sortBy` (type: `string`):

Sort by relevance or posting date.

## `limit` (type: `integer`):

Maximum number of unique jobs to save.

## `includeDescription` (type: `boolean`):

Fetch each job detail to include the full description and extra employer and job fields.

## `excludeJobIds` (type: `array`):

Job IDs already collected in earlier runs.

## `urlParam` (type: `array`):

Optional supported search filter key and value pairs, such as industryNId. Values override matching standard filters.

## `proxyConfiguration` (type: `object`):

Use the configured proxy for public requests.

## Actor input object example

```json
{
  "keywords": "Data Analyst",
  "location": "New York",
  "easyApply": false,
  "remoteWorkType": false,
  "radius": "25",
  "employerSizes": "",
  "sortBy": "relevant_desc",
  "limit": 100,
  "includeDescription": true,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `jobs` (type: `string`):

All collected jobs in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": "Data Analyst",
    "location": "New York"
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/glassdoor-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": "Data Analyst",
    "location": "New York",
}

# Run the Actor and wait for it to finish
run = client.actor("mlg14/glassdoor-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": "Data Analyst",
  "location": "New York"
}' |
apify call mlg14/glassdoor-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/glassdoor-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/iPI5pODtNjRUmUfDS/builds/KR8gUbMU58Zu24Yuk/openapi.json
