Glassdoor Jobs Scraper avatar

Glassdoor Jobs Scraper

Pricing

from $0.80 / 1,000 job results

Go to Apify Store
Glassdoor Jobs Scraper

Glassdoor Jobs Scraper

Scrape job listings from Glassdoor by search query and location, with optional job detail enrichment.

Pricing

from $0.80 / 1,000 job results

Rating

0.0

(0)

Developer

Farhan Ali

Farhan Ali

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

16 hours ago

Last modified

Share

Glassdoor Jobs Scraper creates a structured dataset of job listings collected from Glassdoor, the employer-review and job-search platform. Each dataset item represents one job listing and can include job title, company, rating, location, posting age, apply link, salary estimate, and a description snippet, with optional detail-page enrichment for the full description, employer profile, skills, and geographic fields. Query the source by search keywords and location, control the result limit with maxItems, and retrieve records through the Apify Dataset API or export them as JSON, CSV, Excel, or another supported format.

Dataset at a glance

PropertyValue
Sourceglassdoor.com (job search; 10 country editions)
Record unitOne job listing
Input methodsKeyword search (searchQueries) + location, or search URLs (startUrls)
Main identifierslistingId, jobUrl
DeliveryApify Dataset and API
Export formatsJSON, CSV, Excel, XML, HTML (Apify dataset exports)
Update modelFresh records per Actor run
Pricing$1.00 per 1,000 jobs; +$2.00 per 1,000 for job details

Coverage and available records

The Actor collects jobs from public Glassdoor search results using one of two entry points:

  • Search-based: Set searchQueries to job keywords (for example python or software engineer) and location to a country, state, or city. Leave searchQueries empty to list all jobs for the location.
  • URL-based: Pass Glassdoor job-search URLs in startUrls to scrape those pages in addition to the search terms.
  • Country editions: baseUrl selects one of ten Glassdoor editions (US, DE, UK, FR, NL, CA, AU, AT, IE, IN).

Record types and limits:

  • Listing records are always collected: job title, company, rating, location, posting age, apply link, and description snippet.
  • Detail fields are conditional: the full job description, employer profile, skills, and geographic coordinates are returned only when scrapeJobDetails is enabled, which fetches each job's detail data.
  • Result cap: maxItems limits the number of jobs collected (0 means unlimited, the default). Collection stops after 500 result pages per search.

Known exclusions: the Actor collects jobs, not standalone company profiles (employer data is a byproduct of detail enrichment); content Glassdoor only shows behind login is not collected; each run captures page state at run time (no historical snapshots).

Data dictionary

Field names below match dataset record JSON properties exactly. Fields marked conditional appear only when scrapeJobDetails is enabled.

FieldTypeNullableDescriptionExample
listingIdintegerNoGlassdoor job identifier; best stable deduplication key1010070316579
jobTitlestringNoJob title as listedJUNIOR GENAI ENGINEER (M/W/D)
normalizedJobTitlestringYesNormalized job titleingenieur (m/w/d)
companyNamestringNoCompany nameReply
companyIdintegerYesGlassdoor company identifier306932
companyRatingnumberYesAverage company rating (0–5)3.8
companyLogoUrlstringYesCompany logo image URLhttps://media.glassdoor.com/sql/306932/reply-squareLogo-1680691862007.png
locationNamestringYesJob location as listedDeutschland
locationIdintegerYesGlassdoor location identifier96
ageInDaysintegerYesPosting age in days134
easyApplybooleanNoWhether the listing supports one-click applyfalse
isSponsoredbooleanNoWhether the listing is sponsoredfalse
expiredbooleanNoWhether the listing has expiredfalse
urlstringNoJob listing URLhttps://www.glassdoor.de/job-listing/...?jl=1010070316579
jobUrlstringNoSame as urlSame as url
applyUrlstringYesExternal apply linkhttps://www.glassdoor.de/partner/jobListing.htm?...
descriptionSnippetstringYesShort description snippetAls Junior GenAI Engineer...
jobAttributesstring[]YesAttribute tags["Azure", "Cloud-Architektur", ...]
sourceQuerystringYesSearch keyword that produced the recordpython
searchQuerystringYesSearch keyword usedpython
searchLocationstringYesLocation used in the searchDeutschland
salaryCurrencystringYesSalary currency codeEUR
salaryPeriodstringYesSalary period (ANNUAL, etc.)ANNUAL
salaryMediannumberYesMedian salary estimate65000
salaryP10numberYes10th-percentile salary estimate52000
salaryP90numberYes90th-percentile salary estimate81000

Detail-enriched fields (conditional — scrapeJobDetails):

FieldTypeNullableDescriptionExample
jobDescriptionstringYesFull job description (HTML)Long-form HTML
jobDescriptionTextstringYesFull job description (plain text)Long-form text
employerIndustrystringYesEmployer industryIT Services
employerSizestringYesEmployer size bucket1000–5000
employerWebsitestringYesEmployer website URLhttps://www.reply.com
employerHeadquartersstringYesEmployer headquartersTurin, Italy
latitudenumberYesJob latitude52.52
longitudenumberYesJob longitude13.40
citystringYesJob cityBerlin
statestringYesJob state/regionBerlin
countrystringYesJob countryDE
skillsstring[]YesSkill tags["Python", "Azure", ...]
educationstring[]YesEducation requirements["Bachelor's Degree"]
yearsOfExperiencestring[]YesExperience requirements["3+ years"]
jobTypestring[]YesEmployment types["Full-time"]
jobDetailsobjectYesRaw detail payload, with nested employerOverview, companyRatings, employerLinks, map, and other objectsSee example record

Salary fields (salaryMedian, salaryP10, salaryP90) are estimates and are 0 or null when Glassdoor has no pay estimate for the listing.

Example dataset record

Real record produced with the input below (searchQueries: ["python"], location: "Deutschland", maxItems: 3, scrapeJobDetails: true).

{
"listingId": 1010070316579,
"jobTitle": "JUNIOR GENAI ENGINEER (M/W/D)",
"normalizedJobTitle": "ingenieur (m/w/d)",
"companyName": "Reply",
"companyId": 306932,
"companyRating": 3.8,
"companyLogoUrl": "https://media.glassdoor.com/sql/306932/reply-squareLogo-1680691862007.png",
"locationName": "Deutschland",
"locationId": 96,
"ageInDays": 134,
"easyApply": false,
"isSponsored": false,
"expired": false,
"url": "https://www.glassdoor.de/job-listing/junior-genai-engineer-m-w-d-reply-JV_IC5424598_KO0,23_KE306932.htm?jl=1010070316579",
"jobUrl": "https://www.glassdoor.de/job-listing/junior-genai-engineer-m-w-d-reply-JV_IC5424598_KO0,23_KE306932.htm?jl=1010070316579",
"applyUrl": "https://www.glassdoor.de/partner/jobListing.htm?jobListingId=1010070316579",
"descriptionSnippet": "Als Junior GenAI Engineer ...",
"jobAttributes": ["Azure", "Cloud-Architektur", "R"],
"sourceQuery": "python",
"searchQuery": "python",
"searchLocation": "Deutschland",
"salaryCurrency": "EUR",
"salaryPeriod": "ANNUAL",
"salaryMedian": 65000,
"salaryP10": 52000,
"salaryP90": 81000,
"jobDescriptionText": "Als Junior GenAI Engineer ...",
"employerIndustry": "IT Services",
"employerSize": "1000–5000",
"employerWebsite": "https://www.reply.com",
"employerHeadquarters": "Turin, Italy",
"latitude": 52.52,
"longitude": 13.4,
"city": "Berlin",
"state": "Berlin",
"country": "DE",
"skills": ["Python", "Azure", "Machine Learning"]
}

The record above was produced with this input:

{
"searchQueries": ["python"],
"location": "Deutschland",
"maxItems": 3,
"scrapeJobDetails": true,
"proxyConfiguration": {
"useApifyProxy": true,
"apifyProxyGroups": ["RESIDENTIAL"]
}
}

Query and input reference

InputTypeRequiredDefaultAccepted valuesDescription
searchQueriesarray (stringList)No[]Free-text job keywordsJob search keywords; empty lists all jobs for the location
locationstringNoDeutschlandFree-text country, state, or cityJob location as shown on Glassdoor
baseUrlstringNohttps://www.glassdoor.deglassdoor.com/.de/.co.uk/.fr/.nl/.ca/.com.au/.at/.ie/.co.inGlassdoor country edition
startUrlsarray (requestListSources)No[]Glassdoor job-search URLsExtra search pages to scrape in addition to searchQueries
scrapeJobDetailsbooleanNotruetrue / falseFetch each job's detail data (charged as job details)
detailConcurrencyintegerNo8150Parallel detail requests
maxItemsintegerNo00 or any positive integerMaximum jobs to collect; 0 = unlimited
proxyConfigurationobjectNoApify proxy, RESIDENTIAL groupApify proxy groups or custom proxiesResidential proxies are recommended

Minimal request:

{ "searchQueries": ["software engineer"], "location": "Berlin" }

Advanced request with detail enrichment:

{
"searchQueries": ["data engineer"],
"location": "Deutschland",
"baseUrl": "https://www.glassdoor.de",
"scrapeJobDetails": true,
"detailConcurrency": 10,
"maxItems": 500,
"proxyConfiguration": { "useApifyProxy": true, "apifyProxyGroups": ["RESIDENTIAL"] }
}

Retrieve the data through the API

The Actor runs on the Apify platform, so there is no server to host and no crawling infrastructure to maintain.

  1. Start the Actor with a JSON input (console or API).
  2. Wait for the run to finish, or use a synchronous endpoint if you want the response inline.
  3. Retrieve items from the run's default dataset.
  4. Paginate or export the dataset.

Python example:

from apify_client import ApifyClient
client = ApifyClient("YOUR-APIFY-TOKEN")
run_input = {
"searchQueries": ["software engineer"],
"location": "Berlin",
"maxItems": 20,
}
run = client.actor("datascrapers/glassdoor-jobs-scraper").call(run_input=run_input)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item["listingId"], item["jobTitle"], item["companyName"])

Apify generates ready-to-run Python, JavaScript, and cURL examples on the Actor's API tab. Do not put a real API token in shared code or URLs.

Data quality and record handling

  • Conditional fields: detail fields are present only when scrapeJobDetails is enabled. Listings alone return a leaner record.
  • Salary fields: salaryMedian, salaryP10, and salaryP90 are estimates and are 0 or null when Glassdoor has no pay estimate for the listing.
  • Source changes: Glassdoor page structure and values can change; unreadable fields are returned as null rather than fabricated.
  • Deduplication: the Actor deduplicates listings within a run by listingId. Use listingId as the stable key and filter repeated runs against previously stored IDs.
  • Detail fallback: detail fetches are best-effort; listing data is always returned even if detail enrichment fails.

Export and pipeline examples

DestinationRecommended methodTypical use
PostgreSQL / SupabaseDataset API poll or webhook consumerStore job listings alongside talent tables
Google SheetsApify Google Sheets integrationShare job shortlists with recruiting teams
CRM / ATS pipelineWebhook on run completionPush new listings into applicant tracking
S3 / cloud storageScheduled export via Apify scheduler + integrationArchival of labor-market snapshots

Pricing and cost examples

The Actor uses pay-per-event pricing with two chargeable events:

EventTriggerRate (per 1,000 jobs)
Job resultEvery job record pushed to the dataset$1.00
Job detailsscrapeJobDetails enabled, detail data fetched$2.00

Example costs:

RecordsConfigurationEstimated base cost
1,000Listing only$1.00
10,000Listing only$10.00
1,000Listing + details$3.00
10,000Listing + details$30.00

Apify paid plans reduce the per-1,000 rate (for example $0.80 per 1,000 jobs at the Gold tier). Compute units consumed by the run are billed by your Apify plan. Estimates depend on the verified pricing model and the options selected for the run.

Limitations and responsible data use

  • The Actor collects publicly accessible data from Glassdoor job-search pages only, across ten supported country editions.
  • The Actor collects jobs, not standalone company profiles; employer data is a byproduct of detail enrichment.
  • Salary fields are estimates and are frequently absent when Glassdoor has no pay estimate.
  • Collection stops after 500 result pages per search.
  • Field availability depends on what Glassdoor renders at run time; some values can be null or missing, and site changes can alter fields.
  • The Actor does not provide historical snapshots unless you store them yourself.
  • You are responsible for compliance with Glassdoor's terms of service, applicable privacy law, and any contractual obligations before using the data.

Dataset questions

What does one dataset item represent?

One Glassdoor job listing, optionally enriched with its full detail data.

Which field should I use as a unique identifier?

listingId is the stable Glassdoor job identifier and is the recommended deduplication key. The jobUrl is a reasonable secondary key.

Are fields nullable or conditional?

Yes. Detail fields (full description, employer profile, skills, coordinates) exist only when scrapeJobDetails is enabled. Salary fields are 0 or null when Glassdoor has no pay estimate. Fields the listing does not render are returned as null.

Can I retrieve the records as CSV or JSON?

Yes. The dataset can be exported as JSON, CSV, Excel, XML, or HTML from the Apify Console, and queried through the Dataset API.

Does the Actor return historical data?

No. Each run captures the state of the listings at run time. To track salary or posting changes, schedule repeated runs and store the outputs yourself.

What counts as a billable result?

Two pay-per-event charges apply: a job-result charge for every job record ($1.00 per 1,000) and a job-details charge for each enriched job when scrapeJobDetails is enabled ($2.00 per 1,000).

Data Scrapers support

Need an additional field, record type, or export workflow? Contact Data Scrapers at stardustspotlight@gmail.com. Include a sample source URL, required fields, expected record volume, and preferred delivery format.