Job Board Aggregator ๐Ÿš€ - Indeed, LinkedIn & Glassdoor Scraper avatar

Job Board Aggregator ๐Ÿš€ - Indeed, LinkedIn & Glassdoor Scraper

Pricing

from $2.50 / 1,000 results

Go to Apify Store
Job Board Aggregator ๐Ÿš€ - Indeed, LinkedIn & Glassdoor Scraper

Job Board Aggregator ๐Ÿš€ - Indeed, LinkedIn & Glassdoor Scraper

Job board aggregator and multi-board scraper: LinkedIn, Indeed, Glassdoor, The Muse + 8 keyless boards in one run. ATS job scraper built in - watch Greenhouse, Lever and Ashby company boards. New-job-alerts mode returns only new postings. Cross-board dedup: one role bills once. CSV, JSON.

Pricing

from $2.50 / 1,000 results

Rating

5.0

(3)

Developer

Flash Scrape

Flash Scrape

Maintained by Community

Actor stats

4

Bookmarked

23

Total users

9

Monthly active users

2 hours ago

Last modified

Share

Job Board Aggregator - Indeed, LinkedIn & Glassdoor Scraper

One search. Up to eight job boards. One deduplicated row per role โ€” so a job posted to LinkedIn, Indeed and Glassdoor bills you once, not three times.

LinkedIn, Indeed and Glassdoor plus nine keyless job APIs (The Muse, Remotive, Jobicy, Himalayas, Hacker News "Who is hiring?", Remote OK, We Work Remotely, Working Nomads, DevITjobs US/UK) merged into a single clean table: title, company, location, salary (min/max/interval/currency plus annualized columns), remote flag, job type, posting date, full description, employer rating and company details. Search by keyword + location, filter by remote / job type / recency / salary present / title keywords / company, and export to CSV, JSON, or Excel.

Watch specific companies' careers pages โ€” list company names in Watch companies and the actor reads each one's own ATS board directly (Greenhouse, Lever or Ashby โ€” all keyless, tried in that order) and merges the live openings into the same 50-column table. Fresher than any job board, straight from the source. Leave the search term untouched and you get each company's whole board; type a term and only matching roles are delivered and billed.

Only pay for NEW jobs on a schedule โ€” turn on Only new jobs (monitoring mode) and the actor remembers every job this search has already delivered to you (in a private store in your own account) and bills only postings it has not seen before. Measured before this existed: two identical runs a day apart shared 47% of their rows โ€” a schedule was re-buying half its data every run. Memory is kept per search for 90 days.

Set up a daily new-jobs feed in 60 seconds โ€” the recurring setup most buyers actually want:

  1. Fill in your search term and location, tick Only new jobs (monitoring mode), hit Save & start once.
  2. On the run page choose Actions โ†’ Schedule, pick daily (or hourly for hot markets).
  3. Done. Every run now delivers โ€” and bills โ€” only postings the previous runs have not seen. The first run seeds the memory; identical reruns bill zero (measured: 15 rows, then 0). Point the schedule's webhook at Slack, a Google Sheet via Zapier/Make, or your ATS, and it becomes a job alert service you own.

Remote jobs across eleven boards in one click โ€” turn on Remote jobs only and hit Save & start: LinkedIn, Indeed, Glassdoor and The Muse are asked for their own remote filters while the seven remote-only boards (Remotive, Jobicy, Himalayas, HN Who-is-Hiring, Remote OK, We Work Remotely, Working Nomads) join automatically. No other job-board actor on the Store combines the big boards' remote filters with the remote-only APIs in one run.

Built for recruiters & staffing agencies, job boards & aggregators, market researchers, and sales teams tracking hiring signals who want one consolidated jobs feed instead of running eight scrapers and reconciling them by hand.

What you get

  • Eight working boards in one run โ€” LinkedIn, Indeed and Glassdoor selected by default, plus five keyless APIs you switch on (four of them join remote searches automatically). One schema, one dataset.
  • Cross-board dedup that shows its work โ€” the same role on several boards becomes one row carrying found_on_sites and duplicate_count. Company names are normalized first, so Wipro and Wipro Limited collapse together. It merges across boards only: five genuinely different openings one employer posted to one board stay five rows, because you paid to scrape all five. You are billed for the surviving row only.
  • Salaries you can actually sort โ€” hourly, weekly and monthly pay is normalized into salary_min_annual / salary_max_annual on every salaried row, no flag to find.
  • Filters that never bill you โ€” salary-required, title / company / city excludes, competitor-only, experience level, agency excludes and max posting age all run before the charge, and RUN_SUMMARY.filter_removed counts each one.
  • Five ready-made Output views โ€” Overview, With salary, Cross-board dedup, Remote & location, Standard table. No column wrangling on your first run (exports still carry every column).
  • The same 50 columns on every row, every run โ€” a field a board did not publish is null, never a missing column. Spreadsheet imports line up, pd.read_csv behaves, positional parsers keep working.
  • Presets and quick-pickers โ€” pick a goal and the matching options switch themselves on, with every change (and every refusal) written to the log.
  • Paste a search URL โ€” already built the search on LinkedIn, Indeed or Glassdoor? Drop the URL in and skip the form.
  • Honest per-board reporting โ€” every run writes a RUN_SUMMARY naming each board's real outcome. Boards that are blocked at the source are labelled in the table below, not sold to you.

Measured, not promised

Every number below comes from a real run of this Actor. Inputs and dates are given so you can reproduce them.

RunResult
Untouched form, pressed Start (3 boards, 20/board) โ€” 2026-08-0860 rows in 27 s โ€” LinkedIn 20, Indeed 20, Glassdoor 20. When Glassdoor is having one of its intermittent blocked spells you get ~40 rows from the other two, the status message says so, and the missing rows are never billed
data analyst ยท New York, NY ยท 30/board, 3 boards โ€” 2026-08-0890 raw listings โ†’ 89 billed rows in 39 s; 1 duplicate merged across boards (Indeed 30, Glassdoor 30, LinkedIn 29)
Salary coverage on that same runsalary_min_annual filled on 74 of 89 rows (83%)
Stress run, 120/board, 3 boards โ€” 2026-08-07258 rows in 226 s โ€” LinkedIn 118, Indeed 112, Glassdoor 28
Remote-only run, the five API boards, 10 requested each โ€” 2026-08-08The Muse 10, Jobicy 10, HN 5, Himalayas 6, Remotive 2 โ€” per-board totals track each board's live inventory for your term and vary day to day
LinkedIn descriptions on a default run โ€” 2026-08-0720 of 20 LinkedIn rows (detail fetching is on by default)

Glassdoor tops out around ~28โ€“30 rows per query no matter the cap โ€” a board-side limit we report rather than hide. Every row of every run carries the same 50 columns; per-field fill rates are further down, under Output fields and how often they are filled.

First run

{
"searchTerm": "data analyst",
"location": "New York, NY"
}

That is a complete run. Everything else already has a working default: LinkedIn + Indeed + Glassdoor, 20 job postings per board, full LinkedIn job details on, Apify datacenter proxy (measured: same rows as residential at a fraction of the cost). Open the Actor, hit Save & start without touching anything, and you get rows.

Presets

Pick a goal in Preset and the matching options are switched on for you. A preset only fills options you left at their default โ€” anything you set yourself always wins โ€” and both the run log and RUN_SUMMARY.preset_applied list exactly what it changed and what it refused to change, with your conflicting value.

PresetWhat it switches onEffect on your bill
Salary researchrequireSalary, enforceAnnualSalary, sortBy: salary_descLower โ€” rows carrying no pay are dropped before billing
Competitor hiring intelincludeCompanyDetailsUnchanged โ€” it adds columns, not rows. It does not narrow the run by itself: add the employers you are watching under Only these companies, or the run stays unfiltered and the log says so
Remote-only sweepisRemoteHigher โ€” with the board list left untouched this also brings in the seven remote-only boards (Remotive, Jobicy, Himalayas, HN, Remote OK, We Work Remotely, Working Nomads), and extra boards add rows
Fresh postings onlymaxAgeDays: 7, plus a 168-hour board-side window when that cannot clash with your job type / easy-apply / remote choicesLower โ€” anything older is dropped before billing
Recruiter lead-genincludeCompanyDetails, excludeAgenciesLower โ€” staffing-agency reposts are dropped before billing

Two quick-pickers sit beside it: Role (quick pick) and City (quick pick). Precedence is searchTerms > roleSelect > searchTerm and locations > citySelect > location โ€” a picker overrides the matching free-text box (and warns you in the log when it does), while the multi-value arrays override the picker. Picking a non-US metro also switches countryIndeed to that country unless you set it yourself. Leave all three shortcuts alone and the Actor behaves exactly as it always has.

What the Output tab looks like

Five curated views ship with the Actor. Pick one above the results table instead of scrolling a wall of 50 raw columns:

ViewColumnsUse it for
Overviewtitle, company, location, board, posted, company page, job linkthe fast read of any run
With salarytitle, company, min/max yearly pay, currency, job linkcomp benchmarking โ€” pay is already annualized
Cross-board deduptitle, company, found on boards, listings for this role, job linkproof of the dedup, and a hiring-urgency signal
Remote & locationtitle, company, remote flag, location, board, job linktelling true remote from remote-ish
Standard tablethe 16 core columnsthe previous default view, field-for-field unchanged

Three things to know, so nothing surprises you. No view shows every column โ€” not even Standard table, which is the original 16-column view kept byte-identical because existing API callers address it as ?view=overview. Every run emits 50 columns. description, emails, id and job_url_direct are listed by no view at all; salary_min_annual / salary_max_annual, duplicate_count and company_url are missing from Standard table but do appear in With salary, Cross-board dedup and Quick view respectively. Views are Console display only โ€” CSV, JSON and Excel exports and the Dataset API always hand you every column regardless of the view you are looking at. And views are column presets, not filters: With salary still lists rows whose board published no pay (those cells are simply blank) and Cross-board dedup still lists roles seen on a single board (listings for this role = 1). To actually remove rows, use requireSalary or isRemote โ€” and rows removed by a filter are never billed.


Board status โ€” read this first

We would rather tell you than let you find out on a paid run. Checked 2026-08-07, re-measured 2026-08-11: all boards work from Apify's default datacenter pool (52 of 54 rows vs residential on the same search, faster and far cheaper). Residential remains available in Proxy configuration if a board starts coming back short:

BoardStatusNotes
LinkedInWorking (default)Full job details fetched by default โ€” descriptions on every LinkedIn row
IndeedWorking (default)Richest company data (size, revenue, addresses)
GlassdoorWorking (default), intermittently blockedSalary on ~100% of rows it returns, plus employer rating. Glassdoor 403s some runs entirely (upstream, comes back on its own within hours) โ€” when that happens the run status names it and the missing board bills nothing. Tops out around ~28-30 rows per query โ€” a board-side cap, not a bug
The MuseWorking (default since 2026-08-14)The only extra board with real city coverage (server-side location filtering โ€” measured 12/12 success from the standard proxy pool, 20 rows/page). Carries no salary data โ€” 0 salary fields in 220 measured rows
RemotiveWorking (remote-only)Auto-joins remote searches; rows may not be re-posted to other job boards (see Licensing below)
JobicyWorking (remote-only)Auto-joins remote searches; credit Jobicy + keep the original apply links
HimalayasWorking (remote-only)Auto-joins remote searches; link back to himalayas.app
HN "Who is hiring?"Working (remote-only)Startup jobs from the monthly Hacker News thread
Remote OKWorking (remote-only)100 newest remote listings, server-side tag filtering. Rows keep their remoteok.com link โ€” required attribution, see Licensing
We Work RemotelyWorking (remote-only)Category RSS feeds (25-100 rows); salary parsed from listing text where published (~56% of programming posts)
Working NomadsWorking (remote-only)Curated remote list (~50 rows); no salary data
DevITjobs USWorking (opt-in)US/Canada tech board, salary on ~100% of rows (annual, currency as published), experience level + tech stack on nearly every row
DevITjobs UKWorking (opt-in)UK tech board, salary on ~100% of rows โ€” same shape as the US board
Google JobsBlocked at sourceGoogle serves a bot-check page instead of results
ZipRecruiterBlocked at sourceAPI returns HTTP 403
Bayt (MENA)Blocked at sourceHTTP 403
BDJobs (Bangladesh)Blocked at sourceRedirects away from results
Naukri (India)Blocked at sourceDemands a CAPTCHA

In a verified run (2026-08-08, python developer, remote only, 10 requested per board) the five API boards delivered: The Muse 10, Jobicy 10, HN "Who is hiring?" 5, Himalayas 6, Remotive 2 rows. Per-board counts track each board's live inventory for the term and move day to day (an earlier run saw Remotive 10 and Himalayas 9).

LinkedIn, Indeed and Glassdoor remain the default selection. The five API boards are opt-in, with one convenience: when Remote jobs only is on and the board list is left at its default trio, the four remote-only boards (Remotive, Jobicy, Himalayas, HN) join the run automatically โ€” RUN_SUMMARY.remote_boards_auto_added records exactly which ones were added. They are deliberately not added to location searches: for a query like nurse in Dallas they contribute nothing, so we don't run them. The Muse is the exception with real US city coverage, but its category-based search means every word of your search term must appear in the job's title or description โ€” keep Muse search terms short.

The blocked five stay selectable in case they recover, and they cost you nothing when they return no rows โ€” but do not plan a project around them today.

Every run writes a RUN_SUMMARY record to its key-value store with the exact per-board outcome, so you always know which board gave you what and why one was quiet. It also carries columns_per_row and the full columns list, so you can diff a run's schema without reading a single row. If every board you selected fails to fetch, the run is marked FAILED rather than quietly succeeding with an empty dataset.

Every successful run also sets a status message in the Console run header โ€” row count, column count, and which preset fired and what it changed. Previously that line only appeared when search terms had been dropped, so a run driven by a preset finished without ever saying which preset had run.


Why use this

  • Eight boards, one run โ€” LinkedIn + Indeed + Glassdoor by default, plus five keyless job APIs (The Muse, Remotive, Jobicy, Himalayas, HN "Who is hiring?") merged into a single dataset.
  • Paste a search URL โ€” already built the search on LinkedIn, Indeed or Glassdoor? Copy the URL into searchUrls and skip the form. Keywords, location, radius (LinkedIn and Indeed only), remote flag, recency, job type, experience level and sort are read straight out of the URL โ€” every honored parameter is listed below, and an unparseable URL is reported instead of quietly scraping something else. Note this replaces the search term, locations and board list only; a preset and your filters still apply on top.
  • Keyword ร— location matrix โ€” locations runs every search term against every location and merges the results, which is how you get past a board's per-query result cap (one query for "United States" returns one capped page; five city queries return five).
  • Deduplicated, and it tells you โ€” the same job on multiple boards becomes one row carrying found_on_sites (every board it appeared on) and duplicate_count. Company names are normalized before matching (legal suffixes stripped), so Wipro on one board and Wipro Limited on another collapse into one billed row. Being live on three boards is a real hiring-urgency signal. Within a single board two rows are merged only when they are the same listing (same job_url, else same id), so a company advertising six distinct Software Engineer II roles on Indeed gives you six rows, not one. Read duplicate_count as how many listings in this run shared this title and company, not as a count of rows we deleted: across boards those extras really were merged into the one billed row, while on a single board they each kept their own row, so the same number can legitimately appear on several rows.
  • Filters that never bill โ€” salary-required, job type, remote-only, title-keyword excludes, company excludes, company includes (competitor hiring watch), experience level, city excludes, staffing-agency excludes, strict keyword match and max posting age all run after dedup and before billing; a filtered row is never charged, and RUN_SUMMARY.filter_removed says exactly how many each filter took.
  • Score every job against your resume โ€” put your skills in resumeKeywords and each row gains matched_keywords and keyword_match_percent. Pure annotation: it never removes a row, never changes the bill, and sortBy: relevance puts the best matches on top.
  • Comparable salaries โ€” hourly, weekly and monthly pay is normalized into salary_min_annual / salary_max_annual (hourly ร—2080, weekly ร—52, monthly ร—12), so one column sorts every salaried row.
  • Resilient โ€” a slow or blocked board is skipped on a timeout so the run still finishes with everything the others returned. No hung runs, no all-or-nothing failures.
  • Honest per-board reporting โ€” RUN_SUMMARY names the board and the actual HTTP error, instead of a silent empty result.
  • Multi-search โ€” pass several search terms in one run; results are merged and deduped, and each row records which term matched it.
  • Clean text โ€” descriptions are converted to real Markdown without the stray backslash escapes (full\-time, Web3 \| NYC) that the underlying library emits by default.
  • Export anywhere โ€” CSV, JSON, Excel, or pipe to Google Sheets / your CRM.

How to use it

  1. (Shortcut) Already have the search open on LinkedIn, Indeed or Glassdoor? Copy the URL into Search URLs and hit Save & start โ€” everything below is filled in from the URL.
  2. Enter a search term (job title/keywords) and a location (or several under Locations, which runs every term against every location).
  3. Keep the default boards, and set Max job postings per board. For remote searches, turning on Remote jobs only auto-adds the four remote-only boards; add The Muse yourself for extra US-city coverage.
  4. Full LinkedIn job details is on by default โ€” that's where LinkedIn descriptions, job_type and company_industry come from. Turn it off only for faster runs.
  5. Optionally filter: remote-only, job type, posted-within-N-hours, distance, country โ€” plus the post-scrape filters below (salary required, title/company excludes, max posting age), which are never billed.
  6. Run โ†’ get a clean, deduplicated jobs table.

Input

The Console form is grouped into sections that go from loudest to quietest โ€” Quick start, What to search, Paste a job-search URL, Filters & limits, Board-specific options, Output, and a deliberately plain Advanced (networking). The optional plumbing lower down is grouped at the bottom on purpose, because a default run needs none of it.

FieldTypeDescription
presetstringOptional goal shortcut: salary_research, competitor_intel, remote_only, fresh_postings, recruiter_leads. A preset only fills options you left at their default โ€” anything you set yourself wins โ€” and everything it changed (and everything it refused to change) is logged and recorded in RUN_SUMMARY.preset_applied. Empty (the default) is a strict no-op.
roleSelectstringOptional quick-pick job title. Precedence: searchTerms > roleSelect > searchTerm. Sent to the boards as plain keywords โ€” it is not a fixed taxonomy. Empty = use searchTerm.
citySelectstringOptional quick-pick metro. Precedence: locations > citySelect > location. Picking a non-US metro also sets countryIndeed to the matching country unless you set that field yourself. Empty = use location.
searchUrlsarrayPaste LinkedIn / Indeed / Glassdoor search URLs and skip the form. Replaces searchTerm, locations and sites when set. See the table of honored parameters below.
searchTermstringJob title / keywords (e.g. software engineer).
searchTermsarrayMultiple searches in one run (merged + deduped); overrides searchTerm. Capped at 5 โ€” extras are reported in the run status, never dropped silently.
locationstringCity / state / country (e.g. New York, NY). Empty = anywhere.
locationsarraySeveral locations in one run; every search term is run against every location. Replaces location. Capped at 10 locations and 25 term ร— location searches; caps are reported in RUN_SUMMARY.
sitesarrayBoards to scrape. Defaults to linkedin, indeed, glassdoor; the five API boards (muse, remotive, jobicy, himalayas, hn_hiring) are opt-in. Unrecognised names are reported, not silently ignored.
maxResultsintegerJob postings requested per board (1โ€“500, clamped). 0 and negatives clamp to 1 rather than falling back to the default.
isRemotebooleanAsks each board for its own remote filter, then drops any row whose own is_remote is false before billing โ€” LinkedIn accepts f_WT and ignores it, so the board-side filter alone is not enough. Rows with no description carry no evidence and are kept. With the board list left at its default it also auto-adds the four remote-only boards, which adds rows and raises the bill.
jobTypestringFull-time / part-time / internship / contract. Only Indeed and Glassdoor filter this board-side (LinkedIn and the five API boards ignore it), so rows whose own job_type contradicts your choice are dropped before billing. A posting with no job_type at all is kept.
hoursOldintegerOnly jobs posted in the last N hours (board-side filter). Above 8760 (one year) the age limit is dropped and your other board-side filters are kept โ€” it does not silently strip easy-apply/job-type/remote the way any other non-zero value must.
maxAgeDaysintegerDrop postings older than N days. Applied by the actor after scraping, so it works on every board; filtered rows are never billed.
requireSalarybooleanKeep only rows carrying a salary figure. Filtered rows are never billed.
excludeTitleKeywordsarrayDrop jobs whose title contains any of these words/phrases (case-insensitive). Filtered rows are never billed.
excludeCompaniesarrayDrop these companies. Names are normalized first, so Wipro also excludes Wipro Limited. Filtered rows are never billed.
targetCompaniesarrayKeep only these employers (competitor hiring watch). Same normalization, plus a parent match: Amazon keeps Amazon Web Services but not Amazonia. Filtered rows are never billed.
experienceLevelarrayinternship / entry / associate / mid_senior / director / executive. Best-effort โ€” see the note below. Filtered rows are never billed.
excludeCitiesarrayDrop jobs whose location mentions any of these places (whole-word, case-insensitive). Filtered rows are never billed.
excludeAgenciesbooleanDrop rows whose company name looks like a staffing agency or recruiter. Heuristic; see the note below. Filtered rows are never billed.
strictKeywordMatchbooleanKeep only rows whose title or description contains your search words. Off by default. Use it for narrow roles โ€” LinkedIn never answers "nothing matched" and returns loosely related postings instead. Filtered rows are never billed.
resumeKeywordsarrayScore every row against your skills โ€” adds matched_keywords and keyword_match_percent. Never filters, never changes the bill.
sortBystringrelevance, date_desc or salary_desc. Applied before the charge cap.
countryIndeedstringCountry for Indeed & Glassdoor (e.g. usa, uk, india). Leave it alone and a location naming a country (Berlin, Germany) points Indeed at that country's site automatically. A spelling we don't recognise falls back to usa with a note in RUN_SUMMARY โ€” it never fails the run.
distanceintegerSearch radius in miles. LinkedIn and Indeed honor it; Glassdoor and the API boards ignore it.
offsetintegerSkip the first N job postings per board, to page past a run you already have. LinkedIn, Indeed and Glassdoor page server-side; on the five API boards the Actor fetches N extra rows and hands you the tail. Boards order results their own way, so it is not a stable cursor.
easyApplybooleanLinkedIn/Indeed direct-apply jobs only.
linkedinFetchDescriptionbooleanOn by default. Fetches each LinkedIn job's own page โ€” the source of LinkedIn descriptions, job type and industry. Turn off for faster runs.
descriptionFormatstringmarkdown or html.
descriptionHtmlbooleanAlso emit description_html (the board's original HTML) alongside description. Both in one run, no extra requests.
includeCompanyDetailsbooleanFill company_ceo, company_banner, company_addresses_all and company_details_filled, which are blank without it. The columns exist either way. No extra requests.
proxyConfigurationobjectProxy settings. Datacenter is the default and what you want (measured 2026-08-11: 96% of the rows at ~1/5 the cost); switch to Residential only if a board starts coming back short.

Example input:

{
"searchTerm": "data analyst",
"location": "Austin, TX",
"sites": ["linkedin", "indeed", "glassdoor"],
"maxResults": 50,
"hoursOld": 168,
"requireSalary": true,
"excludeTitleKeywords": ["senior", "intern"]
}

Paste-a-URL input:

{
"searchUrls": [
"https://www.indeed.com/jobs?q=python+developer&l=Austin%2C+TX&fromage=14&radius=25",
"https://www.linkedin.com/jobs/search?keywords=python%20developer&location=Austin%2C%20Texas&f_TPR=r604800"
],
"maxResults": 25,
"resumeKeywords": ["Python", "AWS", "Kubernetes"],
"sortBy": "relevance"
}

Multi-city + competitor watch input:

{
"searchTerms": ["data analyst", "business analyst"],
"locations": ["Austin, TX", "Chicago, IL", "Denver, CO"],
"maxResults": 30,
"targetCompanies": ["Deloitte", "Accenture", "Capgemini"],
"includeCompanyDetails": true,
"sortBy": "date_desc"
}

Paste a search URL (searchUrls)

Build the search on the board's own site, copy the URL, paste it in. Each URL scrapes the board it came from โ€” a LinkedIn URL runs LinkedIn only โ€” and searchUrls replaces searchTerm / locations / sites for the run. Anything the URL contains that is not in this table (geoId, tracking ids, currentJobId, saved-search ids) is ignored.

BoardURL shapes acceptedParameters honored
LinkedInlinkedin.com/jobs/search?keywords=โ€ฆkeywords โ†’ search term, location, distance (miles), f_WT=2 โ†’ remote only, f_TPR=r<seconds> โ†’ posted-within, f_JT โ†’ job type, f_E โ†’ experience level, f_C โ†’ company ids, sortBy (DD/R), start โ†’ offset
Indeedindeed.com/jobs?q=โ€ฆ (any country sub-domain, e.g. ca., uk., de.)q โ†’ search term, l โ†’ location, radius, fromage (days) โ†’ posted-within, sc attr(DSQF7) โ†’ remote, sc jt(โ€ฆ) โ†’ job type, explvl โ†’ experience level, sort=date, start โ†’ offset, and the country from the sub-domain
Glassdoorglassdoor.com/Job/โ€ฆ-SRCH_IL.<a>,<b>_ICโ€ฆ_KO<c>,<d>.htm or โ€ฆ/Job/jobs.htm?sc.keyword=โ€ฆsc.keyword / keyword (or the KO slug offsets) โ†’ search term, locKeyword / locName (or the IL slug offsets) โ†’ location, fromAge (days), remoteWorkType=1 โ†’ remote, jobTypes, seniorityType โ†’ experience level, sortBy, and the country from the domain suffix. radius is not read โ€” Glassdoor's scraper has no distance parameter, so a radius in the URL changes nothing (only LinkedIn and Indeed honor distance)

A URL we cannot parse is skipped with a reason, never guessed at: the other URLs still run and RUN_SUMMARY.search_url_errors names the URL and what was missing. If none of them parse, the run scrapes nothing and bills nothing.

Each URL's filters stay on that URL. A seniority filter carried in one URL (f_E, explvl, seniorityType) is applied only to the rows that URL returned โ€” pasting an executive-only LinkedIn URL next to an unfiltered Indeed URL does not thin out the Indeed results. Those rows are removed before de-duplication and therefore before billing, and the count shows up in RUN_SUMMARY.filter_removed as experienceLevel (from search URL), with the per-URL levels in RUN_SUMMARY.search_url_experience_levels. Setting the experienceLevel input explicitly overrides every URL's own level filter.

Verified locally on 2026-08-08: an Indeed URL and the equivalent manual inputs (searchTerm + location + hoursOld: 336 + distance: 25) returned the same 10 job ids and the same output columns โ€” the URL path is a front door onto the same pipeline, not a different one. A Glassdoor SRCH_โ€ฆKOโ€ฆ slug URL returned 8 rows; glassdoor.com/Job/jobs.htm (no keyword anywhere) was rejected with a readable reason while the other URL still ran.

Keyword ร— location matrix (locations)

locations runs every search term against every location and merges the results into one deduplicated table. It is the way past a board's per-query cap: one query for a whole country returns one capped page, five city queries return five. Every pair asks each board for maxResults rows, so cost scales with the number of pairs โ€” capped at 10 locations and 25 pairs per run, and both caps are reported (never applied silently). Rows carry matched_search_term and, whenever a run covers more than one location, matched_location.

Live run (2026-08-08, data analyst ร— [Austin, TX, Chicago, IL] ร— LinkedIn + Indeed + Glassdoor at 12/board): 72 raw rows โ†’ 67 after cross-search dedup โ†’ 63 billed after filters, split 28 Austin / 35 Chicago.

Post-scrape filters โ€” filtered rows are never billed

These filters run inside the actor after scraping and deduplication and before billing, so a filtered row is never charged. RUN_SUMMARY.filter_removed records how many rows each one took.

  • requireSalary โ€” keep only rows carrying a salary figure (salary_min or salary_max).
  • jobType โ€” also enforced here, not just board-side. Indeed and Glassdoor filter employment type themselves; LinkedIn's guest search accepts the parameter and returns the same rows anyway, and the five API boards have no such filter at all. So a row whose own job_type contradicts your choice is dropped before billing. A posting the board published with no job_type is kept โ€” silence is not a contradiction. Measured 2026-08-08: jobType: internship on Remotive + Jobicy + Himalayas dropped all 35 mismatching rows and billed 0; on LinkedIn + Indeed it dropped LinkedIn's 10 full-time rows and billed 10 rows that all carry internship.
  • isRemote โ€” also enforced here for the same reason: LinkedIn ignores f_WT. Rows whose own is_remote is false are dropped before billing; rows with no description carry no evidence and are kept. Measured 2026-08-08 on the Remote-only sweep preset: 19 rows dropped, and 0 of the billed rows had is_remote: false (before this, 10 of 10 LinkedIn rows billed under that preset were on-site jobs in Midland TX, El Paso TX and Pittsburgh PA).
  • strictKeywordMatch โ€” opt-in, off by default. Keep only rows whose title or description contains your search words. LinkedIn never answers "nothing matched"; it degrades to loosely related cards. Measured 2026-08-08: underwater welder in Fargo, ND billed 10 rows of marine techs, trenchless engineers and welders with it off, and 0 with it on. Either way RUN_SUMMARY.rows_not_mentioning_search_term counts the off-topic rows per board (that run: {"linkedin": 10}), so you can see the problem before deciding to pay for it.
  • excludeTitleKeywords โ€” drop titles containing any listed word or phrase (case-insensitive substring match).
  • excludeCompanies โ€” drop listed companies; names are normalized before matching, so Wipro also excludes Wipro Limited.
  • targetCompanies โ€” the inverse: keep only the listed employers. Same normalization plus a parent match, so Amazon keeps Amazon Web Services while leaving Amazonia out. This is the competitor-hiring-watch input.
  • maxAgeDays โ€” drop postings older than N days. Unlike hoursOld this is applied by the actor itself, so it works on every board.
  • excludeCities โ€” drop jobs whose location text mentions a listed place. Whole-word and case-insensitive, matched against the location exactly as the board published it โ€” Dallas drops Dallas, TX, Dall drops nothing. Best-effort: boards format locations inconsistently.
  • excludeAgencies โ€” drop rows whose company name matches a staffing/recruiting pattern (staffing, recruit*, headhunt*, talent*, manpower, personnel, placement*, consultanc*, resourcing, workforce, temp agency, employment agency, search partners/group/associates, HR solutions, staff augmentation). It reads the name only, so a re-poster with a neutral name gets through and a genuine employer with Workforce in its name gets dropped โ€” use excludeCompanies when you need exact control.
  • experienceLevel โ€” best-effort, and here is exactly how. When a board publishes a seniority field we use it (LinkedIn does with detail fetching on โ€” measured 19/19 LinkedIn rows; Jobicy, Himalayas and The Muse also publish one). Otherwise we read it off the title: Senior / Sr. / Staff / Principal / Lead โ†’ mid-senior, Junior / Jr. / Graduate / Entry level โ†’ entry, Intern โ†’ internship, Director โ†’ director, VP / Chief / Head of โ†’ executive. A job whose level we cannot determine is kept, not dropped โ€” we only remove rows we can prove don't match. On a live 63-row run the level came from a board field on 19 rows, from the title on 16, and was genuinely unknown on 28.

Verified on live runs: requireSalary removed 7 rows and excludeTitleKeywords 3 (2026-08-07); excludeAgencies removed 4 rows and experienceLevel 1 (2026-08-08), all recorded in RUN_SUMMARY.filter_removed, and a targetCompanies value matching nothing ended the run with 0 rows pushed and nothing billed. excludeAgencies reads the company name only, so treat it as a heuristic rather than a measurement: across the 178 distinct employers returned by the 2026-08-08 local runs it flagged none, which tells you how few agency reposts a given search contains, not that it is precise. Its known failure mode is the workforce token โ€” Texas Workforce Commission and Workforce Software are both flagged, and both are genuine employers posting their own jobs. Use excludeCompanies when you need exact control.

Resume keyword scoring (resumeKeywords) and sortBy

Put your skills in resumeKeywords (Python, Kubernetes, SOC 2, C++, .NET, node.js โ€” case-insensitive substring, so punctuation works as written) and every row gains:

  • matched_keywords โ€” which of your keywords appear in the title or description, in your order
  • keyword_match_percent โ€” 0โ€“100

It is annotation only: no row is removed and the bill is unchanged. sortBy: relevance then ranks by match percent, then by how many boards the job was found on, then by date. sortBy: date_desc and salary_desc sort on posting date and on the annualized salary; rows with no value for the key always land last, and sorting runs before the charge cap so a truncated run keeps your top rows. Measured on a 63-row live run: 8 rows at 100%, 17 at 75%, 15 at 50%, 8 at 25%, 15 at 0%.


Output fields and how often they are filled

Every row has the same 50 columns

One column tuple per run โ€” and the same tuple in every run. Every row the Actor delivers carries all 50 columns below, in the same order, whatever board it came from and whatever options you set. A field the board did not publish is null, never missing. That is what makes the CSV safe to open in a spreadsheet, pd.read_csv without dtype surprises, and parse by column position.

This changed on 2026-08-08. Rows used to have their empty fields stripped out before delivery, so a column existed on the rows one board filled in and silently vanished from the rest โ€” a LinkedIn + Indeed + Glassdoor run of this shape once shipped 10 different column tuples ranging from 20 to 31 columns. Since the fix, every row of a run carries the same 50-column tuple, however many rows come back and whatever boards fill them. A live default LinkedIn + Indeed + Glassdoor run delivered one tuple of 50 columns on every row โ€” the row count is just whatever the boards return that day (57 the first time this was measured, 60 on the latest run) โ€” and separate runs across the five API boards and an Indeed-only run with every option switched on carried the same 50 columns in the same order.

The column order, which is stable and safe to depend on:

id, title, company, location, job_url, job_url_direct, site, job_type, is_remote, date_posted, salary_source, salary_min, salary_max, salary_currency, salary_interval, salary_min_annual, salary_max_annual, job_level, job_function, listing_type, company_industry, company_url, company_url_direct, company_addresses, company_description, company_num_employees, company_revenue, company_rating, company_reviews_count, company_logo, description, emails, skills, experience_range, vacancy_count, work_from_home_type, found_on_sites, duplicate_count, search_term, matched_search_term, scraped_at, description_html, company_ceo, company_banner, company_addresses_all, company_details_filled, matched_keywords, keyword_match_percent, matched_location, source_search_url.

The last nine are toggle-driven: description_html fills only with descriptionHtml: true, the four company extras and the counter only with includeCompanyDetails: true, the two keyword columns only with resumeKeywords, matched_location only on a multi-location run and source_search_url only on a searchUrls run. The column is there either way โ€” null means the toggle was off. A column that appears and disappears with a setting is the same ragged-CSV problem in miniature, so we don't do that. New fields are only ever appended to the end of this list, never inserted, so a positional parser keeps working.

Two things that can still hand you a narrower row, both of them your own choice: ?fields= / ?omit= on the Dataset API, and clean=true (a shortcut for skipHidden=true + skipEmpty=true), which strips empty values back out. Plain format=csv, format=xlsx, format=json and the Console Export button all give you the full 50.

Fill rates

Job boards publish wildly different amounts of detail, so most fields are conditional โ€” the column is always present, but how often it carries a value depends entirely on the board. These are measured fill rates, not aspirations โ€” taken from live runs on 2026-08-07 (58โ€“85 rows for the detail-off column, 67โ€“72 rows with LinkedIn detail fetching on). Ranges span two independent runs on different queries; expect your own numbers to land inside them.

LinkedIn detail fetching is now on by default, so the right-hand column is what a default run delivers โ€” a live default run ({}) returned 60 rows with descriptions on 20 of 20 LinkedIn rows. The left column is what you fall back to if you turn it off for speed.

FieldDetail fetch offDetail fetch on (default)Comes from
id, title, site, job_url, is_remote, date_posted, scraped_at, search_term, matched_search_term, found_on_sites, duplicate_count100%100%all boards
company98%95%all boards
location98%100%all boards
company_url97%94%all boards
description64%100%Indeed + Glassdoor always; LinkedIn only with detail fetching
salary_min, salary_max, salary_interval, salary_currency, salary_source56โ€“60%59โ€“70%Glassdoor 95โ€“100%, Indeed 50โ€“80% (varies by query), LinkedIn none without detail fetching and ~35% with it
company_logo49%86%Glassdoor + Indeed; LinkedIn with detail fetching
listing_type32%31%Glassdoor only (organic / sponsored)
job_url_direct31%33%Indeed only (the employer's own posting URL)
company_rating29%30%Glassdoor only โ€” ~90% of Glassdoor rows
job_type24%65%Indeed ~80%; LinkedIn 100% only with detail fetching; never from Glassdoor
company_url_direct, company_addresses, company_num_employees, company_revenue, company_description17โ€“20%19โ€“23%Indeed only (~55โ€“70% of Indeed rows)
job_level, job_function0%34%LinkedIn only, and only with detail fetching (then 100%)
company_industry2%38%LinkedIn with detail fetching (100% of LinkedIn rows); Indeed ~10%
emails5โ€“8%12โ€“14%scraped out of description text when a job lists one

linkedinFetchDescription is on by default, so job_type, company_industry, job_level and job_function arrive without touching anything. Turn it off and LinkedIn contributes none of them โ€” that trade is yours to make for speed.

Two derived columns normalize pay so every row sorts on one axis: salary_min_annual / salary_max_annual convert hourly (ร—2080), weekly (ร—52) and monthly (ร—12) figures to annual. In a live default run they were present on 57 of 57 rows that carried a salary. Two guards keep them honest. When enforceAnnualSalary is on, the scraping library annualizes description-parsed pay itself but leaves salary_interval reading hourly/monthly โ€” we rewrite that interval to yearly instead of multiplying a second time, which is what used to ship salary_min_annual: 95180800 on a warehouse job and print $45,760 / hourly in the raw columns. And a result that cannot be a wage is left null rather than published; the raw salary_min / salary_max are never touched. That backstop is applied only to USD-scale currencies, so a real 3,000,000 KRW or 10,000,000 IDR monthly salary still annualizes normally. Measured 2026-08-08 on the salary_research preset (warehouse associate, Dallas TX, Indeed + LinkedIn): 0 of 14 billed rows carried an impossible annual figure and 0 claimed an hourly rate above $1,000/hour โ€” the same run shape previously produced 9 of 13. The three salary qualifier columns (salary_interval, salary_currency, salary_source) are blank on any row without an actual amount โ€” a qualifier with nothing to qualify is noise. Verified on the 2026-08-08 runs: 0 of 57 and 0 of 69 rows carried an orphan qualifier.

company_reviews_count, skills, experience_range, vacancy_count and work_from_home_type are only ever published by Naukri, which is currently CAPTCHA-blocked โ€” the columns are there on every row but expect them to be blank. They fill automatically if Naukri recovers.

includeCompanyDetails โ€” what you actually get, per board

Turning it on does one thing and invents nothing: it fills four columns that are otherwise blank. (Column stability is no longer its job โ€” every row of every run carries all 50 columns whether or not this is on.) It adds company_details_filled, a count of how many company fields the row actually got, and it emits three fields that Indeed already sends in its response and the underlying library throws away: company_ceo, company_banner and company_addresses_all (the full address list; company_addresses only ever held the first). No extra requests, no extra cost per row.

Measured on a live 63-row run (2026-08-08, LinkedIn + Indeed + Glassdoor, LinkedIn detail fetching on):

FieldIndeed rowsLinkedIn rowsGlassdoor rowsDataset-wide
company_url96%100%100%98%
company_logo71%100%90%86%
company_industry33%100%0%43%
company_num_employees (size band)75%0%0%29%
company_url_direct (employer's own site)75%0%0%29%
company_addresses / company_addresses_all71%0%0%27%
company_description67%0%0%25%
company_revenue58%0%0%22%
company_ceo (new)54%0%0%21%
company_banner (new)54%0%0%21%
company_rating0%0%95%30%

Read that as: Indeed is the company-data board (size, revenue, description, corporate website, addresses, CEO), LinkedIn contributes the industry (100% of LinkedIn rows, but only with Full LinkedIn job details on โ€” it was 0% in a run with it off), and Glassdoor contributes the employer rating. For the five API boards, which board a row came from decides this entirely, so the per-board answer is the only one worth quoting (measured 2026-08-08 on a 34-row remote-only run across all five):

ColumnJobicyHimalayasThe MuseRemotiveHN
company_industry10/100/70/100/20/5
company_logo10/105/70/102/20/5
job_level10/107/710/100/20/5
job_function0/100/70/102/20/5

In words: Jobicy is the only API board that publishes company_industry; Jobicy, Himalayas and Remotive publish company_logo; Jobicy, Himalayas and The Muse publish job_level; Remotive is the only one that publishes job_function; HN publishes none of them. Everything else on these boards is null. Don't read the dataset-wide percentage as a property of the Actor โ€” it is just each board's share of your particular run, and Jobicy is capped at 50 rows per call. If a board doesn't publish a field, the column stays null and says so; we never fill it from somewhere else.

descriptionHtml โ€” both formats in one run

With descriptionHtml: true every row carries description_html (exactly what the board published) alongside the normal description in your chosen descriptionFormat. There is no second request and no price change: the boards return HTML anyway, so the Markdown is produced from that same HTML locally. Use the HTML when you need the original links, lists and headings for your own site or ATS; use the Markdown for LLM prompts and spreadsheets. Budget for the size: description_html holds the raw HTML that the Markdown description is generated from, so it is necessarily about as large as the column beside it. On a 19-row A/B run (2026-08-08, same query with the switch off and on) description_html was filled on 19/19 rows and ran 111% of the description column, growing the whole dataset 86%. An earlier run measured +167% / +132%. Plan on roughly doubling your dataset, as the input form says โ€” not the "about 13%" an earlier version of this README claimed.

JSON output sample

A complete, unedited row from a live Glassdoor scrape (2026-08-08, software engineer / New York NY / LinkedIn + Indeed + Glassdoor, description truncated for length). This is the whole row โ€” all 50 columns, nulls included. Every other row in that run, from every board, had exactly these 50 keys in exactly this order:

{
"id": "gd-1009849012184",
"title": "Software Development Engineer, ML Systems, Annapurna Labs",
"company": "Annapurna Labs (U.S.) Inc.",
"location": "New York, NY",
"job_url": "https://www.glassdoor.com/job-listing/j?jl=1009849012184",
"job_url_direct": null,
"site": "glassdoor",
"job_type": null,
"is_remote": false,
"date_posted": "2026-02-28",
"salary_source": "direct_data",
"salary_min": 158100.0,
"salary_max": 213800.0,
"salary_currency": "USD",
"salary_interval": "yearly",
"salary_min_annual": 158100,
"salary_max_annual": 213800,
"job_level": null,
"job_function": null,
"listing_type": "organic",
"company_industry": null,
"company_url": "https://www.glassdoor.com/Overview/W-EI_IE7470741.htm",
"company_url_direct": null,
"company_addresses": null,
"company_description": null,
"company_num_employees": null,
"company_revenue": null,
"company_rating": 3.7,
"company_reviews_count": null,
"company_logo": "https://media.glassdoor.com/sql/7470741/amazon-web-services-squareLogo-1680754591940.png",
"description": "**DESCRIPTION** About Amazon Annapurna Labs ...",
"emails": null,
"skills": null,
"experience_range": null,
"vacancy_count": null,
"work_from_home_type": null,
"found_on_sites": "glassdoor",
"duplicate_count": 1,
"search_term": "software engineer",
"matched_search_term": "software engineer",
"scraped_at": "2026-08-08T15:52:57.101661+00:00",
"description_html": null,
"company_ceo": null,
"company_banner": null,
"company_addresses_all": null,
"company_details_filled": null,
"matched_keywords": null,
"keyword_match_percent": null,
"matched_location": null,
"source_search_url": null
}

The trailing nulls are the toggle-driven columns: this run had descriptionHtml, includeCompanyDetails and resumeKeywords off, one location and no searchUrls. Turn those on and the same nine columns carry values โ€” they never move, appear or disappear.

Results render as a sortable table on the Output tab and export to CSV, JSON, or Excel.

Example output

A real sample from a live run (software engineer, New York NY, three boards):

titlecompanysitecompany_ratingsalary_minjob_url
Software Engineer, Systems ML - CompilersMetaglassdoor3.7154003https://www.glassdoor.com/job-listing/j?jl=101021โ€ฆ
Silicon Software LeadNormal Computing Corporationindeedhttps://www.indeed.com/viewjob?jk=76bcfโ€ฆ
AI Modernization Senior Lead Software Eโ€ฆJPMorganChaseindeed133000https://www.indeed.com/viewjob?jk=74806โ€ฆ
DevOps EngineerOnMedlinkedinhttps://www.linkedin.com/jobs/view/โ€ฆ

Use cases

  • Recruiting & staffing โ€” pull every open role for a title/location across boards into one pipeline.
  • Job boards & aggregators โ€” backfill and keep a niche board fresh from multiple sources.
  • Hiring-signal sales โ€” a role live on three boards at once is a company hiring urgently; duplicate_count and found_on_sites surface exactly that.
  • Labor-market research โ€” salary ranges, remote share, and demand by title and location.
  • Personal job hunt โ€” one deduped feed instead of refreshing three sites.

Use with AI agents & automation

Run from the Apify MCP server so AI agents (Claude, ChatGPT, Cursor) can pull jobs as a tool call, schedule runs via Make, n8n, or Zapier to alert on new postings, or sync the dataset to Google Sheets for a live jobs dashboard. Clean flat JSON drops into ATS/CRM pipelines with no glue code.


Licensing & attribution โ€” what you may do with the rows

The API boards publish their feeds with strings attached, and those strings pass through to you:

BoardMay you re-post rows on another job board?Required credit
The MuseYes (attribution requested)Credit The Muse
RemotiveNo โ€” their ToS explicitly forbid itCredit Remotive
JobicyOnly with credit, and apply buttons must link to the original job URLCredit Jobicy + keep the job_url links
HN "Who is hiring?"User-authored HN content; link back to the HN itemCredit Hacker News / Algolia
HimalayasYes, with a link-backLink to himalayas.app
Remote OKOnly with a visible, clickable (dofollow) link back to the row's job_url and "Remote OK" named as the source โ€” their terms suspend API access otherwise. Logo use is forbiddenDofollow link to remoteok.com + name "Remote OK"
We Work RemotelyTheir terms carry no republication bar; a link-back is good practiceLink to the job_url
Working NomadsNo stated bar; link-back is good practiceLink to the job_url
DevITjobs US/UKUndocumented public API โ€” treat rows as pointers and keep the job_url links. ~99% of rows are partner-syndicated (e.g. from Indeed), so salary figures may be partner-estimatedKeep the job_url links

In plain terms: Remotive rows are fine for lead-gen, research and your own job hunt, but must not be republished on another job board (their terms name Jooble, Google Jobs, LinkedIn and the like explicitly). If you build a job site on Jobicy data, keep the job_url apply links and credit Jobicy. The job_url column already points at each board's original posting, which covers the link-back half of these requirements.


API examples

Same input keys as the Console form, from code:

const { ApifyClient } = require('apify-client');
const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('flash_scraper/multi-jobboard-scraper').call({
searchTerm: 'data analyst', location: 'Austin, TX', maxResults: 50, onlyNewJobs: true,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("flash_scraper/multi-jobboard-scraper").call(run_input={
"searchTerm": "data analyst", "location": "Austin, TX", "maxResults": 50, "onlyNewJobs": True,
})
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
curl -X POST "https://api.apify.com/v2/acts/flash_scraper~multi-jobboard-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"searchTerm":"data analyst","location":"Austin, TX","maxResults":50}'

Wire the finished-run webhook into n8n, Make or Zapier to land rows in Slack, Sheets or your ATS automatically - combined with onlyNewJobs on a schedule, that is a self-maintaining job feed.

Only need LinkedIn, as cheap as possible? The same publisher's LinkedIn Jobs Scraper is the single-board budget tool - priced in the Store's cheapest tier.

Pricing

Pay-per-event. You are charged one Result event per row delivered to your dataset, plus a one-off Actor Start event. Because de-duplication happens before the push, you pay once for a job found on three boards, not three times โ€” and that holds across a keyword ร— location matrix too, so the same role returned by two city searches is one billed row. Boards that return nothing cost nothing, rows removed by any post-scrape filter (requireSalary, excludeTitleKeywords, excludeCompanies, targetCompanies, experienceLevel, excludeCities, excludeAgencies, maxAgeDays) are never billed, and failed runs deliver no rows and so bill no results. includeCompanyDetails, descriptionHtml and resumeKeywords fill columns that are otherwise blank; they never add a row, so they cost nothing extra. See the Apify Store page for current prices.

Cost note for locations: each term ร— location pair is a full search on every selected board, so 3 terms ร— 3 locations ร— 3 boards can bill up to 9ร— a single search before de-duplication. The caps (10 locations, 25 pairs) exist for exactly that reason and are reported in RUN_SUMMARY, never applied silently.

If you set a maximum total charge on the run and the scrape exceeds it, the actor delivers as many rows as the budget allows and says so in the run status โ€” it will not spend your budget and then hand back an empty dataset.


FAQ

Where does the data come from? Public job listings on LinkedIn, Indeed and Glassdoor via the open-source JobSpy engine plus our own fixes on top of it, and the public keyless APIs of The Muse, Remotive, Jobicy, Himalayas and Hacker News (Algolia) for the five extra boards.

Do I need proxies? Yes โ€” the actor is preconfigured to use Apify's default datacenter pool, which handles LinkedIn / Indeed / Glassdoor (measured 2026-08-11: 52 rows vs residential's 54 on the same search, at ~1/5 the cost). The five API boards need no special proxy at all.

Why is a board returning nothing? Check the RUN_SUMMARY record in the run's key-value store โ€” it names each board and the actual error. Five of the thirteen selectable boards are currently blocked at the source (see the table at the top).

Why didn't the remote boards run on my location search? By design. Remotive, Jobicy, Himalayas and HN "Who is hiring?" carry remote jobs only โ€” they add nothing to a query like nurse in Dallas, so they only run when you name them in sites, or automatically when Remote jobs only is on and the board list is at its default. RUN_SUMMARY.remote_boards_auto_added tells you when the auto-add happened.

Why does The Muse return few or no rows for long search terms? The Muse is searched by job category, then filtered so that every word of your search term appears in the job's title or description. Short terms match; long multi-word phrases starve. Keep Muse terms to the core skill words.

How many rows can one run deliver? Up to 500 per board (maxResults). A stress run at 120 per board delivered 258 rows in 226 seconds: 118 from LinkedIn (full descriptions on all 118), 112 from Indeed, 28 from Glassdoor. Glassdoor tops out around ~28โ€“30 rows per query no matter the cap โ€” a board-side practical limit, not a bug or a budget issue.

Why are job_type / company_industry empty on LinkedIn rows? They shouldn't be โ€” Full LinkedIn job details is on by default and fills them. If you turned it off, that's why: LinkedIn's search results don't include them.

Can I search multiple titles at once? Yes โ€” use searchTerms (an array). Results are merged and deduplicated, and every row's matched_search_term records which term found it. Capped at 5 per run; any extras are named in the run status message.

Why did a job I expected not appear? Across boards, dedup keys on title + normalized company name (legal suffixes are stripped, so Wipro and Wipro Limited match), so the same role found on LinkedIn and Indeed arrives once with duplicate_count: 2. Within one board nothing is merged unless it is literally the same listing, so same-titled but distinct openings all survive. Otherwise check RUN_SUMMARY.filter_removed โ€” a filter you set (job type, remote-only, experience level, max ageโ€ฆ) may have dropped it, in which case you were not charged for it.

Can I just paste the search URL from the job board? Yes โ€” put it in searchUrls. LinkedIn, Indeed (any country sub-domain) and Glassdoor are supported, and the exact parameters we read out of each URL are listed above. A URL from another site, or one with no keyword in it, is skipped with a written reason in RUN_SUMMARY.search_url_errors rather than silently scraping something else.

Why did experienceLevel keep a job that isn't at that level? Because we could not prove otherwise. Only some boards publish a seniority field (LinkedIn does with detail fetching on); for the rest we read the title, and a plain "Software Engineer" genuinely says nothing about seniority. Rather than drop rows on a guess, we keep the unknowns โ€” dropping them would hide real jobs from you. Add excludeTitleKeywords if you want a hard cut.

Does resumeKeywords change what I pay? No. It only fills two columns (matched_keywords, keyword_match_percent) on rows you were getting anyway โ€” the columns are on every row regardless. Neither it nor includeCompanyDetails nor descriptionHtml adds a request or a row.

Why are so many cells empty? Because the boards did not publish those values, and we would rather show you an honest blank than invent one. Every row carries all 50 columns so your CSV lines up; the Fill rates table above says exactly which board fills which column and how often. Nothing is filled in from a second source.

Can I export to CSV or Google Sheets? Yes โ€” CSV, JSON, or Excel from the Output tab, or sync to Google Sheets via Make, n8n, or Zapier. Exports carry all 50 columns. The one thing that will narrow them is asking for it: ?fields= / ?omit=, or clean=true on the Dataset API (it is a shortcut for skipHidden=true + skipEmpty=true, so it strips empty cells back out and the rows go ragged again).


Other Flash Scrape scrapers