Job Board Aggregator ๐ - Indeed, LinkedIn & Glassdoor Scraper
Pricing
from $2.50 / 1,000 results
Job Board Aggregator ๐ - Indeed, LinkedIn & Glassdoor Scraper
Job board aggregator and multi-board scraper: LinkedIn, Indeed, Glassdoor, The Muse + 8 keyless boards in one run. ATS job scraper built in - watch Greenhouse, Lever and Ashby company boards. New-job-alerts mode returns only new postings. Cross-board dedup: one role bills once. CSV, JSON.
Pricing
from $2.50 / 1,000 results
Rating
5.0
(3)
Developer
Flash Scrape
Maintained by CommunityActor stats
4
Bookmarked
23
Total users
9
Monthly active users
2 hours ago
Last modified
Categories
Share
Job Board Aggregator - Indeed, LinkedIn & Glassdoor Scraper
One search. Up to eight job boards. One deduplicated row per role โ so a job posted to LinkedIn, Indeed and Glassdoor bills you once, not three times.
LinkedIn, Indeed and Glassdoor plus nine keyless job APIs (The Muse, Remotive, Jobicy, Himalayas, Hacker News "Who is hiring?", Remote OK, We Work Remotely, Working Nomads, DevITjobs US/UK) merged into a single clean table: title, company, location, salary (min/max/interval/currency plus annualized columns), remote flag, job type, posting date, full description, employer rating and company details. Search by keyword + location, filter by remote / job type / recency / salary present / title keywords / company, and export to CSV, JSON, or Excel.
Watch specific companies' careers pages โ list company names in Watch companies and the actor reads each one's own ATS board directly (Greenhouse, Lever or Ashby โ all keyless, tried in that order) and merges the live openings into the same 50-column table. Fresher than any job board, straight from the source. Leave the search term untouched and you get each company's whole board; type a term and only matching roles are delivered and billed.
Only pay for NEW jobs on a schedule โ turn on Only new jobs (monitoring mode) and the actor remembers every job this search has already delivered to you (in a private store in your own account) and bills only postings it has not seen before. Measured before this existed: two identical runs a day apart shared 47% of their rows โ a schedule was re-buying half its data every run. Memory is kept per search for 90 days.
Set up a daily new-jobs feed in 60 seconds โ the recurring setup most buyers actually want:
- Fill in your search term and location, tick Only new jobs (monitoring mode), hit Save & start once.
- On the run page choose Actions โ Schedule, pick daily (or hourly for hot markets).
- Done. Every run now delivers โ and bills โ only postings the previous runs have not seen. The first run seeds the memory; identical reruns bill zero (measured: 15 rows, then 0). Point the schedule's webhook at Slack, a Google Sheet via Zapier/Make, or your ATS, and it becomes a job alert service you own.
Remote jobs across eleven boards in one click โ turn on Remote jobs only and hit Save & start: LinkedIn, Indeed, Glassdoor and The Muse are asked for their own remote filters while the seven remote-only boards (Remotive, Jobicy, Himalayas, HN Who-is-Hiring, Remote OK, We Work Remotely, Working Nomads) join automatically. No other job-board actor on the Store combines the big boards' remote filters with the remote-only APIs in one run.
Built for recruiters & staffing agencies, job boards & aggregators, market researchers, and sales teams tracking hiring signals who want one consolidated jobs feed instead of running eight scrapers and reconciling them by hand.
What you get
- Eight working boards in one run โ LinkedIn, Indeed and Glassdoor selected by default, plus five keyless APIs you switch on (four of them join remote searches automatically). One schema, one dataset.
- Cross-board dedup that shows its work โ the same role on several boards becomes one row carrying
found_on_sitesandduplicate_count. Company names are normalized first, so Wipro and Wipro Limited collapse together. It merges across boards only: five genuinely different openings one employer posted to one board stay five rows, because you paid to scrape all five. You are billed for the surviving row only. - Salaries you can actually sort โ hourly, weekly and monthly pay is normalized into
salary_min_annual/salary_max_annualon every salaried row, no flag to find. - Filters that never bill you โ salary-required, title / company / city excludes, competitor-only, experience level, agency excludes and max posting age all run before the charge, and
RUN_SUMMARY.filter_removedcounts each one. - Five ready-made Output views โ Overview, With salary, Cross-board dedup, Remote & location, Standard table. No column wrangling on your first run (exports still carry every column).
- The same 50 columns on every row, every run โ a field a board did not publish is
null, never a missing column. Spreadsheet imports line up,pd.read_csvbehaves, positional parsers keep working. - Presets and quick-pickers โ pick a goal and the matching options switch themselves on, with every change (and every refusal) written to the log.
- Paste a search URL โ already built the search on LinkedIn, Indeed or Glassdoor? Drop the URL in and skip the form.
- Honest per-board reporting โ every run writes a
RUN_SUMMARYnaming each board's real outcome. Boards that are blocked at the source are labelled in the table below, not sold to you.
Measured, not promised
Every number below comes from a real run of this Actor. Inputs and dates are given so you can reproduce them.
| Run | Result |
|---|---|
| Untouched form, pressed Start (3 boards, 20/board) โ 2026-08-08 | 60 rows in 27 s โ LinkedIn 20, Indeed 20, Glassdoor 20. When Glassdoor is having one of its intermittent blocked spells you get ~40 rows from the other two, the status message says so, and the missing rows are never billed |
data analyst ยท New York, NY ยท 30/board, 3 boards โ 2026-08-08 | 90 raw listings โ 89 billed rows in 39 s; 1 duplicate merged across boards (Indeed 30, Glassdoor 30, LinkedIn 29) |
| Salary coverage on that same run | salary_min_annual filled on 74 of 89 rows (83%) |
| Stress run, 120/board, 3 boards โ 2026-08-07 | 258 rows in 226 s โ LinkedIn 118, Indeed 112, Glassdoor 28 |
| Remote-only run, the five API boards, 10 requested each โ 2026-08-08 | The Muse 10, Jobicy 10, HN 5, Himalayas 6, Remotive 2 โ per-board totals track each board's live inventory for your term and vary day to day |
| LinkedIn descriptions on a default run โ 2026-08-07 | 20 of 20 LinkedIn rows (detail fetching is on by default) |
Glassdoor tops out around ~28โ30 rows per query no matter the cap โ a board-side limit we report rather than hide. Every row of every run carries the same 50 columns; per-field fill rates are further down, under Output fields and how often they are filled.
First run
{"searchTerm": "data analyst","location": "New York, NY"}
That is a complete run. Everything else already has a working default: LinkedIn + Indeed + Glassdoor, 20 job postings per board, full LinkedIn job details on, Apify datacenter proxy (measured: same rows as residential at a fraction of the cost). Open the Actor, hit Save & start without touching anything, and you get rows.
Presets
Pick a goal in Preset and the matching options are switched on for you. A preset only fills options you left at their default โ anything you set yourself always wins โ and both the run log and RUN_SUMMARY.preset_applied list exactly what it changed and what it refused to change, with your conflicting value.
| Preset | What it switches on | Effect on your bill |
|---|---|---|
| Salary research | requireSalary, enforceAnnualSalary, sortBy: salary_desc | Lower โ rows carrying no pay are dropped before billing |
| Competitor hiring intel | includeCompanyDetails | Unchanged โ it adds columns, not rows. It does not narrow the run by itself: add the employers you are watching under Only these companies, or the run stays unfiltered and the log says so |
| Remote-only sweep | isRemote | Higher โ with the board list left untouched this also brings in the seven remote-only boards (Remotive, Jobicy, Himalayas, HN, Remote OK, We Work Remotely, Working Nomads), and extra boards add rows |
| Fresh postings only | maxAgeDays: 7, plus a 168-hour board-side window when that cannot clash with your job type / easy-apply / remote choices | Lower โ anything older is dropped before billing |
| Recruiter lead-gen | includeCompanyDetails, excludeAgencies | Lower โ staffing-agency reposts are dropped before billing |
Two quick-pickers sit beside it: Role (quick pick) and City (quick pick). Precedence is searchTerms > roleSelect > searchTerm and locations > citySelect > location โ a picker overrides the matching free-text box (and warns you in the log when it does), while the multi-value arrays override the picker. Picking a non-US metro also switches countryIndeed to that country unless you set it yourself. Leave all three shortcuts alone and the Actor behaves exactly as it always has.
What the Output tab looks like
Five curated views ship with the Actor. Pick one above the results table instead of scrolling a wall of 50 raw columns:
| View | Columns | Use it for |
|---|---|---|
| Overview | title, company, location, board, posted, company page, job link | the fast read of any run |
| With salary | title, company, min/max yearly pay, currency, job link | comp benchmarking โ pay is already annualized |
| Cross-board dedup | title, company, found on boards, listings for this role, job link | proof of the dedup, and a hiring-urgency signal |
| Remote & location | title, company, remote flag, location, board, job link | telling true remote from remote-ish |
| Standard table | the 16 core columns | the previous default view, field-for-field unchanged |
Three things to know, so nothing surprises you. No view shows every column โ not even Standard table, which is the original 16-column view kept byte-identical because existing API callers address it as ?view=overview. Every run emits 50 columns. description, emails, id and job_url_direct are listed by no view at all; salary_min_annual / salary_max_annual, duplicate_count and company_url are missing from Standard table but do appear in With salary, Cross-board dedup and Quick view respectively. Views are Console display only โ CSV, JSON and Excel exports and the Dataset API always hand you every column regardless of the view you are looking at. And views are column presets, not filters: With salary still lists rows whose board published no pay (those cells are simply blank) and Cross-board dedup still lists roles seen on a single board (listings for this role = 1). To actually remove rows, use requireSalary or isRemote โ and rows removed by a filter are never billed.
Board status โ read this first
We would rather tell you than let you find out on a paid run. Checked 2026-08-07, re-measured 2026-08-11: all boards work from Apify's default datacenter pool (52 of 54 rows vs residential on the same search, faster and far cheaper). Residential remains available in Proxy configuration if a board starts coming back short:
| Board | Status | Notes |
|---|---|---|
| Working (default) | Full job details fetched by default โ descriptions on every LinkedIn row | |
| Indeed | Working (default) | Richest company data (size, revenue, addresses) |
| Glassdoor | Working (default), intermittently blocked | Salary on ~100% of rows it returns, plus employer rating. Glassdoor 403s some runs entirely (upstream, comes back on its own within hours) โ when that happens the run status names it and the missing board bills nothing. Tops out around ~28-30 rows per query โ a board-side cap, not a bug |
| The Muse | Working (default since 2026-08-14) | The only extra board with real city coverage (server-side location filtering โ measured 12/12 success from the standard proxy pool, 20 rows/page). Carries no salary data โ 0 salary fields in 220 measured rows |
| Remotive | Working (remote-only) | Auto-joins remote searches; rows may not be re-posted to other job boards (see Licensing below) |
| Jobicy | Working (remote-only) | Auto-joins remote searches; credit Jobicy + keep the original apply links |
| Himalayas | Working (remote-only) | Auto-joins remote searches; link back to himalayas.app |
| HN "Who is hiring?" | Working (remote-only) | Startup jobs from the monthly Hacker News thread |
| Remote OK | Working (remote-only) | 100 newest remote listings, server-side tag filtering. Rows keep their remoteok.com link โ required attribution, see Licensing |
| We Work Remotely | Working (remote-only) | Category RSS feeds (25-100 rows); salary parsed from listing text where published (~56% of programming posts) |
| Working Nomads | Working (remote-only) | Curated remote list (~50 rows); no salary data |
| DevITjobs US | Working (opt-in) | US/Canada tech board, salary on ~100% of rows (annual, currency as published), experience level + tech stack on nearly every row |
| DevITjobs UK | Working (opt-in) | UK tech board, salary on ~100% of rows โ same shape as the US board |
| Google Jobs | Blocked at source | Google serves a bot-check page instead of results |
| ZipRecruiter | Blocked at source | API returns HTTP 403 |
| Bayt (MENA) | Blocked at source | HTTP 403 |
| BDJobs (Bangladesh) | Blocked at source | Redirects away from results |
| Naukri (India) | Blocked at source | Demands a CAPTCHA |
In a verified run (2026-08-08, python developer, remote only, 10 requested per board) the five API boards delivered: The Muse 10, Jobicy 10, HN "Who is hiring?" 5, Himalayas 6, Remotive 2 rows. Per-board counts track each board's live inventory for the term and move day to day (an earlier run saw Remotive 10 and Himalayas 9).
LinkedIn, Indeed and Glassdoor remain the default selection. The five API boards are opt-in, with one convenience: when Remote jobs only is on and the board list is left at its default trio, the four remote-only boards (Remotive, Jobicy, Himalayas, HN) join the run automatically โ RUN_SUMMARY.remote_boards_auto_added records exactly which ones were added. They are deliberately not added to location searches: for a query like nurse in Dallas they contribute nothing, so we don't run them. The Muse is the exception with real US city coverage, but its category-based search means every word of your search term must appear in the job's title or description โ keep Muse search terms short.
The blocked five stay selectable in case they recover, and they cost you nothing when they return no rows โ but do not plan a project around them today.
Every run writes a RUN_SUMMARY record to its key-value store with the exact per-board outcome, so you always know which board gave you what and why one was quiet. It also carries columns_per_row and the full columns list, so you can diff a run's schema without reading a single row. If every board you selected fails to fetch, the run is marked FAILED rather than quietly succeeding with an empty dataset.
Every successful run also sets a status message in the Console run header โ row count, column count, and which preset fired and what it changed. Previously that line only appeared when search terms had been dropped, so a run driven by a preset finished without ever saying which preset had run.
Why use this
- Eight boards, one run โ LinkedIn + Indeed + Glassdoor by default, plus five keyless job APIs (The Muse, Remotive, Jobicy, Himalayas, HN "Who is hiring?") merged into a single dataset.
- Paste a search URL โ already built the search on LinkedIn, Indeed or Glassdoor? Copy the URL into
searchUrlsand skip the form. Keywords, location, radius (LinkedIn and Indeed only), remote flag, recency, job type, experience level and sort are read straight out of the URL โ every honored parameter is listed below, and an unparseable URL is reported instead of quietly scraping something else. Note this replaces the search term, locations and board list only; a preset and your filters still apply on top. - Keyword ร location matrix โ
locationsruns every search term against every location and merges the results, which is how you get past a board's per-query result cap (one query for "United States" returns one capped page; five city queries return five). - Deduplicated, and it tells you โ the same job on multiple boards becomes one row carrying
found_on_sites(every board it appeared on) andduplicate_count. Company names are normalized before matching (legal suffixes stripped), so Wipro on one board and Wipro Limited on another collapse into one billed row. Being live on three boards is a real hiring-urgency signal. Within a single board two rows are merged only when they are the same listing (samejob_url, else sameid), so a company advertising six distinct Software Engineer II roles on Indeed gives you six rows, not one. Readduplicate_countas how many listings in this run shared this title and company, not as a count of rows we deleted: across boards those extras really were merged into the one billed row, while on a single board they each kept their own row, so the same number can legitimately appear on several rows. - Filters that never bill โ salary-required, job type, remote-only, title-keyword excludes, company excludes, company includes (competitor hiring watch), experience level, city excludes, staffing-agency excludes, strict keyword match and max posting age all run after dedup and before billing; a filtered row is never charged, and
RUN_SUMMARY.filter_removedsays exactly how many each filter took. - Score every job against your resume โ put your skills in
resumeKeywordsand each row gainsmatched_keywordsandkeyword_match_percent. Pure annotation: it never removes a row, never changes the bill, andsortBy: relevanceputs the best matches on top. - Comparable salaries โ hourly, weekly and monthly pay is normalized into
salary_min_annual/salary_max_annual(hourly ร2080, weekly ร52, monthly ร12), so one column sorts every salaried row. - Resilient โ a slow or blocked board is skipped on a timeout so the run still finishes with everything the others returned. No hung runs, no all-or-nothing failures.
- Honest per-board reporting โ
RUN_SUMMARYnames the board and the actual HTTP error, instead of a silent empty result. - Multi-search โ pass several search terms in one run; results are merged and deduped, and each row records which term matched it.
- Clean text โ descriptions are converted to real Markdown without the stray backslash escapes (
full\-time,Web3 \| NYC) that the underlying library emits by default. - Export anywhere โ CSV, JSON, Excel, or pipe to Google Sheets / your CRM.
How to use it
- (Shortcut) Already have the search open on LinkedIn, Indeed or Glassdoor? Copy the URL into Search URLs and hit Save & start โ everything below is filled in from the URL.
- Enter a search term (job title/keywords) and a location (or several under Locations, which runs every term against every location).
- Keep the default boards, and set Max job postings per board. For remote searches, turning on Remote jobs only auto-adds the four remote-only boards; add The Muse yourself for extra US-city coverage.
- Full LinkedIn job details is on by default โ that's where LinkedIn descriptions,
job_typeandcompany_industrycome from. Turn it off only for faster runs. - Optionally filter: remote-only, job type, posted-within-N-hours, distance, country โ plus the post-scrape filters below (salary required, title/company excludes, max posting age), which are never billed.
- Run โ get a clean, deduplicated jobs table.
Input
The Console form is grouped into sections that go from loudest to quietest โ Quick start, What to search, Paste a job-search URL, Filters & limits, Board-specific options, Output, and a deliberately plain Advanced (networking). The optional plumbing lower down is grouped at the bottom on purpose, because a default run needs none of it.
| Field | Type | Description |
|---|---|---|
preset | string | Optional goal shortcut: salary_research, competitor_intel, remote_only, fresh_postings, recruiter_leads. A preset only fills options you left at their default โ anything you set yourself wins โ and everything it changed (and everything it refused to change) is logged and recorded in RUN_SUMMARY.preset_applied. Empty (the default) is a strict no-op. |
roleSelect | string | Optional quick-pick job title. Precedence: searchTerms > roleSelect > searchTerm. Sent to the boards as plain keywords โ it is not a fixed taxonomy. Empty = use searchTerm. |
citySelect | string | Optional quick-pick metro. Precedence: locations > citySelect > location. Picking a non-US metro also sets countryIndeed to the matching country unless you set that field yourself. Empty = use location. |
searchUrls | array | Paste LinkedIn / Indeed / Glassdoor search URLs and skip the form. Replaces searchTerm, locations and sites when set. See the table of honored parameters below. |
searchTerm | string | Job title / keywords (e.g. software engineer). |
searchTerms | array | Multiple searches in one run (merged + deduped); overrides searchTerm. Capped at 5 โ extras are reported in the run status, never dropped silently. |
location | string | City / state / country (e.g. New York, NY). Empty = anywhere. |
locations | array | Several locations in one run; every search term is run against every location. Replaces location. Capped at 10 locations and 25 term ร location searches; caps are reported in RUN_SUMMARY. |
sites | array | Boards to scrape. Defaults to linkedin, indeed, glassdoor; the five API boards (muse, remotive, jobicy, himalayas, hn_hiring) are opt-in. Unrecognised names are reported, not silently ignored. |
maxResults | integer | Job postings requested per board (1โ500, clamped). 0 and negatives clamp to 1 rather than falling back to the default. |
isRemote | boolean | Asks each board for its own remote filter, then drops any row whose own is_remote is false before billing โ LinkedIn accepts f_WT and ignores it, so the board-side filter alone is not enough. Rows with no description carry no evidence and are kept. With the board list left at its default it also auto-adds the four remote-only boards, which adds rows and raises the bill. |
jobType | string | Full-time / part-time / internship / contract. Only Indeed and Glassdoor filter this board-side (LinkedIn and the five API boards ignore it), so rows whose own job_type contradicts your choice are dropped before billing. A posting with no job_type at all is kept. |
hoursOld | integer | Only jobs posted in the last N hours (board-side filter). Above 8760 (one year) the age limit is dropped and your other board-side filters are kept โ it does not silently strip easy-apply/job-type/remote the way any other non-zero value must. |
maxAgeDays | integer | Drop postings older than N days. Applied by the actor after scraping, so it works on every board; filtered rows are never billed. |
requireSalary | boolean | Keep only rows carrying a salary figure. Filtered rows are never billed. |
excludeTitleKeywords | array | Drop jobs whose title contains any of these words/phrases (case-insensitive). Filtered rows are never billed. |
excludeCompanies | array | Drop these companies. Names are normalized first, so Wipro also excludes Wipro Limited. Filtered rows are never billed. |
targetCompanies | array | Keep only these employers (competitor hiring watch). Same normalization, plus a parent match: Amazon keeps Amazon Web Services but not Amazonia. Filtered rows are never billed. |
experienceLevel | array | internship / entry / associate / mid_senior / director / executive. Best-effort โ see the note below. Filtered rows are never billed. |
excludeCities | array | Drop jobs whose location mentions any of these places (whole-word, case-insensitive). Filtered rows are never billed. |
excludeAgencies | boolean | Drop rows whose company name looks like a staffing agency or recruiter. Heuristic; see the note below. Filtered rows are never billed. |
strictKeywordMatch | boolean | Keep only rows whose title or description contains your search words. Off by default. Use it for narrow roles โ LinkedIn never answers "nothing matched" and returns loosely related postings instead. Filtered rows are never billed. |
resumeKeywords | array | Score every row against your skills โ adds matched_keywords and keyword_match_percent. Never filters, never changes the bill. |
sortBy | string | relevance, date_desc or salary_desc. Applied before the charge cap. |
countryIndeed | string | Country for Indeed & Glassdoor (e.g. usa, uk, india). Leave it alone and a location naming a country (Berlin, Germany) points Indeed at that country's site automatically. A spelling we don't recognise falls back to usa with a note in RUN_SUMMARY โ it never fails the run. |
distance | integer | Search radius in miles. LinkedIn and Indeed honor it; Glassdoor and the API boards ignore it. |
offset | integer | Skip the first N job postings per board, to page past a run you already have. LinkedIn, Indeed and Glassdoor page server-side; on the five API boards the Actor fetches N extra rows and hands you the tail. Boards order results their own way, so it is not a stable cursor. |
easyApply | boolean | LinkedIn/Indeed direct-apply jobs only. |
linkedinFetchDescription | boolean | On by default. Fetches each LinkedIn job's own page โ the source of LinkedIn descriptions, job type and industry. Turn off for faster runs. |
descriptionFormat | string | markdown or html. |
descriptionHtml | boolean | Also emit description_html (the board's original HTML) alongside description. Both in one run, no extra requests. |
includeCompanyDetails | boolean | Fill company_ceo, company_banner, company_addresses_all and company_details_filled, which are blank without it. The columns exist either way. No extra requests. |
proxyConfiguration | object | Proxy settings. Datacenter is the default and what you want (measured 2026-08-11: 96% of the rows at ~1/5 the cost); switch to Residential only if a board starts coming back short. |
Example input:
{"searchTerm": "data analyst","location": "Austin, TX","sites": ["linkedin", "indeed", "glassdoor"],"maxResults": 50,"hoursOld": 168,"requireSalary": true,"excludeTitleKeywords": ["senior", "intern"]}
Paste-a-URL input:
{"searchUrls": ["https://www.indeed.com/jobs?q=python+developer&l=Austin%2C+TX&fromage=14&radius=25","https://www.linkedin.com/jobs/search?keywords=python%20developer&location=Austin%2C%20Texas&f_TPR=r604800"],"maxResults": 25,"resumeKeywords": ["Python", "AWS", "Kubernetes"],"sortBy": "relevance"}
Multi-city + competitor watch input:
{"searchTerms": ["data analyst", "business analyst"],"locations": ["Austin, TX", "Chicago, IL", "Denver, CO"],"maxResults": 30,"targetCompanies": ["Deloitte", "Accenture", "Capgemini"],"includeCompanyDetails": true,"sortBy": "date_desc"}
Paste a search URL (searchUrls)
Build the search on the board's own site, copy the URL, paste it in. Each URL scrapes the board it came from โ a LinkedIn URL runs LinkedIn only โ and searchUrls replaces searchTerm / locations / sites for the run. Anything the URL contains that is not in this table (geoId, tracking ids, currentJobId, saved-search ids) is ignored.
| Board | URL shapes accepted | Parameters honored |
|---|---|---|
linkedin.com/jobs/search?keywords=โฆ | keywords โ search term, location, distance (miles), f_WT=2 โ remote only, f_TPR=r<seconds> โ posted-within, f_JT โ job type, f_E โ experience level, f_C โ company ids, sortBy (DD/R), start โ offset | |
| Indeed | indeed.com/jobs?q=โฆ (any country sub-domain, e.g. ca., uk., de.) | q โ search term, l โ location, radius, fromage (days) โ posted-within, sc attr(DSQF7) โ remote, sc jt(โฆ) โ job type, explvl โ experience level, sort=date, start โ offset, and the country from the sub-domain |
| Glassdoor | glassdoor.com/Job/โฆ-SRCH_IL.<a>,<b>_ICโฆ_KO<c>,<d>.htm or โฆ/Job/jobs.htm?sc.keyword=โฆ | sc.keyword / keyword (or the KO slug offsets) โ search term, locKeyword / locName (or the IL slug offsets) โ location, fromAge (days), remoteWorkType=1 โ remote, jobTypes, seniorityType โ experience level, sortBy, and the country from the domain suffix. radius is not read โ Glassdoor's scraper has no distance parameter, so a radius in the URL changes nothing (only LinkedIn and Indeed honor distance) |
A URL we cannot parse is skipped with a reason, never guessed at: the other URLs still run and RUN_SUMMARY.search_url_errors names the URL and what was missing. If none of them parse, the run scrapes nothing and bills nothing.
Each URL's filters stay on that URL. A seniority filter carried in one URL (f_E, explvl, seniorityType) is applied only to the rows that URL returned โ pasting an executive-only LinkedIn URL next to an unfiltered Indeed URL does not thin out the Indeed results. Those rows are removed before de-duplication and therefore before billing, and the count shows up in RUN_SUMMARY.filter_removed as experienceLevel (from search URL), with the per-URL levels in RUN_SUMMARY.search_url_experience_levels. Setting the experienceLevel input explicitly overrides every URL's own level filter.
Verified locally on 2026-08-08: an Indeed URL and the equivalent manual inputs (searchTerm + location + hoursOld: 336 + distance: 25) returned the same 10 job ids and the same output columns โ the URL path is a front door onto the same pipeline, not a different one. A Glassdoor SRCH_โฆKOโฆ slug URL returned 8 rows; glassdoor.com/Job/jobs.htm (no keyword anywhere) was rejected with a readable reason while the other URL still ran.
Keyword ร location matrix (locations)
locations runs every search term against every location and merges the results into one deduplicated table. It is the way past a board's per-query cap: one query for a whole country returns one capped page, five city queries return five. Every pair asks each board for maxResults rows, so cost scales with the number of pairs โ capped at 10 locations and 25 pairs per run, and both caps are reported (never applied silently). Rows carry matched_search_term and, whenever a run covers more than one location, matched_location.
Live run (2026-08-08, data analyst ร [Austin, TX, Chicago, IL] ร LinkedIn + Indeed + Glassdoor at 12/board): 72 raw rows โ 67 after cross-search dedup โ 63 billed after filters, split 28 Austin / 35 Chicago.
Post-scrape filters โ filtered rows are never billed
These filters run inside the actor after scraping and deduplication and before billing, so a filtered row is never charged. RUN_SUMMARY.filter_removed records how many rows each one took.
requireSalaryโ keep only rows carrying a salary figure (salary_minorsalary_max).jobTypeโ also enforced here, not just board-side. Indeed and Glassdoor filter employment type themselves; LinkedIn's guest search accepts the parameter and returns the same rows anyway, and the five API boards have no such filter at all. So a row whose ownjob_typecontradicts your choice is dropped before billing. A posting the board published with nojob_typeis kept โ silence is not a contradiction. Measured 2026-08-08:jobType: internshipon Remotive + Jobicy + Himalayas dropped all 35 mismatching rows and billed 0; on LinkedIn + Indeed it dropped LinkedIn's 10 full-time rows and billed 10 rows that all carryinternship.isRemoteโ also enforced here for the same reason: LinkedIn ignoresf_WT. Rows whose ownis_remoteisfalseare dropped before billing; rows with no description carry no evidence and are kept. Measured 2026-08-08 on the Remote-only sweep preset: 19 rows dropped, and 0 of the billed rows hadis_remote: false(before this, 10 of 10 LinkedIn rows billed under that preset were on-site jobs in Midland TX, El Paso TX and Pittsburgh PA).strictKeywordMatchโ opt-in, off by default. Keep only rows whose title or description contains your search words. LinkedIn never answers "nothing matched"; it degrades to loosely related cards. Measured 2026-08-08:underwater welderin Fargo, ND billed 10 rows of marine techs, trenchless engineers and welders with it off, and 0 with it on. Either wayRUN_SUMMARY.rows_not_mentioning_search_termcounts the off-topic rows per board (that run:{"linkedin": 10}), so you can see the problem before deciding to pay for it.excludeTitleKeywordsโ drop titles containing any listed word or phrase (case-insensitive substring match).excludeCompaniesโ drop listed companies; names are normalized before matching, soWiproalso excludesWipro Limited.targetCompaniesโ the inverse: keep only the listed employers. Same normalization plus a parent match, soAmazonkeepsAmazon Web Serviceswhile leavingAmazoniaout. This is the competitor-hiring-watch input.maxAgeDaysโ drop postings older than N days. UnlikehoursOldthis is applied by the actor itself, so it works on every board.excludeCitiesโ drop jobs whose location text mentions a listed place. Whole-word and case-insensitive, matched against the location exactly as the board published it โDallasdropsDallas, TX,Dalldrops nothing. Best-effort: boards format locations inconsistently.excludeAgenciesโ drop rows whose company name matches a staffing/recruiting pattern (staffing,recruit*,headhunt*,talent*,manpower,personnel,placement*,consultanc*,resourcing,workforce,temp agency,employment agency,search partners/group/associates,HR solutions,staff augmentation). It reads the name only, so a re-poster with a neutral name gets through and a genuine employer withWorkforcein its name gets dropped โ useexcludeCompanieswhen you need exact control.experienceLevelโ best-effort, and here is exactly how. When a board publishes a seniority field we use it (LinkedIn does with detail fetching on โ measured 19/19 LinkedIn rows; Jobicy, Himalayas and The Muse also publish one). Otherwise we read it off the title:Senior/Sr./Staff/Principal/Leadโ mid-senior,Junior/Jr./Graduate/Entry levelโ entry,Internโ internship,Directorโ director,VP/Chief/Head ofโ executive. A job whose level we cannot determine is kept, not dropped โ we only remove rows we can prove don't match. On a live 63-row run the level came from a board field on 19 rows, from the title on 16, and was genuinely unknown on 28.
Verified on live runs: requireSalary removed 7 rows and excludeTitleKeywords 3 (2026-08-07); excludeAgencies removed 4 rows and experienceLevel 1 (2026-08-08), all recorded in RUN_SUMMARY.filter_removed, and a targetCompanies value matching nothing ended the run with 0 rows pushed and nothing billed. excludeAgencies reads the company name only, so treat it as a heuristic rather than a measurement: across the 178 distinct employers returned by the 2026-08-08 local runs it flagged none, which tells you how few agency reposts a given search contains, not that it is precise. Its known failure mode is the workforce token โ Texas Workforce Commission and Workforce Software are both flagged, and both are genuine employers posting their own jobs. Use excludeCompanies when you need exact control.
Resume keyword scoring (resumeKeywords) and sortBy
Put your skills in resumeKeywords (Python, Kubernetes, SOC 2, C++, .NET, node.js โ case-insensitive substring, so punctuation works as written) and every row gains:
matched_keywordsโ which of your keywords appear in the title or description, in your orderkeyword_match_percentโ 0โ100
It is annotation only: no row is removed and the bill is unchanged. sortBy: relevance then ranks by match percent, then by how many boards the job was found on, then by date. sortBy: date_desc and salary_desc sort on posting date and on the annualized salary; rows with no value for the key always land last, and sorting runs before the charge cap so a truncated run keeps your top rows. Measured on a 63-row live run: 8 rows at 100%, 17 at 75%, 15 at 50%, 8 at 25%, 15 at 0%.
Output fields and how often they are filled
Every row has the same 50 columns
One column tuple per run โ and the same tuple in every run. Every row the Actor delivers carries all 50 columns below, in the same order, whatever board it came from and whatever options you set. A field the board did not publish is null, never missing. That is what makes the CSV safe to open in a spreadsheet, pd.read_csv without dtype surprises, and parse by column position.
This changed on 2026-08-08. Rows used to have their empty fields stripped out before delivery, so a column existed on the rows one board filled in and silently vanished from the rest โ a LinkedIn + Indeed + Glassdoor run of this shape once shipped 10 different column tuples ranging from 20 to 31 columns. Since the fix, every row of a run carries the same 50-column tuple, however many rows come back and whatever boards fill them. A live default LinkedIn + Indeed + Glassdoor run delivered one tuple of 50 columns on every row โ the row count is just whatever the boards return that day (57 the first time this was measured, 60 on the latest run) โ and separate runs across the five API boards and an Indeed-only run with every option switched on carried the same 50 columns in the same order.
The column order, which is stable and safe to depend on:
id, title, company, location, job_url, job_url_direct, site, job_type, is_remote, date_posted, salary_source, salary_min, salary_max, salary_currency, salary_interval, salary_min_annual, salary_max_annual, job_level, job_function, listing_type, company_industry, company_url, company_url_direct, company_addresses, company_description, company_num_employees, company_revenue, company_rating, company_reviews_count, company_logo, description, emails, skills, experience_range, vacancy_count, work_from_home_type, found_on_sites, duplicate_count, search_term, matched_search_term, scraped_at, description_html, company_ceo, company_banner, company_addresses_all, company_details_filled, matched_keywords, keyword_match_percent, matched_location, source_search_url.
The last nine are toggle-driven: description_html fills only with descriptionHtml: true, the four company extras and the counter only with includeCompanyDetails: true, the two keyword columns only with resumeKeywords, matched_location only on a multi-location run and source_search_url only on a searchUrls run. The column is there either way โ null means the toggle was off. A column that appears and disappears with a setting is the same ragged-CSV problem in miniature, so we don't do that. New fields are only ever appended to the end of this list, never inserted, so a positional parser keeps working.
Two things that can still hand you a narrower row, both of them your own choice: ?fields= / ?omit= on the Dataset API, and clean=true (a shortcut for skipHidden=true + skipEmpty=true), which strips empty values back out. Plain format=csv, format=xlsx, format=json and the Console Export button all give you the full 50.
Fill rates
Job boards publish wildly different amounts of detail, so most fields are conditional โ the column is always present, but how often it carries a value depends entirely on the board. These are measured fill rates, not aspirations โ taken from live runs on 2026-08-07 (58โ85 rows for the detail-off column, 67โ72 rows with LinkedIn detail fetching on). Ranges span two independent runs on different queries; expect your own numbers to land inside them.
LinkedIn detail fetching is now on by default, so the right-hand column is what a default run delivers โ a live default run ({}) returned 60 rows with descriptions on 20 of 20 LinkedIn rows. The left column is what you fall back to if you turn it off for speed.
| Field | Detail fetch off | Detail fetch on (default) | Comes from |
|---|---|---|---|
id, title, site, job_url, is_remote, date_posted, scraped_at, search_term, matched_search_term, found_on_sites, duplicate_count | 100% | 100% | all boards |
company | 98% | 95% | all boards |
location | 98% | 100% | all boards |
company_url | 97% | 94% | all boards |
description | 64% | 100% | Indeed + Glassdoor always; LinkedIn only with detail fetching |
salary_min, salary_max, salary_interval, salary_currency, salary_source | 56โ60% | 59โ70% | Glassdoor 95โ100%, Indeed 50โ80% (varies by query), LinkedIn none without detail fetching and ~35% with it |
company_logo | 49% | 86% | Glassdoor + Indeed; LinkedIn with detail fetching |
listing_type | 32% | 31% | Glassdoor only (organic / sponsored) |
job_url_direct | 31% | 33% | Indeed only (the employer's own posting URL) |
company_rating | 29% | 30% | Glassdoor only โ ~90% of Glassdoor rows |
job_type | 24% | 65% | Indeed ~80%; LinkedIn 100% only with detail fetching; never from Glassdoor |
company_url_direct, company_addresses, company_num_employees, company_revenue, company_description | 17โ20% | 19โ23% | Indeed only (~55โ70% of Indeed rows) |
job_level, job_function | 0% | 34% | LinkedIn only, and only with detail fetching (then 100%) |
company_industry | 2% | 38% | LinkedIn with detail fetching (100% of LinkedIn rows); Indeed ~10% |
emails | 5โ8% | 12โ14% | scraped out of description text when a job lists one |
linkedinFetchDescription is on by default, so job_type, company_industry, job_level and job_function arrive without touching anything. Turn it off and LinkedIn contributes none of them โ that trade is yours to make for speed.
Two derived columns normalize pay so every row sorts on one axis: salary_min_annual / salary_max_annual convert hourly (ร2080), weekly (ร52) and monthly (ร12) figures to annual. In a live default run they were present on 57 of 57 rows that carried a salary. Two guards keep them honest. When enforceAnnualSalary is on, the scraping library annualizes description-parsed pay itself but leaves salary_interval reading hourly/monthly โ we rewrite that interval to yearly instead of multiplying a second time, which is what used to ship salary_min_annual: 95180800 on a warehouse job and print $45,760 / hourly in the raw columns. And a result that cannot be a wage is left null rather than published; the raw salary_min / salary_max are never touched. That backstop is applied only to USD-scale currencies, so a real 3,000,000 KRW or 10,000,000 IDR monthly salary still annualizes normally. Measured 2026-08-08 on the salary_research preset (warehouse associate, Dallas TX, Indeed + LinkedIn): 0 of 14 billed rows carried an impossible annual figure and 0 claimed an hourly rate above $1,000/hour โ the same run shape previously produced 9 of 13. The three salary qualifier columns (salary_interval, salary_currency, salary_source) are blank on any row without an actual amount โ a qualifier with nothing to qualify is noise. Verified on the 2026-08-08 runs: 0 of 57 and 0 of 69 rows carried an orphan qualifier.
company_reviews_count, skills, experience_range, vacancy_count and work_from_home_type are only ever published by Naukri, which is currently CAPTCHA-blocked โ the columns are there on every row but expect them to be blank. They fill automatically if Naukri recovers.
includeCompanyDetails โ what you actually get, per board
Turning it on does one thing and invents nothing: it fills four columns that are otherwise blank. (Column stability is no longer its job โ every row of every run carries all 50 columns whether or not this is on.) It adds company_details_filled, a count of how many company fields the row actually got, and it emits three fields that Indeed already sends in its response and the underlying library throws away: company_ceo, company_banner and company_addresses_all (the full address list; company_addresses only ever held the first). No extra requests, no extra cost per row.
Measured on a live 63-row run (2026-08-08, LinkedIn + Indeed + Glassdoor, LinkedIn detail fetching on):
| Field | Indeed rows | LinkedIn rows | Glassdoor rows | Dataset-wide |
|---|---|---|---|---|
company_url | 96% | 100% | 100% | 98% |
company_logo | 71% | 100% | 90% | 86% |
company_industry | 33% | 100% | 0% | 43% |
company_num_employees (size band) | 75% | 0% | 0% | 29% |
company_url_direct (employer's own site) | 75% | 0% | 0% | 29% |
company_addresses / company_addresses_all | 71% | 0% | 0% | 27% |
company_description | 67% | 0% | 0% | 25% |
company_revenue | 58% | 0% | 0% | 22% |
company_ceo (new) | 54% | 0% | 0% | 21% |
company_banner (new) | 54% | 0% | 0% | 21% |
company_rating | 0% | 0% | 95% | 30% |
Read that as: Indeed is the company-data board (size, revenue, description, corporate website, addresses, CEO), LinkedIn contributes the industry (100% of LinkedIn rows, but only with Full LinkedIn job details on โ it was 0% in a run with it off), and Glassdoor contributes the employer rating. For the five API boards, which board a row came from decides this entirely, so the per-board answer is the only one worth quoting (measured 2026-08-08 on a 34-row remote-only run across all five):
| Column | Jobicy | Himalayas | The Muse | Remotive | HN |
|---|---|---|---|---|---|
company_industry | 10/10 | 0/7 | 0/10 | 0/2 | 0/5 |
company_logo | 10/10 | 5/7 | 0/10 | 2/2 | 0/5 |
job_level | 10/10 | 7/7 | 10/10 | 0/2 | 0/5 |
job_function | 0/10 | 0/7 | 0/10 | 2/2 | 0/5 |
In words: Jobicy is the only API board that publishes company_industry; Jobicy, Himalayas and Remotive publish company_logo; Jobicy, Himalayas and The Muse publish job_level; Remotive is the only one that publishes job_function; HN publishes none of them. Everything else on these boards is null. Don't read the dataset-wide percentage as a property of the Actor โ it is just each board's share of your particular run, and Jobicy is capped at 50 rows per call. If a board doesn't publish a field, the column stays null and says so; we never fill it from somewhere else.
descriptionHtml โ both formats in one run
With descriptionHtml: true every row carries description_html (exactly what the board published) alongside the normal description in your chosen descriptionFormat. There is no second request and no price change: the boards return HTML anyway, so the Markdown is produced from that same HTML locally. Use the HTML when you need the original links, lists and headings for your own site or ATS; use the Markdown for LLM prompts and spreadsheets. Budget for the size: description_html holds the raw HTML that the Markdown description is generated from, so it is necessarily about as large as the column beside it. On a 19-row A/B run (2026-08-08, same query with the switch off and on) description_html was filled on 19/19 rows and ran 111% of the description column, growing the whole dataset 86%. An earlier run measured +167% / +132%. Plan on roughly doubling your dataset, as the input form says โ not the "about 13%" an earlier version of this README claimed.
JSON output sample
A complete, unedited row from a live Glassdoor scrape (2026-08-08, software engineer / New York NY / LinkedIn + Indeed + Glassdoor, description truncated for length). This is the whole row โ all 50 columns, nulls included. Every other row in that run, from every board, had exactly these 50 keys in exactly this order:
{"id": "gd-1009849012184","title": "Software Development Engineer, ML Systems, Annapurna Labs","company": "Annapurna Labs (U.S.) Inc.","location": "New York, NY","job_url": "https://www.glassdoor.com/job-listing/j?jl=1009849012184","job_url_direct": null,"site": "glassdoor","job_type": null,"is_remote": false,"date_posted": "2026-02-28","salary_source": "direct_data","salary_min": 158100.0,"salary_max": 213800.0,"salary_currency": "USD","salary_interval": "yearly","salary_min_annual": 158100,"salary_max_annual": 213800,"job_level": null,"job_function": null,"listing_type": "organic","company_industry": null,"company_url": "https://www.glassdoor.com/Overview/W-EI_IE7470741.htm","company_url_direct": null,"company_addresses": null,"company_description": null,"company_num_employees": null,"company_revenue": null,"company_rating": 3.7,"company_reviews_count": null,"company_logo": "https://media.glassdoor.com/sql/7470741/amazon-web-services-squareLogo-1680754591940.png","description": "**DESCRIPTION** About Amazon Annapurna Labs ...","emails": null,"skills": null,"experience_range": null,"vacancy_count": null,"work_from_home_type": null,"found_on_sites": "glassdoor","duplicate_count": 1,"search_term": "software engineer","matched_search_term": "software engineer","scraped_at": "2026-08-08T15:52:57.101661+00:00","description_html": null,"company_ceo": null,"company_banner": null,"company_addresses_all": null,"company_details_filled": null,"matched_keywords": null,"keyword_match_percent": null,"matched_location": null,"source_search_url": null}
The trailing nulls are the toggle-driven columns: this run had descriptionHtml, includeCompanyDetails and resumeKeywords off, one location and no searchUrls. Turn those on and the same nine columns carry values โ they never move, appear or disappear.
Results render as a sortable table on the Output tab and export to CSV, JSON, or Excel.
Example output
A real sample from a live run (software engineer, New York NY, three boards):
| title | company | site | company_rating | salary_min | job_url |
|---|---|---|---|---|---|
| Software Engineer, Systems ML - Compilers | Meta | glassdoor | 3.7 | 154003 | https://www.glassdoor.com/job-listing/j?jl=101021โฆ |
| Silicon Software Lead | Normal Computing Corporation | indeed | https://www.indeed.com/viewjob?jk=76bcfโฆ | ||
| AI Modernization Senior Lead Software Eโฆ | JPMorganChase | indeed | 133000 | https://www.indeed.com/viewjob?jk=74806โฆ | |
| DevOps Engineer | OnMed | https://www.linkedin.com/jobs/view/โฆ |
Use cases
- Recruiting & staffing โ pull every open role for a title/location across boards into one pipeline.
- Job boards & aggregators โ backfill and keep a niche board fresh from multiple sources.
- Hiring-signal sales โ a role live on three boards at once is a company hiring urgently;
duplicate_countandfound_on_sitessurface exactly that. - Labor-market research โ salary ranges, remote share, and demand by title and location.
- Personal job hunt โ one deduped feed instead of refreshing three sites.
Use with AI agents & automation
Run from the Apify MCP server so AI agents (Claude, ChatGPT, Cursor) can pull jobs as a tool call, schedule runs via Make, n8n, or Zapier to alert on new postings, or sync the dataset to Google Sheets for a live jobs dashboard. Clean flat JSON drops into ATS/CRM pipelines with no glue code.
Licensing & attribution โ what you may do with the rows
The API boards publish their feeds with strings attached, and those strings pass through to you:
| Board | May you re-post rows on another job board? | Required credit |
|---|---|---|
| The Muse | Yes (attribution requested) | Credit The Muse |
| Remotive | No โ their ToS explicitly forbid it | Credit Remotive |
| Jobicy | Only with credit, and apply buttons must link to the original job URL | Credit Jobicy + keep the job_url links |
| HN "Who is hiring?" | User-authored HN content; link back to the HN item | Credit Hacker News / Algolia |
| Himalayas | Yes, with a link-back | Link to himalayas.app |
| Remote OK | Only with a visible, clickable (dofollow) link back to the row's job_url and "Remote OK" named as the source โ their terms suspend API access otherwise. Logo use is forbidden | Dofollow link to remoteok.com + name "Remote OK" |
| We Work Remotely | Their terms carry no republication bar; a link-back is good practice | Link to the job_url |
| Working Nomads | No stated bar; link-back is good practice | Link to the job_url |
| DevITjobs US/UK | Undocumented public API โ treat rows as pointers and keep the job_url links. ~99% of rows are partner-syndicated (e.g. from Indeed), so salary figures may be partner-estimated | Keep the job_url links |
In plain terms: Remotive rows are fine for lead-gen, research and your own job hunt, but must not be republished on another job board (their terms name Jooble, Google Jobs, LinkedIn and the like explicitly). If you build a job site on Jobicy data, keep the job_url apply links and credit Jobicy. The job_url column already points at each board's original posting, which covers the link-back half of these requirements.
API examples
Same input keys as the Console form, from code:
const { ApifyClient } = require('apify-client');const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });const run = await client.actor('flash_scraper/multi-jobboard-scraper').call({searchTerm: 'data analyst', location: 'Austin, TX', maxResults: 50, onlyNewJobs: true,});const { items } = await client.dataset(run.defaultDatasetId).listItems();
from apify_client import ApifyClientclient = ApifyClient("YOUR_APIFY_TOKEN")run = client.actor("flash_scraper/multi-jobboard-scraper").call(run_input={"searchTerm": "data analyst", "location": "Austin, TX", "maxResults": 50, "onlyNewJobs": True,})items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
curl -X POST "https://api.apify.com/v2/acts/flash_scraper~multi-jobboard-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \-H 'Content-Type: application/json' \-d '{"searchTerm":"data analyst","location":"Austin, TX","maxResults":50}'
Wire the finished-run webhook into n8n, Make or Zapier to land rows in Slack, Sheets or your ATS
automatically - combined with onlyNewJobs on a schedule, that is a self-maintaining job feed.
Only need LinkedIn, as cheap as possible? The same publisher's LinkedIn Jobs Scraper is the single-board budget tool - priced in the Store's cheapest tier.
Pricing
Pay-per-event. You are charged one Result event per row delivered to your dataset, plus a one-off Actor Start event. Because de-duplication happens before the push, you pay once for a job found on three boards, not three times โ and that holds across a keyword ร location matrix too, so the same role returned by two city searches is one billed row. Boards that return nothing cost nothing, rows removed by any post-scrape filter (requireSalary, excludeTitleKeywords, excludeCompanies, targetCompanies, experienceLevel, excludeCities, excludeAgencies, maxAgeDays) are never billed, and failed runs deliver no rows and so bill no results. includeCompanyDetails, descriptionHtml and resumeKeywords fill columns that are otherwise blank; they never add a row, so they cost nothing extra. See the Apify Store page for current prices.
Cost note for locations: each term ร location pair is a full search on every selected board, so 3 terms ร 3 locations ร 3 boards can bill up to 9ร a single search before de-duplication. The caps (10 locations, 25 pairs) exist for exactly that reason and are reported in RUN_SUMMARY, never applied silently.
If you set a maximum total charge on the run and the scrape exceeds it, the actor delivers as many rows as the budget allows and says so in the run status โ it will not spend your budget and then hand back an empty dataset.
FAQ
Where does the data come from? Public job listings on LinkedIn, Indeed and Glassdoor via the open-source JobSpy engine plus our own fixes on top of it, and the public keyless APIs of The Muse, Remotive, Jobicy, Himalayas and Hacker News (Algolia) for the five extra boards.
Do I need proxies? Yes โ the actor is preconfigured to use Apify's default datacenter pool, which handles LinkedIn / Indeed / Glassdoor (measured 2026-08-11: 52 rows vs residential's 54 on the same search, at ~1/5 the cost). The five API boards need no special proxy at all.
Why is a board returning nothing? Check the RUN_SUMMARY record in the run's key-value store โ it names each board and the actual error. Five of the thirteen selectable boards are currently blocked at the source (see the table at the top).
Why didn't the remote boards run on my location search? By design. Remotive, Jobicy, Himalayas and HN "Who is hiring?" carry remote jobs only โ they add nothing to a query like nurse in Dallas, so they only run when you name them in sites, or automatically when Remote jobs only is on and the board list is at its default. RUN_SUMMARY.remote_boards_auto_added tells you when the auto-add happened.
Why does The Muse return few or no rows for long search terms? The Muse is searched by job category, then filtered so that every word of your search term appears in the job's title or description. Short terms match; long multi-word phrases starve. Keep Muse terms to the core skill words.
How many rows can one run deliver? Up to 500 per board (maxResults). A stress run at 120 per board delivered 258 rows in 226 seconds: 118 from LinkedIn (full descriptions on all 118), 112 from Indeed, 28 from Glassdoor. Glassdoor tops out around ~28โ30 rows per query no matter the cap โ a board-side practical limit, not a bug or a budget issue.
Why are job_type / company_industry empty on LinkedIn rows? They shouldn't be โ Full LinkedIn job details is on by default and fills them. If you turned it off, that's why: LinkedIn's search results don't include them.
Can I search multiple titles at once? Yes โ use searchTerms (an array). Results are merged and deduplicated, and every row's matched_search_term records which term found it. Capped at 5 per run; any extras are named in the run status message.
Why did a job I expected not appear? Across boards, dedup keys on title + normalized company name (legal suffixes are stripped, so Wipro and Wipro Limited match), so the same role found on LinkedIn and Indeed arrives once with duplicate_count: 2. Within one board nothing is merged unless it is literally the same listing, so same-titled but distinct openings all survive. Otherwise check RUN_SUMMARY.filter_removed โ a filter you set (job type, remote-only, experience level, max ageโฆ) may have dropped it, in which case you were not charged for it.
Can I just paste the search URL from the job board? Yes โ put it in searchUrls. LinkedIn, Indeed (any country sub-domain) and Glassdoor are supported, and the exact parameters we read out of each URL are listed above. A URL from another site, or one with no keyword in it, is skipped with a written reason in RUN_SUMMARY.search_url_errors rather than silently scraping something else.
Why did experienceLevel keep a job that isn't at that level? Because we could not prove otherwise. Only some boards publish a seniority field (LinkedIn does with detail fetching on); for the rest we read the title, and a plain "Software Engineer" genuinely says nothing about seniority. Rather than drop rows on a guess, we keep the unknowns โ dropping them would hide real jobs from you. Add excludeTitleKeywords if you want a hard cut.
Does resumeKeywords change what I pay? No. It only fills two columns (matched_keywords, keyword_match_percent) on rows you were getting anyway โ the columns are on every row regardless. Neither it nor includeCompanyDetails nor descriptionHtml adds a request or a row.
Why are so many cells empty? Because the boards did not publish those values, and we would rather show you an honest blank than invent one. Every row carries all 50 columns so your CSV lines up; the Fill rates table above says exactly which board fills which column and how often. Nothing is filled in from a second source.
Can I export to CSV or Google Sheets? Yes โ CSV, JSON, or Excel from the Output tab, or sync to Google Sheets via Make, n8n, or Zapier. Exports carry all 50 columns. The one thing that will narrow them is asking for it: ?fields= / ?omit=, or clean=true on the Dataset API (it is a shortcut for skipHidden=true + skipEmpty=true, so it strips empty cells back out and the rows go ragged again).
Other Flash Scrape scrapers
- LinkedIn Jobs Scraper โ single-board LinkedIn postings at volume pricing
- ATS Job Scraper โ live openings straight from any company's Greenhouse, Lever or Ashby board
- Remote Job Aggregator โ RemoteOK, WeWorkRemotely + 8 more remote boards, deduplicated
- Google Maps Leads Scraper โ local business leads
- Trustpilot Reviews Scraper โ company reviews