Dice Jobs Scraper - US Tech Jobs & Salary Data avatar

Dice Jobs Scraper - US Tech Jobs & Salary Data

Pricing

from $2.00 / 1,000 job returneds

Go to Apify Store
Dice Jobs Scraper - US Tech Jobs & Salary Data

Dice Jobs Scraper - US Tech Jobs & Salary Data

Scrape Dice US tech job listings into a clean JSON or CSV job dataset - software, developer, data, DevOps and IT roles. Salary parsed to the cent with currency and period, normalised regions, employment type, seniority, and ghost-job scoring on every row.

Pricing

from $2.00 / 1,000 job returneds

Rating

0.0

(0)

Developer

Dave Fergins

Dave Fergins

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

17 hours ago

Last modified

Share

Dice Jobs Scraper

Scrape Dice US tech job listings into a clean JSON or CSV job dataset - software, developer, data, DevOps and IT roles. Salary parsed to the cent with currency and period, normalised regions, employment type, seniority, and ghost-job scoring on every row.

Why this board is different from the others here

Dice states salary properly, and almost nobody else does. Most job boards publish a free-text blob and leave you to guess whether "$60 - $70" is per hour or per year. Dice publishes the currency and the period explicitly — USD 119,000.00 - 146,000.00 per year, USD 27.52 per hour — which makes these the best-specified salaries of any source in this family.

This Actor parses them to the cent. That is not a detail: on an hourly contract rate, rounding 27.52 to 28 overstates a year's pay by roughly a thousand dollars, and filtering on a wrong number is worse than filtering on none. Where Dice renders a rate without stating a period, the numbers are deliberately left unparsed rather than guessed at, and you still get the original text in salaryRaw.

Access is by the paths Dice allows. Dice's robots.txt disallows keyword search (/jobs?q*) and explicitly allows /jobs and /job-detail. This Actor reads only the allowed paths, identifies itself honestly in its User-Agent, and does not spoof a browser or defeat a bot check. That is a deliberate constraint, and it is why there is no keyword parameter that filters at the source — filtering happens on the rows after they arrive.

Measured on live data 2026-08-15: 327 listings at the default page budget, with 100% company, posted-date and description coverage and 96% of rows resolving to a hiring region. Roughly half carry a published salary.

Dice is a US tech board and most roles are on-site or hybrid. Remote is reported from the board's own isRemote flag and workplace types, and a listing counts as remote only when Dice says so — never merely because it failed to say otherwise.


What each row contains

Identitystable id across runs, posting URL, direct applyUrl where published
The roletitle, company, plain-text description, tags, employment type, seniority
Whereremote flag, the board's own location text, and normalised regions (worldwide / usa / canada / latam / uk / europe / apac / africa / middle_east)
Paymin, max, currency and period as numbers, annualised on request
Whenposted date, expiry where published, age in days
Trustfreshness 0-1, ghostRisk low/medium/high, and ghostReason explaining the verdict in plain words

The trust fields

Job boards are full of postings still published but no longer open — filled roles left up for pipeline, evergreen "talent pool" adverts, and listings nothing ever expires. Every row here carries a risk band and the reasons behind it:

{
"title": "Senior Backend Engineer",
"company": "Acme",
"freshness": 0.71,
"ageDays": 10.4,
"ghostRisk": "low",
"ghostReason": ["recent, and nothing contradicts it"]
}

Signals come from what the board actually publishes: age, its own expiry date where there is one, missing application links, and evergreen phrasing. Nothing is inferred by a model. Freshness decays on a 21-day half-life, and a posting with no date scores 0.5 rather than 1.0 — absence of evidence is not evidence of freshness.

Example input

{
"query": ["golang", "backend"],
"regions": ["usa"],
"seniority": ["senior", "lead"],
"maxAgeDays": 21,
"maxGhostRisk": "low",
"maxItems": 200
}

Everything is optional — run it empty and you get the whole board, best-first.

  • Search terms are OR-ed. ["go", "rust"] returns jobs mentioning either.
  • Worldwide jobs match every region filter, because a job open to everyone is open to you.
  • minSalary is annualised first, so hourly and monthly rates compare correctly.
  • maxItems is your cost ceiling — you are billed per job returned.

Output

One dataset item per job, ordered best-first: lowest ghost-job risk, then freshest. Export as JSON, CSV or Excel, or read it from the API like any Apify dataset.

Common uses

  • US tech and engineering hiring data. Dice is a US technology board — software, developer, data, DevOps and IT roles — so it is the source here for American tech job market data rather than global remote listings.
  • Salary benchmarking and compensation data. Dice states pay more explicitly than the other boards here, parsed to the cent, so contract hourly rates and salaried roles compare correctly once minSalary annualises them.
  • Contract vs full-time analysis. Dice publishes an employment type and this Actor normalises it, so contract, part-time and permanent roles can be separated rather than eyeballed from the title.
  • Filtering out expired and stale job postings. ghostRisk, ghostReason and freshness say which listings are probably no longer open, so you can drop them before anyone wastes an application.
  • Job market and hiring data. Run it on a schedule and track how hiring, salary ranges, seniority mix and regions move over time.
  • Building a job board, job feed or job alerts. id is stable across runs, so diffing today's dataset against yesterday's gives you genuinely new jobs rather than a board reshuffle.
  • Recruitment and talent research. Company, title, tags, seniority, employment type and normalised regions on every row.

Export as JSON, CSV or Excel, or read the dataset straight from the Apify API.

How it fetches

Dice publishes no open API, so this Actor reads its server-rendered listing pages — but only the paths robots.txt allows (/jobs, /job-detail), never the ones it disallows (/jobs?q*, /jobsearch/). Requests are paced, budgeted and sent with an honest user agent. No browser spoofing, no bot-check evasion, no personal data. Job adverts only.

How to scrape Dice jobs with this Actor

  1. Open the Actor and press Start. With no input it returns every job Dice currently lists, best-first.
  2. Narrow it when you need to: search terms, hiring regions, seniority, a minimum salary, a maximum ghost-job risk, and maxItems as a hard cost ceiling.
  3. Export the dataset as JSON, CSV or Excel, read it through the Apify API, or schedule the run and attach a webhook or an integration (Google Sheets, Make, Zapier, Slack) so new Dice jobs arrive on their own.

How much does it cost to scrape Dice?

You pay per job returned, on Apify's pay-per-event model: $0.006 per job on the Free and Bronze plans, $0.003 on Silver, $0.002 on Gold and above, with no platform usage charged on top. A run that returns nothing costs nothing. The default 10 pages is roughly 330 jobs, about $1.98 on the Free tier; raise maxPagesPerSource for more. Set maxItems to cap a run: 200 jobs cost at most $1.20, and Apify's free plan includes a monthly usage credit, so a first run costs nothing out of pocket.

All eight share one schema, one ghost-job model and one set of filters, so a query written for this board runs unchanged against any of the others.

FAQ

Does Dice have an official API, and is it allowed to scrape it? Dice publishes no open API, so this Actor reads its server-rendered listing pages — but only the paths robots.txt allows (/jobs, /job-detail), never the ones it disallows (/jobs?q*, /jobsearch/). Requests are paced, budgeted and sent with an honest user agent. No browser spoofing, no bot-check evasion, no personal data. Job adverts only.

How often does Dice post new jobs, and how often should I run this? New listings appear every day. A daily schedule keeps the dataset current, and because id is stable across runs, diffing today's dataset against yesterday's yields only the genuinely new Dice jobs, not a reshuffle.

How do I get only fresh, still-open Dice jobs? Set maxGhostRisk to low and maxAgeDays to 14 or 21. Every row also carries ghostReason, the plain-language grounds for its rating, so you can apply your own rule instead.

What is a ghost job? A posting that is still published but no longer open: a filled role left up for pipeline, an evergreen "talent pool" advert, or a listing on a board that never expires anything. The ghostRisk band and freshness score flag them from what the board itself publishes — age, expiry dates, missing application links, evergreen phrasing — never from a model's guess.

Does it include salaries? Yes, the best-specified salaries in this family: Dice states the currency and the period explicitly, and they are parsed to the cent into salaryMin, salaryMax, salaryCurrency and salaryPeriod. Roughly half of listings carry pay.

Do I need proxies, a login or an API key? No. Dice is read the way it asks to be read, and Apify runs the Actor for you; nothing to install and no credentials to manage.

Can I export Dice jobs to CSV, Excel or Google Sheets? Yes. Every run's dataset downloads as JSON, CSV or Excel from the Apify Console or API, and the Google Sheets integration writes rows straight into a sheet.

Want more than one board?

Remote Jobs Aggregator runs this same pipeline across six boards at once — Remote OK, Remotive, Himalayas, Arbeitnow, We Work Remotely and Jobicy — folding duplicates and recording which boards carry each job.

Measured while building it: those six boards barely overlap. Only 1 of 1,513 company+title pairs appeared on more than one. The aggregator is not about removing duplication — there is almost none — it is about getting six boards' worth of distinct jobs from one call.