# Dice Scraper — Tech Jobs by Keyword & Location (`scrapersdelight/dice-scraper`) Actor

Scrape Dice.com tech job postings by keyword & location: title, company, company profile URL, location, posted date, employment type and salary. Real server-side filtering & pagination. No login, no cookies.

- **URL**: https://apify.com/scrapersdelight/dice-scraper.md
- **Developed by:** [Scrapers Delight](https://apify.com/scrapersdelight) (community)
- **Categories:** Jobs, Lead generation, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.54 / 1,000 per job returneds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## Dice Scraper — Tech Jobs by Keyword & Location

Scrape **Dice.com** tech job postings by keyword and location — **no login, no cookies, no Dice account**. Point it at a keyword (and optionally a location), and it walks Dice's public results pages and returns one clean, deduplicated row per job.

### What you get (one row per job)

| Field | Example |
|-------|---------|
| `title` | `Senior Python Developer with Snowflake` |
| `company` | `PETADATA` |
| `company_url` | `https://www.dice.com/company-profile/…` |
| `location` | `Austin, Texas` / `Remote` / `Hybrid in Newark, New Jersey` |
| `posted_raw` | `Today` / `13d ago` |
| `employment_type` | `Full-time, Third Party` / `Contract` |
| `salary_raw` | `140,000 - 200,000` / `Depends on Experience` |
| `salary_min` / `salary_max` / `salary_currency` / `salary_period` | `140000` / `200000` / `USD` / `year` (parsed when Dice prints a figure) |
| `is_easy_apply` | `true` |
| `job_url` | `https://www.dice.com/job-detail/<guid>` |
| `results_count` | `124` (total matches Dice reports for the query) |
| `search_keywords` / `search_location` / `scraped_at` | your query + capture time |

Rows are **deduplicated by Dice's job GUID** (the id in the `/job-detail/<guid>` URL), so you are never charged twice for the same posting.

### Input

```json
{
  "keywords": "python developer",
  "location": "Austin",
  "maxItems": 50,
  "proxyConfiguration": { "useApifyProxy": true }
}
```

- **keywords** — job title / skill / tech, as you would type it into Dice.
- **location** — city, state or `Remote`. Leave empty to search everywhere.
- **maxItems** — stop after N jobs (`0` = every job Dice serves; Dice returns ~30-36 per page).
- **proxyConfiguration** — Apify Proxy (datacenter) is enough; this actor was measured at 17/17 successful calls on standard datacenter proxies, so residential is **not** required.

### Notes on data quality (measured, not assumed)

- `title`, `company`, `location`, `posted_raw`, `employment_type`, `job_url` and the GUID come back on **~100%** of rows.
- `salary_raw` is present on ~80% of rows, but Dice frequently prints **"Depends on Experience"** instead of a number — so the parsed `salary_min`/`salary_max` fill on roughly half of rows. That is Dice's data, not a scraper bug.
- `company_url` is present on ~96–98% of rows (a few listings carry no company-profile link).

### Legality & responsible use

This actor reads only Dice's **public, server-rendered** job-search pages — no login, no CAPTCHA solving, no anti-bot bypass. Scraping may be restricted by **Dice's Terms of Service** regardless of technical access; whether your use is permitted is your decision. Any **personal data** in the results (for example a recruiter or contact name) is yours to process lawfully and in line with applicable privacy law (GDPR/CCPA). Use the data responsibly and at your own risk.

# Actor input Schema

## `keywords` (type: `string`):

Job title, skill or tech, exactly as you would type it into Dice — e.g. 'python developer', 'devops', 'salesforce'. Leave empty only if you set a location.

## `location` (type: `string`):

City, state or region as Dice shows it — e.g. 'Austin', 'New York', 'Remote'. Leave empty to search everywhere.

## `maxItems` (type: `integer`):

Stop after this many jobs. Set 0 for every job Dice will serve for the query (Dice returns ~30-36 per page).

## `proxyConfiguration` (type: `object`):

Apify Proxy is enough — this scraper was measured at 17/17 successful calls on standard datacenter proxies. Residential is not required.

## Actor input object example

```json
{
  "keywords": "python developer",
  "location": "Remote",
  "maxItems": 50,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `jobs` (type: `string`):

Title, company, company URL, location, posted date, employment type, salary (min/max/currency/period where published), easy-apply flag, job URL and the query's total result count.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keywords": "python developer",
    "location": "Remote",
    "maxItems": 50
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapersdelight/dice-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keywords": "python developer",
    "location": "Remote",
    "maxItems": 50,
}

# Run the Actor and wait for it to finish
run = client.actor("scrapersdelight/dice-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keywords": "python developer",
  "location": "Remote",
  "maxItems": 50
}' |
apify call scrapersdelight/dice-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrapersdelight/dice-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/yebuFidnJkMVxhBVu/builds/qy9G6EyiK2JCidENU/openapi.json
