# Naukri Scraper (`mlg14/naukri-scraper`) Actor

Scrape public Naukri public job listings and job identifiers. Export structured jobs, source URLs, IDs, and available details to CSV, JSON, or Excel for India job-market tracking and vacancy discovery.

- **URL**: https://apify.com/mlg14/naukri-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** Jobs, Automation
- **Stats:** 2 total users, 1 monthly users, 50.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Naukri Scraper

Scrape Naukri public job listings and job identifiers from public pages and export Naukri data to CSV, JSON, or Excel. This Naukri API alternative turns repeatable inputs into structured jobs for India job-market tracking and vacancy discovery.

Each dataset row represents a public job or a documented alternate record type. The output includes source identifiers and links where available, so you can verify a row, compare later runs, and distinguish missing source data from an omitted column.

### What data can you extract from Naukri?

The dataset schema defines every field below. Examples come from one successful published run; `null` means that field was not exposed for that particular record. Availability may differ by input mode, page, and record type.

| Field | Description | Example |
| --- | --- | --- |
| `jobId` | Source job identifier. | `230926010009` |
| `jdURL` | Direct job listing URL. | `https://www.naukri.com/job-listings-staff-software-developer-node-js-react-aw…` |
| `title` | Job title when available from the job endpoint. | `null` |
| `companyName` | Employer name when available. | `null` |
| `location` | Job location when available. | `null` |
| `experience` | Required experience when available. | `null` |
| `salary` | Published salary when available. | `null` |
| `currency` | Salary currency when available. | `null` |
| `createdDate` | Listing creation date when available. | `null` |
| `jobDescription` | Full description when detailed data is available. | `null` |
| `tagsAndSkills` | Published tags and skills when available. | `null` |
| `companyJobsUrl` | Employer jobs page when available. | `null` |
| `companyId` | Source employer identifier when available. | `null` |
| `logoPath` | Employer logo URL when available. | `null` |
| `logoPathV3` | Alternate employer logo URL when available. | `null` |
| `ambitionBoxData` | Public employer rating data when available. | `null` |
| `footerPlaceholderLabel` | Additional listing label when available. | `null` |
| `footerPlaceholderColor` | Display color of the additional label when available. | `null` |
| `isSaved` | Saved status when supplied by the endpoint. | `null` |
| `showMultipleApply` | Multiple-apply display flag when available. | `null` |
| `groupId` | Source group identifier when available. | `null` |
| `isTopGroup` | Top-group flag when available. | `null` |
| `exclusive` | Exclusive-listing flag when available. | `null` |
| `brandingTags` | Listing branding labels when available. | `null` |
| `mode` | Source listing mode when available. | `null` |
| `board` | Source job board when available. | `null` |
| `source` | Origin of the record: jobApi, sitemap, or inputUrl. | `sitemap` |
| `sitemapLastModified` | Last modification timestamp from the sitemap; this is not a posting date. | `2026-09-26T02:24:23.900+05:30` |

Use the identifier and URL fields as your join keys before comparing snapshots. Fields that describe a page, search, category, author, or tournament establish where the record came from; keep them when you export a subset. Numeric values and boolean flags reflect the public page at collection time, not a permanent claim about the underlying item.

### How to scrape Naukri

1. Open the actor input form and choose a narrow public source: a search phrase, page URL, item URL, or ID supported by the input fields below.
2. Set the relevant per-source page or result limit and the total item limit. Start with a small sample to check which optional fields the public source exposes.
3. Run the actor. Its dataset contains one structured row per saved result; inspect the first rows and their source links for the chosen input.
4. Download the dataset as JSON, CSV, or Excel. Preserve identifiers when combining runs so repeated results can be deduplicated.

For this actor, a keyword or search URL discovers jobs when the search endpoint is open. Job IDs or URLs take precedence. When a captcha blocks the endpoint, the sitemap can still return recent job IDs and URLs but not the missing job attributes.

### Input

Use only the parameters relevant to your collection mode. An omitted optional filter uses the schema default shown here; an empty array, zero, and an omitted value can have different meanings, so keep intentional settings in your saved input.

| Parameter | Type | Default | Description |
| --- | --- | --- |
| `keyword` | `string` | No default; example `software developer` | Words to find in job listings. A search URL or job IDs take precedence. |
| `location` | `string` | None | Optional city or region to narrow the search. |
| `searchUrl` | `string` | None | A naukri.com search results URL. Its keyword and location take precedence over the separate fields. |
| `jobIds` | `array` | None | Optional 12-digit job IDs or full job URLs to look up directly. Takes precedence over search. |
| `maxJobs` | `integer` | `100` | Maximum number of unique jobs to save. |
| `fetchDetails` | `boolean` | `false` | Request full job details when the source endpoint is available. Captcha-protected results remain partial. |
| `sortBy` | `string` | `relevance` | Sort search results by relevance or date when the search endpoint is available. |
| `experience` | `string` | `all` | Years of required experience for the search endpoint. Use all for no filter. |
| `freshness` | `string` | `all` | Maximum posting age in days for the search endpoint. Use all for no filter. |
| `workMode` | `array` | `[]` | Optional work modes for the search endpoint. |
| `proxyConfiguration` | `object` | `{"useApifyProxy":true}` | Proxy configuration for fetching source pages. |

Example input based on the published golden run (long URL lists are shortened):

```json
{
  "keyword": "software developer",
  "maxJobs": 40,
  "fetchDetails": false
}
```

The example is a starting shape, not a guarantee of a particular result count. Source inventory and page accessibility change. When you need repeatable comparisons, save the exact input JSON with the run date and inspect the returned source or record-type field.

### Output example

The following is one real item from a successful published dataset. Long text and media arrays are shortened for readability; the actual dataset keeps the original values and all schema fields.

```json
{
  "jobId": "230926010009",
  "jdURL": "https://www.naukri.com/job-listings-staff-software-developer-node-js-react-aws-ai-diligent-bengaluru-8-to-13-years-230926010009",
  "source": "sitemap",
  "sitemapLastModified": "2026-09-26T02:24:23.900+05:30"
}
```

This row illustrates the observed output structure, including its identifiers and public links. Empty values elsewhere in the dataset should be interpreted field by field; a field shown in this example is not promised for every job.

### Use cases

- Recruiting analysts can track newly listed public job URLs and IDs by keyword.
- Labor-market researchers can measure the volume of recent vacancies when source coverage is available.
- Job boards can reconcile known Naukri URLs against their catalog without treating sitemap-only rows as complete vacancies.
- Hiring teams can gather available skills, salary, and employer fields when the detailed endpoint is open.

The strongest analyses keep source context. A field such as price, rating, engagement, or rank has meaning only with its associated item, query, date, and public URL. Keep the raw export and create a separate cleaned view for charts or alerts.

### How much does it cost to scrape Naukri?

The price is **$1.00 per 1,000 saved results**. Platform usage is included. The charge scales with output rows, so a restrictive filter or inaccessible page can produce fewer billable results than the requested maximum.

- 100 jobs: **$0.10**. This is useful for checking a small cohort and confirming which optional fields are present.
- 1,000 jobs: **$1.00**. This is the reference price for a larger export.
- 5,000 jobs: **$5.00**. Reaching this size may require multiple focused sources or scheduled runs, depending on public inventory and source caps.

Compute any other estimate as saved result count × $1.00 / 1,000. A maximum input is a ceiling, not a purchase of that many rows. For planning, use the actual saved-item count from an initial representative run.

### Tips for best results

Start with a narrow keyword and location. For a known role, provide jobIds directly. Enable fetchDetails only when title, employer, and full descriptions are needed; check source on each row before relying on optional attributes.

Collect a small baseline first, record the exact input and date, and inspect both a typical row and a sparse row. Expand by adding focused sources instead of assuming one broad input can reveal the full public inventory. When comparing two runs, match stable IDs or canonical URLs and use the same filters so changes reflect the source rather than a changed query.

Export the full JSON when nested arrays or objects matter. CSV and Excel are convenient for sorting and joins, but nested structures may need flattening before spreadsheet analysis. Keep numeric fields numeric and preserve source URLs as text; do not infer a zero from a null.

### Limits

The successful golden run used the sitemap fallback: it returned identifiers and URLs, while title, company, location, and other job attributes were null. Search filters and detail enrichment may be ineffective during captcha periods.

Public pages can change, disappear, or expose different fields for different records. The schema is a list of possible output columns, not a promise that each column is filled in each row. The `maxItems`-style input limits cap saved results; they do not bypass source pagination, public visibility, or a site-specific result ceiling.

Treat a saved result as a snapshot. If a later run returns fewer rows, first compare the input, source access, and public inventory before concluding that the underlying market changed. If you need an audit trail, retain the source URL, stable identifier, and run timestamp with the export.

### Use with AI agents (MCP)

An agent can supply the documented input JSON, run this actor, and work from its dataset. Ask it to keep source links and identify null values explicitly when summarizing results. Example prompts:

> Run Naukri Scraper for the sample input above. Return the first 20 jobs with their source URLs and the fields needed for India job-market tracking and vacancy discovery. Mark unavailable fields as null.

> Schedule a repeat Naukri collection with the same filters. Compare records by stable ID or canonical URL and report only new, removed, or changed public values.

### FAQ

#### Is it legal to scrape Naukri?

This actor collects public data. Check the site’s terms, applicable law, and the rights of people whose data appears in your export. Follow GDPR and other privacy rules where relevant; do not use personal data for misuse, intrusive profiling, or unauthorized contact.

#### Do I need to configure proxies?

The input includes proxyConfiguration and defaults to Apify Proxy. Usually the default is enough to start. Availability can vary by page and region; changing network settings cannot make private or login-only information public.

#### How fast will a run finish?

Duration depends on the number of inputs, pages, detail requests, and source responses. A small sample is the best way to measure your workload. Increase the limits gradually, then use the observed run duration for scheduling and monitoring.

#### Can I schedule and monitor recurring runs?

Yes. Save the input and schedule recurring runs in Apify. Monitor run status, item count, and any missing-field changes; store the run date alongside exports when comparing snapshots.

#### Can I export to Google Sheets or Excel?

Yes. Download CSV or Excel from the dataset, or pass JSON to a spreadsheet integration. For nested fields, flatten the specific child values you need rather than losing the original JSON.

#### What if a field is empty?

An empty or null field means it was not available from that public source record in that run. Check the source URL, input mode, and record type. Do not replace missing prices, counts, dates, or flags with zero unless your own analysis has a documented rule for doing so.

#### Why are job titles and companies sometimes empty?

A captcha may block the search and job endpoints. The actor then uses the public incremental sitemap, which provides recent job IDs and URLs but not detailed job fields.

### Integrations

Use the Apify API to start runs and retrieve the dataset, webhooks to react when a run finishes, or Zapier, Make, and n8n to route records into other systems. Google Sheets supports lightweight review, while scheduled runs provide repeat snapshots. Keep the raw JSON when downstream workflows need nested fields or exact null values.

### Support

Open an issue on the Issues tab; we reply within 24h and add fields on request.

# Actor input Schema

## `keyword` (type: `string`):

Words to find in job listings. A search URL or job IDs take precedence.

## `location` (type: `string`):

Optional city or region to narrow the search.

## `searchUrl` (type: `string`):

A naukri.com search results URL. Its keyword and location take precedence over the separate fields.

## `jobIds` (type: `array`):

Optional 12-digit job IDs or full job URLs to look up directly. Takes precedence over search.

## `maxJobs` (type: `integer`):

Maximum number of unique jobs to save.

## `fetchDetails` (type: `boolean`):

Request full job details when the source endpoint is available. Captcha-protected results remain partial.

## `sortBy` (type: `string`):

Sort search results by relevance or date when the search endpoint is available.

## `experience` (type: `string`):

Years of required experience for the search endpoint. Use all for no filter.

## `freshness` (type: `string`):

Maximum posting age in days for the search endpoint. Use all for no filter.

## `workMode` (type: `array`):

Optional work modes for the search endpoint.

## `proxyConfiguration` (type: `object`):

Proxy configuration for fetching source pages.

## Actor input object example

```json
{
  "keyword": "software developer",
  "maxJobs": 40,
  "fetchDetails": false,
  "sortBy": "relevance",
  "experience": "all",
  "freshness": "all",
  "workMode": [],
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `jobs` (type: `string`):

Job records in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "keyword": "software developer",
    "maxJobs": 40
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/naukri-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "keyword": "software developer",
    "maxJobs": 40,
}

# Run the Actor and wait for it to finish
run = client.actor("mlg14/naukri-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "keyword": "software developer",
  "maxJobs": 40
}' |
apify call mlg14/naukri-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/naukri-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/QsB03waFtXs8mkPZ1/builds/tKz8Q0pGG82dNWaHf/openapi.json
