# elempleo.com Jobs Scraper — Colombia Job Offers & Salaries (`oswaldocarabano/elempleo-jobs-scraper`) Actor

Scrape elempleo.com, a leading Colombian job board: title, company, city, salary band parsed to COP, contract, work modality, keywords. Optional full job page adds 32 fields: full description, expiry date, education, experience, vacancies, schedule, sector. Filters by city, area, salary. No login.

- **URL**: https://apify.com/oswaldocarabano/elempleo-jobs-scraper.md
- **Developed by:** [Oswaldo Carabano](https://apify.com/oswaldocarabano) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.70 / 1,000 job offers

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## elempleo.com Jobs Scraper: Colombia job offers, salaries and employers

**Scrape job offers from elempleo.com, one of Colombia's largest job boards.**
You get every offer in the country, or only the ones that match a keyword, city,
job area, industry, salary band, contract type, modality or seniority.

Every row carries the title, the employer, the city, the **salary band parsed
into COP numbers**, the contract type, the work modality, the publication date,
keywords and the standard positions elempleo maps the offer to. Turn on the full
job page and you also get **32 more fields**: the complete description, the
salary range from the page's structured data, the expiry date, education level, minimum experience, number of vacancies,
work schedule, professions and the employer's industry and profile.

**No login. No cookies from any account. No browser.** $0.70 per 1,000 job offers,
and $0.74 per 1,000 with the full job page.

***

### What you can use this for

| If you are… | You want |
|---|---|
| Building a **Colombia labour-market dataset** | the nationwide feed: 20,876 live offers on 30 Sep 2026, with city, contract and modality on every row |
| Doing **salary benchmarking** in Colombia | `salary_min` / `salary_max` in COP, parsed from the published band, on 75.9 % of offers; the rest are published as confidential and say so |
| Running **recruitment or B2B lead generation** | who is hiring, where and for what: employer name, ID, logo, industry and profile page |
| Tracking **hiring demand** by city, area or industry | the same filters the site offers, with the site's own counts, on a schedule |
| Feeding a **job aggregator** or alerts product | clean JSON over the Apify API, stable `job_id` for deduplication, `published_at` for freshness |

***

### What actually arrives in a row

Every percentage in this section was measured **on Apify on 30 Sep 2026**: 1,000
of the newest offers nationwide for the search result, and 300 with the full job
page. Each figure comes with its sample size. **We do not list a field whose fill
rate we have not measured.**

A search-result row has a **median of 30 filled fields** (n = 1,000). With the
full job page it has **58** (n = 300). The schema has 69 fields, and every key is
present in every row, as `null` when the site has no value. No duplicates, no
`undefined`, no leftover HTML in any of the 1,300 rows.

#### Always there: 100 % (n = 1,000 offers)

`job_id` · `title` · `title_original` · `url` · `company_display_name` ·
`is_confidential` · `city` · `salary_text` · `salary_currency` · `has_salary` ·
`contract_type` · `work_modality` · `is_remote` · `is_hybrid` · `published_at` ·
`published_label` · `description` · `keywords` · `related_positions`

#### Usually there

| Field | Fill rate | n | Why it is not 100 % |
|---|---:|---:|---|
| `salary_min` (COP) | 75.9 % | 1,000 | the rest publish "Salario confidencial" (23.2 %) |
| `salary_max` (COP) | 76.6 % | 1,000 | null for the open band "Más de $21 millones" |
| `company_name`, `company_id` | 68.5 % | 1,000 | 31.5 % of offers are published as confidential |
| `company_logo_url` | 64.5 % | 1,000 | confidential offers and companies without a logo |

**The salary is parsed, not guessed.** elempleo publishes bands such as
"$2,5 a $3 millones", with a decimal comma. We turn that into `2500000` and
`3000000`, and we checked the result against the salary in the job page's
structured data: **216 of 216 matched** (30 Sep 2026; the page adds 1 peso to
the lower bound, so `salary_exact_min` reads `2500001`).

**Confidential offers stay confidential.** When an employer hides its name, the
site shows a placeholder name and a default logo. We return `company_name: null`,
`company_logo_url: null` and `is_confidential: true`, and never pass the default
image off as the company's logo.

#### With the full job page (n = 300 job pages)

| Field | Fill rate |
|---|---:|
| `valid_through` (expiry date), `date_posted`, `employment_type`, `location_country` | 100 % |
| `position_level`, `position_sublevel`, `min_experience`, `experience_code` | 100 % |
| `vacancies`, `job_areas`, `professions`, `job_sector`, `company_industry` | 100 % |
| `hiring_organization`, `direct_apply`, `related_positions_detail` | 100 % |
| `company_page_url` | 99.0 % |
| `location_locality`, `location_region` | 97.0 % |
| `education_level` | 90.0 % |
| `job_area` | 86.0 % |
| `work_schedule` | 75.0 % |
| `salary_exact_min` / `salary_exact_max` / `salary_unit` | 72.0 % / 73.0 % / 73.0 % |
| `company_sector` | 69.3 % |
| `profession` | 53.3 % |
| `company_description` | 34.7 % |
| `job_location_type` | 16.7 % |
| `inclusive_talent` | 3.0 % |

**The search result cuts long descriptions.** elempleo stops the description at
about 2,000 characters in search results, sometimes in the middle of a word. That
affected **15.9 % of offers** (159 of 1,000). `description_is_truncated` tells you
which ones. With the full job page the description is always complete (0 of 300
truncated).

***

### Full job pages: read this before turning them on

The search returns 100 offers per request. The full job page is **one extra
request per offer**, about 60 KB each. It is off by default. Turn it on with
**Include full job page** when you need the fields above.

- It costs **$0.00004 more per offer** ($0.04 per 1,000), charged only when the
  page was actually read and its data added to the row.
- It is slower. Measured on Apify: 50 offers without pages in 9 s, and 30 offers
  with pages in 23 s.
- If a page cannot be read (the offer expired a minute ago, for example), you
  still get the search-result row, charged as a plain offer. `detail_error` says
  what happened, and the page is not charged. (Offers withdrawn by the employer
  answer "410 Gone"; they are reported as no longer available.)

***

### Contacts inside the offer

Some employers write a recruitment phone number or email inside the job
description. We return them in `contact_phones` and `contact_emails`:

| | Full job page (n = 300) | Search result (n = 1,000) |
|---|---:|---:|
| `contact_phones` | 4.3 % | 2.8 % |
| `contact_emails` | 4.0 % | 0 % |

The search result hides emails, so they only appear with the full job page.

**Whose contacts are these?** On elempleo, only registered company accounts can
publish offers. **494 of 494** offers we measured (23 Sep 2026) came from a company account,
confidential ones included. These are recruitment contacts published by an
employer for applicants. We do not collect anything about candidates or site
users, and we never log in.

***

### How the search works, and why there is a cap

elempleo's search returns **at most 10,000 results per query**. Asking for result
10,001 does not return an empty page: the site answers with a server error. We
never ask for a page beyond the cap.

- If your search matches 10,000 offers or fewer, or you ask for 10,000 or fewer,
  it is one search, paged 100 at a time.
- If you ask for more than 10,000 and your search matches more (the whole country
  had 20,876 offers on 30 Sep 2026, and Bogotá alone 13,072), the
  actor **splits the search by job area** automatically. It spreads your quota
  fairly across areas and removes duplicates, because an offer can belong to more
  than one area.

Every filter in the input was tested against the live site and checked against
the site's own counter. For example, on 23 Sep 2026 Medellín returned 1,431 offers and the
city's counter said 1,431. Two filters the site accepts but silently ignores
(minimum experience and sort order) are **not** offered, because they would give
you the whole country while looking like they worked.

***

### Input

Everything has a working default. An **empty input returns the 50 most recent
offers nationwide** in a couple of seconds.

| Field | What it does |
|---|---|
| `searchQuery` | Free text, like the site's search box. "contador" matched 593 offers on 30 Sep 2026. Letters, numbers and `. , & + # ' ( ) / -` only |
| `cities` | One or more of 60 cities, as elempleo lists them |
| `areas` | One or more of 33 job areas (elempleo's own taxonomy) |
| `industry` | Industry of the offer (the site's sector, `job_sector`) |
| `workModality` | On-site, hybrid or remote |
| `contractType` | Permanent, fixed term, service contract, apprenticeship, per project, other |
| `positionLevel` | Operational, assistant, professional, middle or senior management |
| `publishedWithin` | Today and yesterday, last week, last 2 weeks, last month |
| `salaryRanges` | One or more monthly salary bands, in millions of COP |
| `maxItems` | Hard cap on offers delivered and charged (default 50) |
| `scrapeDetails` | Add the full job page (default off) |
| `maxConcurrency` | Job pages read in parallel (default 3). 5 ran at 5.4 pages/s with no refusals |

If a filter matches nothing, the run finishes **green** with one row that says so,
and nothing is charged.

***

### Output example

```json
{
  "job_id": 1886771176,
  "title": "Gestor de nuevos negocios b2b entidad financiera - bogota",
  "company_name": "SUMMAR PRODUCTIVIDAD S.A.S",
  "is_confidential": false,
  "city": "Bogotá",
  "salary_text": "$2,5 a $3 millones",
  "salary_min": 2500000,
  "salary_max": 3000000,
  "salary_currency": "COP",
  "contract_type": "Por obra o labor",
  "work_modality": "Presencial",
  "published_at": "2026-09-23T00:00:00Z",
  "keywords": ["VENTAS"],
  "related_positions": ["Asesor comercial"],
  "description_is_truncated": false,
  "detail_fetched": false
}
```

Values are returned as elempleo publishes them, in Spanish ("Presencial",
"Indefinido"). Field names and every message from the actor are in English.

***

### Pricing

Pay per event: you only pay for what is delivered.

| Event | Price | When |
|---|---:|---|
| Job offer | **$0.0007** | once per offer row ($0.70 per 1,000) |
| Full job page | **$0.00004** | only when you turn it on and the page was read ($0.04 per 1,000) |
| Actor start | $0.00001 | platform minimum, once per run |

- **Failed requests, error rows and duplicates are never charged.** Within a run,
  an offer is delivered and charged once, even if it shows up in several job
  areas.
- If your maximum charge per run is reached, the run **stops cleanly** and says so
  in its status. You never get rows you were not charged for, and you are never
  charged for rows you did not get.
- An empty input costs about $0.035 (50 offers). 1,000 offers with full job
  pages cost $0.74.

***

### Data freshness

Every row is read live from elempleo.com during your run. `scraped_at` records
when. `published_at` is the day the employer published the offer (the site gives
the day, not the hour). `valid_through` on the full job page says until when it
accepts applications.

***

### Frequently asked questions

**Does it need an elempleo account?** No. It reads what any visitor sees, without
logging in.

**Can I get all offers in Colombia?** Yes. Set `maxItems` above the number of
offers (20,876 on 30 Sep 2026) and the actor splits the search by job area
to get past the site's 10,000 cap.

**Why is `company_name` empty on some rows?** The employer published the offer as
confidential (31.5 % of offers on 30 Sep 2026). `company_display_name` shows what the site shows.

**Why is `salary_min` empty on some rows?** The employer chose not to publish the
salary ("Salario confidencial"). `salary_is_confidential` is `true` on those rows.

**How fast is it?** Measured on Apify with the default 512 MB: 1,000 offers in
109 s, 300 offers with full job pages in 185 s, and the empty input (50 offers)
in 9 s. Peak memory was 99 MB.

**Can I run it every day?** Yes. Schedule it with `publishedWithin` set to "Today
and yesterday" and deduplicate on `job_id`.

**Is the job ID stable?** Yes. `job_id` is elempleo's own offer ID and it is also
in the URL. Use it to deduplicate across runs.

**What if elempleo refuses requests?** The actor switches to a different
connection straight away and carries on. If requests are still refused, it pauses
(60 s, then longer), never more than 4 minutes in total and never more than a
third of the time left in your run, and says so in the status message. Failed
requests are never charged.

**Why was my keyword rejected?** elempleo's firewall blocks searches that look
like web addresses or code (for example `<script>` or `https://…`) and then
refuses every request for minutes. The actor only accepts letters, numbers,
spaces and `. , & + # ' ( ) / -`, which covers real job searches such as `C#`,
`auxiliar (bodega)` or `café & té`. A rejected keyword finishes green with one
row explaining why, and nothing is charged.

***

### Support and data removal

Found a problem or a field that looks wrong? Open an issue on the actor's
**Issues** tab with the run ID and we will look at it. If you are an employer and
want an offer's data removed from our examples, write to us there too.

***

### Disclaimer

This actor collects publicly available job offers. It does not log in, create
accounts or solve CAPTCHAs. You are responsible for using the data in line with
the law that applies to you, including Colombian data protection law (Ley 1581
de 2012) and elempleo.com's terms. The actor is not affiliated with or endorsed
by elempleo.com.

# Actor input Schema

## `searchQuery` (type: `string`):

Free-text search, exactly like the search box on elempleo.com (job title, skill or company). Letters, numbers, spaces and . , & + # ' ( ) / - only: the site's firewall blocks web addresses and code. Leave empty to get every job offer. Example: "contador" matched 593 of 20,876 offers on 30 Sep 2026.

## `cities` (type: `array`):

Only job offers located in these cities. Empty = all of Colombia. City names are shown as elempleo.com publishes them.

## `areas` (type: `array`):

Only job offers in these job areas (elempleo.com's own taxonomy, e.g. "Sistemas y Tecnología"). Empty = every area. An offer can belong to several areas; duplicates are removed.

## `industry` (type: `string`):

Only job offers that elempleo.com classifies in this industry. It is the sector of the offer (job\_sector on the full job page, 40 of 40 matched in a test on 30 Sep 2026), which can differ from the employer's own industry (company\_industry). Empty = every industry.

## `workModality` (type: `string`):

On-site, hybrid or remote.

## `contractType` (type: `string`):

Type of employment contract as published by the employer.

## `positionLevel` (type: `string`):

Seniority of the role as classified by elempleo.com.

## `publishedWithin` (type: `string`):

Only job offers published recently.

## `salaryRanges` (type: `array`):

Only job offers in these monthly salary bands, in millions of Colombian pesos, as published by the employer. Empty = any salary, including confidential.

## `maxItems` (type: `integer`):

Hard cap on job offers delivered and charged in this run. The site returns at most 10,000 results per search; above that the actor splits the search by job area automatically and removes duplicates.

## `scrapeDetails` (type: `boolean`):

Also open each job offer page and add 32 fields: full description (the search result cuts ~21% of descriptions), salary range with its unit, expiry date, education level, minimum experience, vacancies, position level, work schedule, professions, company industry and description. Charged as a separate event. One extra request per offer, so runs take longer (measured: 30 offers with pages in 15 s, 50 without in 2 s).

## `maxConcurrency` (type: `integer`):

How many job pages to open at the same time when "Include full job page" is on. Measured: 5 in parallel ran at 5.4 pages/s with no refusals. Lower values are slower, not safer.

## `proxyConfiguration` (type: `object`):

Optional. Not needed today: elempleo.com answers directly. Only used as a fallback if the site starts refusing requests.

## Actor input object example

```json
{
  "searchQuery": "analista",
  "cities": [],
  "areas": [],
  "industry": "",
  "workModality": "any",
  "contractType": "any",
  "positionLevel": "any",
  "publishedWithin": "any",
  "salaryRanges": [],
  "maxItems": 30,
  "scrapeDetails": true,
  "maxConcurrency": 3,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `jobs` (type: `string`):

One row per job offer: title, company, city, salary band parsed to COP, contract, modality, keywords and, with the full job page, 32 more fields.

## `runSummary` (type: `string`):

Counts of delivered and charged rows, pages fetched and any issues.

## `errors` (type: `string`):

Requests that failed or job pages that could not be read. Never charged. Always present, empty when all went well.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQuery": "analista",
    "maxItems": 30,
    "scrapeDetails": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("oswaldocarabano/elempleo-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchQuery": "analista",
    "maxItems": 30,
    "scrapeDetails": True,
}

# Run the Actor and wait for it to finish
run = client.actor("oswaldocarabano/elempleo-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQuery": "analista",
  "maxItems": 30,
  "scrapeDetails": true
}' |
apify call oswaldocarabano/elempleo-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,oswaldocarabano/elempleo-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3HdbKfL7X0CA1Wzta/builds/rEFrOyqgxpqTJuQWJ/openapi.json
