Laborum Jobs Search Scraper avatar

Laborum Jobs Search Scraper

Pricing

from $2.99 / 1,000 laborum job records

Go to Apify Store
Laborum Jobs Search Scraper

Laborum Jobs Search Scraper

Extract rich, current job listings from Laborum.cl with clean descriptions, job attributes, search context, and stable URLs.

Pricing

from $2.99 / 1,000 laborum job records

Rating

0.0

(0)

Developer

Jobs API

Jobs API

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

3

Monthly active users

11 days ago

Last modified

Share

This Apify Actor retrieves public job listings from Laborum.cl through its configured public HTTPS search and normalized-detail APIs. It enriches search results with detail data and writes complete job records to the default dataset. Requests use direct HTTPS and are processed sequentially; the Actor does not use an Apify proxy, browser automation, stealth, or browser fingerprinting.

Responses are capped at 5 MB. Search and detail records are staged until fetching finishes so an access barrier cannot leave buffered partial job rows in the dataset. Only complete, source-verified job records can be written; diagnostics, request receipts, skips, and run state are saved separately in the key-value store.

Modes

  • search: one keyword search with an optional client-side location filter.
  • searchMultiple: several distinct keyword searches; each query is bounded by maxPages.
  • single: enrich one official Laborum detail URL.
  • multiple: enrich a list of official Laborum detail URLs.
  • startUrls: process a mixture of official search and detail URLs.

Search pagination is bounded per query/search URL. maxItems limits the records returned across the run. Search candidates are enriched sequentially; the Actor may inspect up to four times maxItems candidates to find enough complete records. Detail-only modes do not use search pagination.

Input

The default search input is:

{
"mode": "search",
"query": "desarrollador",
"maxItems": 3,
"maxPages": 3,
"requestTimeoutSecs": 25
}

Supported inputs and runtime bounds:

FieldApplies toDefaultBounds / behavior
modeAllsearchsearch, searchMultiple, single, multiple, or startUrls.
querySearch modesdesarrolladorUp to 120 characters.
queriessearchMultipleFalls back to query1–10 unique queries, each up to 120 characters.
locationSearch modesEmpty (no filter)Up to 100 characters; matching is applied to listing location text.
urlsingle—One official HTTPS Laborum job-detail URL.
urlsmultiple—1–50 official HTTPS Laborum job-detail URLs.
startUrlsstartUrls—1–50 objects of the form { "url": "https://www.laborum.cl/..." }; each must be an official search or detail URL.
maxItemsAll3Integer from 1 to 50.
maxPagesSearch modes3Integer from 1 to 10 pages per query/search URL.
requestTimeoutSecsAll25Integer from 5 to 45 seconds per request.

Search URLs may use the home page, /empleos.html, or /empleos-busqueda-<query>.html. Detail URLs use /empleos/<slug>-<numeric-id>.html. The runtime rejects unsupported input properties; there are no public fixture, debug, proxy, cookie, user-agent, or concurrency controls. Requests are sequential rather than configurable for concurrency.

The checked-in INPUT*.json files are Actor run inputs, not offline fixtures; running one may send requests to Laborum's public APIs.

Example for a search URL plus a detail URL:

{
"mode": "startUrls",
"startUrls": [
{ "url": "https://www.laborum.cl/empleos-busqueda-desarrollador.html" },
{ "url": "https://www.laborum.cl/empleos/desarrollador-de-software-mundo-1118387120.html" }
],
"maxItems": 3,
"maxPages": 2,
"requestTimeoutSecs": 25
}

Dataset quality

Each job row includes a stable laborum-<id> ID, canonical public URL, title, employer, location, search provenance, source dates, and detail fields when published. The parser retains readable description text and sanitized HTML, sections, bullets, requirements, qualifications, responsibilities, benefits, salary, application details, and company attributes where available. Duplicate aliases do not count as separate facts: a record must contain at least 21 distinct populated meaningful-field groups, as well as complete identity, a verified canonical detail URL, and a readable description of at least 80 characters. Missing, blank, and empty values are omitted.

The dataset contract is in .actor/dataset_schema.json. Run state is stored under RUN_SUMMARY, RUN_DIAGNOSTICS, RUN_SKIPS, REQUEST_RECEIPTS, RUN_HEALTH, and RUN_METADATA. Records preserve source provenance separately from the curated normalized fields.

RUN_SUMMARY.status and resultStatus are separate. status is the Apify process result (SUCCEEDED or FAILED); resultStatus is COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED. A run that stored one or more real rows remains platform SUCCEEDED if a later error occurs and is marked LIMITED. Fatal zero-row failures use FAILED / FAILED. Unconfirmed empty responses, transport errors, and parser misses use SUCCEEDED / DEFERRED. NO_DATA is reserved for every requested search returning HTTP 200, an explicit source count of zero, and an empty result array; the summary retains each official URL, count, and SHA-256 response hash. SKIPPED requires an explicit source access barrier and does not publish buffered rows. Diagnostics and test/dummy records are never dataset rows.

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

Access barriers and failures

HTTP 401, 403, 429, or 451, and a visible source-authored denial, CAPTCHA, human-verification, or security-challenge page are terminal barriers. The Actor stops requests to that site, does not retry or switch routes, browsers, identities, proxies, cookies, or fingerprints, does not solve the challenge, discards staged job rows, and records resultStatus: SKIPPED with diagnostics and request receipts. A handled barrier is a successful process exit (status: SUCCEEDED), not a successful extraction.

Generic transport failures, HTTP 407, malformed or unrecognized responses, and parser misses are reported as resultStatus: DEFERRED, not as proof that Laborum blocked access or has no results. A fatal zero-row runtime, input, or persistence failure is FAILED / FAILED; after at least one successful dataset write, a later error preserves the written rows and resolves to SUCCEEDED / LIMITED. A bounded retry may be made only for designated transient statuses (408, 425, 500, 502, 503, or 504); access barriers are never retried. Other extraction or validation failures are not mislabeled as access barriers or empty results.

Local development and safe validation

Install the locked dependencies and run the offline checks from this Actor directory:

npm ci
npm test
npm run lint
npm run check
apify validate-schema

The available npm scripts are start, test, lint, check, and validate. The dataset validator accepts a storage directory as its first argument, or reads APIFY_LOCAL_STORAGE_DIR; without either, it checks the default storage directory. For example, after selecting a fresh isolated path, run npm run validate -- "$storageDir".

For an Actor run, use a new, actor-relative storage directory and --resurrect. Never use --purge for validation. Inspect and preserve any existing storage/ data first. Also check for legacy apify_storage in this directory and its parent: the Apify CLI may migrate that directory into the selected store. If a legacy directory exists, preserve it separately or run from an isolated working copy before proceeding.

PowerShell example (the generated store name is unique; this does not delete or reuse existing data):

$legacyStores = @('.\apify_storage', '..\apify_storage') | Where-Object { Test-Path -LiteralPath $_ }
if ($legacyStores) { throw 'Preserve or isolate the existing legacy apify_storage before running.' }
$storageDir = "storage-laborum-$([guid]::NewGuid().ToString('N'))"
if (Test-Path -LiteralPath $storageDir) { throw 'The selected local storage path already exists.' }
$env:APIFY_LOCAL_STORAGE_DIR = $storageDir
apify run --resurrect --input-file .\INPUT.json
npm run validate -- "$storageDir"

Local Actor storage is separate from Apify Cloud storage. A local run or local validation does not establish Cloud qualification.

Responsible use

Use only information made publicly available by the source. Respect Laborum's terms, published access controls, rate limits, privacy obligations, and applicable law. This Actor is not affiliated with Laborum. Do not attempt to bypass access barriers; report issues with the run ID and relevant summary/diagnostics.