Cutshort Jobs Search Scraper avatar

Cutshort Jobs Search Scraper

Pricing

from $2.99 / 1,000 cutshort job records

Go to Apify Store
Cutshort Jobs Search Scraper

Cutshort Jobs Search Scraper

Extract rich public job, recruiter, salary, skills, description, and company data from Cutshort.io.

Pricing

from $2.99 / 1,000 cutshort job records

Rating

0.0

(0)

Developer

Jobs API

Jobs API

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

11 days ago

Last modified

Share

Cutshort Public Jobs Search Scraper

Extract complete public Cutshort job records. Search mode reads a public category/search page, then verifies each selected job against its official detail page. Single and multiple modes resolve exact official public job URLs.

Input

  • mode: search, single, or multiple; if omitted, it is inferred from direct URLs.
  • query: public job search phrase in search mode.
  • location: optional published location filter.
  • categoryUrl: optional HTTPS cutshort.io home or /jobs category URL for search mode.
  • jobUrl, jobUrls, or startUrls: exact public detail URLs for direct modes. Search mode accepts at most one public home/category startUrl per run so an empty-result claim cannot omit another requested search source.
  • maxItems: maximum complete detail-verified records (1–10).
  • maxPages: bounded search pages (1–3).
  • requestTimeoutSecs: per-page timeout.
  • retries: retries only for transient network/server failures.
  • requestDelayMs: spacing between sequential public requests.

The Actor uses ordinary direct HTTPS requests. It does not accept proxy configuration, browser fingerprint settings, or a custom user agent. Requests are sequential so the source receives a bounded, auditable request pattern.

Output

Each dataset item contains the official job identity and canonical URL, title, full HTML/text/Markdown description, company profile and published company attributes, recruiter, locations and remote policy, experience, salary, skills, job categories, employment types, funding/stage facts, related category links, application metadata, dates, structured JobPosting data, and source payloads. Only complete detail-verified job records are written to the dataset; diagnostics, request receipts, run health, and skipped-listing details are kept in separate key-value-store artifacts.

The runtime writer requires a successful detail receipt, matching official canonical identity and provenance, complete company/location/description data, and more than 20 populated meaningful source-backed fields. The storage-backed semantic validator additionally checks each stored item against the dataset JSON schema.

Access policy

HTTP 401, 403, 429, or 451 and visible source-authored access-denial or human-verification pages stop the crawl without retry or bypass. If no job row was persisted, all buffered job rows are discarded and the result is SKIPPED; a previously persisted prefix remains SUCCEEDED / LIMITED. HTTP 407, timeouts, network failures, and transient server errors are DEFERRED, not source blocks. They are never routed through a proxy or alternate identity. Bounded retries apply only to transient failures.

Run lifecycle

OUTPUT_SUMMARY, RUN_HEALTH, RUN_DIAGNOSTICS, and REQUEST_RECEIPTS report the platform status (SUCCEEDED or FAILED) separately from resultStatus (COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED). A clean extraction with stored jobs is SUCCEEDED / COMPLETE; stored jobs followed by any later limitation are SUCCEEDED / LIMITED. Only successfully awaited sequential Dataset writes count as persisted rows. A later dataset-write, KVS-write, or finalization failure preserves the successfully written prefix as SUCCEEDED / LIMITED; a fatal zero-row runtime or storage failure uses Actor.fail() and reports FAILED / FAILED. A handled source barrier with zero persisted rows is SUCCEEDED / SKIPPED; a transport-only or unrecognized zero-row result is SUCCEEDED / DEFERRED. NO_DATA is used only when the first successful search listing contains an explicit empty Next.js jobs array; the matching 2xx request receipt and empty-array evidence are retained in OUTPUT_SUMMARY and required by the local validator. In the current implementation, details are buffered until retrieval ends, so a barrier encountered before writing clears the entire buffer.

Example

{
"mode": "search",
"query": "developer",
"location": "",
"maxItems": 3,
"maxPages": 1,
"requestTimeoutSecs": 15,
"retries": 1,
"requestDelayMs": 500
}

Local validation

npm test
npm run lint
apify validate-schema

npm run validate -- <storage-dir> is a separate storage-backed check that reads Dataset and key-value-store files from the selected local run. It is not part of offline source, runtime, or schema checks.

Run metadata and diagnostics are written to the default key-value store under OUTPUT, OUTPUT_SUMMARY, RUN_HEALTH, RUN_DIAGNOSTICS, REQUEST_RECEIPTS, and RUN_SKIPS. OUTPUT contains only summary metadata and the dataset count; the dataset is the sole location for job rows.