OCC Mexico Jobs Scraper avatar

OCC Mexico Jobs Scraper

Pricing

from $2.00 / 1,000 results

Go to Apify Store
OCC Mexico Jobs Scraper

OCC Mexico Jobs Scraper

Scrape job listings from OCC Mexico: titles, companies, locations, and salary ranges.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Danilo Frias

Danilo Frias

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Scrapes public job listings from OCC Mundial (occ.com.mx), Mexico's large job board, using plain HTTP requests and HTML parsing (httpx + BeautifulSoup). It runs without login, cookies, JavaScript execution, proxies, or headless browsers.

Who would use it: labor-market researchers, recruiters doing competitive mapping, salary-benchmark builders, and job aggregators who need a simple Mexico-wide listing feed.

How it works

OCC's listing pages (https://www.occ.com.mx/empleos/...) return fully server-rendered HTML (verified 2026-09-25, HTTP 200, no login wall, no CAPTCHA). Each listing card carries:

  • Job title (h2), salary text, benefits/perks bullets
  • Company name (often "Empresa confidencial" when the employer hides its name)
  • Location (e.g. "Jalisco", "Ciudad de México")
  • Posted-date text in Spanish ("Hoy", "Ayer", "Hace N días")
  • A numeric offer id (data-id) used to build the detail URL

The actor paginates with OCC's own ?page=N parameter until max_items is reached, with ~1 request/second politeness and per-page try/except resilience.

Inputs

  • search_query (string, default ""): keyword filter, e.g. "contador". Converted to OCC's semantic URL segment (de-contador). Empty = no keyword filter.
  • location (string, default "Mexico"): location filter, e.g. "Ciudad de Mexico", "Jalisco". Converted to OCC's semantic URL segment (en-ciudad-de-mexico). "Mexico" (default) uses the plain Mexico-wide /empleos/ page.
  • start_urls (array of strings, default derived from the two fields above): explicit public OCC listing URLs. Takes precedence over search_query/location.
  • max_items (integer, default 50, min 1): cap on total records pushed.

Example input for accountants in Mexico City:

{
"search_query": "contador",
"location": "Ciudad de Mexico",
"max_items": 50
}

Sample output record

{
"title": "Contador general",
"company": "Keyperspot",
"location": "Benito Juárez, Ciudad de México",
"salary_text": "$ 25,000 - $ 27,000 Mensual",
"job_url": "https://www.occ.com.mx/empleo/oferta/21350252",
"posted_date": "Ayer",
"description_snippet": null
}

(Taken verbatim from the local test run, 2026-09-26, input contador / Ciudad de Mexico.)

Coverage and limits

  • Listing pages: WORKING (200, server-rendered cards, ~20-22 cards/page, ?page=N pagination returns fresh cards).
  • Job detail pages (https://www.occ.com.mx/empleo/oferta/<id>): BLOCKED for bots - the server returns HTTP 403 with the body "scraping abuse" for plain-HTTP clients (verified 2026-09-25). job_url is still emitted (the URL is valid and opens in a real browser), but description_snippet is limited to the benefits/perks bullets shown on the listing card - the full job description cannot be fetched without browser rendering, which this actor deliberately does not do.
  • posted_date is the raw Spanish text from the card ("Hoy", "Ayer", "Hace N días") - no normalization.
  • OCC pins sponsored cards at the top of every page; the actor dedupes by offer id so they are not duplicated across pages.
  • ~1 req/s politeness between requests; timeouts and HTTP errors end pagination for that start URL gracefully.

Data rules

  • Public pages only: the actor requests only pages that load fully for logged-out visitors. No login flows, no session cookies, no CAPTCHA solving, no proxy services, no headless browsers.
  • No personal data: company names are collected (businesses are fine), but recruiter names, emails, phone numbers, or any other private-individual data are never scraped or stored. Listing cards do not expose personal contacts.