OCC Mexico Jobs Scraper
Pricing
from $2.00 / 1,000 results
OCC Mexico Jobs Scraper
Scrape job listings from OCC Mexico: titles, companies, locations, and salary ranges.
Pricing
from $2.00 / 1,000 results
Rating
0.0
(0)
Developer
Danilo Frias
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Scrapes public job listings from OCC Mundial (occ.com.mx), Mexico's large job board, using plain HTTP requests and HTML parsing (httpx + BeautifulSoup). It runs without login, cookies, JavaScript execution, proxies, or headless browsers.
Who would use it: labor-market researchers, recruiters doing competitive mapping, salary-benchmark builders, and job aggregators who need a simple Mexico-wide listing feed.
How it works
OCC's listing pages (https://www.occ.com.mx/empleos/...) return fully server-rendered HTML (verified 2026-09-25, HTTP 200, no login wall, no CAPTCHA). Each listing card carries:
- Job title (
h2), salary text, benefits/perks bullets - Company name (often "Empresa confidencial" when the employer hides its name)
- Location (e.g. "Jalisco", "Ciudad de México")
- Posted-date text in Spanish ("Hoy", "Ayer", "Hace N días")
- A numeric offer id (
data-id) used to build the detail URL
The actor paginates with OCC's own ?page=N parameter until max_items is reached, with ~1 request/second politeness and per-page try/except resilience.
Inputs
search_query(string, default""): keyword filter, e.g."contador". Converted to OCC's semantic URL segment (de-contador). Empty = no keyword filter.location(string, default"Mexico"): location filter, e.g."Ciudad de Mexico","Jalisco". Converted to OCC's semantic URL segment (en-ciudad-de-mexico)."Mexico"(default) uses the plain Mexico-wide/empleos/page.start_urls(array of strings, default derived from the two fields above): explicit public OCC listing URLs. Takes precedence oversearch_query/location.max_items(integer, default50, min1): cap on total records pushed.
Example input for accountants in Mexico City:
{"search_query": "contador","location": "Ciudad de Mexico","max_items": 50}
Sample output record
{"title": "Contador general","company": "Keyperspot","location": "Benito Juárez, Ciudad de México","salary_text": "$ 25,000 - $ 27,000 Mensual","job_url": "https://www.occ.com.mx/empleo/oferta/21350252","posted_date": "Ayer","description_snippet": null}
(Taken verbatim from the local test run, 2026-09-26, input contador / Ciudad de Mexico.)
Coverage and limits
- Listing pages: WORKING (200, server-rendered cards, ~20-22 cards/page,
?page=Npagination returns fresh cards). - Job detail pages (
https://www.occ.com.mx/empleo/oferta/<id>): BLOCKED for bots - the server returns HTTP 403 with the body "scraping abuse" for plain-HTTP clients (verified 2026-09-25).job_urlis still emitted (the URL is valid and opens in a real browser), butdescription_snippetis limited to the benefits/perks bullets shown on the listing card - the full job description cannot be fetched without browser rendering, which this actor deliberately does not do. posted_dateis the raw Spanish text from the card ("Hoy", "Ayer", "Hace N días") - no normalization.- OCC pins sponsored cards at the top of every page; the actor dedupes by offer id so they are not duplicated across pages.
- ~1 req/s politeness between requests; timeouts and HTTP errors end pagination for that start URL gracefully.
Data rules
- Public pages only: the actor requests only pages that load fully for logged-out visitors. No login flows, no session cookies, no CAPTCHA solving, no proxy services, no headless browsers.
- No personal data: company names are collected (businesses are fine), but recruiter names, emails, phone numbers, or any other private-individual data are never scraped or stored. Listing cards do not expose personal contacts.