Meteojob Search Scraper
Pricing
from $2.99 / 1,000 meteojob job records
Meteojob Search Scraper
Extract rich Meteojob listings with structured descriptions, locations, contracts, experience, qualifications, and publication metadata.
Pricing
from $2.99 / 1,000 meteojob job records
Rating
0.0
(0)
Developer
Jobs API
Maintained by CommunityActor stats
0
Bookmarked
3
Total users
2
Monthly active users
11 days ago
Last modified
Categories
Share
What does Meteojob Jobs Search Scraper do?
This Apify Actor searches public job listings and fetches exact public job details from Meteojob, returning structured job data with source provenance. It is a direct-HTTPS, Cheerio-based alternative to manual Meteojob search; it does not use a browser, proxy, mirror, fingerprint spoofing, or access-control bypass.
Why use Meteojob Jobs Search Scraper?
Use search mode to collect a small, bounded set of current listings, or choose single/multiple mode to fetch known public posting URLs. Detail records combine public JobPosting JSON-LD with Meteojob's embedded job-offer state and preserve readable description sections, source labels, salary where published, and request receipts. You can schedule runs and access their datasets, run history, and monitoring through Apify. This Actor makes one sequential attempt per started request and does not rotate proxies or retry blocked/failed requests.
How to scrape Meteojob
- Open the Actor's Input tab and choose
search,single, ormultiple. - For search, enter a French keyword and optional location. For detail modes, provide only official Meteojob HTTPS posting URLs.
- Set the result and request limits, run the Actor, then inspect the dataset and
RUN_SUMMARY/RUN_DIAGNOSTICS/RUN_REQUESTSvalues.
Modes
searchrequests the official/jobs?what=...&where=...SSR page, parses current offer cards, then fetches selected/jobs/<numeric-id>detail pages.singlefetches one exact official detail URL.multiplefetches several exact official detail URLs.
Detail records combine the official JobPosting JSON-LD with the official candidate-front-state job-offer payload. This captures title, company, rich descriptions, profile and company text, benefits, location coordinates and country data, contracts, salary, job taxonomy, experience, skills/languages, synonyms, telework/travel labels, publication dates, and source provenance. Contact data is redacted from the retained structured source record. Incomplete rows are omitted and diagnostics are stored only in RUN_DIAGNOSTICS.
What data can Meteojob Jobs Search Scraper extract?
Every dataset item is a verified job record. The dataset keeps the complete source-backed record; the table below groups its most useful fields:
| Group | Fields | Notes |
|---|---|---|
| Identity | id, jobId, title, jobTitle, company, companyName, companyLogoUrl, reference | Stable Meteojob posting identity and published employer data; repeated title/company fields are compatibility aliases. |
| Location and job | location, locations, city, region, country, employmentType, contractTypes, industry, occupationalCategory, jobCategories, experienceLevels | Structured location parts, job classification, and employment labels are included when published. |
| Compensation | salary, salaryMin, salaryMax, salaryCurrency, salaryPeriod, salaryRanges | The source display is preserved; numeric values and period are omitted when the public posting does not provide them clearly. |
| Description | description, descriptionHtml, descriptionSections, descriptionItems, companyDescription, profileDescription, responsibilities, requirements, benefits | Full readable text is retained. HTML and section/bullet arrays are optional and controlled by includeDescription. |
| Skills and conditions | skills, qualifications, languages, synonyms, telework, travel, directApply | These are source-published labels or structured posting facts, not inferred from the search phrase. |
| Provenance and quality | jobUrl, canonicalUrl, searchUrl, searchResultRank, requestReceipts, detailVerified, quality.meaningfulFieldCount | Receipts and quality metadata explain where the record came from and whether the detail was verified. The meaningful-field count excludes aliases and provenance. |
Optional fields are omitted when the source does not expose them; the Actor does not fill missing salary, application links, qualifications, or other facts with guesses. The schema permits additional source payload keys in sourceRecord so the official structured records remain available for auditing.
Input
See the input tab for the complete configuration. Use jobUrl / jobUrls for normal inputs; url / urls are compatibility aliases.
{"mode": "search","query": "developpeur","location": "Paris","maxItems": 3,"maxCandidates": 20,"maxRequests": 20,"timeoutMs": 15000,"deadlineMs": 120000,"includeDescription": true}
For single, provide jobUrl; for multiple, provide jobUrls. Compatibility aliases url and urls are accepted. Detail URLs must be official HTTPS https://www.meteojob.com/jobs/<numeric-id> URLs.
Single-detail input (use an active public posting URL):
{"mode": "single","jobUrl": "https://www.meteojob.com/jobs/57063397","maxItems": 1,"maxRequests": 1}
Multiple-detail input (both URLs must be active public postings):
{"mode": "multiple","jobUrls": ["https://www.meteojob.com/jobs/57063397","https://www.meteojob.com/jobs/56471308"],"maxItems": 2,"maxRequests": 2}
| Input | Type | Default | Behavior |
|---|---|---|---|
mode | string | search | search, single, or multiple. |
query, location | string | developpeur, empty | Search term and optional location for search; the configured search uses one listing page. |
jobUrl / jobUrls | string / array | — | Exact detail URL(s) for single / multiple; url / urls are aliases. When both forms are supplied, jobUrl takes precedence in single mode and an array-valued jobUrls takes precedence in multiple mode. |
maxItems | integer, 1–50 | 3 | Maximum complete dataset records. |
maxCandidates | integer, 1–80 | 20 | Maximum search candidates or multiple-mode URLs considered. single mode always uses its one normalized detail URL. |
maxRequests | integer, 1–90 | 20 | Hard cap on started direct requests. Each started request is attempted at most once; later candidates may not be reached when the request budget is used. |
timeoutMs | integer, 1,000–30,000 | 15,000 | Per-request timeout, capped by the remaining request-work deadline. |
deadlineMs | integer, 10,000–180,000 | 120,000 | Request-work deadline measured from run start; final dataset/KVS publication may finish shortly after it. |
includeDescription | boolean | true | When false, keeps required text and omits optional HTML/section/item arrays. |
There are no proxy, custom-header, retry, fixture, or debug input controls. Integer settings are clamped to their documented range; non-integer values use the runtime default. Search candidate, maxItems, request, and deadline limits can result in fewer records than requested. In multiple mode, maxCandidates truncates the supplied URL list before requests begin.
includeDescription: false retains the required text description while omitting optional HTML, section, and bullet arrays. Search currently reads one public listing page. Requests are sequential and each started URL is attempted at most once. If a public access barrier is returned, the run stops with SKIPPED and publishes no job rows. Unverified transport failures are DEFERRED rather than treated as proof that Meteojob blocked access.
Each record must contain more than 20 distinct populated source-fact fields to be considered complete. The quality.meaningfulFieldCount value excludes duplicate aliases, derived summaries, identifiers, and request/provenance metadata.
Meteojob’s current apply control is rendered as a client-side button. The actor therefore preserves directApply and the application method when published by the source, but does not fabricate an applyUrl from the job URL.
Output
You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.
Search records are detail-enriched before they are written. RUN_SUMMARY reports the final status and record/request counts; RUN_REQUESTS keeps one receipt per ordinary public request; and RUN_DIAGNOSTICS explains parse failures or unavailable requests without mixing diagnostic objects into the job dataset. quality.meaningfulFieldCount is computed over canonical populated source facts and each emitted record needs at least 21. The structured sourceRecord is retained for auditing with contact data redacted.
Example (a compact subset of a record from a fresh local Actor run):
{"recordType": "job","id": "57063397","title": "Développeur .Net / React H/F","company": "Teamis","location": "Paris (75)","contractTypes": ["CDI"],"salary": "50 000 € - 75 000 € par an","salaryRanges": [{"min": 50000,"max": 75000,"currency": "EUR","period": "ANNUM"}],"responsibilities": ["Concevoir et développer des services backend à l'aide d'ASP.NET Core et EF Core,","Construire des API basées sur Python et des composants de traitement de données avec FastAPI et Pandas,"],"jobUrl": "https://www.meteojob.com/jobs/57063397","detailVerified": true,"quality": {"complete": true,"meaningfulFieldCount": 36}}
Local verification
npm ci --ignore-scripts --no-audit --no-fundnpm run checknpm testnpm run lintapify validate-schema
INPUT-search-global.json, INPUT-single.json, INPUT-multiple.json, INPUT-no-description.json, and INPUT-negative.json provide reproducible local checks. Run evidence is written to RUN_SUMMARY, RUN_DIAGNOSTICS, RUN_REQUESTS, RUN_HEALTH, and RUN_METADATA in the default local key-value store.
For an Actor run, use a new relative storage directory and preserve it with --resurrect. This PowerShell example generates a fresh name, refuses to proceed if legacy apify_storage contents exist, and passes the same isolated directory to the validator:
$storageDir = "storage-meteojob-validation-$([guid]::NewGuid().ToString('N'))"if (Test-Path -LiteralPath "apify_storage") { throw "Preserve legacy apify_storage contents separately before running." }if (Test-Path -LiteralPath $storageDir) { throw "Choose a fresh storage directory." }$env:APIFY_LOCAL_STORAGE_DIR = $storageDirapify run --resurrect --input-file INPUT.jsonif ($LASTEXITCODE -ne 0) { throw "Actor run failed; preserve the isolated run directory for inspection." }npm run validate -- $storageDirif ($LASTEXITCODE -ne 0) { throw "Dataset validation failed; preserve the isolated run directory for inspection." }Remove-Item Env:APIFY_LOCAL_STORAGE_DIR
The validator accepts only a relative path that resolves inside this Actor folder. Do not reuse a prior run directory; keep earlier local run data intact.
How much will it cost to scrape Meteojob?
Apify compute usage depends on the plan and runtime available to your account; this Actor does not promise a fixed price. Search mode uses one listing request and then one sequential detail request for each candidate needed to fill maxItems, subject to candidate, request, timeout, and request-work-deadline limits. single and multiple modes use one detail request per started URL, subject to maxItems, maxCandidates, maxRequests, and the deadline. Requests are sequential, with no retries or retry-driven request multiplier.
Tips and advanced options
- Raise
maxCandidatesif the first search results do not contain enough complete detail records; the search page is bounded to one listing page. - In
multiplemode, setmaxCandidateshigh enough for the portion ofjobUrlsyou want considered, and setmaxItemsto the maximum number of complete rows you want returned. - Use
maxRequestsanddeadlineMsto cap request work. These limits can leave later candidate URLs unrequested; requests that start are single-attempt and sequential. - Set
includeDescriptiontofalseonly when you need the required readable description but do not need HTML, sections, or bullet arrays.
FAQ, support, and responsible use
NO_RESULTSmeans the run ended with zero records and zero diagnostics. If a search page loads but contains no parseable cards, the Actor recordsSOURCE_PARSE_FAILEDand returnsFAILEDinstead.FAILEDwithSOURCE_PARSE_FAILEDmeans the expected public listing/detail structure was not parsed; inspect the diagnostic and receipts rather than treating it as a clean empty search.SKIPPEDmeans the ordinary public request received HTTP 401, 403, 429, or 451, or a clear source-authored access challenge. The Actor stops, discards staged job rows, and preserves the diagnostic and receipts. Do not try another route, browser identity, proxy, or cookies to get around that barrier.DEFERREDmeans the source could not be verified because of a generic transport failure or an unverified response such as HTTP 407. The Actor does not retry or label this as a source block.- For other failures, inspect
RUN_SUMMARY,RUN_DIAGNOSTICS, andRUN_REQUESTSin the run's key-value store and include the run ID when contacting support through the Actor's Issues tab.
This Actor reads public job postings only and is not affiliated with Meteojob. It does not extract private account data, and contact details in retained source payloads are redacted, but public postings may still contain personal or professional information. Use results for a lawful, legitimate purpose, follow Meteojob's terms, robots guidance, and published access limits, and observe applicable privacy and data-use laws. Use the Actor's API tab for run/API access and the Issues tab for bugs or feature requests; include a run ID and summary when possible.