Meteojob Search Scraper avatar

Meteojob Search Scraper

Pricing

from $2.99 / 1,000 meteojob job records

Go to Apify Store
Meteojob Search Scraper

Meteojob Search Scraper

Extract rich Meteojob listings with structured descriptions, locations, contracts, experience, qualifications, and publication metadata.

Pricing

from $2.99 / 1,000 meteojob job records

Rating

0.0

(0)

Developer

Jobs API

Jobs API

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

11 days ago

Last modified

Share

What does Meteojob Jobs Search Scraper do?

This Apify Actor searches public job listings and fetches exact public job details from Meteojob, returning structured job data with source provenance. It is a direct-HTTPS, Cheerio-based alternative to manual Meteojob search; it does not use a browser, proxy, mirror, fingerprint spoofing, or access-control bypass.

Why use Meteojob Jobs Search Scraper?

Use search mode to collect a small, bounded set of current listings, or choose single/multiple mode to fetch known public posting URLs. Detail records combine public JobPosting JSON-LD with Meteojob's embedded job-offer state and preserve readable description sections, source labels, salary where published, and request receipts. You can schedule runs and access their datasets, run history, and monitoring through Apify. This Actor makes one sequential attempt per started request and does not rotate proxies or retry blocked/failed requests.

How to scrape Meteojob

  1. Open the Actor's Input tab and choose search, single, or multiple.
  2. For search, enter a French keyword and optional location. For detail modes, provide only official Meteojob HTTPS posting URLs.
  3. Set the result and request limits, run the Actor, then inspect the dataset and RUN_SUMMARY / RUN_DIAGNOSTICS / RUN_REQUESTS values.

Modes

  • search requests the official /jobs?what=...&where=... SSR page, parses current offer cards, then fetches selected /jobs/<numeric-id> detail pages.
  • single fetches one exact official detail URL.
  • multiple fetches several exact official detail URLs.

Detail records combine the official JobPosting JSON-LD with the official candidate-front-state job-offer payload. This captures title, company, rich descriptions, profile and company text, benefits, location coordinates and country data, contracts, salary, job taxonomy, experience, skills/languages, synonyms, telework/travel labels, publication dates, and source provenance. Contact data is redacted from the retained structured source record. Incomplete rows are omitted and diagnostics are stored only in RUN_DIAGNOSTICS.

What data can Meteojob Jobs Search Scraper extract?

Every dataset item is a verified job record. The dataset keeps the complete source-backed record; the table below groups its most useful fields:

GroupFieldsNotes
Identityid, jobId, title, jobTitle, company, companyName, companyLogoUrl, referenceStable Meteojob posting identity and published employer data; repeated title/company fields are compatibility aliases.
Location and joblocation, locations, city, region, country, employmentType, contractTypes, industry, occupationalCategory, jobCategories, experienceLevelsStructured location parts, job classification, and employment labels are included when published.
Compensationsalary, salaryMin, salaryMax, salaryCurrency, salaryPeriod, salaryRangesThe source display is preserved; numeric values and period are omitted when the public posting does not provide them clearly.
Descriptiondescription, descriptionHtml, descriptionSections, descriptionItems, companyDescription, profileDescription, responsibilities, requirements, benefitsFull readable text is retained. HTML and section/bullet arrays are optional and controlled by includeDescription.
Skills and conditionsskills, qualifications, languages, synonyms, telework, travel, directApplyThese are source-published labels or structured posting facts, not inferred from the search phrase.
Provenance and qualityjobUrl, canonicalUrl, searchUrl, searchResultRank, requestReceipts, detailVerified, quality.meaningfulFieldCountReceipts and quality metadata explain where the record came from and whether the detail was verified. The meaningful-field count excludes aliases and provenance.

Optional fields are omitted when the source does not expose them; the Actor does not fill missing salary, application links, qualifications, or other facts with guesses. The schema permits additional source payload keys in sourceRecord so the official structured records remain available for auditing.

Input

See the input tab for the complete configuration. Use jobUrl / jobUrls for normal inputs; url / urls are compatibility aliases.

{
"mode": "search",
"query": "developpeur",
"location": "Paris",
"maxItems": 3,
"maxCandidates": 20,
"maxRequests": 20,
"timeoutMs": 15000,
"deadlineMs": 120000,
"includeDescription": true
}

For single, provide jobUrl; for multiple, provide jobUrls. Compatibility aliases url and urls are accepted. Detail URLs must be official HTTPS https://www.meteojob.com/jobs/<numeric-id> URLs.

Single-detail input (use an active public posting URL):

{
"mode": "single",
"jobUrl": "https://www.meteojob.com/jobs/57063397",
"maxItems": 1,
"maxRequests": 1
}

Multiple-detail input (both URLs must be active public postings):

{
"mode": "multiple",
"jobUrls": [
"https://www.meteojob.com/jobs/57063397",
"https://www.meteojob.com/jobs/56471308"
],
"maxItems": 2,
"maxRequests": 2
}
InputTypeDefaultBehavior
modestringsearchsearch, single, or multiple.
query, locationstringdeveloppeur, emptySearch term and optional location for search; the configured search uses one listing page.
jobUrl / jobUrlsstring / array—Exact detail URL(s) for single / multiple; url / urls are aliases. When both forms are supplied, jobUrl takes precedence in single mode and an array-valued jobUrls takes precedence in multiple mode.
maxItemsinteger, 1–503Maximum complete dataset records.
maxCandidatesinteger, 1–8020Maximum search candidates or multiple-mode URLs considered. single mode always uses its one normalized detail URL.
maxRequestsinteger, 1–9020Hard cap on started direct requests. Each started request is attempted at most once; later candidates may not be reached when the request budget is used.
timeoutMsinteger, 1,000–30,00015,000Per-request timeout, capped by the remaining request-work deadline.
deadlineMsinteger, 10,000–180,000120,000Request-work deadline measured from run start; final dataset/KVS publication may finish shortly after it.
includeDescriptionbooleantrueWhen false, keeps required text and omits optional HTML/section/item arrays.

There are no proxy, custom-header, retry, fixture, or debug input controls. Integer settings are clamped to their documented range; non-integer values use the runtime default. Search candidate, maxItems, request, and deadline limits can result in fewer records than requested. In multiple mode, maxCandidates truncates the supplied URL list before requests begin.

includeDescription: false retains the required text description while omitting optional HTML, section, and bullet arrays. Search currently reads one public listing page. Requests are sequential and each started URL is attempted at most once. If a public access barrier is returned, the run stops with SKIPPED and publishes no job rows. Unverified transport failures are DEFERRED rather than treated as proof that Meteojob blocked access.

Each record must contain more than 20 distinct populated source-fact fields to be considered complete. The quality.meaningfulFieldCount value excludes duplicate aliases, derived summaries, identifiers, and request/provenance metadata.

Meteojob’s current apply control is rendered as a client-side button. The actor therefore preserves directApply and the application method when published by the source, but does not fabricate an applyUrl from the job URL.

Output

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

Search records are detail-enriched before they are written. RUN_SUMMARY reports the final status and record/request counts; RUN_REQUESTS keeps one receipt per ordinary public request; and RUN_DIAGNOSTICS explains parse failures or unavailable requests without mixing diagnostic objects into the job dataset. quality.meaningfulFieldCount is computed over canonical populated source facts and each emitted record needs at least 21. The structured sourceRecord is retained for auditing with contact data redacted.

Example (a compact subset of a record from a fresh local Actor run):

{
"recordType": "job",
"id": "57063397",
"title": "Développeur .Net / React H/F",
"company": "Teamis",
"location": "Paris (75)",
"contractTypes": ["CDI"],
"salary": "50 000 € - 75 000 € par an",
"salaryRanges": [
{
"min": 50000,
"max": 75000,
"currency": "EUR",
"period": "ANNUM"
}
],
"responsibilities": [
"Concevoir et développer des services backend à l'aide d'ASP.NET Core et EF Core,",
"Construire des API basées sur Python et des composants de traitement de données avec FastAPI et Pandas,"
],
"jobUrl": "https://www.meteojob.com/jobs/57063397",
"detailVerified": true,
"quality": {
"complete": true,
"meaningfulFieldCount": 36
}
}

Local verification

npm ci --ignore-scripts --no-audit --no-fund
npm run check
npm test
npm run lint
apify validate-schema

INPUT-search-global.json, INPUT-single.json, INPUT-multiple.json, INPUT-no-description.json, and INPUT-negative.json provide reproducible local checks. Run evidence is written to RUN_SUMMARY, RUN_DIAGNOSTICS, RUN_REQUESTS, RUN_HEALTH, and RUN_METADATA in the default local key-value store.

For an Actor run, use a new relative storage directory and preserve it with --resurrect. This PowerShell example generates a fresh name, refuses to proceed if legacy apify_storage contents exist, and passes the same isolated directory to the validator:

$storageDir = "storage-meteojob-validation-$([guid]::NewGuid().ToString('N'))"
if (Test-Path -LiteralPath "apify_storage") { throw "Preserve legacy apify_storage contents separately before running." }
if (Test-Path -LiteralPath $storageDir) { throw "Choose a fresh storage directory." }
$env:APIFY_LOCAL_STORAGE_DIR = $storageDir
apify run --resurrect --input-file INPUT.json
if ($LASTEXITCODE -ne 0) { throw "Actor run failed; preserve the isolated run directory for inspection." }
npm run validate -- $storageDir
if ($LASTEXITCODE -ne 0) { throw "Dataset validation failed; preserve the isolated run directory for inspection." }
Remove-Item Env:APIFY_LOCAL_STORAGE_DIR

The validator accepts only a relative path that resolves inside this Actor folder. Do not reuse a prior run directory; keep earlier local run data intact.

How much will it cost to scrape Meteojob?

Apify compute usage depends on the plan and runtime available to your account; this Actor does not promise a fixed price. Search mode uses one listing request and then one sequential detail request for each candidate needed to fill maxItems, subject to candidate, request, timeout, and request-work-deadline limits. single and multiple modes use one detail request per started URL, subject to maxItems, maxCandidates, maxRequests, and the deadline. Requests are sequential, with no retries or retry-driven request multiplier.

Tips and advanced options

  • Raise maxCandidates if the first search results do not contain enough complete detail records; the search page is bounded to one listing page.
  • In multiple mode, set maxCandidates high enough for the portion of jobUrls you want considered, and set maxItems to the maximum number of complete rows you want returned.
  • Use maxRequests and deadlineMs to cap request work. These limits can leave later candidate URLs unrequested; requests that start are single-attempt and sequential.
  • Set includeDescription to false only when you need the required readable description but do not need HTML, sections, or bullet arrays.

FAQ, support, and responsible use

  • NO_RESULTS means the run ended with zero records and zero diagnostics. If a search page loads but contains no parseable cards, the Actor records SOURCE_PARSE_FAILED and returns FAILED instead.
  • FAILED with SOURCE_PARSE_FAILED means the expected public listing/detail structure was not parsed; inspect the diagnostic and receipts rather than treating it as a clean empty search.
  • SKIPPED means the ordinary public request received HTTP 401, 403, 429, or 451, or a clear source-authored access challenge. The Actor stops, discards staged job rows, and preserves the diagnostic and receipts. Do not try another route, browser identity, proxy, or cookies to get around that barrier.
  • DEFERRED means the source could not be verified because of a generic transport failure or an unverified response such as HTTP 407. The Actor does not retry or label this as a source block.
  • For other failures, inspect RUN_SUMMARY, RUN_DIAGNOSTICS, and RUN_REQUESTS in the run's key-value store and include the run ID when contacting support through the Actor's Issues tab.

This Actor reads public job postings only and is not affiliated with Meteojob. It does not extract private account data, and contact details in retained source payloads are redacted, but public postings may still contain personal or professional information. Use results for a lawful, legitimate purpose, follow Meteojob's terms, robots guidance, and published access limits, and observe applicable privacy and data-use laws. Use the Actor's API tab for run/API access and the Issues tab for bugs or feature requests; include a run ID and summary when possible.