# Yacht Crew Jobs Scraper: Yacrew, Meridian, Crew Network (`solalab_digital/yacht-crew-jobs-scraper`) Actor

Superyacht crew vacancies from Yacrew, Meridian and The Crew Network in one normalised dataset: department, seniority tier, yacht length, rotation, start month and monthly-pay restatement. Pick boards, exclude expired listings by default, get only what's new. Each run writes an HTML dashboard.

- **URL**: https://apify.com/solalab\_digital/yacht-crew-jobs-scraper.md
- **Developed by:** [Sankov Vadim](https://apify.com/solalab_digital) (community)
- **Categories:** Jobs
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Superyacht Crew Jobs Scraper — Yacrew, Meridian & The Crew Network

Yacht crew vacancies are split across boards that publish the same job in three different shapes. This Actor reads **Yacrew**, **Meridian** and **The Crew Network** in one run and returns a single normalised dataset — every board's department, contract type and vessel length on the same scale, whichever board a row came from.

> **Disclaimer.** This is an **unofficial, independent job aggregator**. It is **not affiliated with, endorsed by, sponsored by, or connected to** Yacrew, Meridian, The Crew Network, or any employer or yacht listed on those sites. Their names are used here only to identify the publicly accessible websites this Actor reads. All job content belongs to its respective publishers.

### What this Actor does

A recruiter or crew agency covering the superyacht market has to check three sites with three different data models: one publishes an expiry date that means what it says, one computes a fake one on every page load, and one never removes an old posting at all. This Actor absorbs those differences so you don't have to.

1. It reads each selected board through the method that actually fits it — Yacrew's sitemap, Meridian's paginated listing pages, The Crew Network's public JSON API — and opens the individual job page or record where the board publishes richer data there.
2. One board blocked or slow never fails the run: each source is crawled independently, and a broken source is reported `FAILED` in the run summary while the others finish and are billed normally.
3. It maps all three sources onto **one shared record shape** (44 columns), so `department`, `contractType` and `vesselLengthM` mean the same thing in every row. A field only one board publishes is never thrown away — it survives under `siteSpecificExtra`.
4. It works out whether a listing is actually open using **the signal that is true for that board**, not a one-size-fits-all date check — see "How liveness works" below.
5. It **enriches** each record with derived fields no board publishes: a monthly wage in USD and EUR, a seniority tier and a days-since-posted count.
6. Every string field is scrubbed of emails, phone numbers and messaging handles before it is saved — this Actor collects job postings, not the personal contact details sometimes left inside them.
7. It offers a cross-board `possibleDuplicateKey`, because the same berth is sometimes posted through more than one channel.
8. When the run finishes it builds an **HTML run report** and saves it to the run's key-value store.

#### The three sources

| Source | Site | Approx. active vacancies | Extraction | Notes |
|---|---|---|---|---|
| Yacrew | `yacrew.com` | A few hundred, out of 23,600+ pages in its sitemap | Sitemap → individual job pages | The only source with a real archive: `includeArchive` unlocks it here specifically |
| Meridian | `meridiango.com` | ~550 | 56 paginated listing pages → job pages for richer records | No posting date field at all — see `firstSeenAt` below |
| The Crew Network | `crewnetwork.com` | ~75 | Public JSON API, no login | The only source that publishes the yacht's own name |

Use the **`sources`** input to pick which boards actually run. Every source you include adds to the record count, and therefore the cost, of a run — deselecting a board you don't need is the main cost control here.

#### How liveness works — two different signals, on purpose

Yacrew publishes a real, checkable expiry signal: the page's own Apply button disappears exactly when a listing closes. Meridian and The Crew Network do not — Meridian's "valid through" field is computed fresh on every page load (`today + 30 days`) rather than stored, and neither board removes old postings from its listing at all (Meridian entries from 2023 were still listed at the time this was built).

So this Actor reads each board on its own terms instead of forcing one rule onto all three:

- **Yacrew** — a record is `isArchived: true` or `false` based on that real Apply-button/expiry signal. With `includeArchive` off (the default), you get only currently open listings; turn it on to pull the archive too, useful as salary and demand history.
- **Meridian and The Crew Network** — every record is `isArchived: false`. Neither board publishes an expiry, and neither removes old postings, so "currently in the board's own listing" is the signal, full stop — that's how these two boards work, not a gap in the parser. Want to know when a listing disappeared from one of them? Compare two runs' `sourceListingId` sets.

#### `firstSeenAt` — for boards that publish no reliable posting date

Meridian publishes no posting date on most listings (a usable date exists on some job pages, not the fast listing view), and a small share of its individual pages have none at all. For any record that arrives with no date, this Actor records the moment **it** first saw that listing as `firstSeenAt`, with the record flagged `FIRST_SEEN_BASELINE` on the very first run of a given `stateKey` (because on that run "first seen" really means "seen for the first time by this Actor", not "posted today"). The `postedWithinDays` filter uses `firstSeenAt` as a fallback for exactly these records, and excludes them by default when it cannot tell whether they are within the window.

#### How it treats the source sites

- `robots.txt` is fetched and honoured for every source on every request; personal-data zones (candidate search, CV pages, recruiter directories) are never requested regardless of what robots.txt allows.
- Each site's own published rate limit wins where it is stricter than your `maxRequestsPerMinute`. The Crew Network publishes a 10-second crawl delay and is never crawled faster than 6 requests per minute regardless of the input.
- Only publicly accessible pages are read. There is **no login, no paywall circumvention and no bot-protection evasion** of any kind.
- Every text field is scrubbed for emails, phone numbers and common messaging-app handles before it reaches the dataset or the run report. This is pattern-based, not a guarantee against every possible format, and it does not remove a person's name on its own when no contact detail sits next to it.

### Input parameters

Every parameter is optional. With no input at all the Actor pulls a small 60-record sample so a plain run finishes quickly — raise `maxItems` once you know how many records your filters actually return.

| Parameter | Type | Default | Description |
|---|---|---|---|
| `sources` | array | all three | Which boards to crawl: `yacrew`, `meridian`, `crewnetwork`. |
| `includeArchive` | boolean | `false` | Also return closed listings. Applies to **Yacrew only** — see "How liveness works" above. |
| `maxItems` | integer | `60` | Hard cap on records saved across all sources combined, counted after the filters below. `0` for no cap. Enforced exactly. |
| `onlyNew` | boolean | `false` | Skip listings an earlier run with the same `stateKey` already saved. Detects by listing id; an edited listing already saved is not "new". |
| `categories` | array | all | Keep only these departments: `DECK`, `ENGINEERING`, `GALLEY`, `INTERIOR`, `SPECIALIST`, `SHORE_BASED`, `OTHER`. |
| `positionKeywords` | array | none | Keep only vacancies whose position or title contains one of these words (case-insensitive). Up to 20 entries, 60 characters each. |
| `regions` | array | none | Keep only vacancies whose location text contains one of these words. Same limits as above. Listings with no location text are excluded while this is set. |
| `postedWithinDays` | integer | none | Keep only vacancies from the last N days, by `datePosted`/`dateUpdated`, or `firstSeenAt` where a board has no date. |
| `minVesselLengthM` / `maxVesselLengthM` | integer | none | Yacht length bounds in metres. Listings with no known length are excluded while either is set. |
| `includeJobDescription` | boolean | `true` | Include the plain-text description (contact details always removed). Turn off for a compact export. |
| `stateKey` | string | `default` | Names the saved-listing set `onlyNew` reads and writes. Change it to keep two schedules apart. |
| `maxRequestsPerMinute` | integer | `40` | Courtesy rate limit per source. A site's own stricter limit still wins. |
| `proxyConfiguration` | object | disabled | Usually unnecessary — all three sources serve plain HTML or a public JSON API with no challenge. |

#### Example input

```json
{
  "sources": ["yacrew", "meridian", "crewnetwork"],
  "categories": ["DECK", "INTERIOR"],
  "minVesselLengthM": 40,
  "maxItems": 500
}
```

### Output

One dataset item per vacancy, on a **stable 44-column set that does not depend on which sources a run included**. Fields a given board does not publish are `null` rather than omitted.

#### Core fields

| Field | Type | Description |
|---|---|---|
| `sourceSite` | string | `yacrew`, `meridian` or `crewnetwork`. Part of the public contract; these values will not be renamed. |
| `sourceUrl` | string | A page a person can open for this listing. |
| `sourceListingId` | string | The board's own id. Unique **per source**, not globally. |
| `dataCompleteness` | string | `FULL` (built from a detail page/record) or `LISTING_ONLY` (built from a listing card, when `includeJobDescription` allows the faster path). |
| `scrapedAt` | string | ISO 8601 timestamp of the extraction. |
| `isArchived` | boolean | See "How liveness works" above. Always `false` for Meridian and The Crew Network. |
| `title` / `positionLabel` | string | Posting title, and the role on its own where the board publishes them separately. |
| `department` | string | One of the values above. Normalised from each board's own category or, where a board publishes none, inferred from the position text and flagged `DEPARTMENT_INFERRED_FROM_POSITION`. |
| `seniorityTier` | string | Coarse level derived from the role text — a heuristic, not an official grade. |
| `companyName` / `agencyName` | string | `companyName` is `null` on every record — none of the three boards publishes the actual employer, and a placeholder would be a lie. `agencyName` holds the crewing/recruitment agency where a board publishes one; The Crew Network's own name (`TCN`) is kept as published, not filtered out. |
| `locationText` | string | As published; the three boards give this at different granularity (region, city, or only inside the description). |
| `yachtName` | string | The vessel's name — **published only by The Crew Network**, and not always a real yacht (some are agency office listings; see FAQ). |
| `vesselType` | string | `MOTOR` or `SAILING`, where a board's field actually means that. Meridian's own "Vessel Type" field means the yacht's use (private/charter), not this, so it stays `null` there — the raw value is kept under `siteSpecificExtra`. |
| `vesselLengthM` / `vesselLengthFt` | number | Yacht length. `vesselLengthSourceUnit` records which unit the board actually published. |
| `salaryMin` / `salaryMax` / `salaryCurrency` / `salaryPeriod` | number/string | Wage bounds as published. `salaryPeriod` is only ever filled when the board states it explicitly — see "No guessed salary period" below. |
| `salaryTerms` | string | Free-text pay terms the boards attach (e.g. "DOE", net/gross), kept as published. |
| `rotationRaw` / `rotationOn` / `rotationOff` | string/integer | Rotation schedule as published; `On`/`Off` day counts only when the board states a ratio like `5:1`. |
| `contractType` | string | One of `PERMANENT`, `TEMPORARY`, `SEASONAL`, `CHARTER_SHORT`, `DAY_WORK`, `DELIVERY`, `FREELANCE`, mapped from each board's own wording. `null` where a board's field means something else (e.g. Meridian's "Full-Time" is an employment mode, not a contract type) — the raw value is always kept under `siteSpecificExtra`. |
| `contractLengthValue` / `contractLengthUnit` | number/string | Contract duration, where stated as a fixed term. |
| `startDateText` / `startDateIso` / `startMonth` | string | Requested start date as published, and resolved forms where the text parses safely. |
| `datePosted` / `dateUpdated` / `validThrough` | string | As published, where a board's field is a genuine fact rather than a computed placeholder (see "How liveness works"). |
| `firstSeenAt` | string | See above. `null` for any record that arrived with its own date. |
| `daysSincePosted` | integer | From `datePosted` (or `firstSeenAt` where that is the only date available) to today. |
| `descriptionText` | string | Plain text only — contact details removed, original HTML never returned. |
| `siteSpecificExtra` | object | Everything one board publishes and the others do not, or that a shared column's strict definition excludes. See below. |
| `possibleDuplicateKey` | string | Cross-board cross-post signal. See below. |
| `extractionFlags` | array | Machine-readable notes on how a value was produced — see "Reading `extractionFlags`" below. |

#### `siteSpecificExtra`

Examples: Yacrew's raw multi-department category string when a listing spans two departments; Meridian's vessel-use wording (`Private`/`Charter and Private`) and employment-mode text; The Crew Network's public apply link and raw job id. A parser gaining a new field never silently loses it — anything that does not fit a shared column lands here instead of being dropped.

#### Derived fields

| Field | Type | Description |
|---|---|---|
| `monthlyEquivalentUsd` / `monthlyEquivalentEur` | number | The published wage restated as a monthly figure, in whichever of the two currencies the wage was published in — daily rates × 30, weekly × 52/12, annual ÷ 12, monthly unchanged. `null` when no wage is published, the currency is neither USD nor EUR, or the period is not stated (see below). |
| `salaryNormalization` | string | Why the wage above is or is not present: `NORMALIZED`, `NO_SALARY_PUBLISHED`, `PERIOD_NOT_STATED`, `CURRENCY_MISSING` or `UNSUPPORTED_PAY_PERIOD`. |
| `daysSincePosted` | integer | See above. |

#### No guessed salary period

Some listings publish a figure ("3500-3800 USD") with no period attached. This Actor does **not** assume "per month" — `salaryPeriod` and the monthly-equivalent fields stay `null` for these, flagged `PERIOD_NOT_STATED`, and the raw text is preserved under `siteSpecificExtra`. Inventing a period would be a guess dressed up as data; the run report shows what share of a run this affected so you know how much of the wage picture is missing rather than wrong.

#### Cross-posting: `possibleDuplicateKey`

The same berth is sometimes advertised through more than one channel. `possibleDuplicateKey` hashes the identifying parts every board can publish — the role, normalised, and the yacht's length rounded to a 5-metre bucket — so cross-posts collide. Fields that only some sources publish (yacht name, agency, location, salary, posting date) are deliberately **not** in the hash: including them would tie the key to whichever source happens to publish it, not to the vacancy itself.

**This Actor never de-duplicates on it.** Two matching keys are evidence, not proof — two different "Deckhand, 50 m" berths on two boards will also collide. The key is exported and the decision is yours. It is `null` when either signal is missing, and **always `null` on an archived record** — an old, closed Yacrew listing cannot be the same open vacancy as a live one, so tagging it would only add noise on the one board large enough for that noise to matter.

#### Dataset views

The dataset ships with views selectable in the Console's dataset tab and in exports, covering an overview, compensation fields, cross-posting signals, and source/completeness provenance.

### Run report (HTML dashboard)

Every run writes a self-contained HTML report to the run's key-value store under the key **`DASHBOARD`**. The Actor logs its URL at the end of the run log, and you can also reach it from the run's **Storage → Key-value store** tab. Open it in a browser — no download, no spreadsheet, no notebook.

The report shows run summary tiles (records scraped, wage coverage, distinct employers and vessel types), a wage distribution histogram, top departments and vessel types, and a source/completeness breakdown per board. Every chart has a plain data table underneath it, follows your system light/dark setting, and is readable on a phone.

### Why this rather than a bare-JSON scraper

Most job-board Actors stop at one site: whatever fields it renders, whatever shape it uses, no idea whether a listing has actually expired. The work of making that usable — reading the other boards, reconciling their field names, working out which "valid through" date is real and which is computed on every page load, converting wage periods, spotting cross-posts, then building a chart to see what you actually got — lands on you, every run.

| | Typical single-board Actor | This Actor |
|---|---|---|
| Coverage | One board | Three boards, selectable per run |
| Field names | Whatever the site used | One shared 44-column shape across all sources |
| Is it still open? | Unstated, or trusts whatever date the board shows | The signal that's actually true for each board — see "How liveness works" |
| One-board-only fields | Dropped, or left in an ad-hoc blob | Preserved under `siteSpecificExtra` |
| Wage fields | Whatever string the site printed | Parsed bounds and currency, **plus** a monthly USD/EUR equivalent — never a guessed period |
| Department | Whatever category (or none) the board used | Normalised to one vocabulary, inferred from the role where a board publishes no category, and flagged when it is |
| Cross-posts between boards | Invisible | `possibleDuplicateKey`, exported and left to you |
| Contact details | Left in the description text, if the board didn't remove them | Scrubbed from every field before it's saved |
| Missing data | Field absent, or a silent `0` | Explicit `null` with a reason code (`extractionFlags`) |
| Seeing your results | Open the JSON | An HTML report linked from every run |

### Who this is for

#### Crewing and recruitment agencies

An agency placing interior and deck crew needs to know what's actually open **today**, not a board's stale listing — filter to `categories: ["INTERIOR"]`, `minVesselLengthM: 40`, and you get open Yacrew listings, Meridian's whole current board and The Crew Network's whole current board in one table, wage already restated to a monthly figure so a day-rate posting doesn't look cheap next to a monthly one. Because `department` is normalised, "which boards are we missing candidates on" is one `GROUP BY sourceSite`, not three separate reads.

#### Independent yacht crew recruiters and placement consultants

Running a schedule with `onlyNew` turns this into a live feed of new berths instead of a manual daily check of three bookmarks — set `positionKeywords` to the ranks you place and get only the postings that match, with `possibleDuplicateKey` flagging the ones a competing agent has also posted.

#### Maritime market analysts

Track how department mix, vessel length and wage move over a season across three boards at once, or enrich an existing yacht-crew feed with a source that arrives already normalised and already flagged for cross-posting.

### Reading `extractionFlags`

Each record carries a list of short codes noting how a value was produced, so a filtered or derived column is never silently wrong. Examples: `DEPARTMENT_INFERRED_FROM_POSITION` (department guessed from the role text, not published by the board), `PERIOD_NOT_STATED` (salary figure with no period attached), `FIRST_SEEN_BASELINE` (this run is the first time this Actor has seen the listing, so `firstSeenAt` is a lower bound, not a posting date), `LIVENESS_WEAK` (a rare case on The Crew Network where a listing appeared in the board's own results but its individual record could not be confirmed). An unrecognised flag never breaks a run — anything not in the published list is kept under `siteSpecificExtra.unnormalisedValues` instead.

### Pricing

This Actor uses Apify's **pay-per-event** model. One event is charged once per job record saved to the dataset — one price for every record, live or archived, from every source. You pay nothing for pages that yield no data, nothing for a blocked source, and nothing for retries.

- **You are never billed twice for the same listing.** Records are de-duplicated by `sourceListingId` within each source during the crawl.
- **`maxItems` is a hard spend cap**, enforced exactly across all sources combined, guarded against concurrent source crawlers.

### Frequently asked questions

#### Why does a Meridian or Crew Network record never show `isArchived: true`?

Because neither board removes old postings from its own listing, and neither publishes a field that reliably means "this has expired" — see "How liveness works" above. `isArchived` is `false` for every record from these two, by design, not because old listings were filtered out.

#### Why is `yachtName` sometimes not a real yacht?

The Crew Network is the only source that publishes it, and a few of its listings are shore-based roles at an agency's own office rather than aboard a vessel — the field is published as the board wrote it, including those cases, rather than filtered.

#### Why do some records have less data than others?

Check `dataCompleteness`. `FULL` records were built from an individual job page or API record; `LISTING_ONLY` records come from a faster listing-card path used when `includeJobDescription` is off and a board's own listing view already carries enough fields.

#### Why is the salary period missing on some jobs?

Because the board did not state one. See "No guessed salary period" above — this Actor does not assume "per month".

#### How current is the data?

Every run reads the live sites. `datePosted`/`dateUpdated` come from the posting where a board publishes them, `firstSeenAt` covers the rest, and `scrapedAt` records when the extraction happened.

#### Can I run only one board?

Yes. Set `sources` to just that one. It is also the cheapest way to sample: one board, `maxItems: 20`.

#### Does it need a proxy?

Normally no. All three sources have been reached from an Apify platform run with no proxy configured — two serve plain server-rendered HTML and one a public JSON API, none behind a bot-protection challenge.

#### Can I run this on a schedule?

Yes. Use `onlyNew` with a fixed `stateKey` to get just what changed since the last run.

### Legal and responsible use

This Actor collects only publicly available job postings from employers and agencies. It does not collect candidate resumes, CVs or personal contact details — every text field is scrubbed of emails, phone numbers and messaging handles before it is saved, and paths that lead to candidate search or CV pages on these sites are never requested at all. It does not access any account-protected area and does not bypass authentication or bot protection on any of the three sources.

Job listing content remains the property of the publishers. You are responsible for ensuring your use of the data complies with applicable law — including data protection rules in your jurisdiction — and with each source site's own terms. If you intend to republish scraped listings, review those terms first.

# Actor input Schema

## `sources` (type: `array`):

Which yacht crew boards to crawl. All three are included by default. Each board is crawled independently: one blocked or broken site never fails the run, it is reported as FAILED in the run summary while the others finish. Rough sizes: Yacrew — 23,641 job pages in its sitemap, of which only the newest few hundred are still open; Meridian — about 556 listings across 56 list pages; The Crew Network — 75 listings reported by its public jobs API. Leave a source out to cut both run time and cost.

## `includeArchive` (type: `boolean`):

Also return vacancies that no longer accept applications. Applies to Yacrew only — it is the one board that keeps its expired pages online, and they outnumber the open ones by more than an order of magnitude. Turning this on therefore multiplies both the run time and the number of billed records, so raise maxItems deliberately rather than by accident. Expired records are marked isArchived = true and are useful as salary history; for a current hiring picture leave this off.

## `maxItems` (type: `integer`):

Hard cap on records saved across all sources combined, counted after the filters below. Set 0 for no cap. The cap is shared, not per source, and it is enforced exactly: the run never saves — or bills for — one record more. Leave it capped for a sample run; raise it once you know how many records your filters actually return.

## `onlyNew` (type: `boolean`):

Skip listings that an earlier run of this Actor already saved, so a scheduled run returns just what appeared since last time. State is kept between runs under the state key in the advanced section. It detects new listings by their ID: an edit to a listing already saved is not a new listing and will not come back. Changing the filters resets the part of the state that depends on them, so a filter change can never hide a listing permanently.

## `categories` (type: `array`):

Keep only vacancies in these departments. Leave empty for all. The department is normalised from each board's own category wording, so DECK means the same thing on all three. A listing whose category does not map to any of these values is returned only when the filter is empty — its raw category is always kept in siteSpecificExtra.

## `positionKeywords` (type: `array`):

Keep only vacancies whose position or job title contains one of these words (case-insensitive, OR logic). Example: chef, bosun, engineer. Matching the title as well as the position matters here: some adverts list two roles at once and then publish no single position value, so the role is only in the title. At most 20 entries of 60 characters each — the same limits the run enforces internally, listed here so a run over the limit fails at the form, not after it has started.

## `regions` (type: `array`):

Keep only vacancies whose location text contains one of these words (case-insensitive, OR logic). Example: Mediterranean, Caribbean, Fort Lauderdale. The three boards publish location at different granularity — a region on one, a city on another, only inside the description text on the third — so this is a substring match, not a place database. Listings with no location text are excluded while this filter is set. At most 20 entries of 60 characters each, the same as position keywords above.

## `postedWithinDays` (type: `integer`):

Keep only vacancies published in the last N days, measured from datePosted (or dateUpdated when that is what the board publishes). Leave empty for no date filter. Meridian publishes no dates at all: its listings are judged by firstSeenAt, the moment this Actor first saw them, and the ones whose firstSeenAt is only a baseline from the very first run are excluded while this filter is set.

## `minVesselLengthM` (type: `integer`):

Keep only vacancies on yachts at least this long, in metres. Leave empty for no lower bound. Listings whose length could not be determined are excluded while this filter is set — a length filter that silently passed unchecked records would be worse than no filter. The run summary counts them separately under filtered.unknownVesselLength, and fieldCoverage.vesselLength reports how much of the run had a length at all.

## `maxVesselLengthM` (type: `integer`):

Keep only vacancies on yachts at most this long, in metres. Leave empty for no upper bound. Same rule as the minimum: listings with an unknown length are excluded while this filter is set.

## `includeJobDescription` (type: `boolean`):

Include the free-text description of each vacancy. Turn it off for a compact, spreadsheet-friendly table. The description is always plain text with contact details removed; the original HTML is never returned.

## `stateKey` (type: `string`):

Names the set of already-seen listings that 'Only listings not seen before' reads and writes. Change it only to keep two schedules apart: two schedules with different filters sharing one state key would consume each other's new listings. It also drives firstSeenAt for boards that publish no posting date, so a fresh key means those listings are seen for the first time again. Letters, digits and !\_.'()- only, up to 64 characters — this becomes part of an Apify key-value store key, and anything else would fail that source at storage time instead of at the form.

## `maxRequestsPerMinute` (type: `integer`):

Courtesy rate limit, applied to each source separately rather than shared. Whatever a site's own robots.txt declares still wins where it is stricter — The Crew Network publishes a 10-second crawl delay and is never requested faster than 6 times a minute regardless of this setting. Lower it if you run the Actor often.

## `proxyConfiguration` (type: `object`):

Usually not needed: all three sources serve plain HTML or a public JSON API with no challenge. Add a proxy only if a source starts returning no records from Apify's own IP ranges — the run summary reports a blocked source explicitly instead of silently returning an empty dataset.

## Actor input object example

```json
{
  "sources": [
    "yacrew",
    "meridian",
    "crewnetwork"
  ],
  "includeArchive": false,
  "maxItems": 60,
  "onlyNew": false,
  "categories": [],
  "positionKeywords": [],
  "regions": [],
  "includeJobDescription": true,
  "stateKey": "default",
  "maxRequestsPerMinute": 40,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `jobs` (type: `string`):

Every scraped vacancy as a normalised 44-column record, one row per listing across Yacrew, Meridian and The Crew Network. Download as JSON, CSV or XLSX, or read it through the API.

## `dashboard` (type: `string`):

Standalone HTML report for the run: vacancies by source, department and seniority tier, the pay and yacht-length distributions, and a table under every chart. Open it in a browser.

## `runSummary` (type: `string`):

What each source actually did: status and reason, pages fetched, robots.txt skips, live versus archived counts, records filtered out by each filter, and the share of records carrying each optional field (fieldCoverage). Read it to tell an empty result apart from a blocked source.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sources": [
        "yacrew",
        "meridian",
        "crewnetwork"
    ],
    "maxItems": 60,
    "stateKey": "default",
    "maxRequestsPerMinute": 40
};

// Run the Actor and wait for it to finish
const run = await client.actor("solalab_digital/yacht-crew-jobs-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sources": [
        "yacrew",
        "meridian",
        "crewnetwork",
    ],
    "maxItems": 60,
    "stateKey": "default",
    "maxRequestsPerMinute": 40,
}

# Run the Actor and wait for it to finish
run = client.actor("solalab_digital/yacht-crew-jobs-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sources": [
    "yacrew",
    "meridian",
    "crewnetwork"
  ],
  "maxItems": 60,
  "stateKey": "default",
  "maxRequestsPerMinute": 40
}' |
apify call solalab_digital/yacht-crew-jobs-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,solalab_digital/yacht-crew-jobs-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/84e4ezixLS3eSsa99/builds/kty0PHVCDhEREAIV5/openapi.json
