Welcome to the Jungle (WTTJ) Jobs Scraper: Salary and RAG Data avatar

Welcome to the Jungle (WTTJ) Jobs Scraper: Salary and RAG Data

Pricing

from $0.59 / 1,000 jobs

Go to Apify Store
Welcome to the Jungle (WTTJ) Jobs Scraper: Salary and RAG Data

Welcome to the Jungle (WTTJ) Jobs Scraper: Salary and RAG Data

Scrape WTTJ (Welcome to the Jungle) job listings with salary, location, contract type, and company data across France, UK, US, Germany, and 50+ countries. No login required. Export raw JSON for spreadsheets and CRMs, or RAG-ready chunked text for vector databases and AI recruiting pipelines.

Pricing

from $0.59 / 1,000 jobs

Rating

0.0

(0)

Developer

GetAScraper

GetAScraper

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

0

Monthly active users

7 days ago

Last modified

Share

๐Ÿ’ Welcome to the Jungle (WTTJ) Jobs Scraper: Salary and RAG Data

Extract Welcome to the Jungle tech jobs as raw JSON or RAG-ready chunks. Salary ranges, remote policy, and contract type across France, UK, US, Germany, and 50+ countries, ready for spreadsheets or vector databases.
Global technology and remote hiring ย ย โ€ขย ย Hacker News Hiring, GoFractional, NoFluffJobs, and Welcome to the Jungle tech roles
ย HN Hiring
Structured jobs from HN threads
ย GoFractional
Fractional jobs, direct ATS links
ย NoFluffJobs
European tech jobs and salaries
ย WTTJ
โžค You are here
Welcome to the Jungle (WTTJ) jobs: salary, location, and RAG-ready data
Job listings with salary ranges, remote policy, contract type, and company details across France, UK, US, Germany, and 50+ countries. Export clean JSON or pre-chunked text for vector databases.
๐Ÿ”€ Dual output modes
Raw JSON for pipelines, RAG-ready chunks for AI agents
๐Ÿงฉ Framework ready
Drops into LangChain, LlamaIndex, Pinecone, or Qdrant
๐ŸŒ Europe plus US coverage
France, UK, Germany, Spain, and 50+ countries
๐Ÿ”“ No auth required
Public Algolia API, no login or API key needed

Extract job listings from Welcome to the Jungle (WTTJ) with two output modes: raw JSON or RAG-ready chunks. Get salary ranges, company data, locations, contract type, and remote policy across France, UK, US, Germany, and 50+ countries. Drop raw output into spreadsheets or pipelines. Use RAG-ready chunks directly in Qdrant, Pinecone, Weaviate, LangChain, or LlamaIndex for AI-powered job matching and resume screening. No login or API key needed.

๐Ÿ” What does Welcome to the Jungle Scraper do?

This Actor queries WTTJ's public Algolia search API to extract structured job data. It supports two output modes:

  1. Raw JSON - Full job listings with title, company, salary, locations, contract type, remote policy, summary, and sectors. Ready for spreadsheets, CRMs, or data warehouses.

  2. RAG-Ready - Job descriptions tokenized into fixed-size chunks using tiktoken cl100k_base. Configurable chunk size (default 512 tokens) with overlap. Drop straight into any vector database for LLM-based job matching, resume screening, or market intelligence.

The scraper covers 10,000+ companies across Europe and the US. No login, no API key, no proxy required.

๐Ÿ’ก Why use Welcome to the Jungle Scraper for RAG?

  • Dual output modes. Raw JSON for traditional pipelines, RAG-ready chunks for AI agents and vector databases. One Actor handles both workflows.
  • RAG-ready chunks. Job descriptions pre-split for LLM ingestion. No LaTeX stripping, no HTML parsing, no custom chunking logic on your side. Works with Qdrant, Pinecone, Weaviate, pgvector, and Chroma.
  • Framework-ready. Raw output drops into Google Sheets, BigQuery, n8n, or any CRM. RAG chunks drop into LangChain, LlamaIndex, Haystack, or custom LLM pipelines.
  • Europe plus US coverage. France, UK, Germany, Spain, and 50+ countries in one query. Compare salary ranges across markets.
  • Rich job metadata. Salary ranges (EUR, GBP, USD), remote policy, contract type, experience level, company size, and hiring velocity.
  • No auth required. Uses WTTJ's public Algolia API. No cookies, no API key, no account.

๐Ÿš€ How to use Welcome to the Jungle Scraper

STEP 1
Set your search
Enter a query like data engineer, product manager, or sales.
STEP 2
Choose output mode
Pick raw for standard JSON or rag for chunked text.
STEP 3
Filter and run
Add country, contract type, or remote policy filters, then export.

๐Ÿ“ฅ Input

FieldTypeRequiredDescription
querystringYesJob title, skill, or keyword to search.
localeenumNoLanguage index to search: en or fr. Default: en.
countryCodesarray of stringsNoFilter by country codes (e.g., FR, GB, US).
locationsarray of stringsNoFilter by city or region (e.g., Paris, London).
contractTypesarray of stringsNoFilter by contract type (full_time, part_time, internship, freelance, apprenticeship, vie).
remotePoliciesarray of stringsNoFilter by remote policy (full, partial, punctual, no).
categoriesarray of stringsNoFilter by job category (tech, data, sales, marketing).
includeJobDetailsbooleanNoFetch the full candidate profile text for each listing. Default: false.
outputModeenumYesraw for full JSON or rag for chunked text. Default: raw.
chunkSizeintegerNoTarget tokens per chunk in RAG mode. Default: 512.
chunkOverlapintegerNoToken overlap between chunks in RAG mode. Default: 50.
maxItemsintegerYesMaximum jobs to return. Default: 100.
proxyConfigurationobjectNoProxy settings for the run. Defaults to Apify Proxy.

๐Ÿ“ค Output

Raw Mode Example

{
"jobId": "2d4fbe25-352d-47a4-8280-bd6d4642cfb1",
"slug": "data-analytics-engineer-visian-xxx",
"title": "Data analytics engineer",
"contractType": "full_time",
"remotePolicy": "partial",
"publishedAt": "2026-06-04T09:39:54Z",
"experienceMin": 3,
"educationLevel": "master",
"profession": {
"name": "Data analyst",
"category": "Data"
},
"company": {
"name": "Visian",
"slug": "visian",
"size": "50-249 employees",
"employeeCount": 200,
"description": "Consulting et execution technique pour projets data.",
"sectors": ["Artificial Intelligence / Machine Learning", "IT / Digital"]
},
"locations": [
{ "city": "Courbevoie", "country": "France", "countryCode": "FR", "region": "Ile-de-France" }
],
"primaryLocation": "Courbevoie, France",
"salary": {
"hasSalary": true,
"min": 42000,
"max": 55000,
"yearlyMinimum": 42000,
"currency": "EUR",
"period": "yearly"
},
"summary": "Rejoignez Visian, une societe de conseil specialisee en innovation...",
"benefits": ["Remote work", "Health insurance"],
"applyUrl": "https://www.welcometothejungle.com/en/jobs/data-engineer-visian-xxx/apply",
"url": "https://www.welcometothejungle.com/en/jobs/data-engineer-visian-xxx",
"scrapedAt": "2026-06-06T10:00:00.000Z"
}

RAG Mode Example

{
"jobId": "2d4fbe25-352d-47a4-8280-bd6d4642cfb1",
"title": "Data analytics engineer",
"company": { "name": "Visian", "slug": "visian" },
"primaryLocation": "Courbevoie, France",
"salary": { "hasSalary": true, "min": 42000, "max": 55000, "currency": "EUR" },
"chunks": [
{ "idx": 0, "text": "Job Title: Data analytics engineer\nCompany: Visian\nLocation: Courbevoie, France\n\nRejoignez Visian...", "tokens": 256 },
{ "idx": 1, "text": "...SQL, Databricks, Git et Power BI...", "tokens": 128 }
],
"scrapedAt": "2026-06-06T10:00:00.000Z"
}

RAG mode drops the summary field and replaces it with chunks. Every other field from raw mode is still present.

๐Ÿ“Š Data Table

FieldTypeDescription
jobIdstringWTTJ unique job identifier.
slugstringJob URL slug.
titlestringJob title as posted.
contractTypestringContract type code (full_time, part_time, internship, freelance, apprenticeship, vie).
remotePolicystringRemote work policy (full, partial, punctual, no).
publishedAtstringISO date the job was published.
experienceMinintegerMinimum years of experience required, when disclosed.
educationLevelstringMinimum education level required, when disclosed.
professionobjectJob classification: name and category.
companyobjectEmployer details: name, slug, url, logo, size, employeeCount, description, sectors.
locationsarrayLocation objects with city, country, countryCode, region.
primaryLocationstringPrimary location as a formatted string.
salaryobjectSalary details: hasSalary, min, max, currency, period.
summarystringShort job summary. Raw mode only.
profilestringCandidate profile or requirements text, when includeJobDetails is enabled.
benefitsarrayListed job benefits, when disclosed.
applyUrlstringDirect application URL, when available.
urlstringCanonical job URL on welcometothejungle.com.
chunksarrayRAG mode only: array of { idx, text, tokens } chunks replacing summary.
scrapedAtstringISO timestamp of extraction.

๐Ÿ’ฐ Pricing

Pay per result. You are billed only for job listings successfully saved to your dataset. An empty run costs nothing, and there is no subscription or minimum spend. See the Actor's Pricing tab in Apify Console for the current rate.

โญ Enjoying the WTTJ Jobs Scraper?

โญ โญ โญ โญ โญ
Getting job data pre-chunked for a vector database instead of writing your own tokenizer and cleanup logic?
A 5-star rating takes 10 seconds and helps other recruiters and RAG pipeline builders find it. Your feedback also tells us what to build next.
โ˜…ย ย Rate this Actor on Apify

๐Ÿ› ๏ธ Tips

  • Use RAG mode for AI agents. Feed chunks directly into a vector database for job matching, resume screening, or market intelligence. Compatible with LangChain, LlamaIndex, Qdrant, Pinecone, and Weaviate.
  • Filter by country first. WTTJ has different job densities per market. France and UK have the most listings.
  • Combine with salary filters. Many jobs disclose salary ranges. Filter salary.hasSalary: true downstream for benchmarking.
  • Schedule weekly for market tracking. Compare job counts and salary ranges over time to detect hiring trend shifts.

โ“ FAQ, Disclaimers, and Support

Is scraping Welcome to the Jungle legal? This Actor uses WTTJ's public Algolia search API, which requires no authentication. It collects publicly visible job listing data for research, market analysis, and recruiting intelligence. Users are responsible for ensuring their use complies with applicable laws and WTTJ's terms of service.

Why are some salary fields empty? Many jobs on WTTJ do not disclose salary. The Actor only saves salary values that the source exposes. Filter salary.hasSalary: true downstream if salary data is critical.

What is RAG mode? RAG (Retrieval-Augmented Generation) mode splits job descriptions into fixed-token chunks using tiktoken cl100k_base encoding. These chunks are ready to embed and store in a vector database for LLM-based job matching or resume screening.

How is this different from other WTTJ scrapers? This is the only WTTJ scraper with dual output modes: raw JSON for traditional pipelines and RAG-ready chunks for AI agents. No other scraper offers tiktoken-based chunking optimized for LLM ingestion.

Support: Open an issue on the Actor's Issues tab in Apify Console for bug reports, feature requests, or custom-run quotes.


Built with Apify + Crawlee + TypeScript. Part of the actorstack portfolio.

๐Ÿ”— Other actors