Website Job Extractor (Browser) avatar

Website Job Extractor (Browser)

Pricing

from $8.00 / 1,000 job extracted (browser)s

Go to Apify Store
Website Job Extractor (Browser)

Website Job Extractor (Browser)

Extract job listings from JavaScript-rendered career pages (React, Vue, Angular) using AI + Playwright. Companion to the HTTP-only Website Job Extractor. Use it for the ~28% of company sites that need a real browser. Same output format, same quality, same LLM fallback chain.

Pricing

from $8.00 / 1,000 job extracted (browser)s

Rating

0.0

(0)

Developer

NanoScrape

NanoScrape

Maintained by Community

Actor stats

0

Bookmarked

38

Total users

4

Monthly active users

11 hours

Issues response

19 minutes ago

Last modified

Share

Extract job listings from JavaScript-rendered career pages (React, Vue, Angular SPAs) using AI + Playwright. This is the specialist tool in a three-tier extraction chain — use it as a fallback, not a default.

Which tool should I use? (Read this first)

Running a browser + LLM is inherently more expensive than parsing static HTML. Try these in order:

  1. website-job-extractor (HTTP + AI) — try this first. Handles ~70% of company career pages. About 5× cheaper per company than this browser variant. Set enableBrowserFallback: true on its input and it will automatically hand off JS-only sites to this actor.

  2. career-site-jobs-scraper — if you know the URL is hosted on a supported ATS platform (Lever, Greenhouse, Workable, Ashby, Teamtailor, Personio, SmartRecruiters, BambooHR, Workday, iCIMS, Recruitee, JOIN, Pinpoint, Rippling, JazzHR, Comeet). Deterministic, HTTP-only, no LLM cost. Cheapest of the three.

  3. This actor — when the HTTP extractor returns nothing and the URL isn't a known ATS. Renders the page with Playwright, extracts with an LLM. Slower and pricier, but works on anything.

When to use this actor

  • Career pages built with React, Vue, Angular, or other JS frameworks
  • Pages that return empty/skeleton HTML without JavaScript execution
  • Companies flagged by the HTTP actor's JS-rendering detection
  • Auto-chained via enableBrowserFallback on the HTTP actor (was enablePlaywrightFallback, deprecated)

Input shape: website_url vs career_urls

These two fields serve different purposes:

FieldWhat it meansWhen to use it
website_urlCompany homepage. The actor runs discovery from here: probes subdomain patterns, reads navigation links.When you only have the company's main website and want the actor to find the career page for you.
career_urlsSpecific job listing URLs, SERP pages, or detail pages. The actor extracts directly from these without discovery.When you already know the career page URL, or the URL is already a search results page (e.g., amazon.jobs/en/search?base_query=...).

If your input URL is already a SERP or job listing, pass it via career_urls, not website_url. Homepage discovery on a SERP page extracts nav links, which are all rejected as junk category pages.

As of v1.1.8, the actor performs best-effort SERP detection: if website_url looks like a search results page (contains ?base_query=, ?q=, /search?, etc.) and career_urls is empty, it treats website_url as a direct extraction target automatically. Explicit is still better.

Use with AI Agents (MCP)

Connect this actor to any MCP-compatible AI client — Claude Desktop, Claude.ai, Cursor, VS Code, LangChain, LlamaIndex, or custom agents.

Apify MCP server URL:

https://mcp.apify.com?tools=santamaria-automations/website-job-extractor-browser

Example prompt once connected:

"Use website-job-extractor-browser to process data with website job extractor browser. Return results as a table."

Clients that support dynamic tool discovery (Claude.ai, VS Code) will receive the full input schema automatically via add-actor.

How it works

  1. Playwright renders the full page (waits for network idle + text content)
  2. Career page discovery from homepage navigation (same as HTTP actor)
  3. ATS detection for 19 systems (Personio, Greenhouse, Softgarden, etc.)
  4. LLM extraction using Gemini Flash / Groq / OpenRouter
  5. Validation with confidence scoring and deduplication
  6. Pagination follow-up for multi-page listings

Same extraction pipeline as the HTTP actor — same output format, same quality.

Input

Same input format as the HTTP actor. Typically auto-chained:

{
"companies": [
{
"company_id": "abc-123",
"company_name": "TechCorp AG",
"website_url": "https://techcorp.ch"
}
],
"llmProvider": "gemini",
"geminiApiKey": "YOUR_KEY"
}

Output

Each job is a dataset item with browser_extraction: true:

{
"company_id": "abc-123",
"company_name": "TechCorp AG",
"title": "Senior Frontend Developer (m/w/d)",
"location": "Zürich",
"employment_type": "Vollzeit",
"department": "Engineering",
"job_url": "https://techcorp.ch/jobs/senior-frontend-developer-zurich",
"job_id": "senior-frontend-developer-zurich",
"application_url": "https://techcorp.ch/jobs/apply/123",
"confidence": 0.85,
"browser_extraction": true,
"extracted_at": "2026-03-09T10:00:00.000Z"
}

Memory requirements

  • Minimum: 1024 MB (Playwright + Chrome)
  • Recommended: 2048 MB for 5+ companies
  • Maximum: 4096 MB

LLM API keys -- BYOK vs Managed mode

This actor works in two modes:

Managed mode (default -- no key needed)

Leave all API key fields blank. The actor uses a shared OpenRouter account with deepseek/deepseek-chat-v3-0324 as the extraction model. No sign-up required.

Budget caps by tier:

TierDefault capMaximum cap
Free Apify account$0.30 per run (hard ceiling)$0.30 (cannot raise)
Paying Apify account$10.00 per run$50.00 (via managedLlmBudgetUsd)

When the cap is reached mid-run, a managed_llm_run_capped sentinel row is pushed to the dataset and the actor exits cleanly with a message telling you how to continue. No crash.

Managed mode PPE events (charged per extracted job in managed mode):

EventPrice
Company processed$0.05
Job extracted (managed)$0.024

The 3x premium on managed jobs vs BYOK jobs ($0.024 vs $0.008) covers the deepseek-v3 LLM cost plus a margin. Browser compute overhead means the premium is in absolute terms the same multiplier as HTTP but at a higher base price due to Camoufox/Playwright costs.

BYOK mode (bring your own key)

Supply any of geminiApiKey, groqApiKey, or openrouterApiKey. The actor uses your quota, your rate limits, and the standard PPE event prices:

EventPrice
Company processed$0.05
Job extracted (BYOK)$0.008

Managed mode input

{
"companies": [
{
"company_id": "stripe",
"company_name": "Stripe",
"career_urls": ["https://stripe.com/jobs/search"]
}
]
}

No API key fields required. The actor runs immediately.

Raising the managed budget (paying users only)

{
"companies": [...],
"managedLlmBudgetUsd": 25
}

Free-tier users cannot raise the cap above $0.30 regardless of this input.

Sentinel rows

The dataset may also contain sentinel rows with a _type field. These are not job listings but machine-readable events:

_type valueWhen emittedExtra fields
job_extraction_doneWhen all pages for a company are processedjob_count, jobs_extracted, pages_fetched, pagination_source (css_selector, url_pattern, button_click, or none)
llm_extraction_failedWhen every LLM call failed (all providers exhausted)error_category, provider, model, attempts, user_message
managed_llm_run_cappedWhen the managed LLM budget cap is reached mid-run (managed mode only)error_category (free_tier_budget_reached or paying_tier_budget_reached), tier, spent_usd, cap_usd, user_message

All sentinel rows include the four required dataset fields: company_id, company_name, source_url, extracted_at.

Pricing

Browser-based extraction costs more than the HTTP actor due to Chrome overhead. PPE events reflect the actual compute done:

EventListed rateWhen it fires
browser-company-enriched$0.05 / companyOnce per company on the first page render
browser-page-processed$0.005 / pageSee the smart-pricing note below
browser-job-result$0.008 / jobEach validated, non-duplicate job (BYOK mode)
browser-job-result-managed$0.024 / jobEach validated, non-duplicate job (managed mode)

Smart per-page pricing (why your actual bill is lower than the listed rate)

The actor decides internally which proxy pool to use for each URL, defaulting to a fast primary pool and only escalating to residential when a URL gets blocked (403, 429, or repeated timeouts). Because the primary pool costs a fraction of what residential does, we pass the saving on rather than pocket it:

  • browser-page-processed fires only on the 1st, 5th, 9th, ... render that succeeds on the primary pool (1 in 4).
  • Renders that fell back to residential bill 1:1, since that's where our real cost sits.

Result: on typical company career pages, you pay for roughly one page in four, and the actor's structural proxy strategy handles the "hard" URLs invisibly.

Cost examples

Small test (1 company, 3 primary-pool pages, 5 jobs): $0.05 + 1 × $0.005 + 5 × $0.008 = **$0.095** (first page bills, next two are absorbed by the 1-in-4 cadence)

Bulk run (900 companies, mostly primary pool, ~1,850 jobs): Roughly $65-75/run in practice, vs the ~$130 the flat listed rate would imply.

Historical context (why the pricing model is what it is)

An earlier flat per-page charge introduced 2026-07-27 hit deep-crawl users disproportionately. Investigation showed our real cost was almost entirely residential-proxy bandwidth, not compute. Fixing that (making the actor decide proxy strategy per-URL) removed the underlying cost driver, and the smart per-page cadence passes the saving on to you.

Pagination stops when maxResultsPerCompany (default 200) jobs have been extracted per company, or when the 20-page hard safety cap fires. The proxyConfiguration input has been removed; if you send one, it's silently ignored.

Troubleshooting: LLM key errors

Runs that fail with LLM provider ... rejected the supplied API key or LLM model ... is unavailable to your ... key almost always fall into one of these three buckets:

  • Key has referer / IP restrictions. Common on Gemini keys minted in Google AI Studio. Google returns API_KEY_INVALID because the actor runs on Apify's AWS us-east-1 IPs, which aren't on your allow-list. Fix: remove restrictions on the key in Google AI Studio, or add apify.com as an allowed referer.
  • Quota exhausted. Free tiers on Gemini and Groq have daily and per-minute caps. Google returns RESOURCE_EXHAUSTED, Groq returns HTTP 429. The exact error is in the log. Fix: wait for the reset, or move to a paid tier.
  • Model was retired. Vendors deprecate faster than their docs update. Gemini 2.5 Flash was restricted to existing users only. Groq retired Llama 3.3 70B Versatile and Llama 3.1 8B Instant. Fix: set llmModel to a currently-available model. Actor defaults are refreshed when we detect a retirement, but an explicit llmModel overrides the default.

One-line workaround: use OpenRouter. Set llmProvider: "openrouter" and llmApiKey: <your OpenRouter key>. OpenRouter accepts calls from any IP, aggregates spend on one balance, and lets you switch model providers via llmModel (for example google/gemini-3.6-flash, openai/gpt-oss-120b). One key, no per-provider restrictions.

Auto-chaining

The HTTP actor can automatically trigger this browser actor for JS-flagged companies:

  1. Run the HTTP actor with enablePlaywrightFallback: true
  2. Companies with js_rendering_suspected are collected
  3. A browser actor run starts automatically (fire-and-forget)
  4. The browser run ID is saved in the key-value store as BROWSER_FALLBACK_RUN_ID

LLM fallback chain

Like the HTTP actor, this actor supports automatic provider fallback. Just provide API keys for the providers you want to use:

{
"geminiApiKey": "YOUR_GEMINI_KEY",
"llmApiKey": "YOUR_GROQ_KEY",
"openrouterApiKey": "YOUR_OPENROUTER_KEY"
}

The system auto-discovers available providers and builds a fallback chain (e.g. Gemini → Groq → OpenRouter). If one provider's quota runs out, it instantly falls back to the next.

End-to-end pipeline

This actor is part of a 5-actor enrichment suite:

ActorPurposeMemoryLink
Google Maps ScraperFind companies by location~80MBView
Website Job ExtractorExtract jobs (HTTP)~128MBView
Website Job Extractor (Browser)Extract jobs from JS pages~1-4GBThis actor
Website Contact ExtractorExtract contacts (HTTP)~256MBView
Website Contact Extractor (Browser)Extract contacts from JS pages~1-4GBView

Limitations

  • Higher memory usage (~1GB vs ~128MB for HTTP)
  • Slower execution (page rendering + wait times)
  • Higher cost per result (2x HTTP rates)
  • Use the HTTP actor first — only fall back to browser when needed

Why a company returned 0 jobs (reason field)

Every company gets one job_extraction_done row. When job_count is 0 the row also carries a machine-readable reason and a short human reason_message. When jobs were found both are null.

reasonMeaning
dns_errorThe domain does not resolve (typo or retired domain)
connection_errorThe page could not be fetched (timeout, TLS, connection reset)
blockedThe site answered 401/403/429 or showed a bot challenge
http_error:<code>The site answered with another error status, for example http_error:404 or http_error:503
parked_domainThe domain is parked, expired or redirects to a domain seller
not_a_career_pageThe URL is not a careers or jobs page (for example an encyclopedia article)
parse_errorThe page was fetched but the extraction step failed
no_jobs_foundA careers page was read successfully and lists no open positions

Example:

{
"_type": "job_extraction_done",
"company_id": "example-ag",
"company_name": "Example AG",
"job_count": 0,
"reason": "dns_error",
"reason_message": "The domain could not be resolved (DNS lookup failed). Check the URL for typos or a retired domain."
}