Website Job Extractor (Browser)
Pricing
from $8.00 / 1,000 job extracted (browser)s
Website Job Extractor (Browser)
Extract job listings from JavaScript-rendered career pages (React, Vue, Angular) using AI + Playwright. Companion to the HTTP-only Website Job Extractor. Use it for the ~28% of company sites that need a real browser. Same output format, same quality, same LLM fallback chain.
Pricing
from $8.00 / 1,000 job extracted (browser)s
Rating
0.0
(0)
Developer
NanoScrape
Maintained by CommunityActor stats
0
Bookmarked
38
Total users
4
Monthly active users
11 hours
Issues response
19 minutes ago
Last modified
Categories
Share
Extract job listings from JavaScript-rendered career pages (React, Vue, Angular SPAs) using AI + Playwright. This is the specialist tool in a three-tier extraction chain — use it as a fallback, not a default.
Which tool should I use? (Read this first)
Running a browser + LLM is inherently more expensive than parsing static HTML. Try these in order:
-
website-job-extractor(HTTP + AI) — try this first. Handles ~70% of company career pages. About 5× cheaper per company than this browser variant. SetenableBrowserFallback: trueon its input and it will automatically hand off JS-only sites to this actor. -
career-site-jobs-scraper— if you know the URL is hosted on a supported ATS platform (Lever, Greenhouse, Workable, Ashby, Teamtailor, Personio, SmartRecruiters, BambooHR, Workday, iCIMS, Recruitee, JOIN, Pinpoint, Rippling, JazzHR, Comeet). Deterministic, HTTP-only, no LLM cost. Cheapest of the three. -
This actor — when the HTTP extractor returns nothing and the URL isn't a known ATS. Renders the page with Playwright, extracts with an LLM. Slower and pricier, but works on anything.
When to use this actor
- Career pages built with React, Vue, Angular, or other JS frameworks
- Pages that return empty/skeleton HTML without JavaScript execution
- Companies flagged by the HTTP actor's JS-rendering detection
- Auto-chained via
enableBrowserFallbackon the HTTP actor (wasenablePlaywrightFallback, deprecated)
Input shape: website_url vs career_urls
These two fields serve different purposes:
| Field | What it means | When to use it |
|---|---|---|
website_url | Company homepage. The actor runs discovery from here: probes subdomain patterns, reads navigation links. | When you only have the company's main website and want the actor to find the career page for you. |
career_urls | Specific job listing URLs, SERP pages, or detail pages. The actor extracts directly from these without discovery. | When you already know the career page URL, or the URL is already a search results page (e.g., amazon.jobs/en/search?base_query=...). |
If your input URL is already a SERP or job listing, pass it via career_urls, not website_url. Homepage discovery on a SERP page extracts nav links, which are all rejected as junk category pages.
As of v1.1.8, the actor performs best-effort SERP detection: if website_url looks like a search results page (contains ?base_query=, ?q=, /search?, etc.) and career_urls is empty, it treats website_url as a direct extraction target automatically. Explicit is still better.
Use with AI Agents (MCP)
Connect this actor to any MCP-compatible AI client — Claude Desktop, Claude.ai, Cursor, VS Code, LangChain, LlamaIndex, or custom agents.
Apify MCP server URL:
https://mcp.apify.com?tools=santamaria-automations/website-job-extractor-browser
Example prompt once connected:
"Use
website-job-extractor-browserto process data with website job extractor browser. Return results as a table."
Clients that support dynamic tool discovery (Claude.ai, VS Code) will receive the full input schema automatically via add-actor.
How it works
- Playwright renders the full page (waits for network idle + text content)
- Career page discovery from homepage navigation (same as HTTP actor)
- ATS detection for 19 systems (Personio, Greenhouse, Softgarden, etc.)
- LLM extraction using Gemini Flash / Groq / OpenRouter
- Validation with confidence scoring and deduplication
- Pagination follow-up for multi-page listings
Same extraction pipeline as the HTTP actor — same output format, same quality.
Input
Same input format as the HTTP actor. Typically auto-chained:
{"companies": [{"company_id": "abc-123","company_name": "TechCorp AG","website_url": "https://techcorp.ch"}],"llmProvider": "gemini","geminiApiKey": "YOUR_KEY"}
Output
Each job is a dataset item with browser_extraction: true:
{"company_id": "abc-123","company_name": "TechCorp AG","title": "Senior Frontend Developer (m/w/d)","location": "Zürich","employment_type": "Vollzeit","department": "Engineering","job_url": "https://techcorp.ch/jobs/senior-frontend-developer-zurich","job_id": "senior-frontend-developer-zurich","application_url": "https://techcorp.ch/jobs/apply/123","confidence": 0.85,"browser_extraction": true,"extracted_at": "2026-03-09T10:00:00.000Z"}
Memory requirements
- Minimum: 1024 MB (Playwright + Chrome)
- Recommended: 2048 MB for 5+ companies
- Maximum: 4096 MB
LLM API keys -- BYOK vs Managed mode
This actor works in two modes:
Managed mode (default -- no key needed)
Leave all API key fields blank. The actor uses a shared OpenRouter account with deepseek/deepseek-chat-v3-0324 as the extraction model. No sign-up required.
Budget caps by tier:
| Tier | Default cap | Maximum cap |
|---|---|---|
| Free Apify account | $0.30 per run (hard ceiling) | $0.30 (cannot raise) |
| Paying Apify account | $10.00 per run | $50.00 (via managedLlmBudgetUsd) |
When the cap is reached mid-run, a managed_llm_run_capped sentinel row is pushed to the dataset and the actor exits cleanly with a message telling you how to continue. No crash.
Managed mode PPE events (charged per extracted job in managed mode):
| Event | Price |
|---|---|
| Company processed | $0.05 |
| Job extracted (managed) | $0.024 |
The 3x premium on managed jobs vs BYOK jobs ($0.024 vs $0.008) covers the deepseek-v3 LLM cost plus a margin. Browser compute overhead means the premium is in absolute terms the same multiplier as HTTP but at a higher base price due to Camoufox/Playwright costs.
BYOK mode (bring your own key)
Supply any of geminiApiKey, groqApiKey, or openrouterApiKey. The actor uses your quota, your rate limits, and the standard PPE event prices:
| Event | Price |
|---|---|
| Company processed | $0.05 |
| Job extracted (BYOK) | $0.008 |
Managed mode input
{"companies": [{"company_id": "stripe","company_name": "Stripe","career_urls": ["https://stripe.com/jobs/search"]}]}
No API key fields required. The actor runs immediately.
Raising the managed budget (paying users only)
{"companies": [...],"managedLlmBudgetUsd": 25}
Free-tier users cannot raise the cap above $0.30 regardless of this input.
Sentinel rows
The dataset may also contain sentinel rows with a _type field. These are not job listings but machine-readable events:
_type value | When emitted | Extra fields |
|---|---|---|
job_extraction_done | When all pages for a company are processed | job_count, jobs_extracted, pages_fetched, pagination_source (css_selector, url_pattern, button_click, or none) |
llm_extraction_failed | When every LLM call failed (all providers exhausted) | error_category, provider, model, attempts, user_message |
managed_llm_run_capped | When the managed LLM budget cap is reached mid-run (managed mode only) | error_category (free_tier_budget_reached or paying_tier_budget_reached), tier, spent_usd, cap_usd, user_message |
All sentinel rows include the four required dataset fields: company_id, company_name, source_url, extracted_at.
Pricing
Browser-based extraction costs more than the HTTP actor due to Chrome overhead. PPE events reflect the actual compute done:
| Event | Listed rate | When it fires |
|---|---|---|
browser-company-enriched | $0.05 / company | Once per company on the first page render |
browser-page-processed | $0.005 / page | See the smart-pricing note below |
browser-job-result | $0.008 / job | Each validated, non-duplicate job (BYOK mode) |
browser-job-result-managed | $0.024 / job | Each validated, non-duplicate job (managed mode) |
Smart per-page pricing (why your actual bill is lower than the listed rate)
The actor decides internally which proxy pool to use for each URL, defaulting to a fast primary pool and only escalating to residential when a URL gets blocked (403, 429, or repeated timeouts). Because the primary pool costs a fraction of what residential does, we pass the saving on rather than pocket it:
browser-page-processedfires only on the 1st, 5th, 9th, ... render that succeeds on the primary pool (1 in 4).- Renders that fell back to residential bill 1:1, since that's where our real cost sits.
Result: on typical company career pages, you pay for roughly one page in four, and the actor's structural proxy strategy handles the "hard" URLs invisibly.
Cost examples
Small test (1 company, 3 primary-pool pages, 5 jobs):
$0.05 + 1 × $0.005 + 5 × $0.008 = **$0.095**
(first page bills, next two are absorbed by the 1-in-4 cadence)
Bulk run (900 companies, mostly primary pool, ~1,850 jobs): Roughly $65-75/run in practice, vs the ~$130 the flat listed rate would imply.
Historical context (why the pricing model is what it is)
An earlier flat per-page charge introduced 2026-07-27 hit deep-crawl users disproportionately. Investigation showed our real cost was almost entirely residential-proxy bandwidth, not compute. Fixing that (making the actor decide proxy strategy per-URL) removed the underlying cost driver, and the smart per-page cadence passes the saving on to you.
Pagination stops when maxResultsPerCompany (default 200) jobs have been extracted per company, or when the 20-page hard safety cap fires. The proxyConfiguration input has been removed; if you send one, it's silently ignored.
Troubleshooting: LLM key errors
Runs that fail with LLM provider ... rejected the supplied API key or LLM model ... is unavailable to your ... key almost always fall into one of these three buckets:
- Key has referer / IP restrictions. Common on Gemini keys minted in Google AI Studio. Google returns
API_KEY_INVALIDbecause the actor runs on Apify's AWS us-east-1 IPs, which aren't on your allow-list. Fix: remove restrictions on the key in Google AI Studio, or addapify.comas an allowed referer. - Quota exhausted. Free tiers on Gemini and Groq have daily and per-minute caps. Google returns
RESOURCE_EXHAUSTED, Groq returns HTTP 429. The exact error is in the log. Fix: wait for the reset, or move to a paid tier. - Model was retired. Vendors deprecate faster than their docs update. Gemini 2.5 Flash was restricted to existing users only. Groq retired Llama 3.3 70B Versatile and Llama 3.1 8B Instant. Fix: set
llmModelto a currently-available model. Actor defaults are refreshed when we detect a retirement, but an explicitllmModeloverrides the default.
One-line workaround: use OpenRouter. Set llmProvider: "openrouter" and llmApiKey: <your OpenRouter key>. OpenRouter accepts calls from any IP, aggregates spend on one balance, and lets you switch model providers via llmModel (for example google/gemini-3.6-flash, openai/gpt-oss-120b). One key, no per-provider restrictions.
Auto-chaining
The HTTP actor can automatically trigger this browser actor for JS-flagged companies:
- Run the HTTP actor with
enablePlaywrightFallback: true - Companies with
js_rendering_suspectedare collected - A browser actor run starts automatically (fire-and-forget)
- The browser run ID is saved in the key-value store as
BROWSER_FALLBACK_RUN_ID
LLM fallback chain
Like the HTTP actor, this actor supports automatic provider fallback. Just provide API keys for the providers you want to use:
{"geminiApiKey": "YOUR_GEMINI_KEY","llmApiKey": "YOUR_GROQ_KEY","openrouterApiKey": "YOUR_OPENROUTER_KEY"}
The system auto-discovers available providers and builds a fallback chain (e.g. Gemini → Groq → OpenRouter). If one provider's quota runs out, it instantly falls back to the next.
End-to-end pipeline
This actor is part of a 5-actor enrichment suite:
| Actor | Purpose | Memory | Link |
|---|---|---|---|
| Google Maps Scraper | Find companies by location | ~80MB | View |
| Website Job Extractor | Extract jobs (HTTP) | ~128MB | View |
| Website Job Extractor (Browser) | Extract jobs from JS pages | ~1-4GB | This actor |
| Website Contact Extractor | Extract contacts (HTTP) | ~256MB | View |
| Website Contact Extractor (Browser) | Extract contacts from JS pages | ~1-4GB | View |
Limitations
- Higher memory usage (~1GB vs ~128MB for HTTP)
- Slower execution (page rendering + wait times)
- Higher cost per result (2x HTTP rates)
- Use the HTTP actor first — only fall back to browser when needed
Why a company returned 0 jobs (reason field)
Every company gets one job_extraction_done row. When job_count is 0 the row also carries a machine-readable reason and a short human reason_message. When jobs were found both are null.
reason | Meaning |
|---|---|
dns_error | The domain does not resolve (typo or retired domain) |
connection_error | The page could not be fetched (timeout, TLS, connection reset) |
blocked | The site answered 401/403/429 or showed a bot challenge |
http_error:<code> | The site answered with another error status, for example http_error:404 or http_error:503 |
parked_domain | The domain is parked, expired or redirects to a domain seller |
not_a_career_page | The URL is not a careers or jobs page (for example an encyclopedia article) |
parse_error | The page was fetched but the extraction step failed |
no_jobs_found | A careers page was read successfully and lists no open positions |
Example:
{"_type": "job_extraction_done","company_id": "example-ag","company_name": "Example AG","job_count": 0,"reason": "dns_error","reason_message": "The domain could not be resolved (DNS lookup failed). Check the URL for typos or a retired domain."}


