Google Jobs Search Scraper
Pricing
from $2.99 / 1,000 google jobs job records
Google Jobs Search Scraper
Scrape job listings from Google Jobs (via Google Search) by keyword and location with country and language targeting. Extracts job title, company, location, salary range, job type, date posted, and application URL — perfect for job market research and recruitment intelligence.
Pricing
from $2.99 / 1,000 google jobs job records
Rating
0.0
(0)
Developer
Jobs API
Maintained by CommunityActor stats
0
Bookmarked
28
Total users
14
Monthly active users
11 days ago
Last modified
Categories
Share
Google Jobs Public Search Scraper
This Apify Actor reads the public Google Jobs panel in Google Search and opens each visible job card's detail modal before emitting a record. It is designed for source-grounded job research, not for bypassing Google access controls.
Source and access policy
The Actor uses public Google Jobs search URLs such as:
https://www.google.com/search?q=software+engineer+in+United+States&gl=us&hl=en&udm=8&ibp=htl%3Bjobs
It runs one ordinary direct Chrome request at a time, with no proxy, stealth, fingerprint injection, alternate identity, cookie injection, request retry, or Yahoo/source-page fallback. The visible detail modal is read from the same Google Jobs result surface. An affirmative Google access barrier (HTTP 401/403/429/451, /sorry/, or a visible challenge/denial) stops the route and discards buffered rows as resultStatus: SKIPPED. Transport, parser, and controller failures are DEFERRED, never inferred as a source block or empty search; previously persisted verified rows are preserved as SUCCEEDED / LIMITED.
Run summaries distinguish the Apify process lifecycle (status: SUCCEEDED or FAILED) from extraction completeness (resultStatus: COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED). NO_DATA is allowed only when every requested query has a successful Google search receipt, a zero result count, a response hash, and explicit source text confirming no jobs. This Actor currently has no verified empty-state detector, so an unrecognized zero-row response remains DEFERRED.
Input
{"query": "software engineer","location": "United States","gl": "us","hl": "en","maxItems": 3,"maxPages": 1}
The public input schema contains only functional search and result filters:
| Field | Type | Default | Description |
|---|---|---|---|
query | string | software engineer | Main job search phrase. |
queries | string array | — | Optional additional phrases processed in order. |
location | string | United States | Search location. |
maxItems | integer | 50 | Complete records to emit, from 1 to 100. |
gl | string | us | Two-letter Google country code. |
hl | string | en | Google language code. |
maxPages | integer | 3 | Search-page limit, from 1 to 5. |
experienceLevel | enum | — | Google Jobs experience filter. |
salaryMin, salaryMax | integer | — | Salary filters when supported by the public surface. |
remoteFilter | enum | — | remote, hybrid, or onsite. |
educationLevel | enum | — | Education filter. |
Proxy, retry, timeout, user-agent, session, fixture, raw-payload, and debug controls are intentionally excluded from the public contract.
Dataset
Each record follows .actor/dataset_schema.json and contains more than 20 meaningful source-backed leaves where the detail modal exposes them, including:
- stable Google job identity, result position, title, employer, provider, logo, location, applicant-location signal, and work arrangement;
- visible posting age, salary/stipend, employment type, deadline, and application links;
- readable full detail description, normalized responsibilities, qualifications, skills, education/experience requirements, benefits, sections, and character/word counts;
- Google search, locale, requested-filter, card-evidence, detail-evidence, canonical/application URLs, direct HTTP response receipt, and quality provenance.
Raw JobPosting JSON-LD, raw HTML snapshots, temporary searchQuery/searchLocation context fields, fixtures, cookies, and debug artifacts are not emitted.
Local validation
From this actor directory:
npm testnpm run checknpm run validate:datasetnpx apify validate-schema
The dataset validator reconciles generated storage with run summaries, rejects blocked/partial records, checks direct-only provenance and response hashes, enforces unique identities/URLs, and verifies the more-than-20-meaningful-field threshold.
Responsible use
Use public pages in accordance with Google's terms, robots guidance, rate limits, and applicable law. The Actor stops when the source indicates that ordinary access is blocked.