GaijinPot Search Scraper
Pricing
from $2.99 / 1,000 gaijinpot job records
GaijinPot Search Scraper
Extract complete current GaijinPot jobs, including salary, languages, visa/residency requirements, functions, descriptions, and employer details.
Pricing
from $2.99 / 1,000 gaijinpot job records
Rating
0.0
(0)
Developer
Jobs API
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
1
Monthly active users
11 days ago
Last modified
Categories
Share
GaijinPot Jobs Search Scraper
Searches public GaijinPot Jobs pages and enriches every accepted listing from its canonical detail page. It uses lightweight HTTP requests rather than a browser and does not require credentials or a proxy.
Input
query(required): title, function, skill, or keyword.location: optional Japanese city or region.maxItems: complete records to save, from 1 to 100.maxPages: hard listing-page limit, from 1 to 50.
The Actor uses a fixed ordinary direct HTTPS configuration: sequential requests, a bounded timeout and request budget, ordinary headers, no proxy, no browser automation, no stealth or fingerprint changes, and no retries. These transport controls are intentionally not public inputs. Raw JobPosting JSON-LD and fixture/debug/context fields are never emitted.
{ "query": "engineer", "location": "Tokyo", "maxItems": 20, "maxPages": 5 }
Output
Portable fields include jobId, title, companyName, locationText, descriptionText, salaryText, postedAt, employmentType, applyUrl, and url. Rich fields preserve language, visa/residency, employer, benefits, industry, function, application availability, company details, listing/detail receipts, and source-page metadata. The OUTPUT key-value-store record summarizes pages, requests, skips, diagnostics, records, access status, and runtime.
GaijinPot can mix featured or loosely related listings into search results. The Actor verifies query and location against enriched records and excludes incomplete/nonmatching items. Sparse searches can return fewer than maxItems when maxPages is reached. The summary separates platform status (SUCCEEDED or FAILED) from resultStatus (COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED). storedRecords and the compatibility field emittedRecords count only rows whose dataset writes completed. Diagnostics and candidate skips are kept in key-value artifacts, never in the dataset. Only a receipt-matched first-party HTTP 401/403/429/451 or visible verification/access-denied challenge establishes SKIPPED; requests stop and all buffered rows are discarded. Redirects are not followed. HTTP 407, generic HTTP errors, parser uncertainty, and empty results without authoritative first-party evidence are DEFERRED, not source blocks or NO_DATA. A fatal zero-row runtime or Dataset failure uses the Actor failure lifecycle. No bypass is attempted.
Local QA
npm installnpm testapify run --purge --input-file qa-inputs/gaijinpot/local-search.jsonnpm run validate:dataset