GaijinPot Search Scraper avatar

GaijinPot Search Scraper

Pricing

from $2.99 / 1,000 gaijinpot job records

Go to Apify Store
GaijinPot Search Scraper

GaijinPot Search Scraper

Extract complete current GaijinPot jobs, including salary, languages, visa/residency requirements, functions, descriptions, and employer details.

Pricing

from $2.99 / 1,000 gaijinpot job records

Rating

0.0

(0)

Developer

Jobs API

Jobs API

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

1

Monthly active users

11 days ago

Last modified

Share

GaijinPot Jobs Search Scraper

Searches public GaijinPot Jobs pages and enriches every accepted listing from its canonical detail page. It uses lightweight HTTP requests rather than a browser and does not require credentials or a proxy.

Input

  • query (required): title, function, skill, or keyword.
  • location: optional Japanese city or region.
  • maxItems: complete records to save, from 1 to 100.
  • maxPages: hard listing-page limit, from 1 to 50.

The Actor uses a fixed ordinary direct HTTPS configuration: sequential requests, a bounded timeout and request budget, ordinary headers, no proxy, no browser automation, no stealth or fingerprint changes, and no retries. These transport controls are intentionally not public inputs. Raw JobPosting JSON-LD and fixture/debug/context fields are never emitted.

{ "query": "engineer", "location": "Tokyo", "maxItems": 20, "maxPages": 5 }

Output

Portable fields include jobId, title, companyName, locationText, descriptionText, salaryText, postedAt, employmentType, applyUrl, and url. Rich fields preserve language, visa/residency, employer, benefits, industry, function, application availability, company details, listing/detail receipts, and source-page metadata. The OUTPUT key-value-store record summarizes pages, requests, skips, diagnostics, records, access status, and runtime.

GaijinPot can mix featured or loosely related listings into search results. The Actor verifies query and location against enriched records and excludes incomplete/nonmatching items. Sparse searches can return fewer than maxItems when maxPages is reached. The summary separates platform status (SUCCEEDED or FAILED) from resultStatus (COMPLETE, LIMITED, SKIPPED, DEFERRED, NO_DATA, or FAILED). storedRecords and the compatibility field emittedRecords count only rows whose dataset writes completed. Diagnostics and candidate skips are kept in key-value artifacts, never in the dataset. Only a receipt-matched first-party HTTP 401/403/429/451 or visible verification/access-denied challenge establishes SKIPPED; requests stop and all buffered rows are discarded. Redirects are not followed. HTTP 407, generic HTTP errors, parser uncertainty, and empty results without authoritative first-party evidence are DEFERRED, not source blocks or NO_DATA. A fatal zero-row runtime or Dataset failure uses the Actor failure lifecycle. No bypass is attempted.

Local QA

npm install
npm test
apify run --purge --input-file qa-inputs/gaijinpot/local-search.json
npm run validate:dataset