Company Job Posting Feed Builder avatar

Company Job Posting Feed Builder

Under maintenance

Pricing

from $1.00 / 1,000 jobs

Go to Apify Store
Company Job Posting Feed Builder

Company Job Posting Feed Builder

Under maintenance

Collect public job postings with structured salary, location and date fields.

Pricing

from $1.00 / 1,000 jobs

Rating

0.0

(0)

Developer

Danial Maqbool

Danial Maqbool

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

2 days ago

Last modified

Categories

Share

Collect public job postings with structured salary, location and date fields.

Example.com inputs are placeholders, not verified compatible content. Replace them with your permitted source. Run npm run smoke for an account-free deterministic example. JSON Feed, RSS and table examples use the separately recorded small live-test sources; they do not imply compatibility with every site.

What it does

Produces job records for recruitment researchers and job-feed builders. Values are extracted deterministically from public pages; unavailable fields remain null.

Common use cases

Collect public job postings with structured salary, location and date fields. Use the resulting dataset in scheduled tasks, API pipelines, spreadsheets, or agent workflows.

Input

FieldTypeDescription
startUrlsarrayPublic HTTP(S) URLs. No credentials or local-network targets.
maxPagesintegerMaximum pages scheduled in this bounded run.
maxJobsintegerMaximum visible job records returned.
renderJavaScriptbooleanOpt into Chromium for JavaScript content. More expensive; no access-control bypass.
respectRobotsTxtbooleanHonor robots rules and supported crawl delays. Unavailable rules fail closed.
maxDepthintegerMaximum followed-link depth. Explicit seeds start at zero.
includePatternsarrayAllowed absolute-URL patterns using * as wildcard.
excludePatternsarrayExcluded absolute-URL patterns using * as wildcard.
keywordsarrayKeywords. See README for extraction behavior.
locationKeywordsarrayLocation Keywords. See README for extraction behavior.
followPaginationbooleanFollow Pagination. See README for extraction behavior.
itemSelectorstringItem Selector. See README for extraction behavior.
titleSelectorstringTitle Selector. See README for extraction behavior.
nextPageSelectorstringNext Page Selector. See README for extraction behavior.
concurrencyintegerMaximum concurrent page handlers. Per-origin delays still apply.
requestDelayMsintegerMinimum spacing between request starts to the same origin.
timeoutSecsintegerMaximum individual request duration in seconds.
retriesintegerRetries for page handlers. Access blocks are not retried.
proxyConfigurationobjectOptional Apify or public custom proxy. Direct connections are the default.

Output

FieldType
titlestring or null
companystring or null
locationstring or null
employmentTypestring or null
workplaceTypestring or null
salaryMinnumber or null
salaryMaxnumber or null
salaryCurrencystring or null
salaryPeriodstring or null
descriptionstring or null
qualificationsstring or null
datePostedstring or null
validThroughstring or null
jobUrlstring or null
sourceUrlstring or null
extractedAtstring or null

Results are in the default Dataset. RUN_SUMMARY, FAILURES and DIAGNOSTICS are available in the default key-value store. Diagnostics are not billed entity events. JSON, CSV and spreadsheet exports use Apify Dataset.

Example

Replace example.com with a public source relevant to this Actor. Plain example.com has no products, jobs or opportunities; an empty feed there is expected. Run the included local smoke test for a deterministic working example.

{
"startUrls": [
"https://example.com/"
],
"maxPages": 1,
"renderJavaScript": false,
"maxJobs": 10
}

Example output

The following records come from synthetic local fixtures, not a live website or claimed customer data.

[
{
"title": "Data Engineer",
"company": "Fixture Labs",
"location": "Karachi, PK",
"employmentType": "FULL_TIME",
"workplaceType": "remote",
"salaryMin": 75000,
"salaryMax": 95000,
"salaryCurrency": "USD",
"salaryPeriod": "YEAR",
"description": "Build reliable pipelines.",
"qualifications": "SQL and TypeScript",
"datePosted": "2026-09-01",
"validThrough": "2026-12-31",
"jobUrl": "https://fixture.example/job",
"sourceUrl": "https://fixture.example/index",
"extractedAt": "2026-09-29T23:32:58.495Z"
}
]

Pricing model

One job event per visible result record. Event prices are configured in Apify Console, never in extraction code. The SDK enforces the run's maxTotalChargeUsd. Failed pages and diagnostic-only messages are not billed. Do not configure an additional automatic default-dataset-item event.

How it works

Crawlee BasicCrawler manages bounded requests and retries. Cheerio parses HTTP HTML. A DNS-pinned transport validates every redirect and checks robots rules. Chromium is opt-in and its page requests pass through the same transport. No external paid API is required.

Limits

  • Only explicit JobPosting data or recognizable job cards are extracted.
  • Applicant records and authenticated job portals are not supported.
  • Public GET/HEAD content only; no login, form submission, CAPTCHA solving or browser stealth.
  • Responses are capped at 2 MiB decoded. Browser mode caps requests per page and blocks media, fonts, downloads, service workers and WebSockets.
  • Missing/blocked robots rules fail closed. Robots crawl delays above 60 seconds are not supported.
  • Start at 512 MB for static runs; use 2 GB for browser runs and measure representative targets.

Responsible use

Use only public content you are authorized to access. Respect source terms, licenses and applicable law. No private-account extraction, personal-data enrichment, credential input or access-control bypass is provided. Cookie values are never returned.

Local development

Requires Node.js 24 LTS. All commands below run from this Actor directory.

npm ci --ignore-scripts
npm run input:example
# Edit storage/key_value_stores/default/INPUT.json with your public URLs.
npm run start:dev
npm test
npm run typecheck
npm run lint
npm run validate
npm run build
npm start
npm run smoke
npx apify run
docker build -t job-posting-feed-builder .

Browser testing: set PLAYWRIGHT_BROWSERS_PATH to an Actor-local .cache/browsers directory, then run npx playwright install chromium. Docker installs its own matching browser.

Deployment

Deploy only after validating the Docker build and a representative target. Deployment is not performed by installation or tests.

npx apify login
npx apify push

Then configure the exact event in Console, set pricing and spending limits, test a private run, complete PUBLICATION_CHECKLIST.md, and deliberately enable Store visibility.

API usage

Replace YOUR_USERNAME with your Apify account name. Keep the token in an environment variable rather than input JSON.

$curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~job-posting-feed-builder/runs?maxTotalChargeUsd=1" -H "Authorization: Bearer $APIFY_TOKEN" -H "Content-Type: application/json" --data-binary @examples/basic-input.json

Monitor the returned run ID and read its default Dataset. Store/API names and price configuration are account-side settings.

Verified platform example

The saved Console input performs a bounded HTTP run with one page and at most two result records. The source is a public scraping demonstration site; these records are sample content, not customer or current hiring data.

Use the saved default input in Apify Console, or the repository examples/basic-input.json. Replace its source with your own permitted URL and keep limits small for the first run.

The following unmodified record was returned by a successful Apify platform run on 2026-09-30:

[
{
"title": "Senior Python Developer",
"company": "Payne, Roberts and Davis",
"location": "Stewartbury, AA",
"employmentType": null,
"workplaceType": null,
"salaryMin": null,
"salaryMax": null,
"salaryCurrency": null,
"salaryPeriod": null,
"description": "Stewartbury, AA",
"qualifications": null,
"datePosted": null,
"validThrough": null,
"jobUrl": "https://www.realpython.com/",
"sourceUrl": "https://realpython.github.io/fake-jobs/",
"extractedAt": "2026-09-30T11:47:59.981Z"
}
]

Introductory pricing: $1.00 per 1,000 job results, plus Apify platform usage. The standard Actor start event costs $0.00005 per GB (minimum one event). Empty output and duplicate records incur no primary result event; the start fee and any platform usage still apply. The SDK enforces the event spending limit; platform usage is billed separately.