Company Job Posting Feed Builder
Under maintenancePricing
from $1.00 / 1,000 jobs
Company Job Posting Feed Builder
Under maintenanceCollect public job postings with structured salary, location and date fields.
Pricing
from $1.00 / 1,000 jobs
Rating
0.0
(0)
Developer
Danial Maqbool
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
2 days ago
Last modified
Categories
Share
Collect public job postings with structured salary, location and date fields.
Example.com inputs are placeholders, not verified compatible content. Replace them with your permitted source. Run npm run smoke for an account-free deterministic example. JSON Feed, RSS and table examples use the separately recorded small live-test sources; they do not imply compatibility with every site.
What it does
Produces job records for recruitment researchers and job-feed builders. Values are extracted deterministically from public pages; unavailable fields remain null.
Common use cases
Collect public job postings with structured salary, location and date fields. Use the resulting dataset in scheduled tasks, API pipelines, spreadsheets, or agent workflows.
Input
| Field | Type | Description |
|---|---|---|
| startUrls | array | Public HTTP(S) URLs. No credentials or local-network targets. |
| maxPages | integer | Maximum pages scheduled in this bounded run. |
| maxJobs | integer | Maximum visible job records returned. |
| renderJavaScript | boolean | Opt into Chromium for JavaScript content. More expensive; no access-control bypass. |
| respectRobotsTxt | boolean | Honor robots rules and supported crawl delays. Unavailable rules fail closed. |
| maxDepth | integer | Maximum followed-link depth. Explicit seeds start at zero. |
| includePatterns | array | Allowed absolute-URL patterns using * as wildcard. |
| excludePatterns | array | Excluded absolute-URL patterns using * as wildcard. |
| keywords | array | Keywords. See README for extraction behavior. |
| locationKeywords | array | Location Keywords. See README for extraction behavior. |
| followPagination | boolean | Follow Pagination. See README for extraction behavior. |
| itemSelector | string | Item Selector. See README for extraction behavior. |
| titleSelector | string | Title Selector. See README for extraction behavior. |
| nextPageSelector | string | Next Page Selector. See README for extraction behavior. |
| concurrency | integer | Maximum concurrent page handlers. Per-origin delays still apply. |
| requestDelayMs | integer | Minimum spacing between request starts to the same origin. |
| timeoutSecs | integer | Maximum individual request duration in seconds. |
| retries | integer | Retries for page handlers. Access blocks are not retried. |
| proxyConfiguration | object | Optional Apify or public custom proxy. Direct connections are the default. |
Output
| Field | Type |
|---|---|
| title | string or null |
| company | string or null |
| location | string or null |
| employmentType | string or null |
| workplaceType | string or null |
| salaryMin | number or null |
| salaryMax | number or null |
| salaryCurrency | string or null |
| salaryPeriod | string or null |
| description | string or null |
| qualifications | string or null |
| datePosted | string or null |
| validThrough | string or null |
| jobUrl | string or null |
| sourceUrl | string or null |
| extractedAt | string or null |
Results are in the default Dataset. RUN_SUMMARY, FAILURES and DIAGNOSTICS are available in the default key-value store. Diagnostics are not billed entity events. JSON, CSV and spreadsheet exports use Apify Dataset.
Example
Replace example.com with a public source relevant to this Actor. Plain example.com has no products, jobs or opportunities; an empty feed there is expected. Run the included local smoke test for a deterministic working example.
{"startUrls": ["https://example.com/"],"maxPages": 1,"renderJavaScript": false,"maxJobs": 10}
Example output
The following records come from synthetic local fixtures, not a live website or claimed customer data.
[{"title": "Data Engineer","company": "Fixture Labs","location": "Karachi, PK","employmentType": "FULL_TIME","workplaceType": "remote","salaryMin": 75000,"salaryMax": 95000,"salaryCurrency": "USD","salaryPeriod": "YEAR","description": "Build reliable pipelines.","qualifications": "SQL and TypeScript","datePosted": "2026-09-01","validThrough": "2026-12-31","jobUrl": "https://fixture.example/job","sourceUrl": "https://fixture.example/index","extractedAt": "2026-09-29T23:32:58.495Z"}]
Pricing model
One job event per visible result record. Event prices are configured in Apify Console, never in extraction code. The SDK enforces the run's maxTotalChargeUsd. Failed pages and diagnostic-only messages are not billed. Do not configure an additional automatic default-dataset-item event.
How it works
Crawlee BasicCrawler manages bounded requests and retries. Cheerio parses HTTP HTML. A DNS-pinned transport validates every redirect and checks robots rules. Chromium is opt-in and its page requests pass through the same transport. No external paid API is required.
Limits
- Only explicit JobPosting data or recognizable job cards are extracted.
- Applicant records and authenticated job portals are not supported.
- Public GET/HEAD content only; no login, form submission, CAPTCHA solving or browser stealth.
- Responses are capped at 2 MiB decoded. Browser mode caps requests per page and blocks media, fonts, downloads, service workers and WebSockets.
- Missing/blocked robots rules fail closed. Robots crawl delays above 60 seconds are not supported.
- Start at 512 MB for static runs; use 2 GB for browser runs and measure representative targets.
Responsible use
Use only public content you are authorized to access. Respect source terms, licenses and applicable law. No private-account extraction, personal-data enrichment, credential input or access-control bypass is provided. Cookie values are never returned.
Local development
Requires Node.js 24 LTS. All commands below run from this Actor directory.
npm ci --ignore-scriptsnpm run input:example# Edit storage/key_value_stores/default/INPUT.json with your public URLs.npm run start:devnpm testnpm run typechecknpm run lintnpm run validatenpm run buildnpm startnpm run smokenpx apify rundocker build -t job-posting-feed-builder .
Browser testing: set PLAYWRIGHT_BROWSERS_PATH to an Actor-local .cache/browsers directory, then run npx playwright install chromium. Docker installs its own matching browser.
Deployment
Deploy only after validating the Docker build and a representative target. Deployment is not performed by installation or tests.
npx apify loginnpx apify push
Then configure the exact event in Console, set pricing and spending limits, test a private run, complete PUBLICATION_CHECKLIST.md, and deliberately enable Store visibility.
API usage
Replace YOUR_USERNAME with your Apify account name. Keep the token in an environment variable rather than input JSON.
$curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~job-posting-feed-builder/runs?maxTotalChargeUsd=1" -H "Authorization: Bearer $APIFY_TOKEN" -H "Content-Type: application/json" --data-binary @examples/basic-input.json
Monitor the returned run ID and read its default Dataset. Store/API names and price configuration are account-side settings.
Verified platform example
The saved Console input performs a bounded HTTP run with one page and at most two result records. The source is a public scraping demonstration site; these records are sample content, not customer or current hiring data.
Use the saved default input in Apify Console, or the repository examples/basic-input.json. Replace its source with your own permitted URL and keep limits small for the first run.
The following unmodified record was returned by a successful Apify platform run on 2026-09-30:
[{"title": "Senior Python Developer","company": "Payne, Roberts and Davis","location": "Stewartbury, AA","employmentType": null,"workplaceType": null,"salaryMin": null,"salaryMax": null,"salaryCurrency": null,"salaryPeriod": null,"description": "Stewartbury, AA","qualifications": null,"datePosted": null,"validThrough": null,"jobUrl": "https://www.realpython.com/","sourceUrl": "https://realpython.github.io/fake-jobs/","extractedAt": "2026-09-30T11:47:59.981Z"}]
Introductory pricing: $1.00 per 1,000 job results, plus Apify platform usage. The standard Actor start event costs $0.00005 per GB (minimum one event). Empty output and duplicate records incur no primary result event; the start fee and any platform usage still apply. The SDK enforces the event spending limit; platform usage is billed separately.