Python Jobs Scraper (python.org)
Pricing
$2.00 / 1,000 jobs
Python Jobs Scraper (python.org)
Python jobs scraper for the python.org Job Board: export every Python job listing (title, company, location, categories, date, link) to JSON, CSV or Excel. Public pages only, no login.
Pricing
$2.00 / 1,000 jobs
Rating
0.0
(0)
Developer
COMPASSLAB
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 hours ago
Last modified
Categories
Share
Get Python Jobs Scraper (python.org) job listings as clean JSON, CSV or Excel, from the Apify API or on a schedule. No login, about $2.00 for 1,000 jobs.
What does Python Jobs Scraper (python.org) do?
Python Jobs Scraper (python.org) extracts structured data from python.org. Collect job listings from the official Python.org Job Board so recruiters, job seekers and analysts can track Python hiring demand by company, location and job type. It works as an API for python.org data: run it from Apify Console, on a schedule, or from your own code, and get clean, typed JSON with numbers as numbers and dates in ISO 8601.
What you get
| Data | 7 fields per item: url, title, company, location, categories, postedAt, ... |
| Formats | JSON, CSV, Excel, HTML, or the Apify API |
| Price | $2.00 per 1,000 jobs, pay per result |
| Access | Public python.org data only: no login, no cookies, robots.txt respected |
Why use Python Jobs Scraper (python.org)?
- Recruiters and staffing agencies: see which companies are hiring, for which roles and where, and reach out first.
- Job aggregators and job boards: feed fresh, de-duplicated listings into your own board on a schedule.
- Market research and HR analytics: track hiring trends, salaries (where published), locations and remote share.
- AI agents and RAG: give an LLM live, structured job data instead of stale web pages.
Main features:
- Follows pagination up to
maxPagespages per start URL and stops atmaxItemsresults. - Filters:
keywords(keywords),excludeKeywords(exclude keywords),location(location),remoteOnly(remote only),postedAfter(posted after),onlyNewSinceLastRun(only new jobs since the last run), so you only get (and pay for) the results you need. - Polite by default: respects robots.txt, at most
maxConcurrencyparallel requests and a delay between requests. - Checks every result against field validators, so layout changes show up as clear data-quality warnings.
- Runs on the Apify platform: scheduling, API access, integrations, monitoring and datasets you can export.
What data can Python Jobs Scraper (python.org) extract?
| Field | Type | Description |
|---|---|---|
url | string | Absolute URL of the job detail page on python.org |
title | string | Job title |
company | string | Hiring company name |
location | string | Job location as displayed (city, region, country or remote) |
categories | array | Job types/categories tags shown on the listing (e.g. Web development, Data science) |
postedAt | ISO 8601 date | Posting date in ISO 8601 (YYYY-MM-DD), parsed from the listing's time element datetime attribute |
sourcePage | string | List page URL the job was found on |
How to scrape python.org
- Open Python Jobs Scraper (python.org) in Apify Console and go to the Input tab.
- Enter what to scrape (see the Input section below), for example the start URLs.
- Set Max items to the number of results you need.
- Click Start and wait for the run to finish.
- Download the results from the Output tab, or fetch them with the API.
How much will it cost to scrape python.org?
This Actor is priced per result: $2.00 per 1,000 results, with no extra charge for platform usage. That is about $2.00 for 1,000 jobs: 100 results cost $0.20 and 10,000 results cost $20.00. Set a maximum cost per run and the Actor stops when it is reached. Filters (companies, keywords, location, remote only, posted after, only new jobs) run before charging: you never pay for jobs you filtered out.
Input
See the Input tab for full configuration options.
| Field | Type | Required | Description |
|---|---|---|---|
keywords | array | no | Keep jobs whose title (and description, when 'Include description' is on) contains any of these words. Case-insensitive. |
excludeKeywords | array | no | Drop jobs whose title (and description, when on) contains any of these words. |
location | string | no | Keep jobs whose location contains this text, e.g. 'Berlin' or 'United States'. |
remoteOnly | boolean | no | Keep only remote jobs (location says Remote/Anywhere, or the source marks the job as remote). |
postedAfter | string | no | Keep jobs posted on or after this date (YYYY-MM-DD). Jobs without a date are kept. |
onlyNewSinceLastRun | boolean | no | For scheduled runs: skip jobs this same input already returned. The IDs are kept in a named key-value store; delete it to start over. |
maxItems | integer | no | Maximum number of items to return (0 = unlimited). |
startUrls | array | no | Python.org Job Board list pages to scrape (e.g. https://www.python.org/jobs/ or https://www.python.org/jobs/location/remote/). |
maxPages | integer | no | Maximum listing pages to follow per start URL (pagination). |
maxConcurrency | integer | no | Maximum parallel requests (politeness; 1-10). |
requestDelayMs | integer | no | Minimum delay between requests, in milliseconds (at least 250). |
Example input:
{"startUrls": [{"url": "https://www.python.org/jobs/"}],"maxItems": 30,"maxPages": 3,"maxConcurrency": 2,"requestDelayMs": 1000}
Output
You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. Example results from a real run:
[{"url": "https://www.python.org/jobs/8139/","title": "Senior Staff Engineer - Origination & New Products","company": "tem","location": "Remote (UK / EU), Remote (UK / EU), Remote (UK / EU)","categories": ["Back end","Cloud","Lead"],"postedAt": "2026-09-18","sourcePage": "https://www.python.org/jobs/"},{"url": "https://www.python.org/jobs/8137/","title": "Django Developer","company": "The Developer Society","location": "Birmingham, United Kingdom","categories": ["Back end","Cloud","Database","Web"],"postedAt": "2026-09-16","sourcePage": "https://www.python.org/jobs/"},{"url": "https://www.python.org/jobs/8136/","title": "ML Engineer","company": "Micro1","location": "Remote, Worlwide, Worldwide","categories": ["Back end","Machine Learning"],"postedAt": "2026-09-15","sourcePage": "https://www.python.org/jobs/"}]
Integrations and API
- Apify API: start a run and get the results in one HTTP request:
curl -X POST "https://api.apify.com/v2/acts/compass_lab~python-job-board-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \-H "Content-Type: application/json" -d '{"startUrls": [{"url": "https://www.python.org/jobs/"}], "maxItems": 30}'
- Python (
pip install apify-client):
from apify_client import ApifyClientclient = ApifyClient("<YOUR_APIFY_TOKEN>")run = client.actor("compass_lab/python-job-board-scraper").call(run_input={"startUrls": [{"url": "https://www.python.org/jobs/"}], "maxItems": 30})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
- JavaScript (
npm install apify-client):
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });const run = await client.actor('compass_lab/python-job-board-scraper').call({"startUrls": [{"url": "https://www.python.org/jobs/"}], "maxItems": 30});const { items } = await client.dataset(run.defaultDatasetId).listItems();
- Make, Zapier, n8n, Google Sheets, webhooks: use the Apify integrations (Integrations tab) to send each run's results where you need them, or to start a run from your workflow.
- Schedules: run it hourly, daily or weekly from Apify Console (Schedules) and always have fresh jobs.
Tips and advanced options
- Keep Max items and Max pages as low as you need: fewer pages means a faster, cheaper run.
- Raise Delay between requests if the site responds slowly; keep Max concurrency low to stay polite.
- Missing values are
null. Fields that often come back empty are listed in the run log as data-quality warnings. - Companies: type company names (
Stripe,Acme Inc) or paste job board URLs. Names are matched the way the platform writes them (Stripe Inc->stripe); a company that can't be found is named in the log and skipped. - Daily monitoring: schedule the Actor with Only new jobs since the last run on. Each run returns only jobs
it hasn't returned before for the same input. The IDs are kept in a named key-value store called
<actor-name>-seen-<id>(the run log prints its name); delete that store in Storage > Key-value stores, or change the input, to start over. - Filters before charging: keywords, exclude keywords, location, remote only and posted after are applied before a job is saved, so filtered-out jobs cost nothing.
FAQ, disclaimers and support
Is it legal to scrape python.org?
Checked 2026-09-30. robots.txt (https://www.python.org/robots.txt) only blocks a few named crawlers (HTTrack, puf, MSIECrawler, Nutch) and the paths /~guido/orlijn/ and /webstats/; /jobs/ and ?page=N are allowed for all other user agents. The python.org legal page (https://www.python.org/about/legal/) asserts PSF copyright over site content but contains no clause prohibiting automated access or scraping. Job ads are third-party content: this actor extracts listing metadata and links back to python.org; do not republish full job descriptions. Keep request rates low (defaults: 2 concurrent, 1 s delay).
Our Actors are ethical and do not extract any private user data, such as email addresses, gender, or location. They only extract what the user has chosen to share publicly. We therefore believe that our Actors, when used for ethical purposes by Apify users, are safe. However, you should be aware that your results could contain personal data. Personal data is protected by the GDPR in the European Union and by other regulations around the world. You should not scrape personal data unless you have a legitimate reason to do so. If you're unsure whether your reason is legitimate, consult your lawyers.
How many results can I get?
Up to maxItems per run (0 means no limit), as many as the source lists. Each result is one dataset item, and
you are only charged for items that are saved.
Can I run it on a schedule or from my own code?
Yes. Schedule it in Apify Console (Schedules), or call it with the run-sync-get-dataset-items endpoint or the
Python/JavaScript clients shown in Integrations and API above.
What are the limitations?
- The python.org job board is small (a few dozen open jobs), so runs are short and results limited to what is listed.
- Results come from the listing pages; details only shown on a job's own page (such as contact details) are not collected.
- If python.org changes its page layout, results may be missing until the Actor is updated; the run log warns when fields stop matching.
- Some listings have no posted date; it is
nullthen.
Where can I get help?
Report problems or ideas on the Issues tab. To call this Actor from your own code, see the API tab.
Related actors
Job Boards Suite: the same clean, typed output across sources, so you can combine them in one dataset.
| Actor | What it scrapes | Price |
|---|---|---|
| Breezy HR Jobs Scraper | Job listings from Breezy HR | $1.00 / 1,000 |
| Greenhouse Jobs Scraper | Job listings from Greenhouse | $1.60 / 1,000 |
| Lever Jobs Scraper | Job listings from Lever | $1.60 / 1,000 |
| Recruitee Jobs Scraper | Job listings from Recruitee | $1.00 / 1,000 |
| We Work Remotely Jobs Scraper | Job listings from We Work Remotely | $2.50 / 1,000 |
| Working Nomads Jobs Scraper | Job listings from Working Nomads | $2.50 / 1,000 |