# InfoJobs Scraper — Spain Jobs, Company & City (`noahadler/infojobs-spain-jobs`) Actor

Scrape InfoJobs Spain jobs from infojobs.net: title, company, city, contract, teleworking and offer URL. Salary only when InfoJobs publishes it on the card. Search URLs or keyword + geo. HTTP export, not a multi-board dump.

- **URL**: https://apify.com/noahadler/infojobs-spain-jobs.md
- **Developed by:** [Noah Adler](https://apify.com/noahadler) (community)
- **Categories:** Jobs, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 jobs

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## InfoJobs Spain Jobs — InfoJobs scraper for ofertas

**InfoJobs scraper** for Spain’s #1 job board (`infojobs.net`). Export **InfoJobs ofertas** and **InfoJobs Spain jobs** as one clean Dataset row per listing: title, company, city, contract, teleworking, salary when published, offer URL.

HTTP-only (no Playwright): the public search page hydrates `window.__INITIAL_PROPS__` with structured `offers[]`. Use **Store-like search URLs** or `keyword` + `geo`. Built as an **InfoJobs API**-style export for **empleo Spain scraper** workflows — not a bloated Indeed/LinkedIn multi-board.

> InfoJobs returns ~22 ofertas per search page (`code`, `title`, `description`, `city`, `companyName`, `companyLink`, `contractType`, `workday`, `teleworking`, `publishedAt`, `link`). This Actor paginates with `?page=N` until `maxItems`. **ES residential proxy is optional** locally; enable it on Apify if you hit 403/captcha.

**Why this Actor:** Clean schema · company + salary when the card has them · search URLs you can paste from the site · cheap PPE per Job · no Playwright · no email-harvesting claims

**Store icon:** official InfoJobs mark (see Actor icon in Console).

***

### Table of contents

1. [What this Actor does](#what-this-actor-does)
2. [What you get](#what-you-get)
3. [Features](#features)
4. [Input](#input)
5. [Output example](#output-example)
6. [Output fields](#output-fields)
7. [Quick start (API)](#quick-start-api)
8. [Use cases](#use-cases)
9. [Limitations](#limitations)
10. [FAQ](#faq)
11. [Keywords](#keywords)

***

### What this Actor does

| Step | Action |
|------|--------|
| 1 | Accept `searchUrls` and/or `keyword` + `geo` + `maxItems` |
| 2 | Map geo (Madrid, Barcelona, …) to InfoJobs `provinceIds` when known |
| 3 | `GET` the search HTML (`list.xhtml` or your pasted URL) |
| 4 | Parse `window.__INITIAL_PROPS__` → `offers[]` |
| 5 | Map each oferta → Dataset row (`title`, `companyName`, `city`, `url`, …) |
| 6 | Paginate `?page=N` until `maxItems` or an empty page |

**Input:** InfoJobs search URLs, keyword, geo, max items, proxy.\
**Output:** one Dataset row per job.

***

### What you get

**Job**

- `title`, `description`, `offerId`
- `city`, `contractType`, `workday`, `teleworking`

**Company / pay**

- `companyName`, `companyUrl`, `companyLogoUrl`
- `salary` when InfoJobs shows it on the card (otherwise `null`)

**Links / meta**

- `url` (absolute `https://www.infojobs.net/…`)
- `publishedAt`, `keyword`, `geo`, `searchUrl`, `scrapedAt`

***

### Features

| Capability | Detail |
|------------|--------|
| **Search URLs** | Paste the same InfoJobs search you use in the browser |
| **Keyword + geo** | e.g. `python` in `Madrid` → `provinceIds=33` |
| **No browser** | `curl_cffi` impersonate `chrome124` + `browserforge` headers |
| **Hydrated JSON** | `__INITIAL_PROPS__` offers — no CSS card scraping |
| **Proxy-aware** | Apify RESIDENTIAL + **ES** if antibot kicks in |
| **PPE-friendly** | One **Job** event per row — cheaper volume than a $3/1k blob |

***

### Input

| Field | Required | Description |
|-------|----------|-------------|
| `searchUrls` | No\* | `infojobs.net` search pages only (Indeed/LinkedIn rejected) |
| `keyword` | No\* | e.g. `python`, `comercial`, `enfermero` |
| `geo` | No | City or province (`Madrid`, `Barcelona`, `Valencia`) |
| `maxItems` | No | Cap (default 40, max 500) |
| `proxyConfiguration` | No | Default RESIDENTIAL + ES on Store Try it |

\*Provide **searchUrls** and/or **keyword** (geo optional).

#### Example — keyword + Madrid

```json
{
  "keyword": "python",
  "geo": "Madrid",
  "maxItems": 20,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": ["RESIDENTIAL"],
    "apifyProxyCountry": "ES"
  }
}
```

#### Example — paste an InfoJobs search URL

```json
{
  "searchUrls": [
    "https://www.infojobs.net/jobsearch/search-results/list.xhtml?keyword=python&provinceIds=33"
  ],
  "maxItems": 40
}
```

***

### Output example

```json
{
  "title": "Desarrollador Senior Python (Híbrido)",
  "description": "Como Desarrollador Senior Python, formarás parte de un equipo ágil…",
  "companyName": "DEVOTEAM",
  "companyUrl": "https://devoteam.ofertas-trabajo.infojobs.net",
  "city": "Madrid",
  "contractType": "Contrato indefinido",
  "workday": "Jornada completa",
  "teleworking": "Híbrido",
  "salary": null,
  "publishedAt": "2026-09-17T08:16:28Z",
  "url": "https://www.infojobs.net/madrid/desarrollador-senior-python-hibrido/of-i694a872eb84b0ba839fa86d91335d9",
  "offerId": "694a872eb84b0ba839fa86d91335d9",
  "keyword": "python",
  "geo": "Madrid",
  "scrapedAt": "2026-09-24T16:00:00Z",
  "error": false,
  "errorMessage": null
}
```

***

### Output fields

| Field | Notes |
|-------|--------|
| `title`, `companyName` | Must differ — title is the role, not the employer |
| `city` | City label from the card (not a salary string) |
| `url` | Absolute InfoJobs offer link (`code` in the path) |
| `salary` | Present only when InfoJobs publishes pay on the listing |
| `description` | Card/teaser text from `offers[]` (not a second detail crawl) |
| `error`, `errorMessage` | Hard failure vs empty search |

***

### Quick start (API)

```bash
curl "https://api.apify.com/v2/acts/noahadler~infojobs-spain-jobs/runs?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "keyword": "python",
    "geo": "Madrid",
    "maxItems": 20
  }'
```

Then fetch the default Dataset items URL from the run object. Schedule it for daily **InfoJobs ofertas** monitoring.

***

### Use cases

- **Recruiters** tracking open **InfoJobs Spain jobs** by skill and city.
- **Salary / contract research** when listings show `salary`, `contractType`, `teleworking`.
- **Lead gen without spam claims** — company name + company URL + offer URL (no email scrape).
- **Empleo Spain scraper** pipelines that already pull Fotocasa/Wallapop-style ES portals.
- **Challenger export** vs a generic $3/1k InfoJobs actor when you want a predictable schema.

***

### Limitations

- Public **search-list** fields only — does not open every JD apply form.
- Does **not** harvest emails, phones, or CVs.
- `salary` is often missing on InfoJobs cards (“Salario no disponible”).
- Structure depends on `__INITIAL_PROPS__`; a front-end rewrite may need an Actor update.
- 403/captcha → use Apify **RESIDENTIAL** + **ES**. After repeated WAF, the run fails honestly (no fake rows).
- Spain InfoJobs only (`infojobs.net`). Not InfoJobs Brasil, Indeed, or LinkedIn.

***

### FAQ

**Do I need Playwright?**\
No. v0.1 is HTTP + Chrome TLS impersonation against the public search HTML.

**Can I paste the URL from my browser?**\
Yes. `searchUrls` must be `infojobs.net` search pages.

**What does geo=Madrid do?**\
It sets `provinceIds=33` (InfoJobs Madrid). Barcelona is `9`, Valencia `49`, Sevilla `43`, etc.

**Why is salary null?**\
Many InfoJobs ofertas hide pay. The field is filled only when the hydrated JSON includes it.

**What happens on hard failure?**\
One Dataset row with `error: true` and an actionable `errorMessage` (often “enable RESIDENTIAL + ES”).

***

### Keywords

infojobs scraper, infojobs ofertas, infojobs spain jobs, infojobs api, empleo spain scraper, infojobs.net jobs, Madrid ofertas, Barcelona empleo, job listings Spain, recruiting lead generation, Apify Actor jobs

# Actor input Schema

## `searchUrls` (type: `array`):

Full infojobs.net search pages (e.g. https://www.infojobs.net/jobsearch/search-results/list.xhtml?keyword=python\&provinceIds=33 or /ofertas-trabajo/madrid). Ignored hosts (Indeed, LinkedIn) are rejected.

## `keyword` (type: `string`):

Search term as on InfoJobs (e.g. python, comercial, enfermero). Used when searchUrls is empty, or as metadata.

## `geo` (type: `string`):

Spanish city or province (e.g. Madrid, Barcelona, Valencia). Mapped to InfoJobs provinceIds when known.

## `maxItems` (type: `integer`):

Maximum job rows to collect (1–500). InfoJobs search pages return ~22 ofertas each.

## `proxyConfiguration` (type: `object`):

InfoJobs often works over HTTP without proxy. If you see 403/captcha, enable RESIDENTIAL and country ES.

## Actor input object example

```json
{
  "searchUrls": [
    "https://www.infojobs.net/jobsearch/search-results/list.xhtml?keyword=python&provinceIds=33"
  ],
  "keyword": "python",
  "geo": "Madrid",
  "maxItems": 40,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "ES"
  }
}
```

# Actor output Schema

## `jobs` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchUrls": [
        "https://www.infojobs.net/jobsearch/search-results/list.xhtml?keyword=python&provinceIds=33"
    ],
    "keyword": "python",
    "geo": "Madrid"
};

// Run the Actor and wait for it to finish
const run = await client.actor("noahadler/infojobs-spain-jobs").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "searchUrls": ["https://www.infojobs.net/jobsearch/search-results/list.xhtml?keyword=python&provinceIds=33"],
    "keyword": "python",
    "geo": "Madrid",
}

# Run the Actor and wait for it to finish
run = client.actor("noahadler/infojobs-spain-jobs").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchUrls": [
    "https://www.infojobs.net/jobsearch/search-results/list.xhtml?keyword=python&provinceIds=33"
  ],
  "keyword": "python",
  "geo": "Madrid"
}' |
apify call noahadler/infojobs-spain-jobs --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,noahadler/infojobs-spain-jobs"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/jR25h7thUAyph48vu/builds/UVhtMb3Inb582hoEB/openapi.json
