# OpenAIRE Research Search Scraper (`searchapi/openaire-research-scraper`) Actor

Search OpenAIRE research outputs and export persistent IDs, titles, creators, dates, access status and descriptions.

- **URL**: https://apify.com/searchapi/openaire-research-scraper.md
- **Developed by:** [Search API](https://apify.com/searchapi) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.99 / 1,000 search results

This Actor is paid per event and usage. You are charged both the fixed price for specific events and for Apify platform usage.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## OpenAIRE Research Search Scraper

This Actor searches publications and datasets through the official public OpenAIRE Graph API v3. It supports full-text or title-only search, publication-date and access filters, publication review status, official sort modes, bounded pagination, and fair global limits across selected result types.

Only normalized research products are written to the dataset. Empty searches and failures use the fixed `OUTPUT` summary, so diagnostic placeholders never appear beside results.

### Example input

```json
{
  "query": "climate change",
  "searchField": "all",
  "resultTypes": ["publications", "datasets"],
  "fromPublicationDate": "2024-01-01",
  "openAccessOnly": true,
  "peerReviewedOnly": false,
  "sortBy": "relevance",
  "sortDirection": "desc",
  "maxItems": 20,
  "pageSize": 10,
  "maxPages": 5,
  "maxConcurrency": 2,
  "maxRequestRetries": 2,
  "requestTimeoutSecs": 30
}
```

`titleQuery` remains as a backward-compatible alias that automatically selects title-only search. It cannot be combined with `query`.

The Actor maps inputs directly to documented OpenAIRE parameters:

- `query` → `search` or `mainTitle`
- `resultTypes` → separate `type=publication` / `type=dataset` requests
- `fromPublicationDate`, Open Access, and peer-review filters
- `sortBy`, direction, `page`, and `pageSize`

Pagination metadata and returned product types must match the request before records are accepted. Each response is also checked for HTTP status, final official host/path, JSON content type, maximum size, valid JSON, and documented payload shape. Temporary network, 429, and 5xx failures use bounded exponential backoff.

### Output

Each row is a compact `openaire-research-product` record. It can include:

- OpenAIRE ID, product type, DOI or another persistent identifier
- title, subtitle, creators, compact description, and publication date
- access status, OA color, publisher, language, subjects, and countries
- citation count, funding/review flags, collection sources, and landing URL
- stable source request URL, source page/status, and scrape timestamp

Raw API objects, inventory metadata, internal actor metadata, signed URLs, credentials, and empty optional values are not stored. Records are deduplicated by OpenAIRE ID and buffered atomically. Multi-type results are interleaved page by page so publications cannot consume the global cap before datasets are considered.

### Current API references

- [OpenAIRE Graph API research products](https://graph.openaire.eu/docs/apis/graph-api/research-products/)
- [OpenAIRE paging](https://graph.openaire.eu/docs/apis/graph-api/paging/)
- [OpenAIRE filtering](https://graph.openaire.eu/docs/apis/graph-api/searching-entities/filtering-search-results/)

### Local verification

```text
npm ci
npm test
apify run --purge --input-file .actor/input.json
```

A clean no-result run leaves the dataset empty and sets `status: NO_RESULTS`. Invalid input fails before any API call.

# Actor input Schema

## `query` (type: `string`):

Words or phrase to search in research-product content or titles.

## `titleQuery` (type: `string`):

Backward-compatible title-only query. Do not combine with query.

## `searchField` (type: `string`):

Search all indexed content or only the main title.

## `resultTypes` (type: `array`):

Research product types to collect: publications, datasets, or both.

## `fromPublicationDate` (type: `string`):

Optional inclusive lower publication-date bound in YYYY-MM-DD format.

## `openAccessOnly` (type: `boolean`):

Return only products labeled Open Access by OpenAIRE.

## `peerReviewedOnly` (type: `boolean`):

Apply OpenAIRE's peer-reviewed filter to publication requests; dataset requests are unaffected.

## `sortBy` (type: `string`):

Official OpenAIRE sort field.

## `sortDirection` (type: `string`):

Ascending or descending result order.

## `maxItems` (type: `integer`):

Global record cap, fairly shared across selected types.

## `pageSize` (type: `integer`):

Official API page size for each selected type.

## `maxPages` (type: `integer`):

Maximum pages fetched for each selected result type.

## `maxConcurrency` (type: `integer`):

Concurrent type requests; at most two types are supported.

## `maxRequestRetries` (type: `integer`):

Retries for temporary network, rate-limit, and 5xx failures.

## `requestTimeoutSecs` (type: `integer`):

Timeout for each official API request.

## Actor input object example

```json
{
  "query": "climate",
  "searchField": "all",
  "resultTypes": [
    "publications",
    "datasets"
  ],
  "openAccessOnly": false,
  "peerReviewedOnly": false,
  "sortBy": "relevance",
  "sortDirection": "desc",
  "maxItems": 20,
  "pageSize": 10,
  "maxPages": 5,
  "maxConcurrency": 2,
  "maxRequestRetries": 2,
  "requestTimeoutSecs": 30
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "climate"
};

// Run the Actor and wait for it to finish
const run = await client.actor("searchapi/openaire-research-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "query": "climate" }

# Run the Actor and wait for it to finish
run = client.actor("searchapi/openaire-research-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "climate"
}' |
apify call searchapi/openaire-research-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,searchapi/openaire-research-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/z98lnVcHakiAmngec/builds/cY898PeBJOIYqR6k8/openapi.json
