# Wiley Analytical Science Scraper — 26 Journals, Full Abstracts (`trev0n/wiley-analytical-science-scraper`) Actor

Scrape Wiley's Analytical Science portal: 26 journals covering mass spectrometry, separation science, spectroscopy, proteomics and bioanalysis. Get full abstracts, complete author lists with ORCIDs, keywords, DOIs, volume, issue, dates and exact citation counts. Plus incremental monitoring.

- **URL**: https://apify.com/trev0n/wiley-analytical-science-scraper.md
- **Developed by:** [Paweł](https://apify.com/trev0n) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## 🔬 Wiley Analytical Science Scraper

🎯 **Pull thousands of analytical chemistry papers — with full abstracts, complete author lists, keywords and citation counts — from 26 Wiley journals in minutes.**

This scraper collects publication records from Wiley's Analytical Science collection: mass spectrometry, separation science, chromatography, spectroscopy, proteomics, bioanalysis, biotechnology and toxicology. For every paper you get the title, the **complete abstract**, every author (with affiliations and ORCID identifiers where they are on record), subject keywords, DOI, journal, volume, issue, publication dates and an exact citation count.

### 🚀 What Does It Do?

This scraper automatically searches Wiley's analytical science journals and collects **structured, ready-to-use data** for every matching paper. No manual browsing, no copy-pasting from search pages — set your filters and hit Start.

💡 **Three modes of operation:**

1. **🔍 Search Mode** — search by keyword across the whole collection or inside specific journals, with year and open-access filters
2. **📡 Watch Mode** — grab just the newest articles from the journals you follow, one quick pass per journal
3. **📋 Direct Link Mode** — hand it specific article links, DOIs or journal pages and it collects exactly those

### 👥 Who Is This For?

| 🏢 Use Case                                  | 💬 How It Helps                                                                                                       |
| -------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- |
| 🧪 **Analytical chemists & lab groups**      | Track every new method paper in your technique — LC-MS, Raman, NMR, electrophoresis — without living in a search page |
| 📚 **Research librarians**                   | Build and refresh subject bibliographies across 26 journals in one run, exportable straight into your catalogue       |
| 🎓 **PhD students & post-docs**              | Assemble a complete literature corpus for a review or thesis chapter, abstracts included, in a single afternoon       |
| 💊 **Pharma & biotech R\&D**                  | Monitor competitor and academic output in bioanalysis, proteomics and drug testing as it publishes                    |
| 🤖 **AI & data teams**                       | Feed a clean, abstract-rich scientific corpus into retrieval, embedding or literature-mining pipelines                |
| 📈 **Research analysts & scientometricians** | Measure output, citation impact and topic trends per journal, per year, per author                                    |
| 🏭 **Instrument vendors**                    | See which techniques and instruments are gaining traction, and who is publishing with them                            |

### ✨ Features

- 📄 **Full abstracts included** — every record ships with the complete abstract, not a truncated teaser
- 👥 **Complete author lists** — every author, plus affiliations and ORCID identifiers where they are on record
- 🏷️ **Subject keywords & topics** — automatic topic classification so you can slice a corpus by theme
- 📊 **Exact citation counts** — real citation and reference numbers, not rounded badges
- 📚 **26 journals, one run** — search the whole collection at once or narrow to a single title
- 🎛️ **Smart Filters** — keyword search, per-journal selection, year ranges, open-access-only, minimum citations, must-contain and must-not-contain word lists
- 🔓 **Open access flagging** — know instantly which papers your readers can actually open
- 📅 **Year-by-year collection** — split a journal's back catalogue into clean annual slices
- 🔄 **Incremental Monitoring** — on scheduled runs, get only what is new or changed, cutting the cost of a daily alert by 80–95%
- 🔔 **Instant Alerts** — push new papers straight to Slack, Discord, Telegram or your own endpoint
- 🗜️ **Compact mode** — a slim record shape for lightweight and AI workflows
- 🌐 **No proxy needed** — runs clean and fast out of the box, nothing extra to configure
- ⚡ **Fast & Scalable** — hundreds of complete records per minute, thousands per run
- 🧹 **Deduplication** — the same paper found twice is merged into one enriched record
- 📤 **Export Anywhere** — download results as JSON, CSV, Excel, or push to Google Sheets, Zapier, Make, or your CRM

### 🎛️ Filters & Options

| Option                                | What It Does                                                                       |
| ------------------------------------- | ---------------------------------------------------------------------------------- |
| 🔍 **Search query**                   | One or more phrases to search for; add several to run them all in one go           |
| 📚 **Journals**                       | Pick any of the 26 journals by name, or leave empty to search the whole collection |
| 📅 **Published from / to**            | Restrict to a year range — also the cleanest way to slice a large back catalogue   |
| 🔓 **Open access only**               | Keep only papers that are free to read                                             |
| ↕️ **Sort by**                        | Standard order, most relevant first, or oldest first                               |
| 📈 **Minimum citations**              | Drop papers below a citation threshold                                             |
| ✅ **Must contain**                   | Keep a paper only if these words appear in its title, abstract or keywords         |
| 🚫 **Must not contain**               | Drop papers containing any of these words                                          |
| 📊 **Exact citation counts**          | Add precise citation and reference numbers, plus author affiliations and ORCIDs    |
| 🏷️ **Keywords, topics & open access** | Add subject keywords, topic labels and open-access status                          |
| ✂️ **Abstract max length**            | Trim long abstracts to a set number of characters                                  |
| 📡 **Watch newest articles only**     | Skip searching and just pull each selected journal's latest articles               |
| 🧭 **Deep back-catalogue mode**       | Reach further back than a standard search when you need a journal's full history   |
| 🔄 **Incremental monitoring**         | Emit only new or changed papers across scheduled runs                              |
| 🔔 **Notifications**                  | Send new findings to Slack, Discord, Telegram or a webhook                         |
| 🗜️ **Compact output**                 | Emit only the essential fields                                                     |
| 🧹 **Drop empty fields**              | Leave blank fields out of the records entirely                                     |
| 🔢 **Max Results**                    | Control how many papers to extract per run                                         |
| 🔗 **Direct URLs**                    | Optionally provide specific article links, DOIs or journal pages to scrape         |

### 📦 What You Get (Output Fields)

Every paper includes:

#### Identification

| Field           | Example                                                                            |
| --------------- | ---------------------------------------------------------------------------------- |
| doi             | `10.1002/rcm.70138`                                                                |
| title           | `Hybrid CFD-DSMC Simulation of Ion Transport in Photoionization Mass Spectrometry` |
| url             | `https://analyticalsciencejournals.onlinelibrary.wiley.com/doi/10.1002/rcm.70138`  |
| source          | `wiley`                                                                            |
| discoverySource | `search-feed`                                                                      |
| searchQuery     | `mass spectrometry`                                                                |

#### Journal & Issue

| Field       | Example                                     |
| ----------- | ------------------------------------------- |
| journalName | `Rapid Communications in Mass Spectrometry` |
| journalId   | `10970231`                                  |
| issn        | `["1097-0231"]`                             |
| publisher   | `Wiley`                                     |
| volume      | `40`                                        |
| issue       | `19`                                        |
| pages       | `null`                                      |
| contentType | `article`                                   |
| articleType | `RESEARCH ARTICLE`                          |
| isEarlyView | `false`                                     |

#### Dates

| Field         | Example      |
| ------------- | ------------ |
| publishedDate | `2026-07-25` |
| onlineDate    | `2026-07-25` |
| coverDate     | `2026-10-15` |
| year          | `2026`       |

#### Authors

| Field                | Example                                                                                                                                                          |
| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| authors              | `[{ "name": "Zhiwei Wen", "given": "Zhiwei", "family": "Wen", "affiliation": "University of Science and Technology of China", "orcid": "0000-0002-1825-0097" }]` |
| authorCount          | `11`                                                                                                                                                             |
| correspondingAuthors | `[]`                                                                                                                                                             |

#### Content

| Field          | Example                                                                                                                           |
| -------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| abstract       | `ABSTRACT Rationale The ion motion within a mass spectrometer is governed by the coupling of gas dynamics and electric fields...` |
| abstractLength | `1369`                                                                                                                            |
| keywords       | `["Tandem mass spectrometry", "Electrospray ionization", "Ion transport"]`                                                        |
| topics         | `["Mass Spectrometry Techniques and Applications"]`                                                                               |
| subjects       | `[]`                                                                                                                              |
| language       | `en`                                                                                                                              |

#### Access & Impact

| Field          | Example                                       |
| -------------- | --------------------------------------------- |
| isOpenAccess   | `true`                                        |
| oaStatus       | `hybrid`                                      |
| license        | `http://creativecommons.org/licenses/by/4.0/` |
| citationCount  | `12`                                          |
| citationSource | `crossref`                                    |
| referenceCount | `58`                                          |
| scrapedAt      | `2026-07-28T09:41:06.000Z`                    |

### 📊 Example Output

```json
{
    "doi": "10.1002/rcm.70138",
    "title": "Hybrid CFD-DSMC Simulation of Ion Transport in Photoionization Mass Spectrometry: From Atmospheric Pressure to High Vacuum",
    "url": "https://analyticalsciencejournals.onlinelibrary.wiley.com/doi/10.1002/rcm.70138",
    "publisher": "Wiley",
    "source": "wiley",
    "discoverySource": "search-feed",
    "searchQuery": "mass spectrometry",
    "contentType": "article",
    "articleType": "RESEARCH ARTICLE",
    "journalName": "Rapid Communications in Mass Spectrometry",
    "journalId": "10970231",
    "issn": ["1097-0231"],
    "isbn": [],
    "volume": "40",
    "issue": "19",
    "pages": null,
    "publishedDate": "2026-07-25",
    "onlineDate": "2026-07-25",
    "coverDate": "2026-10-15",
    "year": 2026,
    "authors": [
        {
            "name": "Zhiwei Wen",
            "given": "Zhiwei",
            "family": "Wen",
            "affiliation": "National Synchrotron Radiation Laboratory, University of Science and Technology of China, Hefei China",
            "affiliations": [
                "National Synchrotron Radiation Laboratory, University of Science and Technology of China, Hefei China"
            ],
            "orcid": "0000-0002-1825-0097",
            "email": null,
            "isCorresponding": false
        },
        {
            "name": "Yang Pan",
            "given": "Yang",
            "family": "Pan",
            "affiliation": "National Synchrotron Radiation Laboratory, University of Science and Technology of China, Hefei China",
            "affiliations": [
                "National Synchrotron Radiation Laboratory, University of Science and Technology of China, Hefei China"
            ],
            "orcid": null,
            "email": null,
            "isCorresponding": false
        }
    ],
    "authorCount": 11,
    "correspondingAuthors": [],
    "abstract": "ABSTRACT\nRationale\nThe ion motion within a mass spectrometer is governed by the coupling of gas dynamics and electric fields. Therefore, a comprehensive understanding of ion transport from atmospheric pressure to the high-vacuum mass analyzer is crucial.\nMethods\nIn this work, a hybrid computational approach combining computational fluid dynamics and the direct simulation Monte Carlo method was developed.",
    "abstractLength": 1369,
    "keywords": ["Tandem mass spectrometry", "Ion transport", "Photoionization"],
    "topics": ["Mass Spectrometry Techniques and Applications"],
    "subjects": [],
    "language": "en",
    "isOpenAccess": true,
    "oaStatus": "hybrid",
    "license": "http://creativecommons.org/licenses/by-nc/4.0/",
    "pdfUrl": null,
    "citationCount": 12,
    "citationSource": "crossref",
    "referenceCount": 58,
    "references": [],
    "accessesCount": null,
    "altmetricScore": null,
    "isEarlyView": false,
    "scrapedAt": "2026-07-28T09:41:06.000Z"
}
```

### 📋 Dataset Views

The Apify Console gives you **6 ready-made table views** to quickly browse your results:

| View                          | What It Shows                                                                        |
| ----------------------------- | ------------------------------------------------------------------------------------ |
| 📊 **Overview**               | Title, journal, publication date, author count, citations, open-access flag and link |
| 📝 **Abstracts & keywords**   | Title, full abstract, keywords, topics and journal — the reading list view           |
| 👥 **Authors & affiliations** | Every author with their affiliation and ORCID, plus the paper they wrote             |
| 📈 **Citations & issue**      | Volume, issue, article type, citation and reference counts side by side              |
| 🔄 **Monitoring changes**     | What is new, updated or gone since the previous run                                  |
| 📋 **Full Details**           | Every single field — the complete dataset                                            |

### ❓ FAQ

**🤔 Which journals does it cover?**
All 26 titles in Wiley's Analytical Science collection — including Rapid Communications in Mass Spectrometry, Journal of Mass Spectrometry, Mass Spectrometry Reviews, Electrophoresis, Journal of Separation Science, Journal of Raman Spectroscopy, Magnetic Resonance in Chemistry, X-Ray Spectrometry, Biomedical Chromatography, Luminescence, Surface and Interface Analysis, Drug Testing and Analysis, PROTEOMICS, Biotechnology and Bioengineering and more.

**🤔 Do I really get the full abstract?**
Yes — the complete abstract comes with every record, for paywalled and open-access papers alike. Abstracts are public information; the scraper never touches full text behind a paywall.

**🤔 How do I collect a journal's entire back catalogue?**
Run it one year at a time using the year filters — that is the most reliable way to work through a large archive — and switch on deep back-catalogue mode. Each run adds to the same dataset, and duplicates are merged automatically.

**🤔 Why are keywords sometimes empty on brand-new papers?**
Subject keywords come from an open scholarly index that takes a few days to catalogue a freshly published paper. Titles, abstracts, authors and dates are there immediately; re-run later and the keywords fill in.

**🤔 Can I get an alert when something new is published?**
Yes. Turn on watch mode plus incremental monitoring, schedule the run daily, and add a Slack, Discord, Telegram or webhook destination — you will only be told about papers you have not seen before.

**🤔 Can I export the data?**
Yes — JSON, CSV, Excel, XML, HTML, RSS. You can also push data directly to Google Sheets, Zapier, Make, or any webhook endpoint.

**🤔 How often should I run this?**
For fresh data, run daily or weekly. You can schedule automatic runs on Apify with just a few clicks.

**🤔 Does it work with proxies?**
It does not need one — this scraper runs perfectly without any proxy, which keeps your costs down. If you prefer to route traffic through Apify's proxy service for very large runs, that option is there.

### 🛠️ Need Custom Filters or Features?

**I'm happy to customize this scraper for your specific needs!** 🤝

Whether you need:

- 🎯 Additional filters (specific authors, institutions, countries, funding bodies, instrument or technique mentions)
- 📊 Extra data fields or custom output formats
- 🔄 Integration with your CRM, reference manager, Google Sheets, or database
- ⏰ Scheduled scraping with automatic deduplication
- 🌐 Scraping from other scientific publishing platforms alongside Wiley

👉 **Don't hesitate to reach out via private message** — I respond quickly and I'm always open to building exactly what you need. No request is too small or too specific!

### ⚖️ Legal & Ethical Use

This scraper collects **only publicly available information** from Wiley Online Library. It does not access private data, bypass authentication, log in, or retrieve paywalled full-text content — titles, abstracts and bibliographic details are published openly by the publisher. Please use the data responsibly and in compliance with applicable laws and platform terms of service.

# Actor input Schema

## `searchQuery` (type: `array`):

What to search for across the Analytical Science journals. Add several queries to run them in one go (e.g. "mass spectrometry imaging", "capillary electrophoresis"). Leave empty to collect everything from the journals you pick below.

## `journalIds` (type: `array`):

Pick the journals to search. Leave empty to search the whole portal at once. Choosing one journal per run is the most reliable way to collect a large back catalogue.

## `yearFrom` (type: `integer`):

Only include publications from this year onwards. Setting a single year (same value in both boxes) is also the way to split a big journal into runs that fit. A year range is also how you reach the back catalogue, right down to the 1990s.

## `yearTo` (type: `integer`):

Only include publications up to and including this year. A year range is also how you reach the back catalogue, right down to the 1990s.

## `openAccessOnly` (type: `boolean`):

Keep only freely readable publications.

## `sortBy` (type: `string`):

Result order. To collect older material, set a year range instead — that is what reaches the archive.

## `minCitations` (type: `integer`):

Drop publications cited fewer times than this. Needs citation counts enabled (they are by default).

## `includeKeywords` (type: `array`):

Keep a publication only if its title, abstract or keywords contain at least one of these words.

## `excludeKeywords` (type: `array`):

Drop publications whose title, abstract or keywords contain any of these words.

## `startUrls` (type: `array`):

Scrape specific items — article links, journal pages, plain DOIs (10.1002/…), or a feed link you already have.

## `maxResults` (type: `integer`):

Stop after this many publications. Set 0 for no limit.

## `maxPages` (type: `integer`):

Stop each search after this many result pages. 0 means no page limit.

## `includeCitationCounts` (type: `boolean`):

Add exact citation and reference counts, plus author ORCIDs and affiliations where they are on record. Recommended.

## `includeOpenAlexData` (type: `boolean`):

Add subject keywords, topic classifications and open-access status. Keep this on — keywords are only available from this source. Very recently published articles may not have keywords yet.

## `abstractMaxLength` (type: `integer`):

Trim abstracts to this many characters. 0 keeps them in full.

## `monitorMode` (type: `boolean`):

Instead of searching, pull just the newest articles from each journal you picked. One quick request per journal — the cheapest way to stay current.

## `incrementalMode` (type: `boolean`):

Across scheduled runs, emit only publications that are new or changed since last time. Cuts the cost of daily literature alerts by 80–95%.

## `stateKey` (type: `string`):

Name for this monitoring baseline. Leave empty to derive one from your settings. Use different keys for different watch lists.

## `emitUnchanged` (type: `boolean`):

Also emit publications that haven't changed since the previous run.

## `emitExpired` (type: `boolean`):

Also emit publications that were in the previous run but are no longer found.

## `useCrossrefBackstop` (type: `boolean`):

Also enumerate each journal through an unlimited open index. Turn this on when you need a journal's complete history rather than a recent slice.

## `crossrefMailto` (type: `string`):

Optional. A contact address sent with metadata requests to get the faster service tier.

## `compact` (type: `boolean`):

Emit only the essential fields (title, authors, abstract, keywords, citations, link).

## `excludeEmptyFields` (type: `boolean`):

Leave empty fields out of the records entirely.

## `webhookUrl` (type: `string`):

POST the run results to this URL when monitoring finds something.

## `telegramBotToken` (type: `string`):

Send new publications to a Telegram chat.

## `telegramChatId` (type: `string`):

Chat that receives the Telegram messages.

## `discordWebhookUrl` (type: `string`):

Send new publications to a Discord channel.

## `slackWebhookUrl` (type: `string`):

Send new publications to a Slack channel.

## `notifyOnlyChanges` (type: `boolean`):

Stay silent when a monitoring run finds nothing new.

## `notificationLimit` (type: `integer`):

How many publications to list in a notification message.

## `includeRunSummary` (type: `boolean`):

Add the run statistics to webhook payloads.

## `proxyConfiguration` (type: `object`):

Proxy settings. This scraper does not need a proxy — leave it off unless you are running at very high volume.

## Actor input object example

```json
{
  "searchQuery": [
    "mass spectrometry"
  ],
  "openAccessOnly": false,
  "sortBy": "default",
  "minCitations": 0,
  "maxResults": 100,
  "maxPages": 0,
  "includeCitationCounts": true,
  "includeOpenAlexData": true,
  "abstractMaxLength": 0,
  "monitorMode": false,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "useCrossrefBackstop": false,
  "compact": false,
  "excludeEmptyFields": false,
  "notifyOnlyChanges": true,
  "notificationLimit": 10,
  "includeRunSummary": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "searchQuery": [
        "mass spectrometry"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("trev0n/wiley-analytical-science-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "searchQuery": ["mass spectrometry"] }

# Run the Actor and wait for it to finish
run = client.actor("trev0n/wiley-analytical-science-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "searchQuery": [
    "mass spectrometry"
  ]
}' |
apify call trev0n/wiley-analytical-science-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,trev0n/wiley-analytical-science-scraper"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/wAy5GtdfeU5pZiVZ4/builds/DkxJpGIh2Edh7p1Mb/openapi.json
