DOAJ Scraper: Open Access Journals & Articles avatar

DOAJ Scraper: Open Access Journals & Articles

Pricing

from $0.37 / 1,000 record scrapeds

Go to Apify Store
DOAJ Scraper: Open Access Journals & Articles

DOAJ Scraper: Open Access Journals & Articles

Scrape the Directory of Open Access Journals: journal title, ISSN, publisher, subject, licence, article processing charges and peer-review process. For OA policy analysis.

Pricing

from $0.37 / 1,000 record scrapeds

Rating

0.0

(0)

Developer

Arman Hossain

Arman Hossain

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

9 days ago

Last modified

Share

DOAJ Scraper: Journal metadata and per-currency APC pricing, or article abstracts and DOIs, DOAJ's own query syntax passed through

DOAJ Scraper pulls vetted open-access journals out of the Directory of Open Access Journals, title, ISSN, publisher, country, subject classification, licence, article processing charges and peer-review process. Flip one switch and it searches DOAJ's article index instead, returning abstracts, DOIs and full-text links.

APC data is the reason people come here. Very few sources publish article processing charges in structured, per-currency form; DOAJ does, and this Actor hands it to you as a number and a currency code rather than a sentence buried in a PDF.

Agent skill: SKILL.md

https://api.apify.com/v2/key-value-stores/t7YoTxpZEJOWvw4Ug/records/doaj-journals-scraper.md

What you get

Journals (resourceType: "journals": the default)

Output fieldMeaning
idDOAJ internal journal ID
titleJournal title
issn, eissnPrint ISSN and electronic ISSN (either may be null)
publisher, countryPublisher name and its ISO-2 country code
subjectsLibrary of Congress subject labels, e.g. ["Physics"]
languagePublication languages as ISO codes
licenseLicence types, e.g. ["CC BY", "CC BY-NC-ND"]
apcAmount, apcCurrencyHeadline article processing charge and its currency
apcPricesEvery published price/currency pair, not just the headline one
hasApctrue / false, does the journal charge at all
peerReviewProcesse.g. ["Anonymous peer review"]
oaStartYearYear the journal became fully open access
urlThe journal's own homepage
lastUpdatedWhen DOAJ last revised the record
resourceType, query, scrapedAtWhich index, which query produced the row, and the run timestamp

Articles (resourceType: "articles")

id, title, doi, issn, eissn, journalTitle, publisher, country, subjects, language, keywords, authors, abstract, year, month, volume, issue, url (full-text link), lastUpdated, plus the same resourceType / query / scrapedAt.

A RUN_SUMMARY record in the key-value store holds per-run counts, the filters used, any query that failed, and how many queries hit DOAJ's 1000-record API ceiling.

Common use cases

1. Compare publishing costs across journals. Pull everything that charges, then sort on apcAmount.

{
"searchQueries": ["bibjson.apc.has_apc:true AND bibjson.subject.term:\"Medicine\""],
"maxResults": 1000
}

2. Check OA compliance for a funder mandate. Funders that require CC BY and no reader-side paywall can be checked directly against license and hasApc.

{
"searchQueries": ["bibjson.publisher.name:\"Elsevier\""],
"subjects": ["Medicine"],
"maxResults": 500
}

3. Analyse the open-access landscape. Diamond OA, free to both author and reader, is hasApc: false.

{
"searchQueries": ["bibjson.apc.has_apc:false"],
"maxResults": 1000
}

Quick start

Simplest possible run:

{
"searchQueries": ["machine learning"]
}

Two subject-filtered queries, capped:

{
"searchQueries": ["deep learning", "bibjson.subject.term:\"Physics\""],
"resourceType": "journals",
"subjects": ["Physics", "Computer science"],
"maxResults": 300
}

Article metadata for a RAG corpus:

{
"searchQueries": ["climate adaptation"],
"resourceType": "articles",
"maxResults": 1000
}

Input

FieldTypeDefaultNotes
searchQueriesarray-Required. One search per entry. Plain words or DOAJ field syntax. Each entry is paginated independently, so three queries can return up to 3000 records.
resourceTypestringjournalsjournals or articles. One run searches one index, the output shape differs between them.
subjectsarray[]Client-side filter on the subjects field, case-insensitive substring. Empty = keep everything.
maxResultsinteger200Total cap across all queries. 0 = no cap, but see the 1000-record API ceiling below.

Which combinations make sense. subjects is a post-filter, so it can only narrow what a query already returned, if you want subject-scoped results from the server, put it in the query itself as bibjson.subject.term:"Physics" and DOAJ will do the filtering before pagination. Use the input filter when you want one broad query sliced several ways. maxResults is applied across queries in order, so put your most important query first.

DOAJ query syntax

QueryFinds
*Everything (capped at 1000)
bibjson.title:"Nature"Title match
bibjson.publisher.name:"Elsevier"All journals from a publisher
bibjson.apc.has_apc:falseJournals with no author-side charge
bibjson.subject.term:"Physics"Subject-classified journals
issn:2731-3395Lookup by ISSN
bibjson.apc.has_apc:true AND bibjson.publisher.country:GBBoolean combination

Output example

{
"resourceType": "journal",
"query": "bibjson.apc.has_apc:true",
"id": "e5b4b3d3f1a04a.",
"title": "Communications Engineering",
"issn": null,
"eissn": "2731-3395",
"publisher": "Nature Portfolio",
"country": "GB",
"subjects": ["Engineering (General). Civil engineering (General)"],
"language": ["EN"],
"license": ["CC BY", "CC BY-NC-ND"],
"apcAmount": 2290,
"apcCurrency": "USD",
"apcPrices": [
{ "price": 1990, "currency": "EUR" },
{ "price": 2290, "currency": "USD" },
{ "price": 1650, "currency": "GBP" }
],
"hasApc": true,
"peerReviewProcess": ["Anonymous peer review"],
"oaStartYear": 2022,
"url": "https://www.nature.com/commseng/",
"lastUpdated": "2026-01-15T10:33:48Z",
"scrapedAt": "2026-08-06T12:00:00.000Z"
}

RUN_SUMMARY:

{
"resourceType": "journals",
"queriesRequested": 2,
"queriesFailed": 0,
"failures": [],
"recordsSaved": 300,
"queriesTruncatedByApiCeiling": 1,
"apiResultCeiling": 1000,
"filters": {
"searchQueries": ["deep learning", "bibjson.subject.term:\"Physics\""],
"resourceType": "journals",
"subjects": ["Physics"],
"maxResults": 300
},
"finishedAt": "2026-08-06T12:00:04.512Z"
}

Limits and behaviour

  • DOAJ caps the API at 1000 records per query. Ask for record 1001 and DOAJ answers HTTP 400 and points you at its public data dump. The Actor stops one page short of that boundary, logs a warning naming the query and its true total, and counts it in RUN_SUMMARY.queriesTruncatedByApiCeiling. To go deeper, split one broad query into several narrower ones, three subject-scoped queries fetch 3000 records where one broad query fetches 1000. For a full mirror of DOAJ, use their public data dump instead of any API.
  • A bad query doesn't kill the run. Malformed query syntax gets HTTP 400 from DOAJ; that query is recorded in RUN_SUMMARY.failures and the rest continue. The Actor only errors out if every query fails.
  • Partial results are kept. If a query fails on page 4, the 300 records from pages 1-3 stay in the dataset and the failure is still reported.
  • Transient errors are retried. 429 and 5xx get three attempts with linear backoff. Malformed queries and 404s are treated as fatal immediately, since retrying them cannot help.
  • APC prices are per-currency. DOAJ publishes a list, so apcAmount/apcCurrency carry the preferred single price (USD, then EUR, then GBP, then whatever is first) and apcPrices carries the complete list. Never compare apcAmount across journals without checking apcCurrency.
  • Public data only. No authentication, no personal data, no access-control bypass.

API example

curl -X POST "https://api.apify.com/v2/acts/arman-bd~doaj-journals-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"searchQueries": ["bibjson.subject.term:\"Physics\""],
"resourceType": "journals",
"maxResults": 100
}'

JavaScript example

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('arman-bd/doaj-journals-scraper').call({
searchQueries: ['bibjson.apc.has_apc:true'],
subjects: ['Medicine'],
maxResults: 500,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
const charging = items.filter((j) => j.hasApc).sort((a, b) => b.apcAmount - a.apcAmount);
for (const j of charging.slice(0, 10)) {
console.log(`${j.apcAmount} ${j.apcCurrency}, ${j.title} (${j.publisher})`);
}

FAQ

Do I need a proxy? No. Proxy configuration is not required to run this Actor.

Do I need a DOAJ account or API key? No. You supply no credentials.

What happens if DOAJ is unavailable? 429s and 5xx errors are retried with backoff. If a query still fails. It is recorded in RUN_SUMMARY.failures and the run continues with the remaining queries.

Can I schedule it? Yes. DOAJ records carry lastUpdated, so a scheduled run plus a diff on id + lastUpdated gives you a clean change feed.

Why is issn null on some journals? Many journals are electronic-only and have no print ISSN. Use eissn, one of the two is always present.

Why is apcAmount null when hasApc is true? A handful of DOAJ records flag a charge without publishing the figure. apcPrices will be an empty array in that case; the journal's own url is the place to look.

Can I get more than 1000 records for one search? Not through the API, that's DOAJ's limit, not this Actor's. Split the search into narrower queries, or use DOAJ's public data dump.

Can I integrate it with something else? Yes, Apify API, client libraries, webhooks, scheduled runs, dataset exports (JSON/CSV/Excel) or MCP. Output is structured JSON.