DOAB Open Access Books Scraper avatar

DOAB Open Access Books Scraper

Pricing

from $2.00 / 1,000 book scrapeds

Go to Apify Store
DOAB Open Access Books Scraper

DOAB Open Access Books Scraper

Scrape public scholarly open-access book metadata from the Directory of Open Access Books REST catalog.

Pricing

from $2.00 / 1,000 book scrapeds

Rating

0.0

(0)

Developer

Muhammad Afzal

Muhammad Afzal

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

13 hours ago

Last modified

Categories

Share

Scrape public scholarly open-access book metadata from the Directory of Open Access Books (DOAB) REST catalog. Use it for research discovery, library collection analysis, open-access bibliographies, and academic book lead lists.

What it extracts

FieldDescription
title, subtitleBook title metadata
authorsAuthors and editors
publisher, publicationYearPublication metadata
description, subjects, languagesDiscovery metadata
isbn, doi, handlePersistent identifiers
license, fullTextUrlRights and public access link when indexed
recordUrlDOAB record URL

Input

Use query for full-text or fielded DOAB syntax, such as history AND migration or dc.title:"open science". Use handle for one exact record. maxResults, maxPages, and pageSize bound the run. The Console prefill requests 5 books from one page so Apify's automated quality test remains inexpensive and comfortably inside five minutes; API defaults remain 25 books and 3 pages.

Output

One normalized book object is written per unique result to the default dataset. Run diagnostics are written to SUMMARY, including empty results, API failures, pages completed, warnings, and delivered count.

{
"title": "Open Science",
"authors": ["Jane Doe"],
"publisher": "Open Press",
"publicationYear": 2024,
"subjects": ["Research"],
"handle": "20.500.12854/123",
"recordUrl": "https://directory.doabooks.org/handle/123",
"source": "DOAB"
}

Pricing

The primary event is book-scraped at $0.002 per unique book record, plus the configured $0.001 actor-start event. A one-result run therefore costs up to $0.003 in event charges before any platform usage or optional account charges.

Reliability and scope

The actor uses DOAB's public REST search endpoint and hydrates records through bounded metadata requests instead of asking the catalog to expand a large search response. It does not access restricted content, bypass challenges, or download book files. DOAB metadata and links depend on what publishers and repositories have indexed. A zero-result run is reported as empty; HTTP or transport failures are reported in SUMMARY and as a failed run.

Use the metadata and linked resources according to DOAB, publisher, repository, and applicable license terms. This actor indexes public catalog metadata; it does not grant rights to copyrighted content.

Use cases

  • Supply structured public data to AI, RAG, enrichment, or evaluation workflows.
  • Schedule repeatable collection and export results to downstream workflows.
  • Run a one-off research job and export the structured result as JSON, CSV, Excel, XML, or RSS from Apify.
  • Schedule the same input to monitor changes over time and send completed datasets to a webhook or integration.
  • Feed schema-shaped records into a database, spreadsheet, BI tool, or AI workflow with the source URL retained for verification.

Run DOAB Open Access Books Scraper with the Apify API

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('muhammadafzal/doab-open-access-books-scraper').call({
"query": "open access",
"handle": "",
"maxResults": 25,
"maxPages": 3,
"pageSize": 25,
"includeMetadata": true
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);

You can also run the Actor from Apify Console, schedules, webhooks, the REST API, Make, Zapier, n8n, or the hosted Apify MCP server.

Support

When reporting a problem, include the Actor run ID, a redacted input, the expected result, and a small public example URL when applicable. Do not post API tokens, cookies, credentials, or personal data in an issue.

Frequently asked questions

Can I schedule DOAB Open Access Books Scraper?

Yes. Use an Apify schedule to run the same saved input at a chosen interval, then connect a webhook or integration to process the dataset when the run finishes.

How should I test a new input?

Begin with the prefilled example or a small limit. Confirm that the output fields, source coverage, runtime, and live charges match your workflow before increasing the scope.

How do I export the results?

Open the run's default dataset in Apify Console and export JSON, CSV, Excel, XML, or RSS. Applications can retrieve the same records through the Apify API client or REST dataset endpoint.

Can an AI agent call this Actor?

Yes. Add muhammadafzal/doab-open-access-books-scraper through the hosted Apify MCP server or call it through the API. The Actor's input and dataset schemas help agents construct valid requests and interpret returned records.

  1. Define the smallest useful scope. Choose a representative public URL, query, identifier, or filter and keep the first result limit low.
  2. Run and inspect. Check the run log, dataset item count, field coverage, source URLs, and live event or usage charges.
  3. Validate downstream assumptions. Confirm nullable fields, deduplication keys, timestamps, and any locale-specific formats before importing records into another system.
  4. Scale gradually. Increase limits or scheduling frequency only after the small run behaves as expected. Use Apify's maximum-cost and timeout controls to bound large jobs.
  5. Monitor changes. Keep a small known-good input as a canary. If the source layout or API changes, compare the new dataset with a previously validated run and report the run ID when requesting support.

For recurring workflows, store the exact Actor input with your pipeline configuration. This makes runs reproducible and helps distinguish a source-data change from an input change.