OSTI DOE Research Scraper avatar

OSTI DOE Research Scraper

Pricing

from $4.52 / 1,000 results

Go to Apify Store
OSTI DOE Research Scraper

OSTI DOE Research Scraper

Searches OSTI.gov by keyword and optional document type, then returns each publication as a flat row with title, authors, date, DOI, and abstract. Supports up to 1,000,000 records per run.

Pricing

from $4.52 / 1,000 results

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

10 days ago

Last modified

Share

ParseForge

OSTI DOE Research Scraper

Scrape DOE-funded research publications from OSTI.gov by keyword, document type, and sort order, up to a million per run. Every record includes the title, author, publication date, DOI, and abstract. No API key or registration. Export to CSV, JSON, Excel, or XML.

The U.S. Department of Energy's OSTI.gov repository holds millions of technical reports, journal articles, patents, and datasets, but browsing the web interface page by page is slow and manual. This Actor searches the public OSTI.gov database directly, applies your filters, and returns every matching publication in a structured dataset. It is built for researchers, analysts, and librarians who need to gather DOE-funded science output in bulk.

Who uses itWhat they scrape OSTI.gov for
Energy policy analystsTracking which DOE-funded technologies are producing the most published research output each quarter.
Research librariansBuilding a bibliography of all DOE technical reports on a specific topic for a literature review.
Grant administratorsVerifying that funded projects have produced the required public deliverables in OSTI.gov.
Science journalistsFinding recent DOE-funded studies and patents to support an article on clean energy breakthroughs.

What it does

This Actor searches OSTI.gov by keyword and optional document type, then returns each publication as a flat row with its title, authors, date, DOI, and abstract.

  • ๐Ÿ“„ Full-text keyword search: supply any term, from 'solar cell' to 'nuclear fusion', and the Actor searches the entire OSTI.gov corpus.
  • ๐Ÿ“ Document type filter: limit results to Technical Reports, Journal Articles, Conference papers, Datasets, Theses, Patents, or Books.
  • ๐Ÿ“Š Sort control: order results by relevance, newest first, or oldest first to match your workflow.
  • ๐Ÿ“ฆ Bulk collection: set a maximum from 1 to 1,000,000 publications and let the Actor paginate through every result.

Results export to CSV, JSON, Excel, or XML, or straight from the API.

What you can do with OSTI.gov data

๐Ÿ“ˆ Monitor DOE research output by topic.

An energy analyst runs a weekly search for 'battery storage' filtered to Journal Articles and sorted newest first, then charts publication volume over time to spot research trends.

๐Ÿ“š Build a systematic literature review dataset.

A graduate student searches 'perovskite solar cell' across all document types, collects 500 records, and exports the titles and abstracts to a reference manager for screening.

๐Ÿ” Audit grant deliverables.

A program officer searches for a specific DOE award number, filters to Technical Reports, and checks that every required final report appears in the results.

๐Ÿ“ฐ Find DOE-funded patents for a news story.

A journalist searches 'carbon capture' with the Patent filter, sorts by newest first, and reviews the latest DOE-supported inventions for a feature article.

Why choose this scraper

What you get
No API key or registrationOSTI.gov's public search is read directly; you do not need to sign up for a developer account or manage OAuth tokens.
Structured outputEvery publication lands in a fixed schema ready for analysis in Excel, Python, or a database.
Server-side filteringDocument type and sort order are sent to OSTI.gov, so you download only the records you need.
Runs on Apify infrastructureSchedule recurring searches, get webhook notifications, and store results in Apify's cloud without managing your own servers.

How it compares

This Actor and the DOE OSTI.gov Technical Reports Scraper both search the same public repository. The table below compares the capabilities each listing describes.

FeatureParseForgeDOE OSTI.gov Technical Reports Scraper
Keyword searchYesYes
Document type filter (Technical Report, Journal Article, Patent, etc.)YesNot listed
Sort by relevance, newest, or oldestYesNot listed
Up to 1,000,000 records per runYesNot listed
Abstract text in outputYesNot listed

Configure the run

Drive the Actor with a keyword query and an optional document type filter. Sorting and the maximum item count control the shape of the output, and every filter is applied on the server side so only matching publications reach your dataset. The Input tab lists every parameter.

A first run with the defaults:

{
"query": "renewable energy",
"maxItems": 10
}

A larger pull:

{
"query": "renewable energy",
"maxItems": 200
}

Pricing

Pay-per-result: $0.005 per result collected. You pay only for the results written to your dataset.

Results collectedApproximate cost
100 results$0.50
1,000 results$5.00
10,000 results$50.00

New Apify accounts start with $5 in free credit.

Free users

Free-plan runs return up to 10 results as a preview. Upgrade your Apify plan to collect up to 1,000,000 results per run.

Run it

  1. Create a free Apify account with $5 in credit.
  2. Open the OSTI DOE Research Scraper.
  3. Set your inputs and any filters, then click Start.
  4. Export the results as CSV, Excel, JSON, or XML from the Dataset tab.

Run it programmatically through the Apify API (run-sync-get-dataset-items) or the ApifyClient for JavaScript and Python.

Use with AI agents (MCP)

Give an AI agent live access to OSTI.gov through the Model Context Protocol. Add the Actor to Claude, Cursor, or any MCP client:

$claude mcp add --transport http apify "https://mcp.apify.com?tools=parseforge/osti-doe-research-scraper"

Then prompt it in plain language to run the scraper and read back the results.

Troubleshooting

Why am I getting no results?

Check that your keyword is spelled correctly and is not overly specific. Try removing the Product Type filter and setting it to 'All Types'. A very narrow query combined with a restrictive document type filter can return zero matches.

The run stopped before reaching my max items limit.

OSTI.gov returned fewer total results than your max items setting. The Actor collects every match available for your query and stops when the result set is exhausted.

Some fields are empty in my dataset.

OSTI.gov records vary in completeness. Older technical reports may lack a DOI, and some records may not have an abstract. Empty fields reflect what OSTI.gov provides for that publication.

The Actor is running slowly.

The Actor respects OSTI.gov's servers by pacing requests. Large max items values will take longer. Reduce the max items or run the Actor on a higher-memory Apify plan if you need faster throughput.

I got an error or timeout.

OSTI.gov may be temporarily slow or under maintenance. Retry the run after a few minutes. If the problem persists, check the Apify run log for the specific HTTP status code and contact support.

FAQ

QuestionAnswer
Do I need an OSTI.gov account or API key to use this scraper?No. The Actor reads the public OSTI.gov search interface directly. You do not need to register, obtain a key, or sign any agreement.
What data does each publication row include?Each row returns the publication title, author list, publication date, document type, DOI when available, the abstract or description text, and the OSTI.gov record URL.
Can I search by author name or DOE national lab?The keyword field performs a full-text search across the OSTI.gov index, which includes author names and lab affiliations. Type an author name or lab name into the query field to find their work.
How many publications can I collect in one run?You set the maximum, from 1 up to 1,000,000. The Actor will paginate through OSTI.gov search results until it reaches your limit or exhausts the result set.
What export formats are supported?You can export your dataset as CSV, JSON, Excel, XML, or RSS from the Apify platform, or push it directly to a cloud storage integration.
Can I schedule this scraper to run automatically?Yes. Apify's scheduler lets you run the Actor hourly, daily, or weekly, so you can monitor new DOE publications as they appear.
Does this scraper get the full text of the publication?No. The Actor collects the metadata record (title, authors, abstract, DOI, and URL). The full-text PDF or article is usually linked from the OSTI.gov record page.
Is this an official DOE product?No. This is an independent scraper built on the public OSTI.gov search interface. It is not affiliated with or endorsed by the U.S. Department of Energy.
What happens if my search returns no results?The run completes with an empty dataset. Try broadening your keyword, removing the document type filter, or checking for typos in the query.
Can I filter by date range?The current version does not have explicit date-range inputs, but you can sort by newest or oldest first and set a max items limit to approximate a date window.

Browse the full ParseForge collection for more scrapers.

๐Ÿ†˜ Need help? Email parseforge@protonmail.com with your run ID, your input, and what you expected.

โš ๏ธ Disclaimer. This Actor is unofficial and is not affiliated with, endorsed by, or sponsored by U.S. Department of Energy Office of Scientific and Technical Information. It collects only publicly available data. You are responsible for using the collected data in compliance with the source's terms of service and applicable data-protection laws, including GDPR, CCPA, and PIPL. Do not use it to collect personal data unlawfully.